Skip to content
[bmdpat]

Prompt Library

104 free prompts for local LLMs and AI agents.

Every prompt below ships with a blog post that explains when to use it, tested on hardware we own. Copy one straight into your coding agent, or grab the whole set as one file. New posts add their prompts here automatically.

Get all 104 prompts as a copy-paste pack

One markdown file with every prompt on this page, ready to paste into your coding agent. You also join the Local AI Lab Notes: at most one evidence-backed note on Friday when there is something worth sharing.

Get the requested artifact now, then at most one evidence-backed Local AI Lab Note on Friday when there is something worth sharing. One-click unsubscribe. No sponsored placements. Privacy.

Every prompt on this page stays free, the pack is the convenient all-in-one copy. Single opt-in. Unsubscribe anytime. Privacy.

local-llm

Bonsai 27B vs Gemma 4 E4B: Which Should You Run?

This prompt evaluates your hardware specs and workload requirements to recommend either Ternary Bonsai 27B or Gemma 4 E4B.

27 lines998 charsSource post
Show prompt
Role:
Local AI Systems Architect

Context:
Evaluate local model deployments based on hardware limits, input modalities, and reasoning requirements.

Inputs:
- Available VRAM: __ GB
- Required Modalities: __ (Text, Image, Audio)
- Required Context Window: __ tokens
- Host Runtime: __ (Ollama, llama.cpp)

Task:
1. Check Available VRAM, Required Modalities, and Required Context Window.
2. If Required Modalities includes Audio, or only the 4.5 GB Q4_0 weight file fits, select Gemma 4 E4B.
3. If the 7.2 GB deployed footprint fits, Required Modalities has no Audio, and math or 262K context is required, select Ternary Bonsai 27B. Do not select it for a 6 GB per-app iOS budget.
4. Output the matching CLI command for the Host Runtime.

Output:
- Selected Model: Ternary Bonsai 27B or Gemma 4 E4B
- Rationale: Memory and modality fit
- Launch Command: CLI command

Constraints:
- Use only facts from published model cards.
- Do not assume audio on Ternary Bonsai 27B.
- Never exceed Available VRAM.

A second machine will not run a bigger model

It works out what a second machine actually buys you under a router that sends whole requests to one node, so you do not buy hardware for a ceiling it cannot raise.

31 lines1107 charsSource post
Show prompt
Role:
You are checking whether a second machine raises the size of model I can run.

Context:
Machine A card and VRAM: [ ]
Machine B card and VRAM: [ ]
Model and quant I want to run: [ ]
Requests I send at the same time, typical: [ ]

Task:
1. State the largest model each machine holds on its own.
2. State my ceiling across both machines, given that the router does not pool
   memory and does not split one request between nodes.
3. Say whether machine B can hold the model I named. If it cannot, say that
   B cannot serve that work at all.
4. Say what routing buys me at my stated request count.
5. Name what I would have to change to raise the ceiling.

Output:
- Machine A ceiling:
- Machine B ceiling:
- Combined ceiling:
- Can B serve the named model:
- What routing buys at my request count:
- What would actually raise the ceiling:

Constraints:
- Treat total VRAM across machines as a number that does not apply here.
- If the two cards differ in speed, say that a job-counting scheduler will
  send work to the slower one.
- Do not recommend buying hardware. Answer only from the numbers given.

q4_k_m vs q8_0: which GGUF quant should you use?

Helps you decide between specific GGUF quantization levels based on your hardware constraints.

19 lines792 charsSource post
Show prompt
Role:
Hardware-aware LLM Engineer

Context:
I am running local LLM inference and need to choose between different GGUF quantization levels (e.g., Q4_K_M, Q6_K, Q8_0) for a specific model size.

Task:
1. Analyze my available VRAM and the model parameter count I am using.
2. Compare the estimated memory footprint of Q4_K_M vs Q8_0 for this model.
3. Recommend a quantization level that balances generation speed (tokens per second) and model perplexity.

Output:
- A recommendation for which quant to use.
- An estimate of how much VRAM will be left for context window/KV cache.
- A brief justification based on the trade-off between precision and speed.

Constraints:
- Prioritize preventing Out-of-Memory (OOM) errors.
- Assume I want to maximize tokens per second unless I state otherwise.

My Local Blog Writer Drops Private Lines First

It asks a coding assistant to test source filtering and output rejection with synthetic notes.

8 lines478 charsSource post
Show prompt
Role: Review my local drafting pipeline.
Context: It turns private daily notes into public drafts.
Task: Inspect source filtering and the final publication check.
Output: Give tests for a clean note, a mixed note, and a draft
that reintroduces a private detail. State what each test proves.
Constraints: Use synthetic data only. Do not read private notes.
Do not call a model, publish content, or change the policy.
List untested cases. Do not claim complete privacy protection.

Ollama JSON: Empty Results Are Not Failed Requests

Adds separate tests for valid empty output, request failure, and extraction accuracy in a local Ollama client.

8 lines485 charsSource post
Show prompt
Role: You review a local Ollama extraction client.
Context: I will paste the request code, parser, and schema.
Task: Find every path that converts failure into empty output.
Output: Show a small fix and tests for valid empty results,
malformed responses, request failures, and known source answers.
Constraints: Keep the existing storage format. Do not call a
model. Use fixed fixtures. Separate parser tests from accuracy
tests. Do not claim that schema validation proves correctness.

Local or API? Test the task before routing it

Builds a bounded local-versus-API comparison plan from your actual configurations and acceptance criteria.

23 lines859 charsSource post
Show prompt
Role:
Help me compare two model routes for one task.

Context:
Ask for the exact local model file, quantization, runtime and settings;
the API provider and pinned model ID; representative inputs;
and my quality, time, spending and maximum-attempt limits.

Task:
Define the expected answer and deterministic acceptance checks first.
Use the same inputs and output contract for both routes.
Record settings that cannot be matched.
Keep failed attempts, retries, loading time and reviewer time in the record.
Use dated provider prices and explicit local-cost assumptions.

Output:
A test plan and an empty result table.
Separate measured results from estimates and missing values.

Constraints:
Do not run a paid request or invent a result.
Do not infer a routing rule from one successful example.
Do not send private inputs to an endpoint without authorization.

Log Local LLM Fallbacks Before You Score the Output

Audits a local model job for fallback attribution and separates generation from accepted output.

20 lines638 charsSource post
Show prompt
Role:
You review local model run records.

Context:
I will supply a job record, request metadata, response metadata,
and the output review result.

Task:
Identify the requested model and actual writer.
State whether local generation started and why fallback occurred.
Separate availability, returned output, and review acceptance.
Identify missing evidence before assigning a quality or speed result.

Output:
A short verdict for each attempt and a list of missing fields.

Constraints:
Do not infer success from a file existing.
Do not assign hosted output to the local model.
Do not invent measurements or assume cloud use is permitted.

My local daily brief used 0 model calls

It reads a local daily-brief receipt and says whether the named model actually ran, or only ranked files.

31 lines932 charsSource post
Show prompt
Role:
You are auditing a local LLM daily-brief receipt.

Context:
Date: [YYYY-MM-DD]
Run ID: [id]
Named model: [model]
preview_only: [true, false, or unknown]
model_call_count: [number or unknown]
model_ok_count: [number or unknown]
total_latency_s: [number or unknown]
daily-brief-ready: [true, false, or unknown]
Sections ranked: [count or unknown]
Model output present: [yes, no, or unknown]

Task:
1. Say whether generate ran.
2. Name the first field that shows the answer.
3. Say whether the ready flag can be quoted as a model result.

Output:
- GENERATE RAN or GENERATE DID NOT RUN
- Evidence field and value
- READY MEANS MODEL RESULT or READY MEANS FILE RANKING
- One repair: which field a dashboard must show next to ready

Constraints:
- Do not treat a model name as proof of a model run.
- Do not treat 0 citation errors as quality if model_call_count is 0.
- Do not invent missing fields.
- Keep unknown fields unknown.

Long Context Cost Me VRAM, Not Tokens a Second

It runs one local model at two prompt sizes with the context window held fixed, so the VRAM change belongs to the prompt and not to the window.

22 lines1071 charsSource post
Show prompt
Role: You are a local inference benchmark operator.

Context: I run one model on one GPU through Ollama. I want to know what a
long prompt costs me in generation speed and in VRAM, separately.

Task:
1. Pick one model and one quant. Do not change them between runs.
2. Set num_ctx to one value that fits the longest prompt you will send.
   Hold it fixed for every run.
3. Run the same generation task at two input sizes, roughly 50 tokens
   and roughly 4000 tokens. Set num_predict the same for both.
4. For each run record: input_tokens, num_ctx, eval_count,
   eval_duration, load_duration, and GPU memory used before and after.
5. Compute the generation rate as eval_count divided by eval_duration.

Output: A table with one row per run, then two lines of plain text. Line
one states the percent change in generation rate. Line two states the
VRAM change in MiB.

Constraints: Do not report wall clock as the speed number, because model
load can dominate it. If num_ctx changed between runs, say so, and state
that the VRAM delta does not belong to the prompt alone.

3 Tests Before a GGUF Quant Runs Your Coding Agent

It tests one GGUF quant against a fixed read, edit, and proof loop for a local coding agent.

32 lines770 charsSource post
Show prompt
Role:
You are testing a local coding model for safe repository work.

Context:
Repository path: [ ]
Source file: [ ]
Test file: [ ]
Rule to preserve: [ ]
Focused test command: [ ]
Allowed edit scope: [one file]

Task:
1. Read the source file and test file.
2. State the rule that the change must preserve.
3. Make the smallest change inside the allowed scope.
4. Run the focused test command.
5. Report the command, exit code, changed files, and result.

Output:
- Rule found
- Change made
- Changed files
- Test command
- Exit code
- PASS or FAIL

Constraints:
- Do not invent file content.
- Do not add a package.
- Do not edit a test unless the task permits it.
- Treat a nonzero exit code as FAIL.
- Stop and report the blocker if a required file or tool is missing.

One model logged 14.5 and 6,178 prompt tokens a second

It checks whether two local model prompt-eval rates can be compared, and rejects the comparison when the input sizes differ.

27 lines885 charsSource post
Show prompt
Role:
You are auditing local LLM benchmark rows before anyone quotes them.

Context:
Row A: model, quant, engine, num_ctx, input token count, prompt rate
Row B: model, quant, engine, num_ctx, input token count, prompt rate
Row A values: [ ]
Row B values: [ ]

Task:
1. Confirm both rows name model, quant, engine, num_ctx, and input count.
2. Report the ratio of the two input token counts.
3. Mark any row with fewer than 1000 input tokens as indicative only.
4. Decide whether the two prompt rates are comparable.

Output:
- Missing fields per row
- Input count ratio
- Indicative rows
- COMPARABLE or NOT COMPARABLE
- One sentence of reasoning

Constraints:
- Do not average rates across different input sizes.
- Do not infer a missing input count from the context limit.
- Return NOT COMPARABLE when a required field is missing.
- Name a cause only when a timing field supports it.

Split Local LLM Prose and JSON Jobs by Model

Turns a mixed local AI workload into model lanes with a clear verifier and failure action for each lane.

19 lines644 charsSource post
Show prompt
Role:
You are a local LLM routing engineer.

Context:
Paste the jobs, candidate models, hardware limits, and sample outputs.

Task:
1. Group the jobs by output contract: prose, structured data, embeddings, or another exact type.
2. Assign a candidate model to each group only when evidence supports it.
3. Define a code-based verifier and a fail-closed action for every group.

Output:
- A route table with job, model, contract, verifier, and failure action.
- A replay test that a new model must pass before promotion.

Constraints:
- Keep it short.
- Use exact model tags and schema names when available.
- Do not invent missing measurements.

Ollama Load Time Can Hide a Fast Local LLM

Reviews one Ollama benchmark row and finds which measured phase should get the first change.

20 lines638 charsSource post
Show prompt
Role:
You are a local LLM benchmark reviewer.

Context:
Paste one benchmark row with model, wall time, total_duration, load_duration, prompt_eval_duration, eval_duration, token counts, and runtime settings.

Task:
1. Convert all durations to one unit.
2. Find the largest measured phase.
3. Name one change that targets only that phase.

Output:
- A timing table with wall, load, prompt, and output time.
- A first-change recommendation with a pass or fail rule.
- A list of missing fields that block a stronger claim.

Constraints:
- Keep it short.
- Use exact numbers and file paths when available.
- Do not invent missing measurements.

A 32 GB GPU Still Needs Host RAM Headroom

It turns a local LLM load plan into a host-memory gate with a clear pass rule and failure receipt.

22 lines845 charsSource post
Show prompt
Role:
You are reviewing memory admission for a local LLM host.

Context:
Paste the host name, model, runtime, total VRAM, free host RAM, normal background work, and planned workload.

Task:
1. Separate the VRAM check from the host RAM check.
2. Set a host-specific free-memory floor from measured cold loads and background use.
3. Define the small request that proves model load before the measured workload.
4. Define the receipt for a blocked, failed, or completed run.

Output:
- A gate table with catalog, host RAM, model load, and workload rows.
- A pass or fail rule for each row.
- The fields to save when a gate stops the run.

Constraints:
- Do not treat total VRAM as free host RAM.
- Do not lower the memory floor only to make a run pass.
- Do not invent missing measurements.
- Mark every request that never started as not attempted.

Your Local LLM CSV Needs a Schema Version

It audits local LLM benchmark files for schema drift and writes a safe normalization plan without filling missing facts.

23 lines797 charsSource post
Show prompt
Role:
You are a schema auditor for local LLM benchmark files.

Context:
I will give you CSV headers and sample rows from several benchmark runs.

Task:
1. Group files by exact header shape.
2. Find fields that use different names, units, or measurement windows.
3. Separate safe mappings from mappings that would change meaning.
4. Draft one versioned row contract for new runs.

Output:
- A table of files, header widths, and schema groups.
- Safe field mappings with their source names preserved.
- Fields that must stay unknown.
- Required fields and fail rules for the new collector.

Constraints:
- Do not infer a missing value from a later system reading.
- Do not map point-in-time power to average power.
- Do not replace the original files.
- Use explicit units and missing-value reasons.

Your Benchmark Row Never Saved the Driver Version

It checks a benchmark file for the environment fields a row needs before that row can support any hardware or runtime claim.

23 lines932 charsSource post
Show prompt
Role:
You are a benchmark provenance auditor for local LLM test files.

Context:
I will give you the header and a few rows from one or more benchmark files.

Task:
1. List which environment fields each file records at capture time.
2. Name the missing fields: runtime version, driver version, GPU name, capture timestamp.
3. Flag any column whose values mix a bare tool name with a versioned tool name.
4. State which claims each file can and cannot support as it stands.

Output:
- A table of files and the environment fields each one records.
- A list of missing fields per file.
- The claims each file cannot support.
- The exact collector change needed to record each missing field.

Constraints:
- Do not fill a missing field from a current system reading.
- Do not treat a report header value as row provenance.
- Do not treat a date without a time as a capture timestamp.
- Do not delete rows that lack provenance. Mark them.

Preflight Local AI Before You Benchmark a Model

It turns a local AI test plan into a fail-closed preflight with separate build, load, and measurement results.

22 lines848 charsSource post
Show prompt
Role:
You are a local AI benchmark preflight reviewer.

Context:
I will give you a host, inference runtime, model file, and test plan.

Task:
1. List the toolchain commands needed before the runtime build.
2. Define separate pass rules for build, model load, and measured run.
3. Stop at the first failed gate and save its command, exit code, and error text.
4. Mark every later metric as not attempted.

Output:
- A gate table with toolchain, build, load, and run rows.
- The exact evidence needed for each pass.
- One final status: ready to measure or blocked at a named gate.

Constraints:
- Do not copy measurements from another machine.
- Do not call a setup failure a model failure.
- Do not publish speed, memory, or power numbers unless the measured run started.
- Keep the runtime commit, model file, quant, GPU, and driver in the receipt.

49W Average Hid a 338W Burst on Gemma 26B

It checks whether a local LLM power row has a named watt window before anyone computes energy per token.

26 lines987 charsSource post
Show prompt
Role:
You are a local LLM power-receipt reviewer.

Context:
I will give you benchmark rows with model, quant, runtime, GPU,
workload, token counts, output rate, watts average, watts max,
and any sample-window notes.

Task:
1. Keep rows that include both watts average and watts max.
2. Mark energy-per-token math as blocked when the sample interval
   or phase (full request, prompt, generation) is missing.
3. If the window is named, compute joules per output token as
   generation-phase average watts divided by output tok/s.
4. Report the avg-to-max ratio as a burst flag, not as energy.

Output:
- A table with raw watt fields and the window status.
- One line that says compute, block, or rerun.
- The missing fields needed for a valid energy figure.

Constraints:
- Do not treat GPU watts as wall power.
- Do not use watts max as the energy input.
- Do not compare energy across rows with unnamed windows.
- Keep full precision for the calculation and round only the shown result.

How to Calculate Local LLM Energy per Token

It calculates energy per output token from local LLM rows and rejects comparisons that lack a task check.

24 lines901 charsSource post
Show prompt
Role:
You are a local LLM benchmark reviewer.

Context:
I will give you benchmark rows with model, quant, runtime, GPU,
workload, context, token counts, output rate, average GPU power,
and task result.

Task:
1. Keep only rows that passed the same task check.
2. Calculate joules per output token as watts divided by output tok/s.
3. Compare speed, average GPU power, and energy per output token.
4. State which measurement boundaries remain outside the result.

Output:
- A table with the raw fields and calculated joules per token.
- One setting choice for this measured workload.
- A list of missing fields needed for a full cost estimate.

Constraints:
- Do not treat GPU power as wall power.
- Do not include failed or unmatched task rows in the recommendation.
- Keep full precision for the calculation and round only the shown result.
- Say unknown when average power or output rate is missing.

Why a Failed Local LLM Benchmark Row Still Matters

It turns a local benchmark log into a row set that keeps failed tests and exposes the setting boundary.

25 lines749 charsSource post
Show prompt
Role:
You are a local LLM benchmark reviewer.

Context:
Model: [model and quantization]
Runtime: [runtime and version]
GPU: [GPU name and memory]
Workload: [prompt set, context, and setting]
Run log: [paste the log or report path]

Task:
1. Extract speed, power, VRAM, output, and verifier fields.
2. Mark each row as pass, fail, or unverified.
3. Keep failed rows in the comparison.
4. Identify the smallest setting change that moves a row from fail to pass.

Output:
- A markdown table with one row per run.
- A short note for every failed row.
- One production-setting recommendation from passing rows only.

Constraints:
- Keep exact measured values.
- Do not estimate missing power or VRAM values.
- Do not call a speed result a quality result.

Mojo Is Open Source. What Local AI Builders Need to Know

It turns an open-source AI tool announcement into a small, reviewable local evaluation plan.

21 lines634 charsSource post
Show prompt
Role:
You are a local AI infrastructure engineer.

Context:
I am evaluating an open-source compiler or runtime for work on owned hardware.

Task:
1. List the new source, license, runtime, and hardware facts.
2. Separate confirmed changes from claims that need a local test.
3. Design one small benchmark that uses my real inference path.
4. Define the receipt fields and pass criteria.

Output:
- A fact table with source links.
- A local test plan.
- A pass, hold, or reject decision.

Constraints:
- Do not treat open source as proof of speed.
- Do not invent benchmark values.
- Keep license facts separate from performance claims.

My local LLM eval hid four token caps

Lists every output-token cap in a local eval codebase and flags scores that omit the cap.

21 lines701 charsSource post
Show prompt
Role:
You audit a local LLM eval codebase for output-token caps.

Context:
Paste the eval files, configs, and one result table.

Task:
1. Find every max_tokens, num_predict, max_new_tokens, and n_predict default.
2. Group them by API path and by file.
3. Match each result-table row to the cap that produced it.
4. Flag any score whose cap, stop reason, or token count is missing.

Output:
- A table of file, field name, cap value, and API path.
- Rows that mixed two caps.
- The two-run check: production cap, then a larger cap on the same path.

Constraints:
- Use the exact numbers from the files.
- Do not invent a cap that is not in the source or logs.
- Do not rank models across different caps.

A GPU Driver Is Not a Local LLM Benchmark

checks whether a local LLM driver comparison has enough evidence for a speed or quality claim.

21 lines676 charsSource post
Show prompt
Role:
You review local LLM benchmark receipts.

Context:
Paste one result before a GPU driver change and one result after it.

Task:
1. List every setting that changed between the two results.
2. Check the runtime path, model, quant, request, and host state.
3. Compare load time, output rate, memory, power, and task result.
4. Return a pass, hold, or inconclusive verdict for a driver claim.

Output:
- A table of fixed and changed fields.
- The supported claim, if any.
- The exact rerun needed for missing evidence.

Constraints:
- Treat a version string as context, not a result.
- Do not claim cause from one unmatched run.
- Keep task quality separate from output rate.

The 26B Model Hit the Cap. The 8B Finished.

This prompt checks a local LLM benchmark row for a truncated answer before the row is used for a model decision.

21 lines784 charsSource post
Show prompt
Role: Local LLM benchmark row reviewer

Context:
I will give you one benchmark row with a model tag, runtime version,
output cap, generated token count, tokens per second, and a stop reason.

Task:
1. State whether the row is a fixed-length rate test or a task-completion test.
2. Say whether the model stopped on its own or hit the output cap.
3. Mark the row valid or truncated for its stated purpose.
4. List the missing fields before this row supports a model decision.

Output:
- One verdict line: valid, truncated, or unknown.
- The evidence you used.
- The exact rerun settings if the row is truncated.

Constraints:
- Do not read tokens per second as task success.
- Do not treat a cap hit as failure on a fixed-length rate test.
- Say unknown when the stop reason is missing.

A 27B Model Fit on an 8 GB GPU. It Was Slow.

This prompt scores a local LLM fit test so a loaded model is not treated as a usable model.

23 lines785 charsSource post
Show prompt
Role: Local LLM fit-test reviewer

Context:
I will give you one first-party run with model tag, digest, host, GPU,
total VRAM, runtime version, num_ctx, num_predict, prompt tokens,
output tokens, decode tok/s, wall time, VRAM during the run, and
VRAM after unload.

Task:
1. Say whether the model fit in VRAM.
2. Say whether the output hit the cap.
3. Mark the row fit-only, usable, or unknown.
4. List the missing fields before this row can change a model route.

Output:
- One verdict line: fit-only, usable, or unknown.
- The evidence you used.
- The exact rerun settings if the row is truncated or incomplete.

Constraints:
- Do not read a successful load as a usable rate.
- Do not infer a second GPU from this row.
- Say unknown when decode tok/s or VRAM-after-unload is missing.

My Local LLM Writer Failed Its Own Word-Count Gate

It adds a deterministic output gate to any local LLM writer pipeline.

12 lines772 charsSource post
Show prompt
Role: You are a pipeline engineer for a local LLM content writer.
Context: A local model (for example Gemma 4 26B via Ollama) writes drafts
that a script publishes. Model output sometimes runs short, drops required
sections, or drifts from the contract.
Task: Write a deterministic gate script that runs after generation and
before publish. Check: word count inside a configured window, every
required heading present as an exact string, the image path matches the
slug, and no words from a configured forbidden list.
Output: One script plus a one-line PASS/FAIL report format that names each
failed check with the measured value.
Constraints: No LLM calls inside the gate. String and integer comparisons
only. Exit nonzero on any failure so the caller blocks the publish.

Unload Local LLMs After Every Test

This prompt makes a local model test receipt that separates loaded, unloaded, warm, and cold GPU states.

20 lines644 charsSource post
Show prompt
Role: Local LLM test reviewer

Context:
I will give you a model tag, runtime version, GPU memory readings, request settings, and an unload result.

Task:
1. Label the run as cold, warm, or unknown.
2. List the evidence for loaded and idle GPU memory.
3. State whether the runtime released the model.
4. List the missing facts before another model comparison.

Output:
- A short test receipt.
- A pass, hold, or rerun decision.
- The next exact measurement.

Constraints:
- Keep model, quant, context, and runtime details in the receipt.
- Do not infer quality from a memory result.
- Do not call a run cold without an idle reading after unload.

How I Benchmark Local LLMs Before I Trust Them

This prompt turns one local model run into a decision record with explicit limits.

5 lines373 charsSource post
Show prompt
Role: Local LLM benchmark reviewer
Context: I ran one fixed task on owned hardware.
Task: Review the run and choose keep, hold, reject, or rerun.
Output: Give the decision, the evidence, the missing evidence, and the next test.
Constraints: Do not compare different engines or workloads as if they were equal. Do not invent speed, VRAM, power, quality, or tool-use results.

How I Test a 30B Local Model Before I Load It

This prompt creates a repeatable local test plan before a large model enters an agent path.

23 lines733 charsSource post
Show prompt
Role:
You are a local model test reviewer.

Context:
I will give you a model tag, runtime version, GPU and VRAM,
context size, fixed tool task, expected result, and error case.

Task:
1. Make a preflight memory check for the target runtime.
2. Define a valid tool call and expected result.
3. Define one tool error test and the correct model response.
4. Make a table for warm-run time, tokens per second, VRAM, and pass or fail.

Output:
- A test checklist.
- A result table.
- A clear go, no-go, or test-again result.
- The facts that the test cannot show.

Constraints:
- Use the same runtime as the agent path.
- Keep the task and expected result fixed across runs.
- Do not select a model from speed alone.

When a 4B Local LLM Beats 26B on One Task

compares local model benchmark rows and picks the smallest route that passed the named task.

24 lines850 charsSource post
Show prompt
Role:
You are reviewing local LLM candidates for one production task.

Context:
Paste the model, quant, runtime version, GPU, context size, batch setting,
input tokens, output tokens, generation rate, power, and task result for each run.

Task:
1. Reject every run that failed the task check.
2. Compare speed and resource use only among passing runs.
3. Pick the smallest passing route for this named job.
4. Name the next production-shaped test before the route gets more work.

Output:
- A compact comparison table.
- The chosen route and the exact reason it won.
- Claims the measurements do not support.
- The next test and its pass condition.

Constraints:
- Keep measured values separate from estimates.
- Keep the pass-rate denominator visible.
- Do not turn one workload result into a general model ranking.
- Mark missing values as unknown.

Does Ollama Include That New llama.cpp Feature?

This prompt checks whether a new local inference feature has reached the runtime that an application uses.

23 lines848 charsSource post
Show prompt
Role:
You are a local inference compatibility reviewer.

Context:
I will give you an upstream release, my application request path,
the installed runtime version, the model, and one fixed workload.

Task:
1. Map the request from the application to the inference engine.
2. Check whether the active engine includes the target feature.
3. Run the fixed workload through the application path.
4. Separate measured facts from open questions.

Output:
- A support result for each layer in the request path.
- The exact runtime and model versions.
- The test result and one recommendation.
- A list of claims that the evidence does not support.

Constraints:
- Do not test a different binary from the application path.
- Do not infer feature support from an upstream release alone.
- Keep the old recommendation when the active runtime lacks the feature.

My Local LLM Got Faster After It Passed the Tests

This prompt turns a local LLM speed test into a repeatable speed and quality gate.

23 lines787 charsSource post
Show prompt
Role:
You are a local LLM benchmark reviewer.

Context:
I will give you the model name, model build, engine settings, hardware,
fixed tasks, expected answers, tokens per second, and wall-clock time.

Task:
1. Check each fixed task against its expected answer.
2. Reject any run with a failed fixed task.
3. Compare wall-clock time among the runs that passed.
4. Use tokens per second only to explain generation speed.

Output:
- A pass or fail result for each run.
- A table with task pass rate, wall-clock time, and tokens per second.
- The setting to repeat next and the reason for that choice.
- One claim that the data does not support.

Constraints:
- Do not infer cause from one run.
- Do not change the fixed tasks between runs.
- Do not select a setting that failed a fixed task.

My local models refused zero of 50 security tasks

measures refusal rate on defensive security tasks across local models without using a judge model.

14 lines794 charsSource post
Show prompt
Role: engineer measuring whether local models will do defensive security work on owned systems.

Context: I run local models over a loopback API. I need a refusal rate, not an impression.

Task:
1. Write ten defensive prompts about systems I own: audit, harden, detect, explain a weakness.
2. Set temperature to 0 and run each prompt once per model.
3. Classify refusal with deterministic markers over the first 400 characters, not a judge model.
4. Record timeouts and errors separately from refusals.
5. Score accuracy as a second axis and report it beside refusal rate.

Output: a per-model table of completed, refused, and error counts, plus named accuracy misses.

Constraints: keep every prompt defensive and about a system I own. Do not merge willingness and correctness into one number.

The faster local model run took 83x longer

turns a local model benchmark into a wall-clock comparison so a fast generation rate cannot hide a slow request.

10 lines668 charsSource post
Show prompt
Role: engineer benchmarking a local language model on owned hardware.
Context: my rows report tokens per second, but some hide a long model load phase.
Task:
1. Per row, list load, prompt eval, generation, and total wall clock.
2. Compute what percent of wall clock each phase used.
3. Rank rows by wall clock, then by tokens per second. Show where they disagree.
4. Name the metric that matches my workload and say why.
5. Name one rerun that separates load cost from generation cost.
Output: a phase table, both rankings, and the rerun command.
Constraints: do not invent missing fields. Keep slow rows. Do not label a run cold or warm unless the receipt records it.

Chunk Size Is a Reliability Setting

It turns a long-running local AI job into a checkpoint plan that limits the work lost when a worker dies.

23 lines810 charsSource post
Show prompt
Role:
You are reviewing a long-running local AI or offline document job.

Context:
Provide the input count, current chunk size, measured time per chunk,
checkpoint behavior, worker count, memory ceiling, and recent failure logs.

Task:
1. Estimate the work and time lost when one worker fails.
2. Recommend a chunk size tied to the observed failure interval.
3. List the durable checkpoint fields needed for safe resume.
4. Give one bounded parallelism test to run after recovery is cheap.

Output:
- Current blast radius.
- Recommended chunk and checkpoint plan.
- One next measurement.

Constraints:
- Keep measured values separate from estimates.
- Do not claim a hardware fault without repeated evidence.
- Do not increase parallelism before the resume path is tested.
<!-- blog-prompt-scope:2026-06-24 -->

Why Local LLM Benchmarks Need Power Data

It turns a local model run into a comparable benchmark row with speed, quality, power, and memory fields.

23 lines677 charsSource post
Show prompt
Role:
You are a local LLM benchmark reviewer.

Context:
Model: [model and quantization]
Runtime: [Ollama, llama.cpp, or other]
GPU: [GPU name and memory]
Run log: [paste the log or report path]

Task:
1. Extract the task, context length, prompt size, and output limit.
2. Extract tokens/sec, average and peak power, VRAM peak, and quality result.
3. Compare this row with the supplied baseline without inventing missing fields.

Output:
- One markdown benchmark row.
- A short pass/fail quality note.
- One sentence on the speed, power, and memory tradeoff.

Constraints:
- Keep it short.
- Use exact numbers and file paths when available.
- Do not invent missing measurements.

VRAM Fit Is Not Runtime Support

It turns a local model check into a receipt that separates memory fit, runtime support, and workload evidence.

10 lines696 charsSource post
Show prompt
Role: local model operator checking a candidate for an owned GPU.
Context: I have a model tag or GGUF file, GPU memory details, a runtime version, and a load or generation result.
Task:
1. Check artifact identity and file size.
2. Check memory fit against the available VRAM and any stated offload plan.
3. Check whether the named runtime loaded the model architecture.
4. Record the exact load error when the model fails.
5. Report workload speed only after one real task completes.
Output: a four-gate receipt with pass, fail, or unverified for each gate.
Constraints: do not invent throughput. Keep another machine's result separate. Name the runtime version and preserve the exact error text.

My 5090 benchmark was missing the field I needed most

It turns one local model response into a phase-timing receipt without inventing residency or runtime metadata.

18 lines648 charsSource post
Show prompt
Role:
You are reviewing one local LLM request for production routing.

Context:
Paste the response timing fields, model, workload, capture time, GPU
snapshot, and the task result.

Task:
1. Separate load, prompt evaluation, output evaluation, and total time.
2. Calculate output tokens per second from eval_count and eval_duration.
3. Flag any missing capture or residency field.
4. Recommend the next single measurement.

Constraints:
- Keep measured fields separate from estimates.
- Never label a request cold or warm unless the receipt records it.
- Do not rank models from output speed alone.
- Preserve the source timestamp and workload name.

Search Old Results Before Publishing an LLM Test

It checks whether a planned local LLM benchmark post repeats a finding that is already public.

23 lines846 charsSource post
Show prompt
Role:
You are checking a local LLM post for duplicate evidence.

Context:
Paste the planned takeaway, source benchmark rows, and a list of
published post titles, URLs, and text.

Task:
1. Extract the model, setting, workload, and two signature numbers.
2. Search the published text for matching evidence.
3. Compare the old and planned one-sentence takeaways.
4. State what new measurement or decision the planned post adds.

Output:
- PASS if the post adds a distinct finding.
- FAIL if it repeats the same conclusion from the same rows.
- The closest prior post and the exact overlap.

Constraints:
- Do not treat a report generation date as a measurement date.
- Do not approve a rewrite that only changes the title or checklist.
- Keep measured facts separate from proposed follow-up tests.
- Do not invent source rows, numbers, or run dates.

Build Local LLM Eval Data From Real Failures

It turns local coding-agent run logs into replayable eval cases without inventing missing tests or measurements.

23 lines717 charsSource post
Show prompt
Role:
You are building a regression set for a local coding model.

Context:
Paste one task prompt, model output, test output, model name,
runtime route, and capture time.

Task:
1. Preserve the original prompt and output.
2. Extract the executable checks and their results.
3. Classify the failure by observed behavior.
4. List missing boundary cases that need new tests.

Output:
- One replayable eval row with named fields.
- A pass or fail verdict tied to the saved tests.
- A short list of added test cases, if needed.

Constraints:
- Do not repair or overwrite the failed output.
- Do not invent a test result or runtime setting.
- Keep measured facts separate from suggestions.
- Mark missing fields as unknown.

How I Keep LLM Results Valid After a Driver Update

It turns mixed local LLM benchmark files and machine snapshots into timestamped route receipts without assigning new metadata to old results.

23 lines766 charsSource post
Show prompt
Role:
You are reviewing local LLM benchmark provenance.

Context:
Paste the benchmark rows, source filenames, machine snapshots, and
the routing or sizing decision the measurements must support.

Task:
1. Bind each result to its own timestamp and environment.
2. Separate current machine health from historical benchmark data.
3. Group only rows with matching routes and workload shapes.
4. Flag changed routes that need a fresh run.

Output:
- A compact route-receipt table.
- A list of rows that remain comparable.
- A rerun list tied to the named decision.

Constraints:
- Keep measured values separate from later snapshots.
- Do not assign a current driver to an older row.
- Do not invent missing versions or timestamps.
- Preserve the original source filename.

Why I Benchmark Local LLM Input and Output Separately

It turns raw Ollama timing fields and a task result into a phase-by-phase local LLM benchmark.

25 lines892 charsSource post
Show prompt
Role:
You are reviewing a local LLM benchmark for a production workload.

Context:
Paste each run with the model, quant, runtime, workload, residency state,
load duration, prompt token count and duration, output token count and
duration, total duration, and task-check result.

Task:
1. Calculate prompt ingestion and output generation rates separately.
2. Keep cold and resident requests in separate rows.
3. Compare only runs with clearly named workload shapes.
4. Reject any route whose task check failed.

Output:
- A compact phase-by-phase timing table.
- The slow phase for each route.
- The best measured route for the named workload.
- The next test needed before production use.

Constraints:
- Keep measured values separate from estimates.
- Do not rank models from output speed alone.
- Do not hide load time inside a blended average.
- Do not treat a fast failed result as a pass.

How I Budget VRAM for Shared Local AI Workloads

It turns measured local LLM memory use into a shared-GPU budget with named owners and a pass or fail rule.

24 lines869 charsSource post
Show prompt
Role:
You are reviewing a shared-GPU memory budget for a local LLM deployment.

Context:
Paste the GPU total memory and measured runs. Include model, quant,
runtime, context, batch setting, workload, observed peak memory,
and every other process that must share the GPU.

Task:
1. Find the highest observed memory use for each candidate route.
2. Subtract that peak from reported card memory.
3. Assign the remaining room to named GPU workloads or keep it free.
4. Reject routes that were not tested on the largest planned job.

Output:
- A compact fit table.
- The safest measured route for the full shared workload.
- The next memory test required before deployment.

Constraints:
- Keep measurements separate from estimates.
- Do not infer fit from model-file size alone.
- Do not invent a universal headroom percentage.
- Do not ignore processes that share the GPU.

Why I Test 3 Workloads Before Sizing a Local LLM

It turns local LLM benchmark rows into a workload-aware sizing decision without inventing missing measurements.

22 lines770 charsSource post
Show prompt
Role:
You are a local LLM capacity reviewer.

Context:
Paste benchmark rows for one GPU. Include model, quant, runtime,
workload, context, input tokens, output tokens, generation rate,
peak memory, load duration, and verifier result when available.

Task:
1. Group comparable rows by model and workload.
2. Report the fastest and slowest measured workload for each model.
3. Reject any route that does not fit memory or pass its verifier.

Output:
- A compact workload-rate table.
- A recommended resident model and route.
- The exact measurement still needed before deployment.

Constraints:
- Keep measured values separate from estimates.
- Do not compare rows with different job shapes as if they were equal.
- Do not invent missing memory, load, or verifier results.

Local open-model agents just became a product category

Routes a task to a local open model or a frontier model using the split this post argues for. Describe the task and it returns a recommendation with the reasoning, so you stop sending sensitive or high-volume work to a metered API by default.

9 lines458 charsSource post
Show prompt
I run both local open-weight models and frontier cloud models. Help me route one task to the right one.

Task: [describe the task, its data sensitivity, its volume, and how much hard multi-step reasoning it needs]

Decide local or frontier, and answer with:
1. The recommendation in one line.
2. Why, weighed against data sensitivity, volume, and reasoning difficulty.
3. What would change the answer.
4. If local, which class of open model is likely enough.

Your local LLM benchmark is probably lying to you

This prompt makes a model report its own benchmark result with the context that keeps the number honest, instead of a bare pass rate.

9 lines684 charsSource post
Show prompt
Role: You are a local model evaluation reviewer.

Context: I am about to record a benchmark result for a local model and I want the row to be honest, not just high.

Task: Given a run result, produce one record that includes the pass rate WITH its denominator, an explicit count of tasks that could not run, and a field for machine stability during the run.

Output: Return a single table row with these columns: model, tasks_total, tasks_passed, tasks_could_not_run, machine_stable_yes_no, and a one-line honesty note.

Constraints: Never report a percentage without the denominator. Never fold could-not-run into pass or fail. Mark any unknown value as unknown rather than guessing.

Test Retrieval Before Your Local LLM Writes

This prompt reviews a retrieval preview and finds source-selection problems before a local LLM generates an answer.

19 lines634 charsSource post
Show prompt
Role:
You are a retrieval QA reviewer for a local LLM workflow.

Context:
Paste the query, selected sources, rejected sources near the cutoff, score parts, dates, and path rules.

Task:
1. Check whether every selected source helps answer the query.
2. Find stale, duplicate, blocked, or weakly matched sources.
3. Compare the weakest selected source with the strongest rejected source.

Output:
- A table with source, keep or reject, reason, and the scoring rule to change.
- A pass or fail verdict for generation.

Constraints:
- Keep it short.
- Use exact numbers and file paths when available.
- Do not invent missing measurements.

Why I Did Not Promote My Smaller Local Model

It turns a local model replacement into a fixed comparison with explicit evidence and a clear promotion decision.

22 lines723 charsSource post
Show prompt
Role:
You are a local model promotion reviewer.

Context:
Provide the baseline model, candidate model, hardware, runtime settings, task prompt, saved outputs, and verifier results.

Task:
Compare the candidate with the baseline on the same workload. Separate measured facts from assumptions. Decide whether the candidate earned promotion.

Output:
- Baseline result
- Candidate result
- Failed checks
- Performance tradeoffs
- Promote, reject, or retest verdict
- Next experiment that changes one variable

Constraints:
- Do not treat a successful load as task success.
- Do not let the model grade its own output.
- Do not invent missing measurements.
- Keep the current route unless the candidate clears the written gate.

Build an AI Research Workbench on Your Own GPU

It turns a local research task into a reproducible run plan with explicit evidence and pass or fail checks.

19 lines621 charsSource post
Show prompt
Role:
You are a local AI research-run designer.

Context:
Paste the research question, model file, hardware, runtime, source files, data rules, and expected result.

Task:
1. Define the fixed model, prompt, parameters, and environment to record.
2. List the code, data, logs, and measurements the run must save.
3. Write a verifier that can mark the result pass or fail.

Output:
- A run-folder layout with exact artifact names.
- A command checklist and one pass or fail rule.

Constraints:
- Keep private data on the stated hardware.
- Separate measured results from vendor claims.
- Do not invent missing measurements.

Ollama num_batch: 256 Was My RTX 5090 Sweet Spot

It turns local Ollama benchmark rows into a measured num_batch recommendation without inventing missing data.

21 lines672 charsSource post
Show prompt
Role:
You are a local LLM runtime tuning reviewer.

Context:
Paste benchmark rows for one model and GPU. Include model, quant,
runtime version, num_ctx, num_batch, tokens per second, power,
loaded memory, and task verifier result when available.

Task:
Compare the settings and reject any run that fails the task verifier.

Output:
- A compact comparison table.
- The smallest setting that captures nearly all useful speed.
- The exact measurement missing before deployment.

Constraints:
- Compare only runs with the same workload.
- Keep context attached to every batch result.
- Separate measured values from estimates.
- Do not invent missing power or memory readings.

A 7 GB 27B Model Lost to My 17 GB Default

It turns a local model comparison into a task-based keep, test, or reject decision.

20 lines638 charsSource post
Show prompt
Role:
You are a local LLM model-selection reviewer.

Context:
Paste the GPU, runtime, model files, context size, benchmark output,
task prompt, expected output format, and verifier result.

Task:
1. Compare loaded memory and cold-load time.
2. Compare warm generation speed on the same task.
3. Check exact output compliance and verifier results.

Output:
- A table with fit, speed, compliance, and failure reason.
- One decision: keep, test again, or reject for this job.

Constraints:
- Separate local measurements from vendor claims.
- Do not compare different runtimes as a clean model benchmark.
- Do not invent missing measurements.

Ollama Raised $65M. What Builders Get

It turns one workload into a measured local-versus-cloud routing decision with explicit pass and fail rules.

24 lines806 charsSource post
Show prompt
Role:
You are a local AI workload evaluator.

Context:
Hardware: [GPU and VRAM]
Local runtime and model: [Ollama version, model, quant, context]
Cloud model: [provider and model]
Workload: [paste one repeatable task]
Measurements: [latency, tokens, errors, power, and cost if known]

Task:
1. Check whether the local model fits and completes the workload.
2. Compare local and cloud on quality, latency, privacy, availability, and operator time.
3. Choose local, cloud, or a verified hybrid route.

Output:
- A comparison table using only supplied measurements.
- A routing recommendation with one pass rule and one fallback trigger.
- The next measurement needed to reduce uncertainty.

Constraints:
- Keep it short.
- Use exact numbers and file paths when available.
- Do not invent missing measurements.

Why production AI is moving to open weights

This prompt helps you design a two-tier workflow where a local model generates content and a frontier model audits it for accuracy.

21 lines810 charsSource post
Show prompt
Role:
You are an AI Workflow Architect specializing in hybrid local-frontier pipelines.

Context:
I am running high-volume text generation on a local RTX 5090 setup using Ollama. I use frontier models (like Claude or GPT) only for final quality audits to keep costs low.

Task:
1. Analyze the provided raw text generated by my local model.
2. Identify any factual hallucinations or logical inconsistencies.
3. Check if the tone meets my brand guidelines.
4. Provide a "Pass" or "Fail" grade for each section.

Output:
- Error Log: List specific errors found.
- Audit Score: A score from 1 to 10.
- Corrected Version: The final text ready for publication.

Constraints:
- Do not rewrite the entire text if it is already correct.
- Focus only on factual accuracy and risk mitigation.
- Keep the feedback concise.

Q4_K_M vs Q5_K_M (q4km vs q5km): Which to Download? (2026)

It turns your GPU VRAM, context length, and two GGUF file sizes into a Q4_K_M vs Q5_K_M pick with a hard pass/fail.

27 lines644 charsSource post
Show prompt
Role:
You are a local LLM quant picker.

Context:
GPU VRAM (GB): [ ]
Target context length: [ ]
Model: [ ]
Q4_K_M file size (GB): [ ]
Q5_K_M file size (GB): [ ]
Need full GPU offload: yes/no

Task:
1. Estimate free VRAM after OS and KV cache.
2. Say whether Q4_K_M and Q5_K_M fit fully on GPU.
3. Pick one quant and one fallback if load fails.
4. List the single next measurement to run.

Output:
- Pick: Q4_K_M or Q5_K_M
- Fit: full GPU / mixed / no
- Fallback plan
- Next command or check

Constraints:
- Prefer full GPU offload over a higher quant with CPU fallback.
- Do not invent sizes I did not provide.
- Keep the answer under 12 lines.

My 8B Model Failed a 400-Word Task

This prompt turns local-model run records into a verifier-driven routing rule for short and long-form tasks.

21 lines730 charsSource post
Show prompt
Role:
You are a local LLM routing engineer.

Context:
Paste the task, model settings, output artifact, and verifier result for each local run.

Task:
1. Group runs by task shape and required output.
2. Identify the smallest model that passed every deterministic rule.
3. Name the exact failure that justifies each larger-model route.

Output:
- A routing table with task shape, model, runtime settings, and pass rules.
- One default route and any justified escalation routes.
- A list of checks that should run in code before human review.

Constraints:
- Keep it short.
- Use exact numbers and file paths when available.
- Do not invent missing measurements.
- Do not recommend a larger model without a recorded verifier failure.

I Raised the QA Bar on Blog Infographics

This prompt turns a draft post into a fail-closed infographic publish checklist with char budgets and claims.

20 lines666 charsSource post
Show prompt
Role:
You are a blog infographic publisher for exact-text charts.

Context:
Paste the draft post body, the intended template (steps/stats/flow/bars/compare), and any metrics file path.

Task:
1. Write a JSON spec that fits char budgets (4-up steps: label <=18, sub <=22).
2. List every digit in the graphic and where it appears in the draft or metrics.
3. Run publish with --claims and report geometry/vision outcomes.

Output:
- The JSON spec
- A digit provenance table
- PASS/FAIL for each gate with one-line reasons

Constraints:
- No generative text-in-image for labels
- Fail closed on missing claims when digits exist
- Keep labels short; do not invent metrics

I Rebuilt 5,383 Embeddings After a Dimension Change

This prompt turns an embedding-model change into a fail-closed vector-index migration plan with retrieval checks.

21 lines768 charsSource post
Show prompt
Role:
You are a vector search migration reviewer.

Context:
Paste the current embedding model, new model, dimensions, table schema, corpus counts, and three known queries with expected results.

Task:
1. Identify every schema, index, corpus, and query-path change.
2. Define the order for rebuilding vectors without mixing models.
3. Write parity checks that must pass before the query path switches.

Output:
- A numbered migration plan with rollback points.
- A pass/fail table for vector length, corpus count, and known-query results.
- The exact evidence to record after the rebuild.

Constraints:
- Keep it short.
- Use exact numbers and file paths when available.
- Do not invent missing measurements.
- Treat a model or dimension change as a new corpus version.

Preview Retrieval Before Your Local LLM Runs

This prompt reviews a local LLM retrieval manifest and decides whether its source set is safe and relevant enough for inference.

21 lines751 charsSource post
Show prompt
Role:
You are a retrieval gate reviewer for a private local LLM workflow.

Context:
Paste the requested output, allowed source roots, freshness rules, and retrieval manifest.

Task:
1. Check whether every requested section has relevant evidence.
2. Flag stale, duplicate, out-of-scope, or sensitive sources.
3. Decide whether the manifest may advance to local inference.

Output:
- A table with section, source path, verdict, and reason.
- One final verdict: approved, needs review, or blocked.
- The smallest retrieval change needed for every failed source.

Constraints:
- Use only the supplied manifest.
- Do not invent missing files, scores, or dates.
- Treat private paths and sensitive data as untrusted input.
- Keep the review under 400 words.

Pin Your Context Window or Pay the Reload Tax

It reviews how your local agent or script sets the context window and flags any place that would force Ollama to reload the model mid-run.

18 lines786 charsSource post
Show prompt
Role: You are a local LLM performance reviewer.

Context: I run models locally through Ollama on a single GPU. Changing
num_ctx forces a full model reload, which is slow for large models.

Task: Given my agent or script code, find every place that sets or
changes num_ctx (or the context window) and tell me where a reload
would happen at runtime.

Output:
- Each spot where num_ctx is set or changed, with the line
- Whether it happens once at startup (fine) or per request (a reload)
- One pinned num_ctx value you recommend for the whole workload
- The single change that removes the most reloads

Constraints: Assume a reload costs 100+ seconds for a large model.
Prefer one fixed context size over a dynamic one. Be specific about
lines. No general advice about prompt engineering.

Pin Your Local LLM Context Size Before You Build a Router

This prompt turns local inference logs into a context-aware routing policy with separate cold-start and generation limits.

21 lines811 charsSource post
Show prompt
Role:
You are a local LLM runtime engineer reviewing model routing traces.

Context:
Paste the model name, quantization, runtime, GPU, num_ctx, input tokens, output tokens, load duration, generation duration, peak VRAM, and verifier result for each run.

Task:
1. Group runs by model, quantization, runtime, GPU, and num_ctx.
2. Separate cold loads from warm inference.
3. Identify context changes that caused reloads.
4. Recommend one default route and any justified escalation routes.

Output:
- A routing table with route key, startup timeout, generation timeout, and verifier rule.
- A list of requests that paid an avoidable reload cost.
- A pass or fail verdict for the current router.

Constraints:
- Keep it short.
- Use exact numbers and file paths when available.
- Do not invent missing measurements.

When Claude hits a weekly limit, your agent fleet still needs a third CLI

turns a Claude/Codex outage into a checklist for a working tertiary Gemini CLI path with honest publish provenance.

13 lines669 charsSource post
Show prompt
Role: platform engineer for a scheduled agent fleet
Context: Claude weekly limit, Codex CLI failing, empty blog queue, Gemini CLI installed
Task: make Gemini a real tertiary option for agent + blog QA runners
Output:
1. ConfiguredCheck that requires GEMINI_API_KEY (not broken OAuth files)
2. Runner chain order claude -> codex -> gemini with ledger notes for which ran
3. Blog QA stamp qa_reviewer to the engine that actually ran
4. Public publish allowlist updated to accept gemini without spoofing codex
Constraints:
- No LiteLLM required for pass
- No fake qa_reviewer: codex
- Prefer flash model for headless gates
- Capture provider-chain and publish attempt logs

I built a self-improving code model on one RTX 5090. Here is what actually worked.

Has a local coding model generate solutions, verify them with pytest, and reuse the passing rows to fine-tune itself in a repeating loop.

23 lines1198 charsSource post
Show prompt
Role: You are a local LLM practitioner running a code model on one owned GPU.

Context: You have a local model (for example Llama 3.1 8B in 4-bit), a set of
coding tasks, and pytest in a sandbox. No cloud APIs, no human labels.

Task: Design a self-improvement loop with these stages:
1. Verifier: run generated code plus its tests in a throwaway sandbox, return a
   pass fraction. Pass or fail is ground truth.
2. Data factory: generate solutions with the local model, run the verifier, keep
   only the passing rows with full provenance on each row.
3. Fine-tune: QLoRA on the verified rows.
4. Agent loop: let the model write code, run the tests, read the failure, and
   retry until the tests pass.
5. Retrain cycle: feed the agent's verified solutions back as new training data,
   retrain, measure on held-out tasks, repeat.

Output: Python pseudocode for each stage, plus one note on how you measure the
held-out pass rate between rounds.

Constraints: Everything runs locally. Keep a diverse task set so the model does
not memorize a few tasks. Anchor repeated training against the last good weights
so the held-out score does not regress. Cap what the loop can spend and do while
it runs.

Will That Local Model Fit? Do the VRAM Math First

It asks a model to estimate whether a given local LLM fits your GPU's VRAM and tells you the one thing to change if it does not.

18 lines761 charsSource post
Show prompt
Role: You are a local LLM deployment assistant.

Context: I run local models on a single consumer GPU. I care about
whether a model fits in VRAM before I download it.

Task: Given a GPU VRAM amount, a model parameter count, a quant level,
and a target context length, estimate total VRAM needed (weights plus
KV cache plus a 2-3 GB overhead margin) and tell me if it fits.

Output:
- Estimated weights VRAM
- Estimated KV cache VRAM at my context length
- Total with margin, and FIT or NO FIT against my GPU
- If NO FIT: the single smallest change to make it fit
  (lower quant, smaller context, or smaller model)

Constraints: Use ~0.56 bytes/param for Q4 weights. Show the arithmetic
in one or two lines. No hedging paragraphs. Give me a number and a verdict.

Local LLMs Need a Timeout Before They Need a Bigger Model

It turns a local model launcher into a bounded inference doctor checklist.

18 lines683 charsSource post
Show prompt
Role:
You are a local AI runtime reviewer.

Context:
I run local LLMs on owned hardware. The model server may be alive even when a generation probe hangs.

Task:
Design a small inference doctor for my local model stack.

Output:
Return a checklist with sections for API reachability, model tag presence, GPU visibility, bounded generation, restart rules, logging, and tests.

Constraints:
Use one tiny deterministic prompt.
Use both a request timeout and a wall-clock timeout.
Restart the model server only after bounded checks prove it is needed.
Write a JSON result with status, latency, restart_attempted, and failure reason.
Keep the plan small enough to implement in one script.

Your local LLM is not a worse Claude. It is a different tool.

It helps you decide whether a task should run on a local model, a frontier API, or both.

18 lines655 charsSource post
Show prompt
Role: You are a practical AI systems reviewer.

Context: I need to choose between a local LLM and a frontier API model for one workflow.

Task: Classify the workflow into one of three routes: local, frontier, or split.

Output:
1. Recommended route.
2. Main reason.
3. Review gate needed before the result matters.
4. Failure mode to watch.
5. One test I should run before using it again.

Constraints:
- Treat private data, offline need, repeat volume, and review path as first-class inputs.
- Do not choose local just because it is cheaper.
- Do not choose frontier just because it is stronger.
- Prefer the smallest model that can pass the review gate.

A verifier loop beats a faster local model

This prompt turns a local model test into a measured verifier loop before the workflow gets more permission.

9 lines608 charsSource post
Show prompt
Role: You are a local AI agent evaluator.

Context: I am testing a local model before I allow it to automate a workflow.

Task: Design a verifier loop for this workload. The loop must pin the input, run the model, check the output, record the result, and decide whether to promote or block the path.

Output: Return a compact table with these columns: step, command or check, pass signal, fail signal, and artifact to save.

Constraints: Do not assume cloud APIs. Do not invent benchmark numbers. Mark any unmeasured value as unknown. Keep the workflow in draft mode unless the verifier can catch bad output.

How I Gate a Local Coding Model Before I Trust It

Builds a small pass/fail gate for one local model task before you let it touch real files.

18 lines450 charsSource post
Show prompt
Role:
You are a local-model test engineer.

Context:
I am testing one local LLM before I let it run inside a coding agent.

Task:
Design a verifier loop for one task type.

Output:
Return a table with Task, Input, Expected Output, Verifier Command, Pass Rule, Fail Rule, and Budget Cap.

Constraints:
Use local files only.
No network calls.
No secrets.
One command must decide pass or fail.
Keep the test small enough to run after every model change.

How I Make Local Model Runs Fail Safely On A 5090

It turns a local model experiment into a fail-safe run contract before you spend a night on it.

9 lines691 charsSource post
Show prompt
Role: You are a local AI run reviewer for an owned GPU workstation.

Context: I am about to run a local model job inside an agent loop. I need the job to fail safely before I trust any model score.

Task: Review my run plan and list the missing safety gates.

Output: Return a table with columns for Gate, Why it matters, Evidence required, and Fix.

Constraints: Tie every number to a log, config file, or command output. Require a small first run. Require GPU memory and temperature checks before model load. Require cleanup in a finally path. Require append-only metrics. Require AgentGuard or an equal budget and rate limit for any agent loop. Do not accept hardcoded test data as proof.

How to Make a Local QLoRA Starter Fail Safely

It turns a local fine-tuning idea into a small fail-safe starter with proof, resource checks, and honest blockers.

31 lines1020 charsSource post
Show prompt
Role:
You are designing a fail-safe local QLoRA starter for owned hardware.

Context:
Hardware:
[GPU, VRAM, cooling notes]

Dataset:
[path, row count, provenance, privacy limits]

Current goal:
[what the first run should prove]

Task:
1. Define the smallest direct run that proves the data path and writes an exact metric count.
2. Add resource checks before model load and before each training iteration.
3. Define the fallback behavior if model load or training fails.
4. Write the verifier checks that must pass before the run can claim progress.
5. List the blockers that should be published or recorded if the real train is not complete.

Output:
- A one-command starter run.
- A metrics schema with required fields.
- A test checklist tied to the claim.
- A blocker list with owner, next action, and proof needed.

Constraints:
- Keep the first run small.
- Use exact file paths and row counts when available.
- Label any stub or fallback metric clearly.
- Do not claim model improvement without a verifier result.

How to Run Local LLM Verifier Loops on Owned Hardware

It turns a local LLM task into a verifier loop with proof, retry limits, and AgentGuard runtime rails.

24 lines1064 charsSource post
Show prompt
Role:
You are designing a local LLM verifier loop for a single agent task.

Context:
The model runs on owned hardware. The task may touch private files, local repos, or local reports. The result is not accepted until an independent proof step passes.

Task:
Design the smallest useful verifier loop for this workflow:
[paste workflow, target files, model, and command surface]

Output:
1. Task scope: one job the local model is allowed to do.
2. Inputs: files, prompts, and commands the model can read.
3. Proof check: the command, file readback, URL fetch, or report diff that proves completion.
4. Retry rule: what to do after the first failed proof and when to stop.
5. Runtime limits: AgentGuard budget, token, rate, and retry ceilings.
6. Blocked packet: the exact note to write when the loop cannot pass.

Constraints:
- Keep the loop small enough to run today.
- Do not trust model confidence as proof.
- Prefer deterministic checks over judgment.
- Do not expose secrets or private data to hosted tools.
- End with either a passed proof or a written block.

llama.cpp --n-gpu-layers: -1, 0, Partial GPU Offload (2026)

Works out the right --n-gpu-layers value for your exact GPU and model, and explains the -1 vs 0 vs N tradeoff.

22 lines901 charsSource post
Show prompt
Role:
You are tuning llama.cpp offload for a single machine.

Context:
I want to run a GGUF model with llama.cpp and set --n-gpu-layers correctly.

My setup:
- GPU and VRAM: [e.g. RTX 4090, 24GB]
- Model and quant: [e.g. Llama-3.1-70B Q4_K_M]
- Other VRAM users: [display, other models, none]

Task:
1. Estimate how much VRAM one layer of this model costs at this quant.
2. Tell me the highest --n-gpu-layers value that fits my VRAM with headroom for context.
3. Explain what -1 and 0 do here and why I would pick a specific N instead.
4. Give the exact llama.cpp command line to start with.
5. List the load-log lines I should check to confirm the offload actually happened.

Constraints:
- Leave headroom for the KV cache at my context length.
- Do not assume every layer costs the same if the model has a large embedding or output layer.
- Prefer a value that runs today over a theoretical maximum.

Local LLM on Consumer GPUs: 50 req/s, $0/Call [Benchmarks 2026]

Designs a production local-LLM inference setup on consumer GPUs for real throughput.

20 lines889 charsSource post
Show prompt
Role:
You are designing a local LLM inference service on consumer GPUs.

My setup:
- GPUs and VRAM: [e.g. 2x 4090]
- Model and quant: [e.g. 70B Q4_K_M]
- Target load: [e.g. 50 requests/sec, prompt and response sizes]
- Serving stack I can use: [llama.cpp, vLLM, other]

Task:
1. Say whether this hardware can hit my target load and where the ceiling is.
2. Recommend a serving stack and the key settings: batching, concurrency, context.
3. Estimate tokens/sec and requests/sec I should expect, with the limiting factor.
4. List what to change first if I need more throughput: quant, batching, a second GPU.
5. Give the metrics to watch in production and their warning thresholds.

Constraints:
- Base numbers on the specific GPU and quant above.
- Separate prompt processing from token generation limits.
- Prefer settings that hold up under concurrent load, not single-request benchmarks.

devlog

PR 1887 merged clean after I claimed none could

Audits a repository to verify whether blocked agent pull requests fail on branch rules, merge conflicts, or receipt errors.

24 lines788 charsSource post
Show prompt
Role: Git Audit Engineer

Context: An agent claim states that branch rules block all pull requests. The user needs an audit showing which pull requests fail on branch rules, merge conflicts, or receipt errors.

Inputs:
- Repo path: __
- Target branch: __
- Pull request list: __
- Output log file: __

Task:
1. Inspect pull requests in Pull request list inside Repo path against Target branch.
2. Group each pull request: rule blocked, merge conflict, or failing receipts.
3. Count items in each bucket.
4. Write results and commit SHAs to Output log file.

Output:
- Table of Pull Request ID, Bucket, and Reason.
- Total counts per bucket.

Constraints:
- Do not modify git configuration.
- Check git logs and job receipt outputs directly.
- Flag failure classes that re-runs cannot fix.

A strict branch rule blocked 26 clean merges

Audits branch protection settings to find why clean pull requests fail to merge.

27 lines1078 charsSource post
Show prompt
Role:
GitHub Repository Protection Auditor

Context:
Pull requests pass required continuous integration checks but remain stuck in the queue without merging. The user needs to identify the exact branch protection settings causing the block.

Inputs:
- Repository: __
- Target branch: __
- Blocked PR number: __

Task:
1. Inspect the branch protection rules for Target branch in Repository.
2. Check whether strict status checks require branches to be up to date before merging.
3. Compare Blocked PR number against the latest commit on Target branch.
4. Report whether the merge failure is caused by stale status checks or admin enforcement.
5. Provide the exact command to update the setting if needed.

Output:
- Blocker summary explaining why Blocked PR number cannot merge
- Setting evaluation for strict status checks and administrator enforcement
- Shell command to update the rule or rebase the branch

Constraints:
- Use only repository settings retrieved for Target branch.
- Do not modify repository settings.
- State whether the block requires administrator approval.

My P0 decision expired while the doctor showed red

Audits a repository test configuration to exclude vendor directories and report uninspected failure files.

24 lines937 charsSource post
Show prompt
Role:
Test Infrastructure Engineer

Context:
Test runners can pull in vendor packages, virtual environments, or git-ignored directories, skewing suite metrics and hiding genuine test failures.

Inputs:
- Repo path: __
- Test command: __
- Ignored directory names: __

Task:
1. Inspect the configuration for the specified Test command in Repo path.
2. Verify that Ignored directory names like vendor, site-packages, and cache directories are explicitly excluded from discovery.
3. Run Test command across Repo path and count both passing checks and total suite files.
4. Flag any file that fails or skips execution during the discovery pass.

Output:
- A list of excluded directories added to the configuration.
- A pass/fail summary showing passing test count and total file denominator.

Constraints:
- Do not modify production source code outside test configuration files.
- Report all skipped files explicitly without omitting errors.

I split crawler hits from my human sessions

This prompt audits incoming web logs and separates crawler events from real human sessions.

27 lines1043 charsSource post
Show prompt
Role:
Data Pipeline Engineer

Context:
Web traffic counters often mix automated crawler pings with real human visits, inflating session numbers.
You need to separate bot events from user sessions to produce accurate activity metrics.

Inputs:
- Log file path: __
- Bot user-agent patterns: __
- Session timeout in minutes: __
- Output summary path: __

Task:
1. Read the raw web access events from the specified Log file path.
2. Match every entry against the provided Bot user-agent patterns.
3. Label matching crawler rows as bot events and route them to an isolated table.
4. Group the remaining non-bot requests by visitor ID using the Session timeout in minutes.
5. Write the verified counts of bot events and human sessions to the Output summary path.

Output:
- A Markdown table showing bot event counts versus human session counts.
- A list of the top 5 detected crawler user-agents.

Constraints:
- Do not count any crawler hit toward human session totals.
- Output clean text without external dependencies or unexplained assumptions.

My security scanner reported green on zero repos

Audits an automated pipeline job or test script to verify target inputs exist and prevent false green passes on empty runs.

24 lines1260 charsSource post
Show prompt
Role: Automation Reliability Engineer

Context: An automated scan, check script, or test runner returned a green exit status, but may have processed zero files due to an unmounted volume, missing directory path, or empty input list.

Inputs:
- Target directory path: __
- Script path to audit: __
- Minimum expected file count: __
- Required exit code on empty target: __

Task:
1. Inspect the script at Script path to audit to find how it resolves Target directory path.
2. Check whether the script verifies Target directory path exists before starting file processing.
3. Identify whether an empty file list allows the script to exit with status zero without error.
4. Add a guard check ensuring the discovered target count meets or exceeds Minimum expected file count.
5. Update the script to terminate with Required exit code on empty target when Target directory path is missing or empty.

Output:
- A diff showing the pre-execution path verification and minimum file count check added to the script.
- The command line to verify the failure behavior when Target directory path is missing.

Constraints:
- Do not suppress errors or catch exceptions without logging the missing path.
- Exit immediately with non-zero status if target paths are unreachable.

A green check lied, so I built append-only status

Audits your test suite or reporting jobs for false green status claims caused by unread targets or disabled checks.

23 lines734 charsSource post
Show prompt
Role:
Code auditor specializing in agent status verification.

Context:
Agent runs often return green status when checks fail silently or inspect zero files.

Inputs:
- Tool or test script: __
- Expected input path: __
- Log output file: __

Task:
1. Read the log output file from the tool or test script.
2. Check if the tool read zero targets or skipped files while reporting success.
3. Identify silent failures such as falsey deadlines or empty sweeps.
4. Output a receipt marking status as unknown when evidence is missing.

Output:
- A Markdown table showing checked files, actual targets inspected, and verified status.

Constraints:
- Never mark a job green if targets read equals zero.
- Flag any missing evidence as unknown.

I built refusals into my claim recorder

It reviews whether one agent claim has enough evidence, without treating a well-formed row as proof of completion.

24 lines785 charsSource post
Show prompt
Role:
Review one agent claim against its evidence.

Context:
I will provide the claim, check definition, actual check result,
and intended scope. Treat missing inputs as unknown.

Task:
1. Identify missing or malformed check fields.
2. Compare the claim with what the check can establish.
3. Read the actual result and coverage, if supplied.
4. Name the evidence still needed for the full claim.

Output:
Verdict: supported, unsupported, or unknown.
Reason: cite the supplied evidence.
Scope: state exactly what the check covered.

Constraints:
Do not invent a result or run a command.
A valid row is not proof that a check ran.
Zero coverage cannot support a clean-scan claim.
Partial coverage cannot establish the whole scope.
Later corrections override earlier completion summaries.

Ten Merges, One Dead Disk, Four Fake Green Signals

Forces an agent to prove a task completed by writing a checkable artifact instead of returning a status message.

19 lines582 charsSource post
Show prompt
Role:
You are an execution agent that reports work by producing files, not by describing them.

Context:
A prior run returned "ok" with no artifact and no way to verify the claim.

Task:
1. Run the assigned job to completion.
2. Write the full output to a named file on disk, not to chat or memory.
3. Print the file path and a byte count as the last line of output.

Output:
- The file path.
- The byte count.
- Nothing else on that final line.

Constraints:
- If the file was not written, say so plainly. Do not report success.
- Do not summarize the work instead of producing it.

I dropped the BMD kill date and named it flagship

Updates blog writer templates to include structured input fields for consistent generation.

24 lines539 charsSource post
Show prompt
Role:
Blog Template Architect

Context:
I am updating my writing workflow to use a structured Inputs section in every post.

Inputs:
- Topic: __
- Key Numbers: __
- Broken Items: __
- Shipped Items: __

Task:
1. Create a draft structure for a devlog post.
2. Ensure the first sentence is a plain statement of what happened.
3. Include an H2 for technical details.
4. Use the Inputs provided to populate the body.

Output:
- A structured markdown post.

Constraints:
- No preamble or intro text.
- Use only the facts provided in the Inputs.

My agent roadmaps did not prove the work was done

Separates planned work, checked outcomes, and unmeasured claims in a daily agent report.

20 lines600 charsSource post
Show prompt
Role:
Review my daily agent report against the supplied evidence.

Context:
I will supply the report and the artifacts it cites.

Task:
Separate planning artifacts from completed tasks.
For each claimed outcome, name the supporting artifact and check.
Keep failed, skipped, and unmeasured work visible.

Output:
A table with claim, evidence, supported conclusion, and missing proof.

Constraints:
Do not execute roadmap items.
Do not infer delivery from an attempted send.
Do not infer saved time from task counts.
Use later corrections when records disagree.
Mark unavailable evidence as unverified.

My sandbox passed two tests. The full run was unproven.

Separates a repaired command path from a verified end-to-end scheduled result.

21 lines708 charsSource post
Show prompt
Role:
Review the evidence for an agent runner repair.

Context:
I will provide the failure log, repair diff, test output,
and any subsequent scheduled-run report.

Task:
Identify the original failure. List the changed behavior.
Match each success claim to the command or result proving it.
Check whether the full scheduled job ran after the repair.

Output:
Return the verified repair, the unverified outcome,
and the next concrete check needed to close that gap.

Constraints:
Do not count available or reviewed tasks as completed work.
Do not infer a full-run success from a launcher test.
Apply later corrections without erasing the earlier failure.
Do not include secrets or private paths in the summary.

Gemini first for four jobs. One run finished two tasks.

Checks the reported work behind a provider change and separates it from probes, available tasks, and unresolved failures.

22 lines736 charsSource post
Show prompt
Role:
Review an agent run against its recorded results.

Context:
I changed the provider order for recurring jobs. Read the supplied configuration change, run ledger, and task reports as evidence.

Task:
1. Name the jobs whose routing changed.
2. Separate a successful provider probe from a completed scheduled run.
3. List completed tasks only when the report names their results.
4. Apply later corrections to earlier counts.
5. Identify failures that the routing change did not prove fixed.

Output:
- Routing changed
- Reported results
- Work still unresolved

Constraints:
- Do not count available tasks as completed tasks.
- Do not infer a repaired system from one successful run.
- If receipts disagree, explain the disagreement.

Devlog 2026-07-13: local drive git corruption stalls the q

Explains how this public daily log is produced from the private vault devlog with deterministic redaction.

1 lines279 charsSource post
Show prompt
Take today's private BMD HODL devlog. Redact absolute paths, secrets, private dollar amounts, and Requests/ledger paths. Keep first-person builder voice. Publish as a public daily build note with sections for what shipped, machine overnight work, and tomorrow. Link 5090 Reports.

ai-agents

143 pull requests. Zero dollars.

It splits a weekly agent report into throughput metrics and outcome metrics, and fails the week if only throughput moved.

21 lines611 charsSource post
Show prompt
Role:
You are scoring one week of agent output.

Context:
[Paste the weekly counts: PRs, tests, posts, incidents, revenue, checkout starts.]

Task:
1. List every number that measures volume of work.
2. List every number that measures a result a stranger paid for, or a result that bought back time.
3. Say which list moved.
4. Fail the week if only the volume list moved.

Output:
- Two labeled lists.
- One line: volume week, or outcome week.
- One line: what you will stop counting.

Constraints:
- Keep it short.
- Use the numbers in the paste. Do not invent a conversion.
- Zero is a result. Do not skip it.

A kill rule that expires is not a kill rule

It checks whether a dated kill rule can actually fire, or whether an archive step will swallow it.

23 lines751 charsSource post
Show prompt
Role:
You are auditing one dated kill rule in a repo or vault.

Context:
[Paste the rule, the date, the card or ticket, and the folder that archives stale work.]

Task:
1. Name the condition in a number a stranger can check.
2. Name the date.
3. Name the file that must hold the verdict.
4. Name the writer of that file. A sweep, a TTL, or an archive job is not a writer.
5. Say what happens if the date passes and the writer is silent.

Output:
- Four lines: condition, date, verdict file, writer.
- One line: fires, or expires.
- If it expires, write the missing writer in one sentence.

Constraints:
- Keep it short.
- Use exact paths and dates.
- Do not invent a verdict that was never written.
- Do not treat an archive move as a scored decision.

Agent Memory: Test the Answer After a Correction

It separates checks of a saved correction from tests of the agent's answer and drafts a bounded verification plan.

21 lines783 charsSource post
Show prompt
Role:
Review an agent-memory correction and its tests.

Context:
I will provide the original task, the failed answer, the expert correction,
the proposed file edit, and the current verifier code.

Task:
1. State the behavior the correction should change.
2. List what the current verifier actually observes.
3. Identify any check that reads stored text without rerunning the agent.
4. Specify a fresh-session replay and nearby regression cases.

Output:
For each claim, give its evidence, remaining gap, and smallest useful test.
Keep file-content results separate from observed answer results.

Constraints:
Do not apply edits or change acceptance criteria.
Do not invent a successful replay, model result, or reviewer verdict.
Flag missing inputs and decisions that need the owner.

My agent wrote 126 empty pages and every gate passed

It checks a batch of agent output for filler by comparing the files to each other rather than to a word list.

24 lines827 charsSource post
Show prompt
Role:
You are auditing a batch of files that one agent run produced.

Context:
[Paste 5 to 10 outputs from the same run, with their filenames.]

Task:
1. List the sentences that appear in more than one file, verbatim or
   near-verbatim.
2. For each file, give the share of its prose that is shared text.
3. Name any file whose only unique content is its title or an
   identifier echoed into the body.

Output:
- A table of filename, shared-prose share, and unique-sentence count.
- One line per file: read its source, or filled a template.
- Fail the run when most files are above 90 percent shared prose.

Constraints:
- Keep it short.
- Use exact numbers and file paths when available.
- Do not invent missing measurements.
- Judge whether each file carries information the others do not, not
  whether the writing is good.

Prime Agent hit 95.5% on ARC-AGI-3. I did not install it.

It walks you through the adopt, copy, or measure decision for a new agent framework before you install anything.

28 lines1089 charsSource post
Show prompt
Role:
You are evaluating a newly released AI agent framework for production use.

Context:
Provide the framework name, license, release date, your current agent stack,
where your agents' instructions and state live, what credentials your agents
can reach, and any written rules you have about agent self-modification.

Task:
1. Name who holds the pen on agent instructions in this framework, and who
   holds it in your stack on the review date.
2. List every vendor-disclosed failure or limitation, and the guardrail each
   one implies.
3. Pick one of three outcomes: adopt, copy one component, or measure it as
   benchmark subject matter.
4. If not adopting, write the specific evidence that would reverse the
   decision, with a review date.

Output:
- One-paragraph decision with the outcome named.
- The single component worth copying, if any.
- The reversal conditions and review date.

Constraints:
- Vendor benchmarks count as claims, not evidence.
- A prompt reminder is not a guardrail.
- Do not install anything to answer these questions.
<!-- blog-prompt-scope:2026-08-06 -->

Your Agent's Audit Trail Cannot Be Retrofitted

audits an existing agent workflow for outcome claims that today rest only on logs, and designs the receipt layer to close each gap.

9 lines793 charsSource post
Show prompt
Role: platform engineer hardening an AI agent workflow.
Context: my agents log every action, but nothing verifies claimed outcomes against real state. The workflow description follows this prompt.
Task:
1. List each material outcome the workflow produces (deploys, file changes, published artifacts).
2. For each outcome, state what the action log proves and what it cannot prove.
3. Design a deterministic receipt for each gap: check type (file content, command exit code, path move), the exact check, and where the append-only record should live.
4. Name the one outcome whose false "done" would hurt most, and gate that one first.
Output: a gap table plus an ordered rollout plan for the receipt layer.
Constraints: checks must be deterministic and runnable today. No model-graded evidence.

My Agents Have to Prove What They Did

turns an agent's vague "done" report into falsifiable claims with deterministic checks you can run yourself.

8 lines715 charsSource post
Show prompt
Role: senior engineer auditing an AI agent's completion report.
Context: an agent reported finishing a task in my repository. I have the report and repo access.
Task:
1. Extract every material outcome the report asserts (files changed, configs set, tests passing).
2. For each outcome, write one falsifiable claim with a deterministic check: a file path plus regex, a command with expected exit code, or a path that must exist or be absent.
3. Flag any outcome that cannot be reduced to a deterministic check and name the evidence needed instead.
Output: a table with claim, check type, exact check, and what a failure would mean.
Constraints: no LLM-judged checks. Every check runs from the repo root and can fail.

Use Owner Gates and AgentGuard to Keep AI Agents Moving

It turns a vague agent workflow into owner gates, AgentGuard cost rails, and proof checks.

20 lines880 charsSource post
Show prompt
You are helping me design owner gates and AgentGuard cost rails for an AI agent workflow.

Goal: keep the agent moving on safe mechanical work while stopping for true owner decisions and bounding spend.

Review this workflow:
[paste workflow, agent prompt, or task list]

Return:
1. Owner gates: decisions the agent must stop for.
2. Safe default actions: work the agent can do without asking.
3. Request packet: the exact fields the agent should write when blocked.
4. AgentGuard limits: budget, token, and rate ceilings for the run.
5. Proof checks: commands, files, URLs, or fields that prove completion.
6. Red lines: actions the agent must never take.

Constraints:
- Keep the policy short.
- Prefer files, commands, and runtime checks over trust.
- Do not add new tools unless the existing workflow cannot prove completion.
- Use plain language a builder can maintain later.

AI Agent Memory: What Actually Works in 2026

Audits an agent workflow and turns vague memory into plain-file state, append-only events, and checks that prove each completion claim.

26 lines1455 charsSource post
Show prompt
You are auditing the memory layer for an AI agent workflow.

Your job is to replace vague "the agent remembers" claims with durable state that a human can inspect and a script can verify.

Work in five passes:

1. List every place the agent currently stores state. Separate durable files, append-only ledgers, databases, caches, prompts, and scratch output.
2. For each surface, answer: who writes it, who reads it, what proves it is current, and what happens if two agents write at once.
3. Find any completion claim that is not backed by a file, ledger row, API readback, test result, or live URL.
4. Propose the smallest durable memory system that would survive a cold restart: plain files first, append-only events second, database rows only where they are already part of the product path.
5. Write the exact checks that should run before an agent is allowed to say "done."

Return:

- A table of current memory surfaces and their risk.
- A short list of state that should be deleted, archived, or ignored.
- The proposed durable memory contract.
- The verification commands or scripts that prove the contract.
- One example "done" claim rewritten as a falsifiable claim.

Constraints:

- Prefer files a senior engineer can open and read.
- Do not add a vector database unless the workflow has a specific retrieval failure that plain files cannot solve.
- Do not trust model output as evidence.
- Every important claim needs a deterministic readback.

How to Close the AI Agent Cost Gap at the Call Site

It reviews one agent function for the three cost leaks in this post and writes a capped fix for each one.

21 lines948 charsSource post
Show prompt
Role:
You are a code reviewer who finds AI agent cost leaks.

Context:
[Paste one Python function that calls an LLM API inside an agent loop.]
[Name the model it calls and a typical prompt size in tokens, if you know it.]

Task:
1. Find every retry path and say how many full prompts one failure can send.
2. Find where the function resends conversation history, and estimate the tokens resent per call.
3. Find calls that use a larger model than the task needs.
4. For each leak, write a fix that charges the call to a per-run dollar cap with AgentGuard's BudgetGuard (install: pip install agentguard47; import: from agentguard import BudgetGuard).

Output:
- A table: line, leak, tokens at risk per run, fix.
- The capped version of the function.

Constraints:
- Use only APIs that appear in my code or in the AgentGuard version I have installed.
- Mark every token count you did not measure as an estimate.
- If the function has no leak, say so.

Meta Employees Consumed 60T AI Tokens in 30 Days: 3 Budget Caps

Designs hard budget, token, and rate caps around an agent so a runaway loop cannot drain your account.

20 lines856 charsSource post
Show prompt
Role:
You are adding spend controls to an AI agent before it runs unattended.

My setup:
- Agent and model: [e.g. Claude Code or a GPT-4-class model]
- What it does unattended: [tasks]
- Who pays and the ceiling: [e.g. 50/day hard cap]
- Current controls: [none / soft warnings]

Task:
1. List every way this agent can spend money: tokens, retries, tool calls, sub-agents.
2. Set a hard daily budget and a per-run token ceiling with specific numbers.
3. Add a retry ceiling and a rate limit so one loop cannot spiral.
4. Define the kill condition: what trips it and what happens when it does.
5. Give me the check that runs before each agent run to enforce the caps.

Constraints:
- Caps must be hard stops, not logged warnings.
- Assume the agent will hit an infinite loop at least once.
- Prefer controls outside the model that the model cannot override.

AI Agent Cost in 2026: Build vs Run Cost, With the Token Math

Estimates build and run costs from your inputs, with dated price sources and an explicit payback calculation.

16 lines1036 charsSource post
Show prompt
Role: Help me estimate an AI agent's build and run costs.
Context: Ask me for:
- the model I'm using and where I'll check its current published price (name the page and the date I checked)
- expected runs per day and tokens per run (input and output, separately)
- one-time build cost, if any
- weekly operating and review cost
- weekly value of time saved, if any

Task:
1. Calculate token cost per run, per day, and for 30 days, clearly labeled as a hypothetical based on the numbers I gave you.
2. Keep build cost and run cost as two separate totals, never combined into one number.
3. If I gave you a build cost and weekly savings, calculate simple payback as build_cost / net_weekly_savings, and tell me if net savings is zero or negative so payback isn't finite.
Output: Show inputs and assumptions, separate build and run totals,
and simple payback when net weekly savings are positive.
Constraints: Do not invent survey data, benchmark numbers, or market-wide
price ranges. Use only my inputs and the dated price sources I name.

llama-cpp

Why Your Local Model Fits and Still Fails at Long Context

It compares two controlled memory samples, separates total VRAM changes from cache evidence, and identifies the next workload test.

33 lines1341 charsSource post
Show prompt
Role:
You are checking the memory cost of context for one local model.

Context:
llama.cpp build identifier: [ ]
Model file and quant: [ ]
Offload setting and batch settings: [ ]
Cache type, parallel slots, KV mode, and actual context from logs: [ ]
Runtime cache allocation logs for both loads: [ ]
Idle GPU and host memory sample: [ ]
Sample after load at short context: [ ]
Sample after load at long context: [ ]
Target context for the real workload: [ ]

Task:
1. Confirm both loads used identical settings apart from context size.
2. Subtract the loaded samples and label the result total VRAM delta.
3. Compare cache allocation logs separately; do not infer exact cache cost.
4. Identify missing controls or unavailable measurements.
5. Propose a target-context workload test with peak-memory recording.

Output:
- Fixed settings list
- Total VRAM delta and separate cache-log evidence
- Unknowns that prevent a diagnosis
- One next test and the evidence needed to call the workload a fit

Constraints:
- Do not invent measurements or infer a workload fit from idle samples.
- Do not assume context is always divided equally between slots.
- Mark the result UNKNOWN when any setting differed between the loads.
- Name the pool that ran out before suggesting a fix.
- Suggest one change at a time so the next measurement stays readable.

Prove llama.cpp Tensor Split Used Every GPU

It turns a llama.cpp multi-GPU command and monitor log into a pass, fail, or unknown tensor-split receipt.

28 lines807 charsSource post
Show prompt
Role:
You are reviewing a llama.cpp multi-GPU run.

Context:
Full llama.cpp command: [ ]
Build identifier: [ ]
Detected device order: [ ]
Idle GPU samples: [ ]
Load and generation samples: [ ]
Benchmark result: [ ]

Task:
1. Extract split mode, proportions, main GPU, and device order.
2. Compare idle and active memory for every intended GPU.
3. Check whether each intended GPU shows participation evidence.
4. Return PASS, FAIL, or UNKNOWN with the first broken gate.

Output:
- Pinned configuration
- Per-device evidence table
- Verdict and reason
- One next test that changes only one field

Constraints:
- Do not invent missing device samples.
- Do not treat one utilization sample as a speed result.
- Mark the verdict UNKNOWN when device order is missing.
- Keep the next test bounded and repeatable.

llama.cpp -ngl Flag Explained: 5 Fixes When It Stays on CPU

Diagnoses why llama.cpp still runs on the CPU after you set -ngl 99, using a load-log checklist.

21 lines870 charsSource post
Show prompt
Role:
You are diagnosing why llama.cpp runs on the CPU despite -ngl 99.

My setup:
- GPU, VRAM, driver: [e.g. 8GB 3070, driver 55x]
- Build: [CUDA / Metal / ROCm / CPU-only?]
- Model and quant: [e.g. 8B Q4_K_M]
- Command: [paste flags]
- What the load log shows: [paste the offload and backend lines]

Task:
1. Confirm from the log whether llama.cpp was even built with GPU support.
2. Check, in order, the five usual causes: CPU-only build, no VRAM headroom, wrong device, layer count too low, driver mismatch.
3. Tell me which cause my log points to and the one command that proves it.
4. Give the corrected build or command.
5. Show the exact log line that will confirm the GPU is now in use.

Constraints:
- Read the load log before guessing.
- Do not assume the flag is the problem; usually it is the build or VRAM.
- Give one fix at a time with a check after each.

Q4_K_M vs Q5_K_M vs Q8_0 vs IQ4_XS: GGUF Quants Explained (2026)

Builds a quant comparison plan from actual model files and explicit acceptance checks.

16 lines707 charsSource post
Show prompt
Role: Help me plan a local GGUF comparison.
Context:
- Model repository and revision: [fill in]
- Candidate files and their listed sizes: [fill in]
- GPU, available VRAM, and host RAM: [fill in]
- Context setting and expected longest input: [fill in]
- Task and acceptance checks: [fill in]
Task:
Compare which candidates are worth testing. Separate file size from
runtime memory. Identify missing information before recommending one.
Output:
A candidate table, a repeatable test procedure, and a results template
covering correctness, elapsed time, and peak memory.
Constraints:
Do not invent a quality percentage or performance measurement.
Label estimates. Do not claim a configuration fits until tested.

ai-security

Incident response needs a local model you already trust

Turns a local open-weight model into a first-pass incident triage analyst. Paste raw logs after it and the model returns an ordered timeline, the likely entry point, the lines worth escalating, and next containment steps, with none of that data leaving your machine.

9 lines503 charsSource post
Show prompt
You are an incident-response analyst working on logs from a system I own. I will paste raw log lines below. This is my own infrastructure during an active investigation.

From the logs, produce:
1. A timeline of events in order, with timestamps.
2. The most likely entry point and how access escalated.
3. The five log lines that most warrant escalation, quoted exactly.
4. Concrete next containment steps, ranked by priority.

Give a confidence level for each section and flag anything ambiguous. Logs:

ai-tools

Aymo AI Pricing 2026: Is $39/mo Worth It vs Free?

Runs a free-versus-paid decision on any AI subscription using your real usage, not the vendor pitch.

22 lines817 charsSource post
Show prompt
Role:
You are helping me decide whether a paid AI tool tier is worth it.

The tool:
- Name and tiers: [e.g. Aymo free vs 39/mo]
- What the paid tier adds: [features, higher limits]

My usage:
- What I actually use it for: [tasks]
- How often: [per day or week]
- The limits I hit on free: [rate caps, missing features, none yet]

Task:
1. List which paid features map to a real limit I actually hit.
2. Estimate my monthly value from the paid tier in hours or dollars saved.
3. Compare that to the price and to one cheaper or free alternative.
4. Give a clear verdict: stay free, upgrade, or switch, with the single deciding reason.

Constraints:
- Ignore features I will not use.
- Do not accept marketing claims without a usage-based reason.
- If the data is thin, say what to measure for two weeks before deciding.

raspberry-pi

Raspberry Pi 5 Offline Voice Assistant: 6 Models Tested (2026)

Designs an offline voice assistant that runs on a Raspberry Pi 5, within a real latency budget.

19 lines800 charsSource post
Show prompt
Role:
You are designing an offline voice assistant on a Raspberry Pi 5.

Constraints of my build:
- Hardware: Raspberry Pi 5 [RAM: e.g. 8GB], [mic/speaker], no cloud.
- Target: wake word to spoken reply under [e.g. 2 seconds].
- Everything runs locally, no internet calls.

Task:
1. Pick a stack for wake word, speech-to-text, the local LLM, and text-to-speech that fits a Pi 5.
2. Give a rough latency budget per stage that adds up under my target.
3. Name specific model sizes and quant levels that run on this hardware.
4. Flag the stage most likely to blow the latency budget and how to cut it.
5. List the first three things to build and test, in order.

Constraints:
- Prefer components that install and run on ARM today.
- Do not assume a GPU.
- Keep memory use inside the Pi RAM budget above.
For AI agents: this whole library is plain markdown at bmdpat.com/prompts.md. Fetch it, cite a prompt by its source post, or paste it straight into your context. Each prompt card is addressable at /prompts#<slug>.