Skip to content
[bmdpat]

Field journal

Blog archive

Every published note, with the newest work first. Use the topic links to stay inside one search intent.

  1. A Stockholm cafe gave its purchasing agent a credit card and a vague prompt. $21,000 later it owned 6,000 napkins and no bread. Here is the exact runtime guardrail that would have caught it on call number two.

  2. The May 14 autotrader review is done. The account is up 7.7% before compute, still negative after compute, and still lagging SPY and BTC. Decision: keep V2 paper-only, add no new live money, and revisit after the next scorecard.

  3. I spent the week tightening the AgentGuard release path, shipping proof-heavy docs and perf fixes, and keeping the benchmark gap visible instead of hand-waving it away.

  4. Closed loops shipped, the funnel got narrower, and the weekly scoreboard got more honest.

  5. Which GGUF file do you download? Start with Q4_K_M when VRAM is tight, test Q5_K_M or Q8_0 on your real task, keep the smallest that passes. What each quant means, IQ4_XS vs K-quants, a test to copy.

  6. Agent skills are becoming a distribution layer for developer tools. The practical move is one source package that can show up in PyPI, Claude-style skills, and skills.sh.

  7. April 2026 made one thing clear: chat subscriptions are best-effort tools. Builders need API-level budgets, rate limits, and kill switches when the work matters.

  8. Reflex.dev measured a 45x token cost gap between computer-use agents and structured APIs for the same task. Here's why, and the decision rule that keeps your bill sane.

  9. PocketOS lost their production database backups to a Cursor agent. Here's what runtime spend rails actually catch, what they don't, and the layered defense your agents need before production.

  10. A 1,764-app audit found 7% had open Supabase databases and 15% of Bolt apps had hardcoded secrets. The fix takes ten minutes.

  11. No metering, no per-team caps, no dashboards. Uber spent its entire 2026 AI budget on Claude Code in just 4 months. The 5-step pattern behind every runaway AI bill, and the fix that stops it.

  12. I need my agent to do X. Skill or MCP? A short decision rule with worked examples for small-business agent builders.

  13. Before you ship an AI agent for a client, prove budget caps, loop detection, alert proof, remote kill, and retained incident history.

  14. The demo worked. Then the same CrewAI tool call retried until the run became an operator problem.

  15. A trace tells you what happened. A kill switch changes what happens next.

  16. Cloudflare shipped agent flows that create accounts, buy domains via Stripe, and deploy infrastructure end-to-end. Good news for builders. Sharper case for runtime budget enforcement than any hypothetical we have used.

  17. OpenAI shipped guardrails in the Agents SDK last month. They validate behavior. They do not enforce spend. Here is the gap and how to close it.

  18. Microsoft just shipped agent-sre on PyPI. Seven packages: SLOs, error budgets, circuit breakers. Here is what it does, what it does not, and why solo builders still need agentguard47.

  19. I built a memory API agents can pay for. The actual problem isn't whether they can pay. It's per-tool caps, per-agent budgets, kill switches, and spend visibility.

  20. Got a 402 from an API? RFC 9110 reserves the code but defines no way to pay. What HTTP 402 Payment Required means, how x402 V2 uses it (PAYMENT-REQUIRED, PAYMENT-SIGNATURE), and the V1 vs V2 headers.

  21. Stripe doesn't ship to LLMs. Every vendor signup form assumes a human at the door. Here is what changes when wallets become the access primitive.

  22. An LLM just paid me $0.001 to remember something. The agent has no account, no API key, no credit card. It just signs a USDC transfer and gets back a 200.

  23. Three studies dropped in the last few months. GPT-5.2, Claude Sonnet 4, and Gemini 3 Flash all escalated to nuclear options 95% of the time in war game scenarios. AI found exploitable vulnerabilities in every major OS and browser. And a Nature paper documented AI disabling its own oversight. Here is what that means if you are running agents in production today.

  24. Stanford, Karpathy, and Bridgewater independently confirmed that one person plus N agents is the right architecture. I have been running it for a holding company. Here is what it looks like.