Skip to content
[bmdpat]

Field journal

Blog archive

Every published note, with the newest work first. Use the topic links to stay inside one search intent.

  1. AI agents need two rails before they can run unattended: owner gates for judgment and AgentGuard for spend. Without both, the operator becomes the fallback.

  2. Most agent memory systems add complexity faster than value. This is the small set that actually compounds for one person running a fleet: files, ledgers, and strict verification.

  3. AI agents report work as done that they never did. Make every completion a falsifiable claim a script can verify before you trust it.

  4. An append-only event log lets you replay exactly what your AI agent did, and catches the crashed runs a status field hides.

  5. Automated recovery only fixes a broken machine. When the real failure is an empty queue, retrying does nothing forever. Two failures, one red box, opposite repairs.

  6. A spend ledger that counts missing billing data as $0 hides exactly the unattended agent spend you built it to catch.

  7. Salesforce shipped roughly 20,000 Agentforce deployments and found 90% of agent work happens after launch. Here is what that means for a solo builder running a small agent fleet.

  8. Anthropic says 80% of its new code is Claude-authored. Here is how solo builders manage the review burden.

  9. A June 2026 Mem0 survey of 8 major agent harnesses found that over half of them leak memory across users. Here is why keyword retrieval is a security risk and how to fix it.

  10. Estimate the VRAM required to run local LLMs like Llama 3 with our interactive calculator. Compare quantization levels like Q4 and Q8 to plan your hardware.

  11. The seat price isn't the real cost anymore. Real 2026 prices for GitHub Copilot, Cursor, and Claude Code, pulled from each vendor's own page.

  12. Agentic coding made writing code free. The slow part is now reviewing a queue of plausible PRs.

  13. Agent count is a vanity metric. It tells you about volume, not value. Here is what I track instead after running a one-person AI fleet.

  14. I run a one-person company on scheduled agents and gave almost none of them memory. They write to files instead. Here is why that wins.

  15. Running a 70B across 2 GPUs and hitting OOM? The llama.cpp multi-GPU reference: --tensor-split ratios, --split-mode layer vs row, --main-gpu, and the VRAM math from a real two-card rig.

  16. The cost gap between what an AI agent could cost and what it does cost is 40%. You close it at the call site, not in a dashboard. Here is how.

  17. A 2026 Mem0 survey found 57-71% cross-user memory contamination across major agent frameworks. Here is why it happens and how to stop it.

  18. JPMorgan turned on AI for 250k people. The quiet line is that the usage racks up fees. Here is how to control the bill before it arrives.

  19. Anthropic filed for IPO at a $47B run-rate while 40% of enterprise customers report under 10% cost savings from Claude. Here is how to close that gap.

  20. A repair agent in my own pipeline failed the same check 27 times in a row. Each try was a paid model call. Here is why uncapped retries quietly burn money, and the two-line fix.

  21. Anthropic banned 832 accounts for AI-enabled attacks. What it means for teams running AI agents.

  22. JPMorgan just switched on AI for 250,000 employees. The headline is workforce shift. The quiet story is enterprise AI cost, and why token spend runs away without controls.

  23. Most AI advice tells you to ship more agents. Here is the honest opposite: the four times a plain script and a human beat an agent, learned running a fleet daily.