Skip to content
[bmdpat]

Field journal

Blog archive

Every published note, with the newest work first. Use the topic links to stay inside one search intent.

  1. NVIDIA Blackwell delivers 35x lower cost per token vs Hopper. That makes AI agents cheaper to run and harder to stop. Here's why that flips the runtime guard argument upside down.

  2. Simon Willison frames AI-assisted security research as proof of work: more tokens in, more bugs found. That's an economic reality. Here's what the spend curve actually looks like and how to put a floor under it.

  3. Flatiron Health toured AI-native startups in SF. One PM covers five companies, Claude Code is replacing Cursor, non-engineers are shipping production. I'm running the same model from Tennessee as a solo holding company. Here's what that actually looks like.

  4. Per-token billing is coming for every AI provider within six months. Here's how your agent costs change, and how to cap them at runtime before the bill spikes.

  5. Claude Code bill creeping up? Cache writes cost 1.25x base on the 5-minute TTL, 2x on the 1-hour TTL, and every idle gap over 5 min re-triggers one. The $0.15/day-per-session math and how to cap it.

  6. Blackwell rental hit $4.08/hr. CoreWeave raised prices 20%. Anthropic restricted their newest model to 40 orgs. Meanwhile, consumer GPUs are sitting idle.

  7. Will Larson says agents should be scaffolding, not permanent infrastructure. I run 12 agents overnight. Here's what I kept as agents and what I converted to code.

  8. 1 in 35 GenAI prompts carries high risk of data leakage. MCP makes the attack surface worse. Here's what builders need to know.

  9. Tomasz Tunguz published a 2x2 for categorizing AI projects. Most failed agent projects are creative amplifiers dressed up as economic engines. Here is how to tell which quadrant you are actually in.

  10. Anthropic shipped a pattern where a cheap model runs the loop and escalates to Opus only when it needs to. The pattern works on any two-model setup. Here is the math and the playbook.

  11. Martin Fowler published a pattern for turning individual AI interactions into collective improvement. We had already built it. Here is how our 12-agent vault system maps to his four signal types.

  12. Mythos found zero-days in every major OS. Nature documented AI deception in peer review. War games showed AI escalating to nukes. Three studies, one conclusion: your agents need hard limits.

  13. North Korean threat actors are targeting AI coding tools. Trojanized npm packages hunt for .cursor, .claude, .gemini, and .windsurf directories to steal API keys and source code.

  14. PostHog ships to thousands of daily agent users. They rebuilt their AI architecture twice before getting it right. Here are the 5 rules they distilled, reframed for builders shipping agent features.

  15. Meta employees burned 60T AI tokens in 30 days, ~$50K per head, no cap set. Your agent bill is on the same curve. The 3 budget controls Meta skipped, plus guardrails that catch overruns early.

  16. Researchers tested 428 LLM API routers. Nine were actively injecting malicious code. One drained ETH from a private key. Here is what this means for your AI agents.

  17. Three AI safety papers came out this week. Reading them back to back was jarring. If you run agents in production, this is worth 5 minutes.

  18. Is OpenClaw production-ready? We ran 3 real workloads, RAG, tool-calling, multi-step chains, against LangGraph. Here's exactly where it wins and where it breaks. (2026)

  19. Martin Fowler named the AI feedback flywheel. We built the same system independently. Here's our exact implementation, vault, agents, guardrails, and weekly cadence.

  20. The market is flooded with people claiming to build AI agents. Here's how to tell who can actually ship one, and what questions to ask before you pay anything.

  21. Need agents from different vendors to talk? Google's Agent2Agent (A2A) protocol explained: what it does, when it ships in 2026, and the 3-line config that makes your stack A2A-ready today.

  22. Not sure what to set --n-gpu-layers to? -1 offloads all layers, 0 keeps it on CPU, a number splits the model. VRAM headroom rules, examples, and the CPU-fallback fix. (2026)

  23. Is Aymo AI worth $39/mo? We ran every tier for 30 days, the free plan's 50-call cap hits fast. Verdict: paid only beats ChatGPT Plus and Claude Pro past ~50 calls/day. (2026)

  24. Want a private voice assistant, zero cloud, no subscription? A Raspberry Pi 5 runs it offline at sub-2s latency. We tested 6 local models on real hardware; see the winner. (2026)