Skip to content
[bmdpat]
LIVE

AI agent cost control with runtime guardrails.

AgentGuard checks budgets, repeated tool calls, retries, and elapsed time in instrumented Python code. Handle its exceptions to stop the next operation. Free, MIT-licensed, and local by default.

Free and open source. The SDK runs inside your process, with no account and no hosted service for the first guarded run.

3,672 downloads / last 180 days648 downloads / last month, includes a release190 MCP downloads / last month4 GitHub starsMITClaude Code, Cursor, and Codex MCPMCP v0.2.2v2.0.0

Observed package downloads, updated hourly. SDK totals use Pepy first, then Pypistats no-mirror overall data when Pepy is unavailable; SDK recent-month counts use Pypistats; MCP recent-month counts use npm. Downloads are a usage signal, not an adoption proof. Tracked copy is shown when metric upstreams are unavailable.

View on GitHubConnect MCPView on PyPI

AgentGuard publishes a verified PyPI package, npm MCP package, official MCP Registry entry, MIT license, source repo, local proof commands, and no required hosted service for the first guarded run.

Source proof
Verified PyPI links
PyPI shows repository, docs, issues, and changelog links for bmdhodl/agent47.
Local proof
No key required
The SDK can run doctor, demo, quickstart, report, and incident commands locally.
Supply chain
MIT, Python 3.9+
The source repo declares the MIT license, and the PyPI package advertises Python 3.9+ support.
Runtime scope
Zero runtime deps
The core SDK keeps guardrails in-process instead of adding a required hosted service.

§ 001 / RUNTIME RISK

Agents stop being toys the first time they touch production.

Developers are right to start with scripts and demos. The risk starts when the same agent gets credentials, API budgets, database access, or a deploy path. AgentGuard is the boring runtime layer that fails closed before the run keeps moving.

Toy scripts become production systems

If an agent can spend money, call tools, or touch real data, it is no longer a toy script. It needs limits before the next run.

The bad run costs real money

If an agent can delete records, send email, merge code, or touch prod, missing runtime limits are a bug. Give the run a stop condition before damage keeps growing.

Prevention beats post-mortems

Observability tells you what happened after the run. AgentGuard raises inside the process before the bad run keeps going.

Runtime is the durable layer

Framework and provider checks are useful, but they move with each stack. AgentGuard keeps the runtime ceiling in your Python code.

1

Install the SDK

Put hard budget, loop, timeout, and rate limits inside the Python agent code that spends money.

2

Open the dashboard

Create a read key, keep event history, and see the traces your team needs later.

Create read key
3

Connect MCP

Let Claude Code, Cursor, Codex, or another MCP client inspect traces, alerts, usage, costs, and budget health.

Not sure what to guard?

Scan your agent idea first.

Describe the workflow, tools, users, and expected usage. The Agent Roadmap Scanner returns likely runtime risks, first guardrails, and where AgentGuard fits.

Run the scanner

MCP setup

Give your coding agent a read-only view of AgentGuard.

The MCP server does not replace the Python SDK. Guards only cover instrumented code and do not cancel in-flight provider calls. The MCP server provides read-only visibility; it does not enforce runtime limits. It lets your editor inspect retained traces, alerts, usage, costs, and budget health after you connect a read key from the dashboard.

Claude Code / Claude Desktop
{
  "mcpServers": {
    "agentguard": {
      "command": "npx",
      "args": [
        "-y",
        "@agentguard47/mcp-server"
      ],
      "env": {
        "AGENTGUARD_API_KEY": "ag_your_read_key_here"
      }
    }
  }
}
Cursor .cursor/mcp.json
{
  "mcpServers": {
    "agentguard": {
      "command": "npx",
      "args": [
        "-y",
        "@agentguard47/mcp-server"
      ],
      "env": {
        "AGENTGUARD_API_KEY": "ag_your_read_key_here"
      }
    }
  }
}
Open raw mcp.json

Command

npx -y @agentguard47/mcp-server

Credential

Read-only ag_... API key

Tools

traces, alerts, usage, costs, and budget health

Order of operations

What happens, in the order it happens.

Every box below is a real constructor from the quickstart. Read it top to bottom and you have read the whole runtime.

  1. Trigger01

    The agent starts a run

    Guards are constructed before the loop, so their ceilings exist before the first call can spend anything.

    Tracer(service=...)

  2. Ceiling02

    TimeoutGuard wraps the whole run

    Checks elapsed time when you call check() and when the context exits. It does not cancel an in-flight request. Configure request timeouts on the provider client too.

    max_seconds

  3. Ceiling03

    BudgetGuard counts each call

    Call-count checks can run before an operation. Provider adapters record returned usage; the request that crosses a cost limit can already have spent money.

    max_cost_usdmax_calls

  4. Ceiling04

    LoopGuard checks for repeats

    The same call with the same arguments trips the repeat ceiling before it can run again.

    max_repeats

  5. Ceiling05

    RateLimitGuard meters the call rate

    This one is handed to the tracer rather than called directly, so every traced call passes through it.

    max_calls_per_minuteTracer(guards=[rate])

  6. If / Else06

    Any ceiling reached raises AgentGuardError

    There is no silent degrade. The exception carries the reason, and your own except block decides what happens next.

    except AgentGuardError

The guards are plain Python objects. Nothing is injected, patched, or inferred, you place them where money and loops actually happen.

Quickstart

Install it, trip a guard, then wire it into your agent.

AgentGuard is not magic middleware. You record usage where money or loops happen. When a limit trips, it raises an AgentGuardError with the reason.

1. Install

Install `agentguard47` from PyPI. Import `agentguard` in Python. That is the public SDK module exposed by the package.

macOS/Linux venv

Use when a project virtual environment is active and you want pip tied to that Python interpreter.

Windows venv

Use the Python launcher on Windows when it selects the environment you run.

uv project

Use in a uv-managed project to add AgentGuard to pyproject.toml.

Short pip

Use when pip already points at the Python environment your agent runs in.

Or install via the Claude Code skill / Codex skill / Markdown install guide.

2. Prove local setup

Run the doctor before wiring your real agent.

agentguard doctor --json

3. Stop a simulated fourth operation

Save this as first_run.py. Run it with the Python interpreter where you installed AgentGuard. If you installed with py or python3, use that launcher in the run command too.

first_run.py
from agentguard import BudgetExceeded, BudgetGuard

budget = BudgetGuard(max_calls=3)

for step in range(1, 6):
    try:
        budget.consume(calls=1)
    except BudgetExceeded:
        print(f"Stopped before operation {step}.")
        break
    print(f"Operation {step} allowed.")

Expected output (exit 0)

Operation 1 allowed.
Operation 2 allowed.
Operation 3 allowed.
Stopped before operation 4.

The guard runs before each simulated operation. It allows three and stops the fourth. The script exits with code 0 because it handles the expected exception. No model requests are made. This example checks call count; it does not measure token usage or cost, and it creates no trace file.

If Python cannot import agentguard, install agentguard47 using the same Python interpreter. Keep budget.consume(calls=1) before the operation you want to protect. Do not catch the exception and continue with that operation.

One call limit across two workers

Two workers share one local key. The example allows one simulated call and stops the other before sending. It needs AgentGuard 1.4.0, makes no model requests, and uses a fresh temporary store on each run.

Open the runnable shared-limit example

Save the file as shared_call_limit.py, then run python shared_call_limit.py with the interpreter where you installed AgentGuard. Look for dispatched: 1, stopped: 1, and fixed: true in the RESULT output. This is a local call limit, not a provider invoice cap or a limit across machines.

A successful demo checks setup. Share an optional redacted demo report. For a real workflow, email optional workflow feedback with the task, guard, and observed result. For repeat use, include a redacted project label and dates at least seven days apart. Review it before posting. Do not include secrets or customer data.

After the first run: wire more guards

The first example below stays local. The OpenAI example requires its SDK, an API key and paid model requests.

The OpenAI adapter records returned usage and blocks subsequent calls once the budget is exhausted. A request can exceed the dollar limit before its cost is known. The client request timeout is separate from the elapsed-time guard.

quickstart.pyPython
openai_wiring.pyPython
1from agentguard import AgentGuardError, BudgetGuard, TimeoutGuard, Tracer2from agentguard.instrument import patch_openai3from openai import OpenAI45budget = BudgetGuard(max_cost_usd=2.00, max_calls=20)6timeout = TimeoutGuard(max_seconds=300)7tracer = Tracer(service="support-agent")8patch_openai(tracer, budget_guard=budget)910client = OpenAI(timeout=30.0, max_retries=0)1112try:13    with timeout:14        timeout.check()15        response = client.chat.completions.create(16            model="gpt-4o-mini",17            messages=[{"role": "user", "content": "Reply with one short greeting."}],18        )19        print(response.choices[0].message.content)20except AgentGuardError as exc:21    print(f"AgentGuard stopped the run: {exc}")

What it does

Put checks around the Python operations you control. Record only the events you choose.

Repeated tool calls

Check tool names and arguments before dispatch. A repeated-call exception lets your application stop the loop.

from agentguard import LoopGuard

loop = LoopGuard(max_repeats=3)
loop.check("search", {"query": "agent budgets"})

Local trace files

Write the events you instrument to a local JSONL file. No hosted account is required.

from agentguard import JsonlFileSink, Tracer

tracer = Tracer(sink=JsonlFileSink("traces.jsonl"))
with tracer.trace("agent.run") as span:
    span.event("tool.call", data={"tool": "search"})

Provider integrations

Instrument Python calls to Anthropic, OpenAI, and local LLMs. Record the usage and tool events your guards need; model choice does not add automatic coverage.

from agentguard import BudgetGuard, Tracer, patch_openai

budget = BudgetGuard(max_cost_usd=5.00)
tracer = Tracer()
patch_openai(tracer, budget_guard=budget)

Budget checks

Account for usage at your instrumented boundary. A limit exception stops subsequent work when your application handles it.

from agentguard import BudgetGuard

budget = BudgetGuard(max_cost_usd=5.00, max_calls=50)
budget.consume(cost_usd=0.02, calls=1)

Offline proof

Check installation and see budget, loop, and retry stops before connecting a model provider.

agentguard doctor
agentguard demo

Local by default

The local demo does not send your trace data. You choose whether to configure a hosted sink or share a redacted feedback report.

agentguard demo --feedback

Built for

Teams who need to prove what their agents did, not just hope they behaved.

Engineering in regulated industries

Healthcare, finance, government. Self-hosted policy + audit that meets your compliance review.

Platform teams shipping internal agents

One governance layer across every agent your developers deploy. MCP-aware out of the box.

Security and compliance reviewers

Approve agent deployments with confidence. Decision logs, payload digests, replayable history.

Use this when

  • Your agent can call paid APIs in a loop.
  • You need limits on recorded cost, calls, elapsed time, or call rate.
  • You are shipping agents across more than one framework.
  • You want local guardrails first and MCP visibility later.

Do not use this when

  • You only need prompt-injection defense.
  • You only need JSON schema or output validation.
  • You need human approval workflows before every action.
  • You need a runtime SDK outside Python.
  • You expect the MCP server to enforce Python runtime guards by itself.

Tried the local demo? Share a first-run result. This voluntary public form asks whether it worked. Inspect your report before sharing; no traces or secrets.

Feedback

Bug or suggestion? Include the version. Leave out private data.

Send AgentGuard feedback

Opens your email app, or write to pat@bmdpat.com.

Get AgentGuard release updates

Start with the local quickstart. Get changes and upgrade instructions when a stable release ships.

Get the AgentGuard quickstart and stable release updates, plus at most one Local AI Lab Note on Friday when there is something worth sharing. No account needed. One-click unsubscribe stops all BMD Pat emails. Privacy.

Why this exists

I wrote AgentGuard because the existing options all point at a different problem. Lakera is about prompt injection. Guardrails AI is about output validation. Platform-native guardrails ship with a single vendor and lock you in. None of them stop an agent from burning $200 of OpenAI credit in a runaway loop at 2 AM.

AgentGuard is runtime only. It sits between your code and the model, counts cost and tool calls and wall-clock time, and raises before you get the surprise bill. It works the same whether you're calling GPT-5, Claude, a local llama.cpp server, or something I haven't heard of yet.

I built it for my own agents. I run it on my autotrader and on the agents that post this blog. If it works for me at 2 AM, it should work for you.

FAQ

What is AI agent cost control?
AI agent cost control checks spend, tool calls, elapsed time, and call rate against limits. AgentGuard raises exceptions at instrumented Python boundaries so your application can stop the next operation. It does not cancel a provider request already in flight.
Do I need the SDK or the MCP server?
Use the SDK when you want hard runtime enforcement inside Python code. Use the MCP server when you want Claude Code, Cursor, Codex, or another MCP client to inspect retained traces, alerts, costs, and budget health from the hosted dashboard.
Why is the install name different from the import name?
Install `agentguard47` from PyPI. Import `agentguard` in Python. That is the public SDK module exposed by the package.
Why not use framework-native guardrails?
Framework guardrails usually handle model policy, tool schemas, or output validation. AgentGuard handles runtime failure: budget, loop, timeout, and rate limits. It is the layer you keep when you swap frameworks.
What data leaves my machine?
The OSS SDK does not need a network call to enforce guards. Trace events go to the sink you choose. If you use a local JSONL sink, they stay local. If you wire the hosted dashboard, send runtime metadata only. Do not put prompts or responses into trace data unless you intend to store them.
What is open source, and is anything paid?
The Python SDK and MCP server are open source and run inside your process. There is no paid plan on top of them.
Is this a one-person product?
AgentGuard is a one-developer product by design: small surface, local-first SDK, public docs, and no custom integration promises. It is maintained as open source, with no paid plan.
Which stacks does this work with?
Raw Python works directly. The SDK also exposes modules for OpenAI, LangChain, LangGraph, and CrewAI. You still decide where tool calls, costs, and events are recorded.

THE LOOP

One of three tools.

AgentGuard caps the spend during the run. showwork checks the claims after it. BMD shows the runs in one Windows workspace. Each tool works on its own.

Install the SDK and guard your first run.

Free, MIT licensed, no account. Read the source before you trust it.

View on GitHub