Skip to main content

Guardrails in practice: approvals, scoped tools and spend policies

field note•2 min read

Autonomy without boundaries is just expensive improvisation.

On this site, “agents” means systems that keep running after you close the laptop — with roles, tools, and memory. The interesting engineering problem is not how to make them act. It is how to keep them scoped, measurable, and interruptible. That is the guardrail layer.

Four controls that actually show up in production

1. Approvals on irreversible actions

Anything externally visible — publish, send, customer-facing reply, spend that cannot be rolled back — stays behind a human gate until the loop has clean telemetry. Draft tools and publish tools are separate. The model may prepare; a person (or a strict policy) releases.

See the approval shape in The Unnamed Roads and the production automation case (AI automation in production).

2. Scoped tools (MCP is not a free pass)

An MCP gateway without least privilege is a keyring with a chat UI. Production lessons: allowlists per agent role, read-only by default, separate credentials downstream, and a kill switch that does not require a redeploy. Full write-up: MCP gateway lessons.

3. Evals before promote

Routing and model choice need fixtures, not vibes. Keep a small battery of cases for the lanes that matter (content voice, SQL safety, tool-call shape) and fail closed when a candidate model regresses. Start here: eval fixture pack.

4. Spend policies per role

Do not share one opaque budget across researcher, implementer, critic, and publisher. Cap unit cost and volume by role; tag spend; define hard-stop behaviour; review weekly. Publisher roles get the strictest caps and mandatory approval. Failure-priced thinking: when cheap models get expensive. Series hub: agent cost & model selection.

What this is not

Guardrails are not a longer system prompt. Prompts do not revoke credentials, stop a runaway loop, or explain last week’s token bill.

They are also not “Company OS” as a product name. Company OS is a coordination thesis — map, ownership, approvals — described in the architecture note. Guardrails are the concrete edges of that map.

Bottom line

If an agent can act, it needs a boundary you can name: who may call which tool, what requires a human, how quality is checked, and how much it may spend. Without those four, you do not have an agent system. You have a demo with a credit card.