Operating Note 002: Right Agent, Right Model, Best ROI
Operating Note 002 (30 Aug 2026): Positioning + practice. I am niching toward agent and model cost optimization with OpenRouter as the default routing layer — right agent role, right model, best ROI. Builds on Operating Note 001, the €35/mo content org, and the hands-on OpenRouter + Cursor setup.
Positioning (locked for now)
Category: Agentic cost engineering — model routing for production agent systems.
One-liner: I design agent setups that route every task through OpenRouter to the cheapest capable model — so you get reliable outcomes without frontier bills on routine work.
For whom: Founders and tech leads who already have (or want) agents, but bleed money or chaos on “one model for everything.”
Not: Generic AI strategy decks. Not model research. Not “we’ll fine-tune GPT.”
Proof I already have: live OpenRouter routing tables, ~€35/mo infra band, multi-agent content/dev loops, Company OS as the operating layer.
Most teams still buy agentic capability the wrong way: one frontier model, one mega-agent, then wonder why the bill and the failure rate both climb. The leverage is routing discipline.
Why OpenRouter is the spine
I standardized on OpenRouter as the model control plane:
- One key → many models (free + paid tiers)
- Swap routes without rewriting agents
- Weekly promote/demote based on quality × cost
- Keep expensive frontier capacity for class C/D work only
Day-to-day coding already runs this way (writeup). The niche is extending the same discipline to every agent role in production — not just the IDE.
Four task classes (before you pick a model)
| Class | Examples | OpenRouter posture |
|---|---|---|
| A — Cheap & frequent | Explain, triage, extract, outlines | Free / tiny models |
| B — Structured production | Code with tests, schema-bound drafts, tool calls | Cheap mid-tier; budget retries |
| C — Judgment / critique | QA another agent, risky diffs, architecture | Stronger / different model than producer |
| D — Irreversible / external | Publish, email, prod deploy | Strong model + human approval |
If everything is C/D, you overpay. If D is treated as A, you underpay and then overpay in incidents.
Agent role ≠ model
Two axes:
- Which agent role owns the task? (researcher, implementer, critic, publisher)
- Which OpenRouter model powers that role?
MCP + approvals keep axis 1 safe (MCP lessons). OpenRouter keeps axis 2 measurable and replaceable. Company OS decides whether the work should run at all (layers).
ROI is not token price
A free model that fails three times and burns an hour of human cleanup is expensive.
Track:
- Infra (~€35/mo band for the studio stack)
- OpenRouter spend per workflow / accepted outcome
- Human minutes on approval and cleanup
- Silent-failure rate
Target: accepted outcomes per euro, under an explicit quality bar.
Practice loop
- Name the task class (A–D)
- Pick the smallest agent role that can own it
- Pick the cheapest model that historically clears the bar (via your router or allowlist)
- Log spend + failure mode
- Promote/demote weekly — ideally from an automatic model watch
That loop is the craft I am doubling down on — and the conversation I want when cost or chaos is the bottleneck.
Next in this lane
Series hub: Agent cost & model selection — start here.
Already shipped:
- Automatic model watch
- When cheap models get expensive
- Routing tables for content and data agents
- Eval fixture pack
- OpenRouter + Cursor example
Still planned: Company OS spend policies per role; domain-specific fixture packs.
Work with me
/agentic-engineer-stockholm · /hire · what is agentic engineering