Model Routing Without a Router: Letting the Environment Pick the Model
Mandible colonies pick a tier - a frontier API, a model on your own GPU, or a shell script - by reading labels, failure marks, and lineage already sitting in the environment. No classifier service, no dispatcher, no retry state, and every routing decision leaves a trail.
Every multi-agent system eventually hits the same question: which model should handle this?
A frontier model for the hard refactor. Something small and local for the typo fix. Something in between for the rest. Retry on something stronger when the cheap one gives up. The usual answer is a router - a classifier service in front of your agents, a config file mapping categories to models, a queue per tier, and a growing pile of special cases. It works, but now you have a scheduler to run, and every agent has to ask it before doing anything.
Mandible doesn't have a router. It has something simpler.
The environment already knows
Mandible is a stigmergy framework: agents coordinate the way ants do, by leaving and sensing signals in a shared environment - a filesystem, a GitHub repo, a Dolt database. No agent talks to another agent. They read the substrate and write to it.
Which means the information a router would need is already there:
- A human put
complexity:highon the GitHub issue - The colony that created the task tagged it
kind:bug - The last attempt failed and left a mark
- The artifact this task replaces was rejected by a critic - that's in the lineage
A colony doesn't need to ask anyone what tier a task deserves. It looks at the signal it's holding and acts.
Letting the signal choose the model
Every action provider takes a model. Usually that's a tier alias - fable, opus, sonnet, haiku - resolved to a current model ID at call time, so upgrading Mandible moves every colony forward without touching colony code:
withClaudeCode({ model: 'sonnet', prompt })
Anything that isn't an alias passes through untouched: a pinned ID like claude-opus-4-6, a Bedrock ID, a gpt- or gemini-prefixed model, or the name of whatever you have loaded in vLLM. Either way, model can be a function of the signal:
withClaudeCode({
model: (signal) => signal.meta.tags?.includes('complexity:high') ? 'opus' : 'sonnet',
prompt: (signal) => `Implement: ${signal.payload.description}`,
})
Same provider, same prompt, same tools. Only the model changes. For anything with more than one rule, selectModel keeps it declarative:
withClaudeCode({
model: selectModel({
rules: [
{ match: (s) => escalationLevel(s) > 0, model: 'opus' },
{ match: (s) => s.meta.tags?.includes('complexity:high'), model: 'opus' },
{ match: (s) => s.meta.tags?.includes('complexity:low'), model: 'haiku' },
],
default: 'sonnet',
}),
prompt,
})
That's most people's model routing, done. No new concepts. The rules read the same whatever the strings are.
Different tiers, different tools
Sometimes the tiers aren't just different models - they're different kinds of work, from different vendors, on different machines. A hard issue wants a full coding agent with file access. An easy one wants a single structured-output call, and there's no reason that call has to leave the building. A chore wants a shell script. withModelRouter dispatches to entirely different handlers:
const local = vllmStructuredProvider({ endpoint: 'http://localhost:8001' });
colony('worker')
.in(githubEnv)
.sense('task:ready', { unclaimed: true })
.retry(3)
.do('work', withModelRouter({
routes: [
byEscalation(1, withClaudeCode({ model: 'opus', prompt })),
byTag('complexity:high', withClaudeCode({ model: 'opus', prompt })),
byTag('complexity:low', withStructuredOutput({ provider: local, model: 'Qwen3-Coder-Next', schema, prompt, route })),
byTag('chore', withBash({ command: 'npm run lint -- --fix', output })),
],
fallback: withClaudeCode({ model: 'sonnet', prompt }),
}))
The handlers there are a hosted coding agent, a Qwen model on a box in the corner, and npm run lint. The router doesn't know the difference between them.
First match wins. The matchers all read the same substrate the colony already senses:
| Matcher | Reads | Example |
|---|---|---|
byTag |
Signal tags (issue labels on GitHub) | complexity:high takes the frontier tier |
byType |
Signal type | task:docs takes the cheap tier |
byPayload |
Payload fields | Route on diff size or file count |
byConcentration |
Signal concentration | Fresh signals get the strong tier |
byEscalation |
Failure marks from earlier attempts | Retry a tier up |
byLineage |
Ancestor signals via caused_by |
Rejected work routes up |
Or your own predicate. On a GitHub environment, issue labels are the tags: complexity:high on the issue sends it to your strongest tier with zero glue code.
The router is just another action handler. The colony DSL didn't change. The runtime didn't change.
The tiers don't have to share a vendor
Because routes are handlers rather than model strings, a tier can be anything Mandible knows how to run:
withClaudeCode- a Claude Code session with real file access, direct or through BedrockwithOpenCode,withOpenHands,withQwenCode- other coding agents, driven the same waywithToolLoop- an agentic tool-calling loop against any OpenAI-compatible endpoint, which is how you get file editing and bash out of a model you host yourselfwithStructuredOutputandwithLLM- one-shot calls, whereprovideris'anthropic','bedrock','openai','gemini','vercel-ai', or a function you writewithBash- no model at all
With provider: 'vercel-ai', the model string picks the SDK: claude* loads @ai-sdk/anthropic, gpt* loads @ai-sdk/openai, gemini* loads @ai-sdk/google. And because the provider is just a value, "local if I have it, hosted if I don't" is a line of code:
const provider = vllmFromEnv() ?? 'anthropic'; // reads VLLM_ENDPOINT, null when unset
withLLM({
provider,
model: provider === 'anthropic' ? 'haiku' : 'Qwen3-Coder-Next',
prompt,
route: 'docs:generated',
})
None of this is routing-specific. Our first SDLC run put four colonies against a Nemotron model on a DGX station with no cloud calls at all, and routing works there unchanged. That's rather the point of a cheap tier: it can be a model on hardware you already paid for, and the tier you escalate to can be whatever you keep a credit card on file for.
What a dispatcher can't do
Here's where reading the environment beats asking a service.
It leaves a trail. Before dispatching, the router writes route:<name> onto the signal. On GitHub, that's a label on the issue. A critic colony reviewing the result can walk the lineage and see which tier produced it, and hold the cheap tier's output to a higher bar than the expensive one's.
Failure is a mark, not a state. When a handler throws, the router writes escalation:1 onto the signal and rethrows. The colony's ordinary .retry(3) re-runs the rule, and byEscalation(1, opus) - placed first - catches it. No retry state inside the router. If the colony process dies mid-retry, the mark is still on the signal, and whoever picks it up next routes correctly.
History routes too. A task whose previous artifact drew review:changes-needed shouldn't go back to the cheap model:
byLineage({ environment: env, type: 'review:changes-needed' }, withClaudeCode({ model: 'opus', prompt }))
That walks caused_by, finds the rejection two hops back, and routes up. No one had to throw. No one had to label anything.
Who decides what a task is
All of the above assumes someone marked the task. On GitHub, a human did. In a pipeline where a scout colony creates tasks, it can set tags. But sometimes a task arrives naked.
So classification is a colony action too:
withModelRouter({
classify: withClassifier({
model: 'haiku',
schema: z.object({ complexity: z.enum(['low', 'medium', 'high']), kind: z.enum(['bug', 'feature', 'chore']) }),
prompt: (s) => `Classify this task for routing:\n\n${s.payload.description}`,
tags: (r) => [`complexity:${r.complexity}`, `kind:${r.kind}`],
}),
routes: [ /* as above */ ],
fallback: sonnet,
})
When no route matches, the router asks the cheap tier once, writes the answer onto the signal as tags, stamps classified:default, and routes again. One classification per unlabeled task, ever. Signals that arrive already labeled never trigger it. And the classification isn't hidden in a router's memory - it's on the issue, visible to every colony and every human afterwards.
withClassifier takes the same provider options as withStructuredOutput, so the thing deciding what your tasks are can be a small local model. Classifying is the cheapest work in the pipeline and the easiest to keep in house.
Watching it work
There's a runnable demo with the LLM faked, so you can see the routing decisions without an API key:
npm run demo:model-routing
[classify] add-auth-middleware → complexity:high
[route] add-auth-middleware → tag:complexity:high
[opus] ✓ add-auth-middleware
[route] flaky-lint → tag:complexity:low
[haiku] ✗ flaky-lint — failed (will escalate)
[route] flaky-lint → escalation>=1
[opus] ✓ flaky-lint
[critic] ✗ rewrite-parser (by haiku) — changes needed, re-queuing
[route] rewrite-parser → lineage:review:changes-needed
[opus] ✓ rewrite-parser
Trail left on each task signal:
add-auth-middleware route=tag:complexity:high escalation=0 classified=true
flaky-lint route=escalation>=1 escalation=1 classified=true
rewrite-parser (att 2) route=lineage:review:changes-needed escalation=0 classified=false
Three tasks, three different paths, and none of them decided centrally. The first was labeled by a classifier. The second escalated because it failed. The third escalated because a critic rejected its ancestor. Every decision is a mark on a signal - ls the signals directory and you can read the routing log without a routing log.
The tier names in that output are labels on fake handlers. Swap tier('opus') for withClaudeCode({ model: 'opus', ... }) to run it against an API, or for withToolLoop({ endpoint: 'http://localhost:8001', ... }) to run the whole thing on your own GPU. The routing decisions come out identical, because they never depended on who serves the model.
Why this shape
The temptation with model routing is to build the smart thing in the middle. Mandible's bet is that the middle doesn't need to exist. The environment carries the state. Colonies are stateless workers that read it. Routing intent - labels, tags, failure marks, history - is just more state in the same substrate, written by the same primitives (ctx.enrich) that colonies already use for everything else.
Swap in a new model: change a string. Move a tier onto your own hardware, or off it: change a handler. Add a new routing signal: add a tag. Retry smarter: put a matcher first. Nothing to deploy, nothing to ask.
Check out the GitHub repo to get started - the full routing guide is at docs/how-to/model-routing.md - or explore Mandible Cloud to deploy colonies.