Back to blog

Model Routing Without a Router: Letting the Environment Pick the Model

Mandible colonies pick a tier - a frontier API, a model on your own GPU, or a shell script - by reading labels, failure marks, and lineage already sitting in the environment. No classifier service, no dispatcher, no retry state, and every routing decision leaves a trail.

Every multi-agent system eventually hits the same question: which model should handle this?

A frontier model for the hard refactor. Something small and local for the typo fix. Something in between for the rest. Retry on something stronger when the cheap one gives up. The usual answer is a router - a classifier service in front of your agents, a config file mapping categories to models, a queue per tier, and a growing pile of special cases. It works, but now you have a scheduler to run, and every agent has to ask it before doing anything.

Mandible doesn't have a router. It has something simpler.

The environment already knows

Mandible is a stigmergy framework: agents coordinate the way ants do, by leaving and sensing signals in a shared environment - a filesystem, a GitHub repo, a Dolt database. No agent talks to another agent. They read the substrate and write to it.

Which means the information a router would need is already there:

  • A human put complexity:high on the GitHub issue
  • The colony that created the task tagged it kind:bug
  • The last attempt failed and left a mark
  • The artifact this task replaces was rejected by a critic - that's in the lineage

A colony doesn't need to ask anyone what tier a task deserves. It looks at the signal it's holding and acts.

Letting the signal choose the model

Every action provider takes a model. Usually that's a tier alias - fable, opus, sonnet, haiku - resolved to a current model ID at call time, so upgrading Mandible moves every colony forward without touching colony code:

withClaudeCode({ model: 'sonnet', prompt })

Anything that isn't an alias passes through untouched: a pinned ID like claude-opus-4-6, a Bedrock ID, a gpt- or gemini-prefixed model, or the name of whatever you have loaded in vLLM. Either way, model can be a function of the signal:

withClaudeCode({
  model: (signal) => signal.meta.tags?.includes('complexity:high') ? 'opus' : 'sonnet',
  prompt: (signal) => `Implement: ${signal.payload.description}`,
})

Same provider, same prompt, same tools. Only the model changes. For anything with more than one rule, selectModel keeps it declarative:

withClaudeCode({
  model: selectModel({
    rules: [
      { match: (s) => escalationLevel(s) > 0,                   model: 'opus'  },
      { match: (s) => s.meta.tags?.includes('complexity:high'), model: 'opus'  },
      { match: (s) => s.meta.tags?.includes('complexity:low'),  model: 'haiku' },
    ],
    default: 'sonnet',
  }),
  prompt,
})

That's most people's model routing, done. No new concepts. The rules read the same whatever the strings are.

Different tiers, different tools

Sometimes the tiers aren't just different models - they're different kinds of work, from different vendors, on different machines. A hard issue wants a full coding agent with file access. An easy one wants a single structured-output call, and there's no reason that call has to leave the building. A chore wants a shell script. withModelRouter dispatches to entirely different handlers:

const local = vllmStructuredProvider({ endpoint: 'http://localhost:8001' });

colony('worker')
  .in(githubEnv)
  .sense('task:ready', { unclaimed: true })
  .retry(3)
  .do('work', withModelRouter({
    routes: [
      byEscalation(1,          withClaudeCode({ model: 'opus', prompt })),
      byTag('complexity:high', withClaudeCode({ model: 'opus', prompt })),
      byTag('complexity:low',  withStructuredOutput({ provider: local, model: 'Qwen3-Coder-Next', schema, prompt, route })),
      byTag('chore',           withBash({ command: 'npm run lint -- --fix', output })),
    ],
    fallback: withClaudeCode({ model: 'sonnet', prompt }),
  }))

The handlers there are a hosted coding agent, a Qwen model on a box in the corner, and npm run lint. The router doesn't know the difference between them.

First match wins. The matchers all read the same substrate the colony already senses:

Matcher Reads Example
byTag Signal tags (issue labels on GitHub) complexity:high takes the frontier tier
byType Signal type task:docs takes the cheap tier
byPayload Payload fields Route on diff size or file count
byConcentration Signal concentration Fresh signals get the strong tier
byEscalation Failure marks from earlier attempts Retry a tier up
byLineage Ancestor signals via caused_by Rejected work routes up

Or your own predicate. On a GitHub environment, issue labels are the tags: complexity:high on the issue sends it to your strongest tier with zero glue code.

The router is just another action handler. The colony DSL didn't change. The runtime didn't change.

The tiers don't have to share a vendor

Because routes are handlers rather than model strings, a tier can be anything Mandible knows how to run:

  • withClaudeCode - a Claude Code session with real file access, direct or through Bedrock
  • withOpenCode, withOpenHands, withQwenCode - other coding agents, driven the same way
  • withToolLoop - an agentic tool-calling loop against any OpenAI-compatible endpoint, which is how you get file editing and bash out of a model you host yourself
  • withStructuredOutput and withLLM - one-shot calls, where provider is 'anthropic', 'bedrock', 'openai', 'gemini', 'vercel-ai', or a function you write
  • withBash - no model at all

With provider: 'vercel-ai', the model string picks the SDK: claude* loads @ai-sdk/anthropic, gpt* loads @ai-sdk/openai, gemini* loads @ai-sdk/google. And because the provider is just a value, "local if I have it, hosted if I don't" is a line of code:

const provider = vllmFromEnv() ?? 'anthropic';   // reads VLLM_ENDPOINT, null when unset

withLLM({
  provider,
  model: provider === 'anthropic' ? 'haiku' : 'Qwen3-Coder-Next',
  prompt,
  route: 'docs:generated',
})

None of this is routing-specific. Our first SDLC run put four colonies against a Nemotron model on a DGX station with no cloud calls at all, and routing works there unchanged. That's rather the point of a cheap tier: it can be a model on hardware you already paid for, and the tier you escalate to can be whatever you keep a credit card on file for.

What a dispatcher can't do

Here's where reading the environment beats asking a service.

It leaves a trail. Before dispatching, the router writes route:<name> onto the signal. On GitHub, that's a label on the issue. A critic colony reviewing the result can walk the lineage and see which tier produced it, and hold the cheap tier's output to a higher bar than the expensive one's.

Failure is a mark, not a state. When a handler throws, the router writes escalation:1 onto the signal and rethrows. The colony's ordinary .retry(3) re-runs the rule, and byEscalation(1, opus) - placed first - catches it. No retry state inside the router. If the colony process dies mid-retry, the mark is still on the signal, and whoever picks it up next routes correctly.

History routes too. A task whose previous artifact drew review:changes-needed shouldn't go back to the cheap model:

byLineage({ environment: env, type: 'review:changes-needed' }, withClaudeCode({ model: 'opus', prompt }))

That walks caused_by, finds the rejection two hops back, and routes up. No one had to throw. No one had to label anything.

Who decides what a task is

All of the above assumes someone marked the task. On GitHub, a human did. In a pipeline where a scout colony creates tasks, it can set tags. But sometimes a task arrives naked.

So classification is a colony action too:

withModelRouter({
  classify: withClassifier({
    model: 'haiku',
    schema: z.object({ complexity: z.enum(['low', 'medium', 'high']), kind: z.enum(['bug', 'feature', 'chore']) }),
    prompt: (s) => `Classify this task for routing:\n\n${s.payload.description}`,
    tags: (r) => [`complexity:${r.complexity}`, `kind:${r.kind}`],
  }),
  routes: [ /* as above */ ],
  fallback: sonnet,
})

When no route matches, the router asks the cheap tier once, writes the answer onto the signal as tags, stamps classified:default, and routes again. One classification per unlabeled task, ever. Signals that arrive already labeled never trigger it. And the classification isn't hidden in a router's memory - it's on the issue, visible to every colony and every human afterwards.

withClassifier takes the same provider options as withStructuredOutput, so the thing deciding what your tasks are can be a small local model. Classifying is the cheapest work in the pipeline and the easiest to keep in house.

Watching it work

There's a runnable demo with the LLM faked, so you can see the routing decisions without an API key:

npm run demo:model-routing
  [classify] add-auth-middleware → complexity:high
  [route]    add-auth-middleware → tag:complexity:high
  [opus]     ✓ add-auth-middleware
  [route]    flaky-lint → tag:complexity:low
  [haiku]    ✗ flaky-lint — failed (will escalate)
  [route]    flaky-lint → escalation>=1
  [opus]     ✓ flaky-lint
  [critic]   ✗ rewrite-parser (by haiku) — changes needed, re-queuing
  [route]    rewrite-parser → lineage:review:changes-needed
  [opus]     ✓ rewrite-parser

Trail left on each task signal:
  add-auth-middleware    route=tag:complexity:high  escalation=0  classified=true
  flaky-lint             route=escalation>=1        escalation=1  classified=true
  rewrite-parser (att 2) route=lineage:review:changes-needed  escalation=0  classified=false

Three tasks, three different paths, and none of them decided centrally. The first was labeled by a classifier. The second escalated because it failed. The third escalated because a critic rejected its ancestor. Every decision is a mark on a signal - ls the signals directory and you can read the routing log without a routing log.

The tier names in that output are labels on fake handlers. Swap tier('opus') for withClaudeCode({ model: 'opus', ... }) to run it against an API, or for withToolLoop({ endpoint: 'http://localhost:8001', ... }) to run the whole thing on your own GPU. The routing decisions come out identical, because they never depended on who serves the model.

Why this shape

The temptation with model routing is to build the smart thing in the middle. Mandible's bet is that the middle doesn't need to exist. The environment carries the state. Colonies are stateless workers that read it. Routing intent - labels, tags, failure marks, history - is just more state in the same substrate, written by the same primitives (ctx.enrich) that colonies already use for everything else.

Swap in a new model: change a string. Move a tier onto your own hardware, or off it: change a handler. Add a new routing signal: add a tag. Retry smarter: put a matcher first. Nothing to deploy, nothing to ask.

Check out the GitHub repo to get started - the full routing guide is at docs/how-to/model-routing.md - or explore Mandible Cloud to deploy colonies.