Radar / Ideas / TraceRoute: the coding-agent router…

TraceRoute: the coding-agent router learned from production traces

daily ideamoderateJEV confidence 0.562026-10-02
Outcomecoding-agent platforms cut inference spend roughly in half with no measurable quality regression

The problem

Coding agents burn frontier-model tokens on every turn - one engineer ran up $150,000 in a single month on Claude Code, and an internal Amazon report showed $1.8M spent on menial Claude tasks, 860% over budget. Teams either overpay for frontier models on trivial tasks or hand-tune brittle routing rules that silently degrade quality.

The idea

An OpenAI-compatible routing proxy purpose-built for coding-agent workloads. A cheap difficulty classifier scores each incoming task; easy work (formatting, simple edits, test scaffolding) dispatches to small cheap models while genuinely hard reasoning stays on the frontier - the ensemble pattern Weave Router 2.0 validated on Terminal Bench 4.0. The routing policy is calibrated on Microsoft's public 301k-session Copilot trace dataset, which captures agent-specific dynamics static benchmarks miss: KV-cache collapse at turn boundaries, 4x compute amplification from retry loops, and cache invalidation after model switches. Every routing decision is logged with predicted versus actual cost saved.

Why now

Microsoft just released the first production-scale public telemetry of coding-agent workloads (301,026 sessions, 9.3M model calls) - previously every router vendor trained on private data or static benchmarks. Weave Router 2.0 simultaneously proved open ensemble routing can match GPT-6 Astra coding quality at roughly half the cost. Open decision models give a near-free difficulty pre-classifier. The combination makes a trace-calibrated router buildable by a small team for the first time.

What it combines

Production agent traces (microsoft-releases-301-000-copilot-coding-agent-traces) + open ensemble routing (weave-router-2-0) + typed difficulty classification (kev-open-trainable-jev-like-family-of-small-decision-models-on-qwen3-5-3-8). Weave proved routing works but calibrated on benchmarks; the traces reveal agent-specific failure economics (retry storms, cache collapse at turn boundaries, model-switch invalidation) that benchmark-trained routers misprice; kev supplies the near-free per-request classifier. Together they yield a router that understands how agents actually spend, not how chat benchmarks assume they do.

MVP

Weekend scope: a proxy that classifies task difficulty, routes easy versus hard work to cheap versus frontier models per a Weave-style policy, and logs per-session cost saved; validate no quality regression on a SWE-bench-lite subset. Deliberately skip: learning from each customer's own traces (v1 calibrates on the public Copilot dataset), multi-provider failover, and fine-tuned routing models.

Distribution

B2B2C: sell a drop-in OpenAI-compatible endpoint to coding-agent IDE and plugin vendors (Cline, Roo Code, Kilo Code, Aider-style tools) whose margins are eaten by frontier API bills. Pricing is gain-share - a percentage of metered savings - so adoption carries no downside for the platform. Who pays: dev-tool companies and agent-platform startups running agents at scale.

Why it wins

Martian, Not Diamond, Unify, and OpenRouter auto-route on prompt text using proprietary models or static evals - none train on production agent-trace dynamics, and Martian is enterprise-priced and closed. OSS options (RouteLLM, Semantic Router) are frameworks you assemble yourself. TraceRoute is agent-workload-specific, calibrated on public production traces, open-source at the core, and priced as a share of measured savings. It also differs from hardware-aware inference gateways: this routes tasks to models, not requests to GPUs.

Risks

Biggest risk is providers repricing or shipping new models that invalidate the routing policy. The MVP de-risks with nightly re-benchmarking against the trace-derived eval set - and gain-share pricing means if the savings disappear, the customer's bill does too.

Build it with

Repo to start from

traceroute - open learned router for coding agents: difficulty classifier plus ensemble dispatch calibrated on the public Copilot trace dataset, OpenAI-compatible.

Evidence

Get the week's best AI launches, plus 3 ideas worth building

One email every Saturday. Ranked by traction, not hype. Free.