TraceRoute: the coding-agent router learned from production traces
The problem
Coding agents burn frontier-model tokens on every turn - one engineer ran up $150,000 in a single month on Claude Code, and an internal Amazon report showed $1.8M spent on menial Claude tasks, 860% over budget. Teams either overpay for frontier models on trivial tasks or hand-tune brittle routing rules that silently degrade quality.
The idea
Why now
Microsoft just released the first production-scale public telemetry of coding-agent workloads (301,026 sessions, 9.3M model calls) - previously every router vendor trained on private data or static benchmarks. Weave Router 2.0 simultaneously proved open ensemble routing can match GPT-6 Astra coding quality at roughly half the cost. Open decision models give a near-free difficulty pre-classifier. The combination makes a trace-calibrated router buildable by a small team for the first time.
What it combines
Production agent traces (microsoft-releases-301-000-copilot-coding-agent-traces) + open ensemble routing (weave-router-2-0) + typed difficulty classification (kev-open-trainable-jev-like-family-of-small-decision-models-on-qwen3-5-3-8). Weave proved routing works but calibrated on benchmarks; the traces reveal agent-specific failure economics (retry storms, cache collapse at turn boundaries, model-switch invalidation) that benchmark-trained routers misprice; kev supplies the near-free per-request classifier. Together they yield a router that understands how agents actually spend, not how chat benchmarks assume they do.
MVP
Weekend scope: a proxy that classifies task difficulty, routes easy versus hard work to cheap versus frontier models per a Weave-style policy, and logs per-session cost saved; validate no quality regression on a SWE-bench-lite subset. Deliberately skip: learning from each customer's own traces (v1 calibrates on the public Copilot dataset), multi-provider failover, and fine-tuned routing models.
Distribution
B2B2C: sell a drop-in OpenAI-compatible endpoint to coding-agent IDE and plugin vendors (Cline, Roo Code, Kilo Code, Aider-style tools) whose margins are eaten by frontier API bills. Pricing is gain-share - a percentage of metered savings - so adoption carries no downside for the platform. Who pays: dev-tool companies and agent-platform startups running agents at scale.
Why it wins
Martian, Not Diamond, Unify, and OpenRouter auto-route on prompt text using proprietary models or static evals - none train on production agent-trace dynamics, and Martian is enterprise-priced and closed. OSS options (RouteLLM, Semantic Router) are frameworks you assemble yourself. TraceRoute is agent-workload-specific, calibrated on public production traces, open-source at the core, and priced as a share of measured savings. It also differs from hardware-aware inference gateways: this routes tasks to models, not requests to GPUs.
Risks
Biggest risk is providers repricing or shipping new models that invalidate the routing policy. The MVP de-risks with nightly re-benchmarking against the trace-derived eval set - and gain-share pricing means if the savings disappear, the customer's bill does too.
Build it with
- Microsoft releases 301,000 Copilot coding-agent tracesFirst public production-scale telemetry of agent workloads - timings, cache behavior, retry patterns - to calibrate routing policy
- Weave Router 2.0Open-source proof that ensemble routing matches frontier coding-agent quality at roughly half the cost
- kev: open, trainable Jev-like family of small decision models on Qwen3.5/3.8Near-free per-request difficulty classifier deciding which tasks genuinely need the frontier model
Repo to start from
traceroute - open learned router for coding agents: difficulty classifier plus ensemble dispatch calibrated on the public Copilot trace dataset, OpenAI-compatible.
Evidence
- RuntimeWire: Microsoft releases 301,000 Copilot agent traces
- Microsoft Research: Agentic Coding in the Wild (trace analysis paper)
- Weave Router (GitHub)
- StartupFortune: how AI model routing works and why it is halving LLM bills
- WebProNews: AI's code-churning frenzy - skyrocketing bills
- AI coding cost crisis: Cursor hid the numbers, Amazon blew $1.8M
Get the week's best AI launches, plus 3 ideas worth building
One email every Saturday. Ranked by traction, not hype. Free.