Respan Span-01: behavior-scoring decision model for AI agent monitoring
Respan listed Span-01 on OpenRouter (announced September 24): a specialized decision model that scores behaviors you define in plain language - hallucination, user frustration, tool misuse - as present, absent, or not observable across conversation traces, at $0.02 per million input tokens with free output and a free Lite tier.
Why it matters
Turns LLM-as-a-judge into a single-forward-pass API: no text to parse, multiple behavior definitions scored at once, free Lite tier; aimed at teams scoring large volumes of production agent traces without paying a text-generating model per check.
What you could build with it
An AI observability platform could add continuous behavior scoring over every production agent trace, flagging hallucinations and tool misuse in near real time without paying a text-generating model per check.
Does it hold up?
Too early to judge: listed days ago, no independent benchmarks published; the economic case (cheaper than LLM-as-a-judge) is plausible but unproven in production workflows.
Learn more
First spotted on article: source.
More AI models
ElevenLabs Eleven v4 + v4 Turbo: new emotive TTS architecture with a 100ms real-time variantElevenLabs released Eleven v4 and Eleven v4 Turbo on September 28: a new text-to-speech architecture with inline…models · JEV 0.72VoiceStudio: open-source, fully-local ElevenLabs alternative with voice cloning, dubbing and dictation in 646 languagesVoiceStudio is an open-source, fully-local voice AI studio: voice cloning from a 3-second sample, voice design, video…models · JEV 0.71DeepSeek V4-Flash official API: public beta with upgraded agent capabilities, Responses API and Codex supportDeepSeek's official V4-Flash is now a public beta API with massively upgraded agent capabilities - benchmark scores…models · JEV 0.69Mistral Forge: enterprise platform for training and continuously improving proprietary modelsMistral AI launched Forge (Sep 28) — a full-lifecycle model training platform (pre-training, SFT, DPO/ODPO, RL…models · JEV 0.69kev: open, trainable Jev-like family of small decision models on Qwen3.5/3.8A family of small decision models (0.8B to 27B) built on Qwen3.5/3.8 that reproduces Jev's prefill-only typed-decision…models · JEV 0.68GPT-Synopsys: OpenAI and Synopsys multi-year deal to build an AI model that operates EDA toolsOpenAI and Synopsys signed a multi-year agreement on Sep 30 to jointly develop GPT-Synopsys, a specialized model that…models · JEV 0.66
Get the week's best AI launches, plus 3 ideas worth building
One email every Saturday. Ranked by traction, not hype. Free.