Underdog-Saluki-27B-1.0: 2-bit Qwen3.8-27B that beats the original at tool calling
2-bit (IQ2-mix) GGUF build of Qwen3.8-27B in 7.89 GB vs 54 GB BF16, tuned for agentic tool calling: 88/120 vs 84 for the full model on the author's BFCL-v4-derived bench, 42 vs 35 on parallel tool calls, and 96% average retention across 9 benchmarks (math takes the hit, ~82-85% on AIME); Apache 2.0, runs in stock llama.cpp, optional vision add-on.
Why it matters
A rare case where a 2-bit quant beats the full-size original at the exact skill agents need (tool calling) while running in stock llama.cpp — deliberate quantization targeted at local agents, not just a shrink.
What you could build with it
A startup shipping private, on-device agent software could build its whole tool-calling stack on this 7.89 GB GGUF in stock llama.cpp, running Qwen3.8-27B-class agents on consumer GPUs with no API costs and no data leaving the machine.
Does it hold up?
Good early signals for a 2-day-old release: the official file runs in stock llama.cpp with a documented quickstart, and an independent builder (Kujira) already repacked it byte-identically with an MTP head — the strongest third-party proof it works as advertised. Tool-calling claims still rest on the author's own benchmark suite pending independent replication.
Built with Underdog-Saluki-27B-1.0
- Kujira MTP repack of Saluki 27Bhf · Community build adding Qwen3.8-27B's MTP head (8.35 GB) for self-speculative decoding in llama.cpp; weights byte-identical to the official file, confirming clean stock llama.cpp compatibility.
- IT之家: Saluki 27B 2-bit release coveragearticle · Chinese tech media (IT之家) coverage on 2026-10-10 noting the 7.89 GB file, agent-tuned design, and stock llama.cpp support.
Learn more
technical deep dive →
First spotted on hf: source.
More AI models
OpenAI publishes 722 AI-generated math manuscripts from an unreleased modelOpenAI released 722 mathematical manuscripts in 372 result families to GitHub under Apache-2.0, produced by an…models · JEV 0.74Mistral Large 4 (Le Chonk): 1.05T-param open-weight MoEMistral's new flagship: a 1.05T-parameter mixture-of-experts multimodal model with 49B active params and 1M-token…models · JEV 0.72ElevenLabs Eleven v4 + v4 Turbo: new emotive TTS architecture with a 100ms real-time variantElevenLabs released Eleven v4 and Eleven v4 Turbo on September 28: a new text-to-speech architecture with inline…models · JEV 0.72VoiceStudio: open-source, fully-local ElevenLabs alternative with voice cloning, dubbing and dictation in 646 languagesVoiceStudio is an open-source, fully-local voice AI studio: voice cloning from a 3-second sample, voice design, video…models · JEV 0.71JetBrains Mellum 2.1: 12B MoE coding model trained for agentsJetBrains released Mellum 2.1, a 12B MoE model (2.5B active parameters, 128K context) under Apache 2.0, trained mainly…models · JEV 0.69DeepSeek V4-Flash official API: public beta with upgraded agent capabilities, Responses API and Codex supportDeepSeek's official V4-Flash is now a public beta API with massively upgraded agent capabilities - benchmark scores…models · JEV 0.69 Get the week's best AI launches, plus 3 ideas worth building
One email every Saturday. Ranked by traction, not hype. Free.