SPARSEUP (Linkup)
Open-source learned sparse embedding model on a 149M-parameter ModernBERT backbone, Apache 2.0, reporting 56.4 nDCG@10 on BEIR-13 as the strongest public sub-150M sparse encoder.
Why it matters
Fills the missing sparse leg beside dense and late-interaction retrievers: readable, vocabulary-aligned embeddings that slot into inverted indexes and handle rare words better than dense vectors, loadable via Transformers today.
What you could build with it
Pair SPARSEUP with a dense embedder in a hybrid RAG pipeline for legal-tech document search, using sparse scores for exact-term recall on statute numbers and dense for semantic recall.
Does it hold up?
Loadable and deployable today via standard HF tooling, with one third-party deployment tutorial already out; no real production retrieval deployments or comparative community benchmarks yet.
Built with SPARSEUP
- Official Hugging Face model pagemodel_page · Vendor-published weights (Apache 2.0, 149M ModernBERT backbone, 56.4 avg nDCG@10 on BEIR-13), loadable via Transformers or Sentence Transformers with trust_remote_code=True; the usable artifact everything else builds on.
- Undercode Testing deployment tutorialwriteup · Third-party deployment guide for SPARSEUP covering logit-shift and case-folding internals, pipeline setup, and latency notes.
- MarkTechPost release coveragewriteup · Independent technical writeup on the release: fixes a stopword-saturation problem in SPLADE-style setups, fills the sparse slot beside LightOn's DenseOn and LateOn on shared data and backbone.
Learn more
First spotted on article: source.
More AI models
ElevenLabs Eleven v4 + v4 Turbo: new emotive TTS architecture with a 100ms real-time variantElevenLabs released Eleven v4 and Eleven v4 Turbo on September 28: a new text-to-speech architecture with inline…models · JEV 0.72VoiceStudio: open-source, fully-local ElevenLabs alternative with voice cloning, dubbing and dictation in 646 languagesVoiceStudio is an open-source, fully-local voice AI studio: voice cloning from a 3-second sample, voice design, video…models · JEV 0.71DeepSeek V4-Flash official API: public beta with upgraded agent capabilities, Responses API and Codex supportDeepSeek's official V4-Flash is now a public beta API with massively upgraded agent capabilities - benchmark scores…models · JEV 0.69Mistral Forge: enterprise platform for training and continuously improving proprietary modelsMistral AI launched Forge (Sep 28) — a full-lifecycle model training platform (pre-training, SFT, DPO/ODPO, RL…models · JEV 0.69kev: open, trainable Jev-like family of small decision models on Qwen3.5/3.8A family of small decision models (0.8B to 27B) built on Qwen3.5/3.8 that reproduces Jev's prefill-only typed-decision…models · JEV 0.68GPT-Synopsys: OpenAI and Synopsys multi-year deal to build an AI model that operates EDA toolsOpenAI and Synopsys signed a multi-year agreement on Sep 30 to jointly develop GPT-Synopsys, a specialized model that…models · JEV 0.66
Get the week's best AI launches, plus 3 ideas worth building
One email every Saturday. Ranked by traction, not hype. Free.