Perplexity pplx-embed-v2-context-9b-preview
Perplexity Research and turbopuffer released a 9B MIT-licensed contextual embedding model for RAG that encodes each chunk with the full document in view and is trained to retrieve answers plus their supporting evidence, with 8x smaller int8 vectors than Voyage.
Why it matters
Changes the RAG training signal from 'one gold passage' to 'answer plus verifiable evidence' — directly attacking the chunking context-loss problem with an open-weights model anyone can self-host.
What you could build with it
An indie developer building a research assistant could self-host this model to ground long-document Q&A with cited evidence passages, cutting vector-storage costs 8x versus float32 embeddings while improving verifiability.
Does it hold up?
Too early to judge — preview release Oct 1; the model card warns weights and interface may change without backward compatibility, and independent RAG evaluations are pending.
Built with Perplexity pplx-embed-v2-context-9b-preview
- MarkTechPost: Perplexity releases pplx-embed-v2-context-9b-previewMarkTechPost · Deep dive on the late-chunking design, the evidence-aware training signal, and deployment requirements (transformers>=5.4, trust_remote_code).
- RuntimeWire: Perplexity releases embeddings designed to retrieve answers and evidenceRuntimeWire · Notes the company-reported benchmarks and the private context-bench.
- AIDailyPost: Perplexity's embedding model beats Voyage at 8x smaller sizeAIDailyPost · Covers the 8KB to 1KB vector-size reduction and the trust_remote_code friction.
Learn more
First spotted on hf: source.
More AI models
ElevenLabs Eleven v4 + v4 Turbo: new emotive TTS architecture with a 100ms real-time variantElevenLabs released Eleven v4 and Eleven v4 Turbo on September 28: a new text-to-speech architecture with inline…models · JEV 0.72VoiceStudio: open-source, fully-local ElevenLabs alternative with voice cloning, dubbing and dictation in 646 languagesVoiceStudio is an open-source, fully-local voice AI studio: voice cloning from a 3-second sample, voice design, video…models · JEV 0.71DeepSeek V4-Flash official API: public beta with upgraded agent capabilities, Responses API and Codex supportDeepSeek's official V4-Flash is now a public beta API with massively upgraded agent capabilities - benchmark scores…models · JEV 0.69Mistral Forge: enterprise platform for training and continuously improving proprietary modelsMistral AI launched Forge (Sep 28) — a full-lifecycle model training platform (pre-training, SFT, DPO/ODPO, RL…models · JEV 0.69kev: open, trainable Jev-like family of small decision models on Qwen3.5/3.8A family of small decision models (0.8B to 27B) built on Qwen3.5/3.8 that reproduces Jev's prefill-only typed-decision…models · JEV 0.68GPT-Synopsys: OpenAI and Synopsys multi-year deal to build an AI model that operates EDA toolsOpenAI and Synopsys signed a multi-year agreement on Sep 30 to jointly develop GPT-Synopsys, a specialized model that…models · JEV 0.66
Get the week's best AI launches, plus 3 ideas worth building
One email every Saturday. Ranked by traction, not hype. Free.