Interfaze-1-lite: open-weight model for OCR, speech and structured extraction
Interfaze released an open-weight model for document OCR, speech-to-text and structured extraction that returns confidence scores and bounding boxes alongside its answers, designed to run on a single 80GB GPU under Apache 2.0.
Why it matters
Outputs you can check: every answer ships with confidence scores and location metadata (bounding boxes), so production software can route uncertain extractions to review instead of trusting fluent output. Beating general-purpose models at deterministic extraction is the explicit bet.
What you could build with it
A startup building an invoice-processing API could drop interfaze-1-lite into its pipeline so every extracted field carries a confidence score and bounding box, letting downstream code auto-accept high-confidence fields and route only uncertain ones to human review.
Does it hold up?
Too early to judge — launched three days ago and no independent evaluations are published yet. The verifiable-output thesis (confidence scores, bounding boxes) is the claim to watch once production trials appear.
Built with Interfaze-1-lite
- Interfaze releases open-weight model for OCR, speech and structured extractionarticle · RuntimeWire launch coverage: San Francisco startup founded by repeat entrepreneur Yoeven Khemlani and CV researcher Harsha Khurdula; weights on Hugging Face under Apache 2.0.
- The model weights on Hugging Facehuggingface · Official model repo: 52 likes within days of release, posted 2026-10-03.
Learn more
First spotted on x: source.
More AI models
OpenAI publishes 722 AI-generated math manuscripts from an unreleased modelOpenAI released 722 mathematical manuscripts in 372 result families to GitHub under Apache-2.0, produced by an…models · JEV 0.74Mistral Large 4 (Le Chonk): 1.05T-param open-weight MoEMistral's new flagship: a 1.05T-parameter mixture-of-experts multimodal model with 49B active params and 1M-token…models · JEV 0.72ElevenLabs Eleven v4 + v4 Turbo: new emotive TTS architecture with a 100ms real-time variantElevenLabs released Eleven v4 and Eleven v4 Turbo on September 28: a new text-to-speech architecture with inline…models · JEV 0.72VoiceStudio: open-source, fully-local ElevenLabs alternative with voice cloning, dubbing and dictation in 646 languagesVoiceStudio is an open-source, fully-local voice AI studio: voice cloning from a 3-second sample, voice design, video…models · JEV 0.71Mistral Forge: enterprise platform for training and continuously improving proprietary modelsMistral AI launched Forge (Sep 28) — a full-lifecycle model training platform (pre-training, SFT, DPO/ODPO, RL…models · JEV 0.69DeepSeek V4-Flash official API: public beta with upgraded agent capabilities, Responses API and Codex supportDeepSeek's official V4-Flash is now a public beta API with massively upgraded agent capabilities - benchmark scores…models · JEV 0.69
Get the week's best AI launches, plus 3 ideas worth building
One email every Saturday. Ranked by traction, not hype. Free.