LightOnOCR-3 — Apache 2.0 open-weight document intelligence models
Paris-based LightOn released LightOnOCR-3, a family of lightweight OCR models in 0.8B, 1B and 4B sizes that transcribe pages, emit labeled bounding boxes, describe images and extract chart data as tables in a single pass.
Why it matters
Open weights under Apache 2.0 with leading open-model results on ParseBench layout/chart extraction and first place on French-language FRBench-pdf2md — at under $0.01 per thousand pages when self-hosted, a fraction of competing OCR pipelines.
What you could build with it
A bookkeeping startup could run LightOnOCR-3-0.8B on its own servers to convert clients' scanned invoices into structured line items overnight, keeping sensitive financial documents off third-party APIs entirely.
Does it hold up?
Immediately usable: Apache-2.0 checkpoints are publicly downloadable from Hugging Face with a technical blog, and third-party guides (RohitAI) already walk through choosing a size for OCR plus layout extraction.
Built with LightOnOCR-3
- RohitAI: LightOnOCR-3 — choosing a model for OCR, layout and chart extractionrohitai.com · Practitioner's guide to picking among the 0.8B, 1B and 4B sizes, noting the models combine transcription, labeled page regions, image descriptions and chart tables — and that the public leaderboard tells a more mixed story than vendor claims.
- TipRanks: LightOn launches LightOnOCR-3 to boost low-cost document intelligencetipranks.com · Market coverage of the release: Apache-2.0 licensing, sub-$0.01-per-1K-pages self-hosted cost claim, and LightOn's positioning as a European alternative for regulated-enterprise document AI.
- ExplainX AI: LightOnOCR-3 — Apache 2.0 OCR in 0.8B, 1B and 4B, ranked on ParseBenchexplainx.ai · Spotted in ExplainX's OCR model roundup: three Apache-2.0 sizes that read pages, return labeled bounding boxes, describe images and turn charts into tables, positioned against Baidu's Unlimited-OCR and Cohere Parse 5.
Learn more
First spotted on article: source.
More AI models
OpenAI publishes 722 AI-generated math manuscripts from an unreleased modelOpenAI released 722 mathematical manuscripts in 372 result families to GitHub under Apache-2.0, produced by an…models · JEV 0.74Mistral Large 4 (Le Chonk): 1.05T-param open-weight MoEMistral's new flagship: a 1.05T-parameter mixture-of-experts multimodal model with 49B active params and 1M-token…models · JEV 0.72ElevenLabs Eleven v4 + v4 Turbo: new emotive TTS architecture with a 100ms real-time variantElevenLabs released Eleven v4 and Eleven v4 Turbo on September 28: a new text-to-speech architecture with inline…models · JEV 0.72VoiceStudio: open-source, fully-local ElevenLabs alternative with voice cloning, dubbing and dictation in 646 languagesVoiceStudio is an open-source, fully-local voice AI studio: voice cloning from a 3-second sample, voice design, video…models · JEV 0.71JetBrains Mellum 2.1: 12B MoE coding model trained for agentsJetBrains released Mellum 2.1, a 12B MoE model (2.5B active parameters, 128K context) under Apache 2.0, trained mainly…models · JEV 0.69DeepSeek V4-Flash official API: public beta with upgraded agent capabilities, Responses API and Codex supportDeepSeek's official V4-Flash is now a public beta API with massively upgraded agent capabilities - benchmark scores…models · JEV 0.69
Get the week's best AI launches, plus 3 ideas worth building
One email every Saturday. Ranked by traction, not hype. Free.