OmniExtractBench
Datalab's open, auditable benchmark and Python toolkit for structured document extraction: 620 documents pooled from four existing suites, one deterministic scorer, and per-value verdicts (matched/misread/unfound/fabricated/invented) that make every score explainable.
Why it matters
One shared, vendor-neutral yardstick for PDF-to-JSON extraction with a transparent, fully specified scoring method and a six-verdict breakdown, plus a one-line harness to rerun vendors (Datalab, Reducto, Extend, LlamaExtract) against the same corpus.
What you could build with it
An indie developer building a document-processing product could use OmniExtractBench to pick their extraction backend: run Datalab, Reducto, and LlamaExtract against the same 620-document corpus with one deterministic scorer, and choose based on auditable per-field verdicts instead of vendor leaderboards.
Does it hold up?
Self-evident from the repo: the scorer installs with score() plus scipy, `oeb benchmark` is resumable, and per-address verdicts ship by default so failures can be argued with, not just quoted. Independent third-party usage is just beginning (15 stars at launch).
Built with OmniExtractBench
- The New Benchmark Challenging AI Extraction Leaderboardsartiverse.ca · In-depth launch-day analysis of the benchmark's design: 620 documents, six per-value verdicts, and why vendor leaderboards are hard to compare.
- omni-extract-bench on PyPIpypi · PyPI listing with install instructions for the scorer (score() + scipy only), vendor harness, and benchmark orchestration.
- LlamaIndex runs its Agentic Plus extraction model on OmniExtractBench, scores 93.94% (X thread, twstalker mirror)x · LlamaIndex ran its 'Agentic Plus' extraction mode on OmniExtractBench with the same auditable scorer, scoring 93.94% and noting a compatibility workaround around null-field handling — real third-party usage and results on the new benchmark.
Learn more
First spotted on article: source.
More AI agents
Agent ReachOpen-source CLI and SKILL.md kit that gives AI agents read/search access to 15+ internet platforms (Twitter, Reddit…agents · JEV 0.8Google Agent SkillsGoogle's official catalog of 150+ agent skills for its products and technologies — Cloud, GKE, BigQuery, Gemini, Agent…agents · JEV 0.74hermes-agent (NousResearch)Self-improving personal AI agent across Telegram, Discord, Slack, and terminal with persistent memory.agents · JEV 0.73Paperclip: the app people use to manage agents at workOpen-source orchestration for teams of AI agents: org charts, budgets, governance, heartbeats and a ticket system in…agents · JEV 0.73AIHOT — open-source framework for building automated industry-news sitesFull-stack framework (Node 24 + PostgreSQL) that ingests RSS/webpages/X/WeChat, dedupes, double-scores, clusters with…agents · JEV 0.73orcaAgent development environment for working with a fleet of parallel agents on your own subscription.agents · JEV 0.68
Get the week's best AI launches, plus 3 ideas worth building
One email every Saturday. Ranked by traction, not hype. Free.