ExplorationBench: AI exploration in verifiable alien worlds
Tencent's Hunyuan team with Fudan and Tsinghua released ExplorationBench, a benchmark measuring whether AI systems can discover hidden rules by actively experimenting in two synthetic sandboxes, AlienCode and AlienLogic.
Why it matters
First benchmark to cleanly separate rule discovery from rule execution: invented 'alien worlds' defeat pretraining recall while an interpreter and a proof checker give exact, LLM-judge-free grading; testing 10 frontier systems exposed huge trajectory variance and rule-application failures.
What you could build with it
An eval-tools startup could productize the ExplorationBench protocol as a service: customers submit models and receive contamination-proof discovery scores under the forked-protocol controls, giving lab buyers a reproducible way to compare agents on exploration skill rather than memorized benchmarks.
Does it hold up?
Usable today as a concept and leaderboard - the paper, project site, and leaderboard are live, but the task set stays private to reduce contamination and the GitHub repo is still listed as forthcoming, so independent replication is not possible yet.
Built with ExplorationBench
- AlphaSignal: deep breakdown of the benchmark protocol and findingsnews · Detailed third-party analysis of the forked evaluation protocol, Best@3 reporting, and the discovery-vs-execution divergence across 10 frontier systems.
- Complete AI Training: Fudan and Tencent launch ExplorationBenchnews · Writeup emphasizing the headline result: active probing drove AlienCode accuracy from under 16% to 89%, while unguided thinking alone stalled below 11%.
Learn more
First spotted on article: source.
More AI agents
hermes-agent (NousResearch)Self-improving personal AI agent across Telegram, Discord, Slack, and terminal with persistent memory.agents · JEV 0.73Paperclip: the app people use to manage agents at workOpen-source orchestration for teams of AI agents: org charts, budgets, governance, heartbeats and a ticket system in…agents · JEV 0.73AIHOT — open-source framework for building automated industry-news sitesFull-stack framework (Node 24 + PostgreSQL) that ingests RSS/webpages/X/WeChat, dedupes, double-scores, clusters with…agents · JEV 0.73orcaAgent development environment for working with a fleet of parallel agents on your own subscription.agents · JEV 0.68Univer 1.0: Office Harness for AI AgentsUniver 1.0 unifies six editors (Sheets, Docs, Slides, Boards, Bases, PDFs) into one programmable, embeddable Office…agents · JEV 0.68Nasiko: developer control plane for AI agentsRust control plane giving developers one dashboard to deploy, monitor, and govern fleets of AI agents.agents · JEV 0.68
Get the week's best AI launches, plus 3 ideas worth building
One email every Saturday. Ranked by traction, not hype. Free.