ExplorationBench: Measuring AI Systems' Exploration in Verifiable Alien Worlds
- Published
- Source
- arXiv
- Paper number
- 1120
- Field
- Evaluation
- arXiv ID
- 2609.30199
Key points
- Built AlienCode (31 discovery targets, 70 tasks) and AlienLogic (24 targets, 70 tasks) sandboxes whose executable rules enable exact answer grading
- The conflicting 'alien' design ensures tasks can only be solved by genuine exploration, not memorization
- Introduced a milestone evaluation structure where systems find rules through self-chosen probes starting from a flawed manual
- Across ten systems evaluated, the strongest acquired and applied unfamiliar rules, but variance across trajectories was large
- Cases were observed where continued exploration stalled or even reversed earlier gains
Paper links
External research summaries. These are not HDATF publications or measured product results.