Apodex 1.1: Scaling Agentic Intelligence for Complex Work

Published
Source
arXiv
Paper number
988
Field
AI / General
arXiv ID
2608.23283

Key points

  • It trained the model's 'ability to do work' using a variety of verifiable task environments involving files, search, code, and more.
  • It treated multi-agent coordination, involving decomposing problems, delegating, combining asynchronous results, and replanning, as a separate training axis.
  • It scored 54.3 on FrontierFinance for financial practice, 63.3 on FrontierScience-Research for scientific research, and 78.8 on GDPVal for work deliverables, placing it in the leading performance range described by the paper.
  • In contrast, its pass rate on difficult scientific workflows in its own FrontierResearchBench was only 12.4%, trailing leading systems. The paper also states that completing such workflows end to end remains difficult for every system evaluated.
  • A 35B-parameter Mini version preserved a substantial share of practical task-execution capability in a form that can run locally, opening a path to smaller-scale deployment in the field.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)