At a cost of $19,000 for the evaluation, the model successfully demonstrated agentic intelligence by autonomously exploring and building internal models of complex environments without explicit instructions.
The results are significant because they mark a breakthrough in action efficiency, a metric that tracks how many interactions a system needs to solve a problem compared to a human.
Astra surpassed the human baseline on 96% of the tested levels, requiring roughly 52% fewer actions per level than the average person.
This suggests that the model is not just solving tasks through brute-force computation, but is developing a refined understanding of underlying mechanics.
By identifying rules and creating its own symbolic language to track state, the model reduces the total number of actions and model calls, which lower the overall operational cost.
Beyond simple task completion, the ARC Prize reported that Astra displayed advanced problem-solving behaviors, such as writing its own software libraries and "algebraic shorthand" to navigate game mechanics.
In more complex scenarios, the model generated custom tools—including maze solvers and game-state predictors—to plan its moves.
While this performance nears the ceiling of the ARC-AGI-3 benchmark, the researchers noted that the system is tested within closed-ended, deterministic environments, meaning further benchmarks will be needed to evaluate how these capabilities translate to the open-ended complexity of the real world.