Nvidia agent scores 100% on ARC-AGI-3 reasoning benchmark

4 sources
  • Nvidia's AVO agent system achieved a perfect 100.00 score on the ARC-AGI-3 interactive reasoning benchmark, completing all 183 levels across 25 environments.
  • AVO uses persistent memory and a supervisor component to sustain long-horizon tasks; without the harness, Claude Opus 5 scored just 30% on the same test.
  • The same architecture previously optimized GPU kernels over a seven-day autonomous run, suggesting the agent loop generalizes across vastly different domains.
Sources (4)
  1. 1 NVIDIA AVO Reaches 100% on ARC-AGI-3, Demonstrating a Frontier-Level General-Purpose Architecture for Long-Horizon Autonomous Agents | NVIDIA Technical Blog developer.nvidia.com
  2. 2 Nvidia just showed that the harness, not the AI model, is now the real hero | TechCrunch techcrunch.com
  3. 3 Nvidia Research: AI Agent Control Beats Raw Model Power www.techbuzz.ai
  4. 4 Nvidia just showed that the harness, not the AI model, is now the real hero tech.yahoo.com