AGI benchmarking and progress metrics
As of mid-2026, the definition of Artificial General Intelligence (AGI) has transitioned from static, knowledge-based performance to metrics focused on agentic and interactive capabilities Verified Answer #1. While there is no single consensus on the exact definition of AGI, researchers generally measure progress through a bundle of milestones across expert knowledge, fluid intelligence, autonomous execution, and long-horizon reliability Verified Answer #3Verified Answer #2.
Fluid Intelligence and Novel Problem Solving
Fluid intelligence is defined as the capacity to efficiently learn and adapt to novel tasks without relying on prior training data Verified Answer #3. The primary threshold for AGI in this category is the ability to solve problems in rule-free environments, a trait measured by the Abstraction and Reasoning Corpus (ARC-AGI) Verified Answer #1.
- ARC-AGI-1: Using test-time compute scaling, GPT-5.5 Pro achieved 96.5% on the static version of this benchmark Verified Answer #3.
- ARC-AGI-3: Released in March 2026, this version requires agents to interact with turn-based environments and infer goals in real-time Verified Answer #1.
- Performance Gap: While humans score approximately 100% on these environments, leading frontier models scored less than 1% upon the release of ARC-AGI-3, indicating a significant deficit in independent exploration Verified Answer #1.
Expert Academic Synthesis
To reach AGI, a system must demonstrate graduate-level synthesis across domains rather than retrieving memorized internet data Verified Answer #3. Humanity's Last Exam (HLE) is a 2,500-question benchmark vetted by domain experts to evaluate this boundary Verified Answer #1.
- Human Baseline: Human experts maintain approximately 90% accuracy on HLE Verified Answer #1.
- Current Model Performance: As of early 2026, frontier systems consistently cluster between 40% and 50% accuracy Verified Answer #1. Specific results include Gemini 3.5 Pro at 45.9% and Grok 4 at 44.4% Verified Answer #3.
- Significance: This ~45 percentage-point gap highlights a barrier in high-reasoning synthesis that exceeds simple retrieval patterns Verified Answer #3Verified Answer #1.
Autonomous Agency and Reliability
Researchers use time-based metrics and specialized benchmarks to measure how well an AI can perform real work without supervision Verified Answer #1Verified Answer #2.
- Long-Horizon Reliability: The METR Time Horizon metric tracks the length of complex tasks (in human-expert hours) that an AI can complete with 50% reliability without failure Verified Answer #1.
- Autonomy Benchmarks: Tools like SWE-bench and GAIA are used to determine if a model can execute tasks and use tools effectively Verified Answer #2.
- Calibration: AGI-like systems must demonstrate calibrated truthfulness to ensure they are trustworthy when operating unsupervised Verified Answer #2.
Limitations of Legacy Benchmarks
Historical benchmarks like MMLU are no longer viewed as definitive indicators of AGI progress because they are routinely saturated by Large Language Models (LLMs) Verified Answer #3. While crossing 80–90% on MMLU indicates broad knowledge and test-taking competence, it does not establish general intelligence, as models may still fail at robust planning or novel abstraction Verified Answer #2. For example, GPT-4 achieved 86.4% on MMLU, yet most researchers do not consider such models to be AGI-like if they remain weak in autonomy and novel problem solving Verified Answer #2.