AI Tech News
By K.T.

Why Artificial Analysis Weights Agents at 34%: The June 2026 Shift from Capability Benchmarking to Execution Reality

The Problem Nobody Was Measuring

This isn't a story about a benchmark getting tweaked. Artificial Analysis released v4.1 on June 15, 2026, and the change cuts to something deeper: the industry finally admitted that raw capability scores stopped meaning much once all the frontier models started clustering at the top.

Version 4.1 marks a broader shift toward agentic workloads. Not because agents are trendy—because they're where the real work happens. When enterprises pick an LLM, they don't care if it scores 90th or 95th percentile on a trivia question. They care whether it can actually complete a task over multiple steps without hallucinating, backtracking, or requiring constant human intervention.

That's an execution problem, not a capability problem. And until June 2026, the index wasn't measuring it seriously.

What Actually Changed in v4.1

The new structure assigns Agents 34% of the total weight—up from 25% in v4.0. That's the headline. But the real shift is visible in the benchmarks themselves.

Artificial Analysis upgraded three evaluations and removed one. Terminal-Bench Hard became Terminal-Bench 2.1, and τ²-Bench Telecom was replaced by τ³-Bench Banking—both moving to harder, more realistic agentic scenarios. IFBench, the instruction-following benchmark, was removed due to saturation. Translation: every frontier model had stopped failing at it.

GDPval-AA was upgraded to GDPval-AA v2, re-baselining Elo to human performance at 1000 and raising the turn limit from 100 to 250 for longer-horizon agent trajectories. That matters. A 100-turn limit tests shallow reasoning. A 250-turn limit catches when models forget context, make logical errors, or fail at planning across time.

Category v4.0 Weight v4.1 Weight What Changed
Agents 25% 34% GDPval-AA v2 (20%), τ³-Banking (14%)
Coding 25% 24% Terminal-Bench 2.1 (16%), SciCode (8%)
Scientific Reasoning 25% 24% Humanity's Last Exam (12%), GPQA Diamond (6%), CritPt (6%)
General 25% 18% AA-LCR (6%), AA-Omniscience accuracy (8%) + non-hallucination (4%)

Why This Matters for Model Selection

The weighting shift reveals what enterprises actually care about. If general knowledge and reasoning were the bottleneck, v4.0's equal weighting made sense. But in production workloads in 2026, the bottleneck is execution: can the model plan a sequence of steps, use tools reliably, recover from mistakes, and deliver a working answer?

Here's where the rubber meets the road. DeepSeek V4 Pro achieves an intelligence index score of 44 at just $0.04 per task, making it over 20x cheaper than GPT-5.5 (xhigh) at $0.99 and over 44x cheaper than Claude Opus 4.8 (max) at $1.78. That cost gap doesn't come from one model being smarter in a vacuum—it comes from different operating costs, quantization choices, and inference stacks.

But v4.1's agentic emphasis surfaces something else: for workloads where multi-step agent performance in the low-to-mid 40s is sufficient, the cost-to-capability ratio is extraordinary. Conversely, if you need reasoning chains that hold up past 200 turns or error recovery across complex workflows, you may not get adequate differentiation from the cheaper models.

The official Intelligence Index v4.1 assigns 34% to Agents, 24% each to Coding and Scientific Reasoning, and 18% to General. That means if your workload is pure code generation, don't optimize against the total score—look at the Coding subcategory. If you're running long-running agents that need to debug and retry, the Agents weighting matters more than whether the model can recite physics facts.

What This Signals About the Evaluation Crisis

The bigger picture: v4.1 exists because traditional benchmarks broke. Not because vendors cheated, but because the gap between "frontier model on a test" and "frontier model in production" had become too wide to ignore.

When every top model scores in the high 90s on instruction-following, you lose signal. When every model can pass MMLU questions, you don't know which one will hallucinate during a 50-step workflow or forget a constraint mid-execution.

Released on June 16, 2026, v4.1 marks a decisive shift away from static question-answer benchmarks toward multi-step, long-horizon agent tasks that reveal how models behave under realistic enterprise workloads.

Artificial Analysis is signaling that capability and execution are no longer synonymous. A model can score 60 on raw reasoning and 30 on agent execution—and which score matters depends entirely on what you're trying to do. The 34% weighting for agents finally makes that trade-off explicit instead of hidden.

What This Means for Your Model Selection

If you're evaluating models for your team or business:

  • Don't optimize purely for the total score. Look at which category weights match your workload. If you're building agents that run autonomously, the Agents subcategory is more predictive than General knowledge.
  • Watch the cost-per-task metrics, not just the price card. The per-task metrics in v4.1 reveal a striking divergence between capability and cost. A cheaper model might score adequate on agents but have different failure modes under concurrent load or adversarial prompts.
  • Check the benchmark details, not just the category scores. Terminal-Bench 2.1 is harder than v2.0 because it tests realistic shell workflows. τ³-Banking tests multi-step financial operations. GDPval-AA's 250-turn limit catches long-horizon drift. If your actual workload is similar, those benchmarks are predictive. If not, they're noise.
  • Plan for benchmark churn. Artificial Analysis revises both benchmark composition and model results. A score from August 2026 might not compare cleanly to a score from next year. That's not a flaw—it's a feature. It means the benchmark is staying harder as models improve.

The v4.1 shift signals that the AI evaluation industry is maturing past generic "smarter or dumber" reasoning and toward "works in production or doesn't." That's the question that actually matters.

Our tracked data

AI Intelligence Index (Top 3 Frontier Models)

01632476305-1706-0106-0807-0607-1307-2007-2708-0308-1008-1708-2408-31Claude Opus 4.7 (Adaptive Reasoning, Max Effort) — Anthropic: 57 (2026-05-17)Claude Opus 4.8 (Adaptive Reasoning, Max Effort) — Anthropic: 61 (2026-06-01)Claude Opus 4.8 (Adaptive Reasoning, Max Effort) — Anthropic: 61 (2026-06-08)Claude Opus 4.8 (Adaptive Reasoning, Max Effort) — Anthropic: 56 (2026-07-06)Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) — Anthropic: 60 (2026-07-13)Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) — Anthropic: 59.9 (2026-07-20)Claude Opus 5 (Adaptive Reasoning, Max Effort) — Anthropic: 61 (2026-07-27)Claude Opus 5 (Adaptive Reasoning, Max Effort) — Anthropic: 61 (2026-08-03)Claude Opus 5 (Adaptive Reasoning, Max Effort) — Anthropic: 63 (2026-08-10)Claude Opus 5 (Adaptive Reasoning, Max Effort) — Anthropic: 63 (2026-08-17)Claude Opus 5 (Adaptive Reasoning, Max Effort) — Anthropic: 63 (2026-08-24)Claude Opus 5 (Adaptive Reasoning, Max Effort) — Anthropic: 63 (2026-08-31)63GPT-5.5 (xhigh) — OpenAI: 60 (2026-05-17)GPT-5.5 (xhigh) — OpenAI: 60 (2026-06-01)GPT-5.5 (xhigh) — OpenAI: 60 (2026-06-08)GPT-5.5 (xhigh) — OpenAI: 55 (2026-07-06)GPT-5.6 Sol (max) — OpenAI: 59 (2026-07-13)GPT-5.6 Sol (max) — OpenAI: 58.9 (2026-07-20)GPT-5.6 Sol (max) — OpenAI: 59 (2026-07-27)GPT-5.6 Sol (max) — OpenAI: 59 (2026-08-03)GPT-5.6 Sol (max) — OpenAI: 61 (2026-08-10)GPT-5.6 Sol (max) — OpenAI: 61 (2026-08-17)GPT-5.6 Sol (max) — OpenAI: 61 (2026-08-24)GPT-5.6 Sol (max) — OpenAI: 61 (2026-08-31)61Gemini 3.1 Pro Preview — Google DeepMind: 57 (2026-05-17)Gemini 3.1 Pro Preview — Google DeepMind: 57 (2026-06-01)Gemini 3.1 Pro Preview — Google DeepMind: 57 (2026-06-08)Gemini 3.1 Pro Preview — Google DeepMind: 46 (2026-07-06)Gemini 3.5 Flash (high) — Google DeepMind: 55 (2026-07-13)Gemini 3.1 Pro Preview — Google DeepMind: 46 (2026-07-20)Gemini 3.6 Flash (high) — Google DeepMind: 50 (2026-07-27)Gemini 3.6 Flash (high) — Google DeepMind: 50 (2026-08-03)Gemini 3.6 Flash (high) — Google DeepMind: 52 (2026-08-10)Gemini 3.7 Flash (high) — Google DeepMind: 56 (2026-08-17)Gemini 3.7 Flash (high) — Google DeepMind: 56 (2026-08-24)Gemini 3.7 Flash (high) — Google DeepMind: 56 (2026-08-31)56
  • Anthropic
  • OpenAI
  • Google DeepMind

Intelligence Index — Trend

Hover over each point to see the specific model version at that date.

Last updated: 2026-08-31 · 12 data points · artificialanalysis.ai

Collected weekly by our editorial team from primary sources.

See the full dataset