Comparison of generative AI models
Appearance
This is a comparison of frontier models in generative AI, according to aggregates of benchmarks.
Large language models
[edit]The Intelligence Index released by benchmarking firm Artificial Analysis aggregates nine benchmarks: GDPval-AA v2, đÂł-Banking, Terminal-Bench v2.1, SciCode, AA-LCR, AA-Omniscience, Humanity's Last Exam, GPQA Diamond, and CritPt.[1] Only the highest "effort" setting for each model is shown below.
Table
[edit]Notes
[edit]- â Artificial Analysis notes its evaluation includes 'fallback' where Fable 5 passes some queries to Opus 5
See also
[edit]References
[edit]- â "Intelligence Benchmarking | Artificial Analysis". artificialanalysis.ai. Retrieved 2026-08-24.
- â "Comparison of AI Models across Intelligence, Performance, and Price | Artificial Analysis". artificialanalysis.ai. Retrieved 2026-08-24.