AI Resource

AI Model Reference Guide

A regularly updated reference covering the most relevant large language models from the major AI providers. Use this guide to quickly compare capabilities, pricing, and use cases when evaluating AI models.

Provider Models

AI Models

Data sourced from official provider documentation and OpenRouter.ai. Pricing reflects list rates as of the last update date. Always verify with the provider before making purchasing decisions.

Loading models...

Compare by Need

Quick Reference

Top picks across common use-case categories.

Last Reviewed: August 10, 2026

Largest context window

Llama 5 (5M trained) · Gemini 3.1 Pro (2M extended) · Gemini 3.5 Pro (2M, limited preview) > Claude Opus 4.8 & Claude Opus 4.7 & GPT-5.5 (1M) > Llama 4 Scout (328K via API; 10M trained) > Claude Sonnet 4.6 (200K) > Claude Opus 4.7-fast (128K) > Claude Haiku 4.5 (200K)

Best for coding / SWE

Claude Fable 5 (80.3% SWE-Bench Pro) · Claude Opus 5 (near-Fable-5 on CursorBench 3.2) · Claude Opus 4.8:thinking (69.2%) · Claude Sonnet 5 (63.2%) · GPT-5.5 · o3-pro / o4-mini

Best overall value (frontier)

Claude Sonnet 5 ($2/$10 intro through Aug 2026) · Gemini 3.1 Pro ($2/$12) · Claude Sonnet 4.6 ($3/$15) · Claude Opus 5 ($5/$25, near-Fable-5 intelligence)

Best budget / high-volume

Llama 4 Scout ($0.08/$0.30) · Gemini 3.1 Flash Lite ($0.25/$0.50) · Claude Haiku 4.5 ($1/$5)

Best multimodal (text + image + video + audio)

Gemini 3.1 Pro (only model natively supporting all four) · GPT-5.4 (text + image + computer use)

Best deep reasoning / math / science

Claude Opus 5 (leads Frontier-Bench v0.1) · Claude Opus 4.8:thinking · o3-pro · Gemini 3.1 Pro (ARC-AGI-2 77.1%) · o4-mini · o1-pro · Claude Opus 4.8-fast:thinking

Best safety & instruction following

Claude Opus 4.7 · Claude Sonnet 4.6 (Constitutional AI; strongest instruction adherence) · Claude 3.7 Sonnet

Best open-source / self-hostable

Llama 5 (frontier-class, 600B+ open-weights) · Gemma 4 31B (Apache 2.0) · Llama 4 Scout (budget) · Phi-4 / Phi-4-mini (edge & on-device)

Best for M365 / enterprise ecosystem

Microsoft 365 Copilot (deep M365 integration with Microsoft Graph grounding)

Best computer use / desktop agents

Claude Opus 4.8 (84% Online-Mind2Web, strongest computer-use benchmark) · Claude Sonnet 4.6 (94% on insurance benchmarks) · GPT-5.5 · Claude Opus 4.8-fast (lower-latency agent deployment)