| GPT-5.x | OpenAI | General purpose | GPT-5.6 Sol/Terra/Luna (Jul 9) is the current flagship API family — Sol ($5/$30) best coding, Terra ($2.50/$15) mid-tier, Luna ($1/$6) lightweight, dynamic routing across sub-models; GPT-5.5 Instant (May 5) is the ChatGPT default for all tiers — leaner variant, ~30% fewer words/lines, optimized for low latency; GPT-Live-1/mini (Jul 8) are full-duplex voice models with agentic delegation to GPT-5.5, now the default ChatGPT voice mode |
| Claude 5.x Family | Anthropic | Coding & reasoning | Claude Opus 5 (Jul 24) — new flagship, 1M context, 128K output (300K on Batch API), $5/$25 per 1M tokens; frontier intelligence of Fable 5 at roughly half the price, the recommended starting model for agentic coding & enterprise work; Fable 5 (Jun 9) — max-capability option, 1M context, 128K output, $10/$50 per 1M tokens, 95% SWE-bench Verified; Sonnet 5 (Jun 30) — default Free/Pro model, near-flagship performance at $2/$10 per 1M (intro); Haiku 4.5 hits ~90% of Sonnet 5 coding at a fraction of the price |
| Gemini 3.x | Google DeepMind | Multimodal | Gemini 3.5 Pro (I/O 2026): has missed three consecutive GA deadlines (most recently July 17); no confirmed launch date, reportedly rebuilt from scratch after recursive tool-calling failures; still limited Vertex enterprise preview; leads 3.1 Pro on reasoning; Gemini 3.6 Flash (Jul 21) is the new default in the Gemini app, superseding 3.5 Flash — 17% fewer output tokens, output price cut to $7.50/M (from $9.00), knowledge cutoff advanced to March 2026; siblings Gemini 3.5 Flash-Lite and Flash Cyber launched same day; Gemini 4 teased, no release date; 3.1 Ultra remains the top reasoning tier |
| Grok 4.x | xAI | Real-time info | Grok 4.3 (Apr 30) is the cost-efficient API model at $1.25/$2.50 per M tokens with always-on reasoning and native video input (5-min clips); Grok 4 Heavy is the premium multi-agent variant (256K context, first model to hit 50.7% on Humanity's Last Exam (text-only with tools)) gated behind the $300/mo SuperGrok Heavy tier; Grok 4.5 launched July 8, 2026 — 1.5T-param V9 MoE, 500K context, $2/$6 per 1M tokens; EU rollout partial as of late July — live via Cursor, direct API console still restricted |
| Llama 4 | Meta | Open-source | 10M token context; fully self-hostable; Behemoth effectively shelved — repeated delays, MoE-routing failures; Meta pivoting to Muse Spark |
| Muse Spark 1.1 | Meta Superintelligence Labs | Agentic coding & reasoning | Released Apr 8, 2026, upgraded to 1.1 on July 9, 2026 — Meta's first closed-weight frontier model and first paid developer API ($1.25/$4.25 per 1M tokens); 1M token context, agentic coding focus (tool use, computer use, multi-agent orchestration); top scores on MCP Atlas, JobBench, Humanity's Last Exam |
| DeepSeek V4 | DeepSeek | Cost efficiency | Released April 24, 2026 — V4-Pro (1.6T MoE, $0.435/1M input, $0.87/1M output after a permanent May 22 price cut) and V4-Flash (284B, $0.14/1M input); MIT, 1M context |
| Mistral 3 Family | Mistral | EU compliance | Large 3, Magistral (reasoning), Devstral (open-source coding agent), Small 4, Voxtral — enterprise-safe with data sovereignty |
| Qwen 3.7 | Alibaba | Multilingual | Qwen 3.7 Max (May 19, closed-weight) — 56.6 AI Intelligence Index, $2.50/$7.50 per 1M tokens; Qwen 3.7 Plus with vision GA June 1; Qwen3.5 open-weight (397B) and Qwen3.6-Plus still available; Qwen3.8-Max previewed Jul 19 (2.4T params, not yet GA) |
| Microsoft MAI | Microsoft | Speech & media AI | MAI-Transcribe-1, MAI-Voice-1, MAI-Image-2, plus Phi-4-reasoning (Apr 10) — Microsoft's own foundation stack on Foundry |
| Amazon Nova 2 | Amazon / AWS | AWS-native enterprise | Released Dec 2025 — Lite (1M context, MCP), Pro (reasoning), Sonic (speech-to-speech), Omni (multimodal); powers Nova Act agent (90%+ task reliability) |
| MiniMax M3 | MiniMax | Cost-efficient coding | Jun 1, 2026 — 1M context, 59.0% SWE-bench Pro; successor to M2.5/M2.7; $0.15/$1.15 per 1M tokens |
| Command A+ | Cohere | Enterprise RAG | Released May 20, 2026 — 218B MoE / 25B active, Cohere's first fully Apache 2.0 frontier model, runs on 2 H100s; supersedes Command A (Apr 7); tuned for retrieval-augmented generation |
| Gemma 4 | Google | On-device open-weight | 31B ranks #3 on Arena AI (1452 Elo); E2B/E4B optimized for Android; Apache 2.0 |
| Kimi K3 | Moonshot AI | Agentic open-source | Released Jul 16, full weights Jul 27 — 2.8T MoE, second only to Fable 5 and GPT-5.6 on most benchmarks; supersedes K2.7 Code; Modified MIT |
| GLM-5.2 | Zhipu AI | Open frontier | Released Jun 13, 2026 (Z.ai / Zhipu AI); 1M context window; MIT license; trained on zero NVIDIA GPUs |
| GPT-5.5-Cyber | OpenAI | Defensive cybersecurity | Released May 7, 2026, superseding GPT-5.4-Cyber; TAC-gated fine-tune; lowered security refusals; binary reverse engineering; partners: CrowdStrike, Cloudflare, Palo Alto, Cisco |
| Claude Mythos 5 | Anthropic | Security research | Graduated from Preview to full release June 9, 2026 alongside Fable 5; Project Glasswing expanded to 150+ organizations across 15+ countries for critical-infrastructure cybersecurity; GPQA 0.9; high-cost pricing retained to gate access |
| Sakana Fugu Ultra | Sakana AI | Orchestration / routing | Released Jun 22, 2026 — meta-model that dynamically routes tasks across frontier models internally; 73.7% SWE-Bench Pro; $5/$30 per 1M tokens standard, $10/$45 above 272K context |
| Apple AFM 3 | Apple | On-device + cloud privacy | AFM 3 family shipped at WWDC 2026 (June 8) — five models spanning on-device (Core, Core Advanced) and cloud (Cloud, Cloud Pro, ADM 3 Cloud for image gen); multimodal image input, Python SDK; on-device tiers handle sensitive tasks locally, Cloud Pro (hosted on Google Cloud/NVIDIA via Private Cloud Compute) for heavier workloads |
| Sonar | Perplexity | Search & research | Search-grounded, citation-first answers at 1,200 tok/s on Cerebras inference |
| Composer 2.5 | Cursor | AI-native coding | May 18, 2026 — built on Kimi K2.5 with Cursor's post-training; 79.8% SWE-bench Multilingual, 69.3% Terminal-Bench 2.0; matches Opus 4.7 at ~10× cheaper ($0.50/$2.50 per M tokens) |
| SubQ | Subquadratic | Architecturally novel | First commercial subquadratic LLM (May 5, 2026) — Subquadratic Sparse Attention (SSA) scales ~linearly; 12M token production context; Claude Opus-level coding at ~1/5th the compute cost, attention up to 52× faster at scale |