| GPT-5.x / GPT-6 | OpenAI | General purpose | GPT-6 Astra (Sep 3) is now the flagship, first model rated Critical for cybersecurity; GPT-5.6 Sol/Terra/Luna remains available; GPT-5.5 Instant is the ChatGPT default |
| Claude 5.x Family | Anthropic | Coding & reasoning | Claude Opus 5 (Jul 24) is the recommended starting model — near-Fable-5 intelligence at about half the price; Fable 5.1 (Sep 1) is the max-capability tier, now with 75% cheaper cache reads |
| Gemini 3.x | Google DeepMind | Multimodal | Gemini 3.8 Flash (Sep 2) is the new default workhorse, same price as 3.7 Flash; Gemini 3.5 Pro remains delayed with no confirmed launch date |
| Grok 4.x | xAI | Real-time info | Grok 4.6 (Aug 12) is the current flagship — 1.5T MoE, $2/$6 per 1M tokens, built for long-running agentic work |
| Llama 4 | Meta | Open-source | 10M token context, fully self-hostable; Behemoth shelved as Meta pivots to closed-weight Muse Spark |
| Muse Spark 1.3 | Meta Superintelligence Labs | Agentic coding & reasoning | Meta's first closed-weight frontier model — Spark 1.3 (Sep 2) improves coding/agentic performance for longer-horizon work, #6 on Artificial Analysis Intelligence Index |
| DeepSeek V4 | DeepSeek | Cost efficiency | V4.1-Flash (Sep 10) now outperforms V4-Pro and has taken over V4-Pro's API traffic at its lower rates — MIT license, 1M context, among the cheapest frontier-class models available |
| Mistral 3 Family | Mistral | EU compliance | Large 3, Medium 3.5, Small 4, Voxtral — enterprise-safe with EU data sovereignty |
| Qwen 3.8 | Alibaba | Multilingual | Qwen3.8-Max (Aug 2) is the new flagship; Qwen3.8-27B (Aug 13) is a separate, genuinely self-hostable Apache 2.0 companion |
| Microsoft MAI | Microsoft | Speech & media AI | MAI-Transcribe-1, MAI-Voice-1, MAI-Image-2, Phi-4-reasoning — Microsoft's own foundation stack on Foundry |
| Amazon Nova 2 | Amazon / AWS | AWS-native enterprise | Lite, Pro, Sonic, Omni — AWS-native family powering the Nova Act agent |
| MiniMax M3 | MiniMax | Cost-efficient coding | 1M context, 59.0% SWE-bench Pro at $0.15/$1.15 per 1M tokens |
| Command A+ | Cohere | Enterprise RAG | 218B MoE, Cohere's first fully Apache 2.0 frontier model, tuned for RAG |
| Gemma 4 | Google | On-device open-weight | 31B ranks #3 on Arena AI; E2B/E4B optimized for on-device Android |
| Kimi K3 | Moonshot AI | Agentic open-source | 2.8T MoE, open-weight — second only to Fable 5 and GPT-5.6 on most benchmarks |
| GLM-5.3 | Zhipu AI | Open frontier | MIT license, 1M context; GLM-5.3-Flash (Aug 26) adds native multimodal at ~1/10th the price, trained on Huawei Ascend chips |
| Atria Dawn Preview | Shanghai AI Lab | Long-horizon research agents | 744B MoE on a GLM-5.2 base, MIT license; vendor-reported top-5-of-16 benchmark claims, not yet independently verified |
| GPT-5.5-Cyber | OpenAI | Defensive cybersecurity | TAC-gated fine-tune for defensive cybersecurity — CrowdStrike, Cloudflare, Palo Alto, Cisco partners |
| Claude Mythos 5.1 | Anthropic | Security research | Cybersecurity research model via Project Glasswing — 150+ partner orgs, same weights as Fable 5.1 |
| Gemini 3.8 Flash Cyber | Google DeepMind | Restricted cybersecurity | Defensive-security variant of Gemini 3.8 Flash — gated to vetted defenders via Google's Fairwind Program, not on public price sheet |
| Sakana Fugu Max / Ultra v2 | Sakana AI | Orchestration / routing | Meta-model that dynamically routes tasks across frontier models internally; Ultra v2 hits 74.3 DeepSWE without Fable 5.1 or GPT-6-Astra in its pool |
| Apple AFM 3 | Apple | On-device + cloud privacy | Five-model family spanning on-device and cloud — privacy-preserving inference for iOS/macOS |
| Sonar | Perplexity | Search & research | Search-grounded, citation-first answers at 1,200 tok/s on Cerebras inference |
| Jev | TypeSafe AI | Structured decisions, not text | Not an LLM — outputs typed, schema-conformant decisions in parallel instead of generating text; 70-500ms, $0.042/M input, output free |
| Composer 2.5 | Cursor | AI-native coding | Built on Kimi K2.5 with Cursor's post-training — matches Opus 4.7 at ~10× cheaper |
| SubQ | Subquadratic | Architecturally novel | First commercial subquadratic LLM — 12M token context at ~1/5th the compute cost of transformers |