New: Build with SOLAR

AI Models
Guide

A quick-reference to the major AI models, who makes them, and what they do best.

Updated September 19, 2026
On This Page
ModelCompanyBest ForKey Differentiator
GPT-5.x / GPT-6OpenAIGeneral purposeGPT-6 Astra (Sep 3) is now the flagship, first model rated Critical for cybersecurity; GPT-5.6 Sol/Terra/Luna remains available; GPT-5.5 Instant is the ChatGPT default
Claude 5.x FamilyAnthropicCoding & reasoningClaude Opus 5 (Jul 24) is the recommended starting model — near-Fable-5 intelligence at about half the price; Fable 5.1 (Sep 1) is the max-capability tier, now with 75% cheaper cache reads
Gemini 3.xGoogle DeepMindMultimodalGemini 3.8 Flash (Sep 2) is the new default workhorse, same price as 3.7 Flash; Gemini 3.5 Pro remains delayed with no confirmed launch date
Grok 4.xxAIReal-time infoGrok 4.6 (Aug 12) is the current flagship — 1.5T MoE, $2/$6 per 1M tokens, built for long-running agentic work
Llama 4MetaOpen-source10M token context, fully self-hostable; Behemoth shelved as Meta pivots to closed-weight Muse Spark
Muse Spark 1.3Meta Superintelligence LabsAgentic coding & reasoningMeta's first closed-weight frontier model — Spark 1.3 (Sep 2) improves coding/agentic performance for longer-horizon work, #6 on Artificial Analysis Intelligence Index
DeepSeek V4DeepSeekCost efficiencyV4.1-Flash (Sep 10) now outperforms V4-Pro and has taken over V4-Pro's API traffic at its lower rates — MIT license, 1M context, among the cheapest frontier-class models available
Mistral 3 FamilyMistralEU complianceLarge 3, Medium 3.5, Small 4, Voxtral — enterprise-safe with EU data sovereignty
Qwen 3.8AlibabaMultilingualQwen3.8-Max (Aug 2) is the new flagship; Qwen3.8-27B (Aug 13) is a separate, genuinely self-hostable Apache 2.0 companion
Microsoft MAIMicrosoftSpeech & media AIMAI-Transcribe-1, MAI-Voice-1, MAI-Image-2, Phi-4-reasoning — Microsoft's own foundation stack on Foundry
Amazon Nova 2Amazon / AWSAWS-native enterpriseLite, Pro, Sonic, Omni — AWS-native family powering the Nova Act agent
MiniMax M3MiniMaxCost-efficient coding1M context, 59.0% SWE-bench Pro at $0.15/$1.15 per 1M tokens
Command A+CohereEnterprise RAG218B MoE, Cohere's first fully Apache 2.0 frontier model, tuned for RAG
Gemma 4GoogleOn-device open-weight31B ranks #3 on Arena AI; E2B/E4B optimized for on-device Android
Kimi K3Moonshot AIAgentic open-source2.8T MoE, open-weight — second only to Fable 5 and GPT-5.6 on most benchmarks
GLM-5.3Zhipu AIOpen frontierMIT license, 1M context; GLM-5.3-Flash (Aug 26) adds native multimodal at ~1/10th the price, trained on Huawei Ascend chips
Atria Dawn PreviewShanghai AI LabLong-horizon research agents744B MoE on a GLM-5.2 base, MIT license; vendor-reported top-5-of-16 benchmark claims, not yet independently verified
GPT-5.5-CyberOpenAIDefensive cybersecurityTAC-gated fine-tune for defensive cybersecurity — CrowdStrike, Cloudflare, Palo Alto, Cisco partners
Claude Mythos 5.1AnthropicSecurity researchCybersecurity research model via Project Glasswing — 150+ partner orgs, same weights as Fable 5.1
Gemini 3.8 Flash CyberGoogle DeepMindRestricted cybersecurityDefensive-security variant of Gemini 3.8 Flash — gated to vetted defenders via Google's Fairwind Program, not on public price sheet
Sakana Fugu Max / Ultra v2Sakana AIOrchestration / routingMeta-model that dynamically routes tasks across frontier models internally; Ultra v2 hits 74.3 DeepSWE without Fable 5.1 or GPT-6-Astra in its pool
Apple AFM 3AppleOn-device + cloud privacyFive-model family spanning on-device and cloud — privacy-preserving inference for iOS/macOS
SonarPerplexitySearch & researchSearch-grounded, citation-first answers at 1,200 tok/s on Cerebras inference
JevTypeSafe AIStructured decisions, not textNot an LLM — outputs typed, schema-conformant decisions in parallel instead of generating text; 70-500ms, $0.042/M input, output free
Composer 2.5CursorAI-native codingBuilt on Kimi K2.5 with Cursor's post-training — matches Opus 4.7 at ~10× cheaper
SubQSubquadraticArchitecturally novelFirst commercial subquadratic LLM — 12M token context at ~1/5th the compute cost of transformers

Building Software

  • Build a full-stack app from scratchClaude Opus 5
  • Debug a complex codebaseClaude Opus 5
  • Generate unit tests and docsGPT-5.x
  • Rapid UI prototypingGPT-5.x
  • Background agents for parallel developmentComposer 2.5 (Cursor)
  • Open-source agentic codingKimi K3

Research & Analysis

  • Analyze a long PDF or contractGemini 3.x
  • Summarize a YouTube videoGemini 3.x
  • Get real-time data on a trending topicGrok 4.x
  • Get sourced answers with citationsSonar (Perplexity)
  • Deep competitive researchClaude Fable 5.1

Creative & Visual

  • Create stylized hero imagesMidjourney v7
  • Generate photorealistic product shotsImagen 4
  • Edit and remix existing imagesNano Banana 2
  • Generate a short video from a promptVeo 3.1

Data & Math

  • Solve complex math problems step-by-stepGrok 4.x
  • Write and optimize SQL queriesGPT-5.x
  • Transparent chain-of-thought reasoningDeepSeek V4
  • Analyze spreadsheet dataGemini 3.x

Self-Hosting & Privacy

  • Run a model on your own infrastructureLlama 4
  • Fine-tune for a domain-specific taskLlama 4
  • Deploy in EU-regulated environmentsMistral Large
  • Budget-friendly open-source alternativeDeepSeek V4
  • MIT-licensed frontier alternativeGLM-5.3

Writing & Communication

  • Write long-form technical contentClaude Sonnet 5
  • Draft emails and business writingGPT-5.x
  • Translate content across 100+ languagesQwen 3.8
  • Summarize meeting transcriptsGemini 3.x

GPT-5.x / GPT-6

OpenAI · San Francisco

The versatile all-rounder with dynamic internal routing.

  • GPT-5.6 Sol / Terra / Luna (July 9, 2026): current flagship API family — 1M context, internal router picks the sub-model per request. Sol (best coding): $5/$30 per 1M, Fast Mode 2.5× faster at 2× price. Terra (balanced): $2/$12 per 1M. Luna (cheap): $0.20/$1.20 per 1M.
  • GPT-5.5 Instant (May 5, 2026): ChatGPT's default for all tiers — leaner, lower-latency variant distinct from the API flagship family
  • Available on Plus/Pro/Business/Enterprise via ChatGPT and Codex; native computer use and tool calling for agentic automation
  • GPT-Live-1 / mini (July 8, 2026): full-duplex voice models with real-time translation — now the default ChatGPT voice mode (Live-1 for Plus/Pro, mini for Free)
  • GPT-6 Astra (Sept 3, 2026): new flagship, supersedes GPT-5.6 Sol — $10/$50 per 1M, 72.6% OSWorld 2.0 (vs 65.7% for Sol); first model rated Critical for cybersecurity under OpenAI's Preparedness Framework — rollout started with enterprise Daybreak-program partners before reaching ChatGPT/API

Claude 5.x Family

Anthropic · San Francisco

The developer favorite for coding, reasoning, and safety.

  • Claude Opus 5 (July 24, 2026): flagship — 1M context, 128K output, $5/$25 per 1M; near-Fable-5 intelligence at half the price; model ID: claude-opus-5
  • Fable 5.1 (Sept 1, 2026): most capable model — 1M context, 128K output, $10/$50 per 1M (cache reads now $0.25/M, down 75%); outperforms Fable 5, Opus 5, and GPT-5.6 Sol; ~25-45% cheaper than Fable 5 for typical/agentic work, 60% fewer cybersecurity false positives
  • Sonnet 5 (June 30, 2026): default on Free/Pro tiers — near-flagship performance, $2/$10 per 1M; model ID: claude-sonnet-5
  • Dynamic Workflows: run parallel subagents on independent subtasks
  • Fast mode across the Opus line: ~2.5× faster at ~2-3× price
  • Haiku 4.5: ~90% of Sonnet 5's coding performance at a fraction of the cost
  • Leads human-preference leaderboards, strong ARC-AGI-2 scores
  • Claude Mythos 5.1 (Sept 1, 2026): cybersecurity research model via Project Glasswing — same weights as Fable 5.1

Gemini 3.x

Google DeepMind · Mountain View

Multimodal powerhouse with top benchmark breadth.

  • Gemini 3.5 Pro (I/O 2026): next flagship — repeatedly delayed, no confirmed launch date; 2M context, Deep Think mode
  • Gemini 3.8 Flash (Sept 2, 2026): new default workhorse, built on 3.7 Flash — beats 3.7 on every benchmark Google published, 1M context, 64K output, same $0.75/$3.75 per 1M pricing through Dec 31, 2026
  • Gemini 3.5 Flash-Lite and Flash Cyber (July 21, 2026): cost-efficient high-volume and defensive-security variants
  • Gemini 3.1 Ultra (May 2026): top reasoning tier — 2M context, natively multimodal
  • Gemini 3.1 Pro (Feb 2026): deployed Pro-tier flagship — 94.3% GPQA Diamond, 1M context, $2/$12 per 1M
  • Gemini 3.1 Flash Live and Flash TTS round out the audio lineup; Omni Flash (May 19, 2026) adds any-to-any input with video output
  • Gemini Spark (May 19, 2026): Google's 24/7 persistent AI agent, running on Gemini 3.7 Flash

Grok 4.x

xAI · Austin

Real-time data meets raw reasoning power.

  • Grok 4.3 (Apr 30, 2026): cost-efficient API flagship — $1.25/$2.50 per 1M, always-on reasoning, native video input up to 5-minute clips
  • Grok 4 Heavy: premium multi-agent variant — 256K context, first model to score 50.7% on Humanity's Last Exam; gated behind the $300/month SuperGrok Heavy tier
  • Grok 4 Heavy also scored 100% on AIME 2025 — often mistakenly attributed to the base Grok 4.x model
  • Real-time integration with X (Twitter) for current events
  • Grok 4.6 (Aug 12, 2026): supersedes Grok 4.5 — same 1.5T-param MoE, $2/$6 per 1M, 500K context; scores 61 on the Artificial Analysis Intelligence Index (near GPT-5.6 Sol); focus on long-running agentic work
  • Grok Voice: voice interface model released June 4, 2026
  • Grok 4.7 (2.1T params) expected within weeks of 4.6; Grok 5 (~6T MoE) now targeted before end of 2026

Sonar (Perplexity)

Perplexity AI · San Francisco

Search-native AI built for grounded, cited answers.

  • Custom Sonar model fine-tuned for search-grounded factuality and citation accuracy
  • Runs at 1,200 tokens/sec on Cerebras inference hardware
  • Model family: Sonar, Sonar Pro, Reasoning Pro, Deep Research
  • Matches GPT-4o on user satisfaction benchmarks

Jev (TypeSafe AI)

TypeSafe AI · San Francisco

Not an LLM — a System One model that outputs typed, schema-conformant decisions instead of generated text.

  • Early access Sept 15, 2026 — founded by Diogo Almeida
  • "Unstructured state in, typed probabilistic decisions out" — built for classification, routing, scoring, and extraction, not chat
  • Outputs generated in parallel rather than token-by-token, so schema conformance is structural, not prompted — can't hallucinate or produce a type error by design
  • 70-500ms end-to-end — 40-200× faster than a comparable frontier LLM on the same task
  • $0.042 per 1M input tokens, output free; trade-off is no text generation — structured decisions only

Composer 2.5 (Cursor)

Cursor · San Francisco

Frontier-class coding at one-tenth the price — built on Kimi K2.5 with Cursor's own post-training pipeline.

  • Released May 18, 2026 — built on Kimi K2.5 with Cursor's own post-training
  • 79.8% SWE-bench Multilingual, 69.3% Terminal-Bench 2.0 — matches Claude Opus 4.7 and GPT-5.5 on coding benchmarks
  • $0.50/$2.50 per 1M — roughly 10× cheaper than Opus 4.7
  • Trained with RL for complex, hundreds-of-steps tasks
  • Fast variant is the default; background agents run tasks autonomously

Microsoft MAI

Microsoft · Redmond

Microsoft's own foundation model stack — independent of OpenAI, built for speech, voice, vision, and reasoning.

  • MAI-Transcribe-1: speech-to-text across 25 languages, outperforms Whisper-large-v3
  • MAI-Voice-1: generates 60s of audio in 1s, supports voice cloning
  • MAI-Image-2: high-quality image generation
  • Phi-4-reasoning (Apr 10, 2026): compact reasoning model, strong math/logic at small-model cost
  • MAI-Thinking-1 (Jun 2026): 1T/35B active MoE reasoning model — flagship of a seven-model MAI launch
  • MAI-Code-1 / Flash (Jun 2026): 5B-param coding model trained on Copilot's production workflows
  • All available on Microsoft Foundry — independent of OpenAI

Meta Muse Spark 1.3

Meta Superintelligence Labs · Menlo Park

Meta's first closed-weight frontier model — a deliberate break from the open Llama strategy, now shipping with Meta's first terminal coding agent.

  • Released April 8, 2026 — led by Alexandr Wang (ex-Scale AI)
  • Multimodal reasoning with thought compression and parallel sub-agent orchestration
  • Scored 52 on the Artificial Analysis Intelligence Index — top 5 globally
  • Muse Spark 1.1 (July 9, 2026): 1M context, Meta's first paid developer API, $1.25/$4.25 per 1M
  • Muse Spark 1.3 (Sept 2, 2026): improved coding/agentic performance for longer-horizon work — #6 on Artificial Analysis Intelligence Index, $1.25/$4.25 per 1M

Apple AFM 3

Apple · Cupertino

Apple's native on-device + cloud model framework — privacy-preserving inference shipping with iOS/macOS developer APIs.

  • AFM 3 family shipped at WWDC 2026 (June 8) — five models spanning on-device (Core, Core Advanced) and cloud (Cloud, Cloud Pro, ADM 3 Cloud for image gen)
  • Multimodal image input added this generation; Python SDK released for developers
  • Cloud Pro: hosted on Google Cloud/NVIDIA via Private Cloud Compute for heavier workloads; on-device tiers handle sensitive tasks locally
  • Ships as native iOS/macOS developer APIs

Amazon Nova 2

Amazon / AWS · Seattle

AWS's enterprise foundation model family — multimodal, agentic, and natively integrated with Bedrock.

  • Released December 2025 at AWS re:Invent — Amazon's flagship in-house model family
  • Nova 2 Lite: 1M context with native MCP support — fast, cost-efficient tier
  • Nova 2 Pro: tuned for complex reasoning and agentic workflows
  • Nova 2 Sonic: speech-to-speech for low-latency voice apps
  • Nova 2 Omni: unified multimodal across text, image, audio, and video
  • Powers Nova Act (agentic browser, 90%+ task reliability) and Nova Forge for custom model training

The biggest story of April 2026: Anthropic and OpenAI each released a cybersecurity-focused model within days of each other — gated, expensive, and limited to enterprise partners. These are not general-purpose models.

GPT-5.5-Cyber

OpenAI · San Francisco

Fine-tuned GPT-5.5 for enterprise defensive cybersecurity — gated behind OpenAI's Trusted Access for Cyber program.

  • Released May 7, 2026 — fine-tune of GPT-5.5 for dual-use security research; supersedes GPT-5.4-Cyber, still available under the same program
  • Lowered refusal thresholds for defensive cybersecurity tasks; native binary reverse engineering without source code
  • Gated behind OpenAI's Trusted Access for Cyber (TAC) program; partners include CrowdStrike, Cloudflare, Palo Alto, Cisco, JPMorgan, Goldman Sachs
  • No public API pricing; $10M in credits committed via the Cybersecurity Grant Program
  • Direct counterpart to Claude Mythos 5 — a matched pair of restricted cyber models

Claude Mythos 5.1

Anthropic · San Francisco

Cybersecurity research model via Project Glasswing — Anthropic's counterpart to GPT-5.5-Cyber. Updated to Mythos 5.1 alongside Fable 5.1 on Sept 1, 2026.

  • Originally launched April 6–7, 2026 as Mythos Preview; graduated to Mythos 5 on June 9, 2026 alongside Fable 5; updated to Mythos 5.1 on Sept 1, 2026 alongside Fable 5.1
  • GPQA 0.9; 93.9% SWE-bench Verified; 97.6% USAMO 2026 (as measured on Mythos 5)
  • Expanded June 2, 2026 from ~12 initial partners to 150+ organizations across 15+ countries for critical-infrastructure cybersecurity
  • Same weights as Fable 5.1 — safety classifiers lifted only for Glasswing partners, not a separate model
  • High-cost pricing retained to gate general use and prevent misuse

Gemini 3.8 Flash Cyber

Google DeepMind · Mountain View

Restricted defensive-security variant of Gemini 3.8 Flash — gated to vetted cyber defenders via Google's Fairwind Program.

  • Released September 2, 2026 alongside general-availability Gemini 3.8 Flash
  • Restricted to governments, critical infrastructure operators, and software maintainers approved through the Fairwind Program — not on general release or a public price sheet
  • 70%+ real-world vulnerability discovery rate; sits on the CWE-Bench Pareto frontier for patching speed
  • Participating orgs must limit access to internal cybersecurity/incident-response/pentest teams and enforce MFA
  • Third major lab (after OpenAI's GPT-5.5-Cyber and Anthropic's Mythos 5.1) to ship a gated defensive-cyber model

Models that present as a single API but dynamically route tasks across multiple frontier models internally — no hardcoded rules. A new category as of mid-2026.

Sakana Fugu Max / Fugu Ultra v2

Sakana AI · Tokyo, Japan

A meta-model that presents as a single API but dynamically routes tasks across a pool of frontier models internally — no hardcoded rules. Now ships in two tiers.

  • Fugu Max and Fugu Ultra v2 (Sept 11, 2026) replace the original Fugu Ultra — same orchestration architecture split into a cost-efficient tier (Max) and a max-capability tier (Ultra v2)
  • Fugu Ultra v2 hits 74.3 on DeepSWE and 48.3 on Chartography (vs. 27.3 for Opus 5, 29.5 for Fable 5) — and does it without Fable 5, Fable 5.1, or GPT-6-Astra in its own agent pool
  • Pricing: Fugu Max $2/$6 per 1M; Fugu Ultra v2 $5/$30 per 1M standard ($0.50/1M cache reads), $10/$45/$1 above 272K context
  • Can substitute models around export controls — built for research, cybersecurity, multi-step patent work
  • No hardcoded routing rules — the meta-model selects and coordinates sub-models per task

Llama 4

Meta · Menlo Park

The leading open-source model family.

  • Llama 4 Scout: industry-leading 10M token context window
  • Llama 4 Maverick: 17B active / 128 experts — outperforms GPT-4o and Gemini 2.0 Flash on key benchmarks
  • Fully open weights; can be self-hosted for complete data control
  • Llama 4 Behemoth (288B active): effectively shelved — repeated delays and MoE-routing failures at 2T scale; Meta pivoting to closed-weight Muse Spark
  • Llama 4.5: announced for 2026 as the next open Llama release, distinct from closed-weight Muse Spark — no release date confirmed

DeepSeek V4

DeepSeek · Hangzhou, China

Cost-redefining open-source frontier — V4.1-Flash (Sept 10, 2026) now outperforms the flagship V4-Pro.

  • V4.1-Flash (Sept 10, 2026): 552B MoE, 8B active on input / 16B active on output — new Causal Encoder-Decoder architecture, native multimodal image understanding, MIT license; beats the larger V4-Pro on cost, speed, and total runtime
  • $0.15/$0.60 per 1M off-peak, doubling to $0.30/$1.20 at peak hours — among the cheapest frontier-class models available
  • Starting Sept 14, 2026: all V4-Pro API requests are answered by V4.1-Flash and billed at its lower rate, until a V4.1-Pro ships; V4-Flash and the August vision-preview model are retired
  • MIT license; native 1M context window; built-in agentic long-context and tool-use
  • Entire V3/V4 lineage trained for under $6M — redefining AI cost efficiency

Mistral 3 Family

Mistral AI · Paris, France

The enterprise-safe European model family, now spanning text, reasoning, code, and speech.

  • Mistral Large 3: EU AI Act-compliant flagship for regulated industries; 675B total MoE
  • Mistral Medium 3.5 (Apr 29, 2026): 128B dense, 256K context — retired Magistral (reasoning), Devstral 2 (coding), and Medium 3.1 (chat) into one unified model with configurable reasoning effort; 77.6% SWE-bench Verified, $1.50/$7.50 per 1M
  • Mistral Small 4 (March 16): 119B/6.5B-active MoE unifying reasoning, vision (Pixtral), and coding in one endpoint
  • Voxtral (March 26): open-source 4B text-to-speech, 9 languages, runs on consumer hardware
  • Strong European data sovereignty guarantees across the full model family

MiniMax M3

MiniMax · Shanghai, China

Cost-efficient open-weight coding leader — 1M context and top SWE-bench Pro score per dollar.

  • Released June 1, 2026 — 1M context window; 59.0% SWE-bench Pro; successor to M2.5 (Feb 12) and M2.7 (late April)
  • $0.15/1M input, $1.15/1M output — among the cheapest frontier-class coding models
  • M2.5 base: 230B MoE / 10B active, 80.2% SWE-Bench Verified; M3 extends context and benchmark lead

Kimi K3

Moonshot AI · Beijing, China

Open-weight frontier model that closes the gap with US labs — 2.8T MoE, second only to Fable 5 and GPT-5.6 on most benchmarks.

  • Released July 16, 2026 (full weights July 27) — 2.8T parameter MoE, successor to K2.7 Code
  • Outperforms every other model except Claude Fable 5 and GPT-5.6 on most benchmarks
  • K2.7 Code (June 12, 2026): still available — 1T MoE coding variant of K2.6, 30% fewer thinking tokens
  • Open-weight, Modified MIT license; weights on Hugging Face
  • Kimi Code CLI agent rivals Claude Code and Gemini CLI

GLM-5.3

Zhipu AI · Beijing, China

Frontier-class model on a MIT license — now with a fast multimodal variant.

  • GLM-5.3 (Aug 14, 2026): scaling-post-training upgrade over GLM-5.2 — same ~744B parameter MoE (40B active), 1M context extended to 128K output; text-only, tuned for code, agents, and cyber-defense
  • GLM-5.3-Flash (Aug 26, 2026): first native multimodal model in the GLM-5 series — 320B/18B active MoE, hybrid sparse+linear attention, 1M context, $0.15/$0.50 per 1M (~1/10th GLM-5.3's price)
  • Released under MIT license; trained entirely on Huawei Ascend chips (zero NVIDIA GPUs)
  • Priced roughly 6x cheaper than comparable proprietary models

Atria Dawn Preview

Shanghai AI Lab · Shanghai, China

Open-weight agentic model for long-horizon research, built on a GLM-5.2 base — appeared on Hugging Face before its own technical report.

  • Released Sept 11, 2026, with no announcement; the 140-author technical report followed three days later
  • 744B-parameter MoE built on the GLM-5.2 foundation; MIT license, open weights on Hugging Face
  • Trained via a "Verifiable Experience Pipeline" — connects tool-mediated interactions to executable environments and externally verified outcomes
  • Vendor claims top score on 5 of 16 benchmarks, incl. 59.6% SWE-bench Pro (vs. 74.7% for Fable 5.1) — none independently verified by Artificial Analysis yet
  • Built for carrying scientific work from a published method to executable experiments and a reproducible report

NVIDIA Nemotron 3

NVIDIA · Santa Clara

NVIDIA's open agentic reasoning stack — Nano, Super, and Ultra 550B sizes on Bedrock.

  • Nemotron 3 Ultra 550B (June 4, 2026): 550B/55B active MoE, hybrid Mamba-Transformer, 1M context; 48/100 AA Intelligence Index — strongest open US model benchmarked
  • Nemotron 3 Nano Omni: 30B MoE omni-modal model unifying vision, audio, and text
  • Nano and Super sizes released March 2026; Super peers with Llama 4 Maverick on benchmarks
  • Weights, data, and recipes public on Hugging Face; available via Amazon Bedrock and NVIDIA NIM

Gemma 4

Google · Mountain View

Open-weight models from Gemini 3 research — optimized for on-device and frontier-class performance.

  • Four Apache 2.0 models: E2B (2.3B), E4B (4.5B), 26B MoE (4B active), 31B dense (Apr 2, 2026)
  • 31B ranks #3 on Arena AI leaderboard at 1452 Elo — outperforms models 20× its size
  • E2B/E4B optimized for on-device Android: up to 4× faster and 60% less battery than prior Gemma
  • All models natively multimodal; larger variants support 256K context

Cohere Command A+

Cohere · Toronto

Cohere's first fully open-licensed frontier model — enterprise-tuned MoE for retrieval-augmented generation.

  • Command A+ (May 20, 2026): 218B/25B active MoE, Apache 2.0, runs on two H100s — supersedes Command A as Cohere's flagship
  • Purpose-built for RAG: strong grounding, citation accuracy, and document comprehension
  • Cost-efficient enterprise tier; competitive on retrieval benchmarks against larger models
  • Available via Cohere API and on major cloud marketplaces

Qwen 3.8

Alibaba Cloud · Hangzhou, China

The multilingual giant — Qwen3.8-Max is now the flagship, succeeding Qwen 3.7 Max; open-weight Qwen3.5 remains available.

  • Qwen3.8-Max (GA Aug 2, 2026): flagship succeeding Qwen 3.7 Max — 2.4T/95B active MoE; hosted API is multimodal, 1M context, $2/$6 per 1M; open weights (Aug 12) under a restricted revenue-share license, text-only, 262K context
  • Qwen3.8-27B (Aug 13–14, 2026): separate, genuinely self-hostable companion — 27B dense, multimodal with vision encoder, 262K context, Apache 2.0
  • Qwen3.5 (Feb 2026): open-weight, 397B params, available for self-hosting; Qwen3.5-Omni adds native audio/video/text (256K context, 113 languages)
  • Qwen3-Coder: 69.6% SWE-Bench Verified; Qwen3 Coder Next is the follow-up

SubQ

Subquadratic · San Francisco

The first commercial subquadratic LLM — a clean break from quadratic-attention transformers.

  • Released May 5, 2026 — first production model built on Subquadratic Sparse Attention (SSA), scaling ~linearly with context length
  • 12M-token production context window
  • Matches Claude Opus-class performance on coding benchmarks at roughly 1/5th the compute cost, with attention up to 52× faster at scale
  • Architecturally distinct from every other model on this page — the most significant new attention mechanism shipped in 2026
  • Backed by $29M seed round (May 2026); positioning as infrastructure for very-long-context agents

Midjourney v7

Midjourney

Artistic, stylized visuals with strong aesthetic control

Midjourney V1 Video

Midjourney

First video model from MJ — 5s clips extendable to 20s; ~25× cheaper than competitors

Imagen 4

Google

Photorealistic composition, spelling, and typography accuracy

Nano Banana 2

Google

Fast AI image editing, remixing, and style transfers; built on Gemini Flash

DALL-E 4

OpenAI

Integrated with ChatGPT; strong prompt adherence

Stable Diffusion 3.5

Stability AI

Open-source; self-hostable; highly customizable

FLUX 3

Black Forest Labs

Unified image/video/audio/robotics model (Jul 23); video in Early Access now, image + open-weight Dev + robotics variant rolling out in stages, no public pricing yet

LTX-2.3

Lightricks

Open-weights video+audio in one pass; 22B params; 4K at 50 FPS, up to 20s; one of the most capable open video models available

Grok Imagine Video 1.5

xAI

~2× faster video generation than prior version, native audio output (Jun 2026)

Reve 2.0

Reve

Image generation with precise layout control and typography — strong alternative to DALL-E 4 / Imagen 4 for design-heavy use cases (Jun 3, 2026)

Gemini 3.1 Flash Lite Image

Google

Specialized fast image generation — Flash variant tuned for high-volume, low-latency image tasks (Jun 23, 2026)

Video generation leaders: Google Veo 3.1 (native 4K + vertical video), Kling 3.0 (native 4K/60fps), Runway Gen-4.5 (creative/cinematic), Runway Gen-4 Turbo (July 2026 — ~60% faster inference at comparable quality to Gen-4), and Seedance 2.0 (ByteDance — notable for Identity Lock, which maintains consistent faces across multi-scene video). Sora 2 (OpenAI) was deprecated April 26, 2026 with an API sunset on September 24, 2026 — do not build new integrations on it.

There is no single "best" model in 2026. The landscape has shifted from a winner-take-all race to specialized excellence. Match the model to the task.

Sourced directly from company websites and documentation. Updated weekly.

Developer Writing Assistant

ESC