Models & tools
Every AI vendor, every capability — in one place.
All 11tracked players — Anthropic, OpenAI, Google, Meta, xAI (Grok), Mistral, DeepSeek, Alibaba (Qwen), NVIDIA, plus Oracle (the enterprise host) and MCP (the tool-use standard) — across six capability classes: models, price, context, coding, tooling & security. Scan the breadth, open any vendor for the full detail, follow every figure to its source.
Built and generated by AIU agents — every source linked.
✓ verified vendor/primary or independently corroborated · △ vendor-claimed the source said so / secondary, not independently confirmed · — unpublished / unsourced this pass (never invented) · N/Athe class doesn't apply (host / protocol).
The model landscape — three ways to use a model
flagship · open weight · local · as of 2026-07-22The strongest models the top labs run for you — you call them over an API and pay per use, and you never hold the weights.
Models whose weights you can download to run, fine-tune, or self-host — no single vendor stands between you and the model.
Open-weight models small enough to run on hardware you control — a laptop, a workstation, or your own server — with nothing leaving your machine.
They overlap on purpose. These describe HOW YOU USE a model, not three separate species. An open-weight model can also run locally — but only if it is small enough: Kimi K3's 2.8-trillion parameters are open-weight yet need a server cluster, not a laptop. A flagship can be open-weight too (Kimi K3 is both a frontier flagship and downloadable). The line that matters for you is simple — do you call it over an API (flagship / hosted), hold the weights yourself (open weight), or run it on hardware you control (local)?
Every figure for the models named here is sourced in the vendor matrix below →
The vendors, at a glance
11 vendors · open a card for its full 6-class detail + sourcesAnthropicclosedClaude Fable 5 · Mythos 5 — the Claude 5 family, Anthropic's current flagship generation (2026); Claude Opus 5 (2026-07-24) is the recommended default for agentic coding and enterprise work at half Fable 5's price; Claude Opus 4.8 is the prior generation1MSWE 88.6%Created MCP · nativeSOC 2 · ISO 27001/42001Open the full 6-class detail + sources →
ModelClaude Fable 5 · Opus 5 · Sonnet 5 (Claude 5 family, closed)
Claude Fable 5 is Anthropic's most capable widely released model and, with Claude Mythos 5, forms the current flagship generation (generally available 2026-06-09; Fable 5 was pulled by a June 12, 2026 US export-control order and redeployed worldwide on July 1, 2026). Claude Opus 5 launched 2026-07-24 as the recommended default for agentic coding and enterprise work; Claude Sonnet 5 is the speed/intelligence balance and Claude Haiku 4.5 the fastest tier. Claude Mythos 5 is invitation-only. Claude Opus 4.8 and 4.7 are prior-generation and still available.
API $/1M (in / out)$10 / $50
Per 1M tokens (input / output), from Anthropic's model documentation: Claude Fable 5 (flagship) $10 / $50 · Claude Opus 5 $5 / $25 · Claude Sonnet 5 $3 / $15 (introductory $2 / $10 through 2026-08-31) · Claude Haiku 4.5 $1 / $5. Claude Opus 4.8 and 4.7 (prior generation) are also $5 / $25. Batch API and prompt-caching discounts apply on top.
Context window1M
1M tokens on Claude Fable 5, Opus 5 and Sonnet 5; 200k on Claude Haiku 4.5. Maximum output 128k tokens on the 1M-context models (64k on Haiku 4.5), rising to 300k output on the batch endpoint under a beta header.
SWE-bench Verified88.6%
SWE-bench Verified 88.6% — an Opus 4.8 (prior-generation) figure; Anthropic has not published a SWE-bench Verified score for Opus 5 or Fable 5, so this is NOT a current-generation number. Independent standing for the current generation: Artificial Analysis' Intelligence Index v4.1 ranks Claude Opus 5 first at 61, ahead of Claude Fable 5 at 60.
Tooling / MCPCreated MCP; native
Created MCP (Nov 2024); native; AAIF co-founder/donor.
Security postureSOC 2 I/II · ISO 27001/42001 · ZDR addendum
SOC 2 Type I & II · ISO 27001:2022 · ISO 42001:2023 · Zero-Data-Retention addendum · BAA/HIPAA on API + Enterprise.
OpenAIclosedGPT-5.5 — closed (GPT-5.5 Pro variant; GPT-5.6 emerging Jul 2026)1.05MSWE 88.7% (#1)Native MCP · AAIFSOC 2 · ISO 27001/27701Open the full 6-class detail + sources →
ModelGPT-5.5 (closed)
GPT-5.5, closed (GPT-5.5 Pro variant; GPT-5.6 emerging Jul 2026).
API $/1M (in / out)$5 / $30
$5.00 in / $30.00 out per 1M (cached in $0.50; >272K prompt -> 2x/1.5x). Pro: $30/$180.
Context window1.05M
1,050,000 tokens; 128K max output.
SWE-bench Verified88.7% (#1)
SWE-bench Verified 88.7% (#1).
Tooling / MCPNative MCP (Mar 2025)
Native MCP (adopted Mar 2025; ChatGPT apps Sep 2025); AAIF co-founder.
Security postureSOC 2 Type 2 · ISO 27001/27701 · ZDR
SOC 2 Type 2 · ISO 27001:2022 · ISO 27701:2019 · ZDR on enterprise/API · BAA/HIPAA.
GoogleclosedGemini 3.1 Pro — closed (Gemini 3.6 Flash cost tier)1M-2MSWE 80.6%MCP in GeminiFedRAMP High · ISO 42001Open the full 6-class detail + sources →
ModelGemini 3.1 Pro (closed)
Gemini 3.1 Pro (closed); Gemini 3.6 Flash (cost tier, 2026-07-21) + Gemini 3.5 Flash-Lite (cheapest tier); 3.5 Flash Cyber restricted to governments/trusted partners.
API $/1M (in / out)$2 / $12
$2.00 in / $12.00 out per 1M (>200K -> $4/$18, 3.1 Pro); 3.6 Flash $1.50/$7.50; 3.5 Flash-Lite $0.30/$2.50 (2026-07-21 release).
Context window1M-2M
1M standard; up to 2M (industry-leading, 3.1 Pro).
SWE-bench Verified80.6%
SWE-bench Verified 80.6%.
Tooling / MCPMCP in Gemini (Apr 2025)
MCP support confirmed for Gemini (DeepMind, Apr 2025).
Security postureSOC 1/2/3 · ISO 42001 · FedRAMP High · HITRUST
SOC 1/2/3 · ISO 42001 · HITRUST · FedRAMP High · PCI DSS v4.0; ZDR must be customer-set (default 24h retention); BAA/HIPAA on covered tier.
Metaopen weightsLlama 4 Maverick — open weights (MoE 400B/17B active)1MSWE — · proxy: MMLU-Pro 80.5MCP via clientSelf-host = own boundaryOpen the full 6-class detail + sources →
ModelLlama 4 Maverick (open weights)
Open weights, MoE 400B total / 17B active (128 experts), natively multimodal; Llama 4 Community License.
API $/1M (in / out)~$0.15 / $0.60 (provider)
Provider-set (open weights) — from ~$0.15 in / $0.60 out per 1M (DeepInfra cheapest); Meta blended ref ~$0.19/1M.
Context window1M
1,048,576 (1M).
SWE-bench Verified— · proxy: MMLU-Pro 80.5
SWE-bench Verified ABSENT (no clean published figure this pass). Proxy: MMLU-Pro 80.5, shown as a labelled proxy — NOT a SWE-bench score.
No clean SWE-bench Verified figure this pass; MMLU-Pro 80.5 shown as a labelled proxy, not a SWE-bench number.
Tooling / MCPOpenAI-compat fn-calling; MCP via client
OpenAI-compatible function calling; MCP via client libraries (no native mcp_servers).
Security postureSelf-host = own boundary (open weights)
Open weights -> self-host = your own security boundary (no per-token provider retention when self-hosted). Hosted-provider posture varies by provider.
xAIclosedGrok 4.5 — closed flagship, launched 2026-07-08500KSWE 86.6%Tool-calling; no native MCP yetGDPR risk · EU investigationOpen the full 6-class detail + sources →
ModelGrok 4.5 (closed)
Grok 4.5 — flagship (coding/agentic/knowledge), launched 2026-07-08, closed.
API $/1M (in / out)$2 / $6
$2.00 in / $6.00 out per 1M (cached $0.50; >200K -> $4/$12).
Context window500K
500K tokens.
SWE-bench Verified86.6%
SWE-bench Verified 86.6%.
Tooling / MCPOpenAI-compat tool-calling; no native mcp_servers yet
OpenAI-compatible tool calling; MCP via wrapping client; no native mcp_servers parameter yet.
Security postureEnt. SOC 2/HIPAA advertised — highest GDPR risk; under EU investigation
Enterprise tier advertises SOC 2 Type II · HIPAA · GDPR · CCPA; BUT highest GDPR regulatory risk in 2026 — no enterprise DPA, no EU data-residency, under active investigation (Ireland DPC · France CNIL · UK ICO); no-train-by-default on Business/Enterprise.
Mistralopen weightsMistral Large 3 (2512) — open weights, EU (Paris)256KSWE — · proxy: HumanEval ~92%MCP via clientEU data-residency defaultOpen the full 6-class detail + sources →
ModelMistral Large 3 (2512, open weights)
Mistral Large 3 (pinned mistral-large-2512); open weights.
API $/1M (in / out)$0.50 / $1.50
$0.50 in / $1.50 out per 1M (caching 10% of input on hit; batch -50%).
Context window256K
262,144 (256K).
SWE-bench Verified— · proxy: HumanEval ~92%
SWE-bench Verified ABSENT (no published number this pass). Proxy sourced: HumanEval ~92% pass@1; LiveCodeBench mid-tier. Shown as a labelled proxy, not SWE-bench.
No published SWE-bench Verified figure this pass; HumanEval ~92% pass@1 shown as a labelled proxy.
Tooling / MCPOpenAI-compat fn-calling; MCP via client
OpenAI-compatible function calling; MCP via client.
Security postureEU-based (Paris); EU data-residency default; ZDR API/Ent
EU-based (Paris) — only EU frontier lab; EU data-residency by default (no cross-border transfer); ZDR on API/Enterprise — structural GDPR advantage.
DeepSeekopen weights · MITDeepSeek V4-Pro — open weights (MIT), MoE 1.6T/49B active1MSWE 83.7% (alt 80.6%)OpenAI + Anthropic compatSelf-host clean; hosted = CN infraOpen the full 6-class detail + sources →
ModelDeepSeek V4-Pro (open weights, MIT)
DeepSeek V4 — open weights (MIT), MoE; V4-Pro 1.6T total / 49B active · V4-Flash 284B / 13B; released 2026-04-24.
API $/1M (in / out)$0.435 / $0.87
V4-Pro $0.435 in / $0.87 out; V4-Flash $0.14 / $0.28 per 1M (permanent 75% cut, May 2026).
Context window1M
1M default; 384K max output.
SWE-bench Verified83.7% (alt 80.6%)
SWE-bench Verified 83.7% (alt source 80.6%; range shown with both).
Tooling / MCPOpenAI + Anthropic-format compat
API speaks both OpenAI ChatCompletions and Anthropic formats; MCP via those.
Security postureSelf-host (MIT) clean; hosted API = CN infra, GDPR-hard, no public SOC 2/ISO
Self-host (MIT weights) = clean, no provider relationship. Hosted API = CN infrastructure, GDPR-hard, no public SOC 2/ISO callouts; Korea PIPC found unconsented prompt transfers.
Alibabaclosed (API-only)Qwen3.8-Max — 2.4T MoE, API GA; open weights announced1MSWE 80.4%Qwen3.6 MCP-trainedCN-infra caveat · certs —Open the full 6-class detail + sources →
ModelQwen3.8-Max (API GA; open weights announced)
Qwen3.8-Max — 2.4T-parameter multimodal MoE flagship, broadly available via API and third-party gateways (Aug 2026); open weights announced as shipping within days, alongside an open Qwen3.8-27B. Qwen 3.7 Max (prior flagship) remains closed/API-only.
API $/1M (in / out)$1.25 / $3.75
Qwen 3.7 Max $1.25 in / $3.75 out; Qwen 3.5 Plus $0.30 / $1.80 per 1M.
Context window1M
1M tokens; 65,536 max output.
SWE-bench Verified80.4%
SWE-bench Verified 80.4%.
Tooling / MCPQwen3.6 trained on MCP benchmarks; OpenAI-compat
Qwen3.6 explicitly trained/evaluated on MCP-based agentic benchmarks; OpenAI-compatible.
NVIDIAopen weights · OpenMDW-1.1Nemotron 3 Ultra — open weights (OpenMDW-1.1), Mamba-Transformer 550B/55B262K (BF16) / 1M (NVFP4)SWE 70.7%OpenAI-compatSelf-host · certs —Open the full 6-class detail + sources →
ModelNemotron 3 Ultra (open weights, OpenMDW-1.1)
Nemotron 3 Ultra — open weights (OpenMDW-1.1, Linux Foundation), MoE hybrid Mamba-Transformer 550B total / 55B active; released 2026-06-04.
API $/1M (in / out)Self-host (open weights) · hosted list —
List $/1M ABSENT (open weights; self-host = infra cost, not per-token). Hosted 'free' tier surfaced (OpenRouter); Together AI hosts.
Open-weight self-host has no per-token list price; a hosted list $/1M was not sourced this pass — shown as self-host + hosted list '—'.
Context window262K (BF16) / 1M (NVFP4)
262K (BF16) / 1M (NVFP4 on Blackwell).
SWE-bench Verified70.7%
SWE-bench Verified 70.7% (65-70.4% across agent frameworks).
Tooling / MCPOpenAI-compat fn-calling
OpenAI-compatible function calling; MCP via those.
Security postureSelf-host = own boundary; NVIDIA AI Enterprise; specific certs —
Self-host = own boundary; NVIDIA AI Enterprise / NIM deployment path; specific certs ABSENT this pass — posture shown, certs '—'.
Self-host posture sourced; specific hosted-provider certifications were not sourced this pass (rendered '—').
OraclehostOCI Generative AI — enterprise multi-model HOSTInfrastructure / host, not a model vendor — routed to the Business & Players desk (R4 IA). Its value is which models it serves + OCI security, NOT a model row.Open its columns + sources →
ModelHosts OpenAI · Gemini · Grok · Llama · Cohere · Mistral (OCI)
Hosts OpenAI · Google Gemini · xAI Grok (added Jun 2025; Grok 4.1 Fast Jan 2026) · Meta Llama (3.3, 4 Scout, 4 Maverick) · Cohere · Mistral; 13 vision-capable models via one library.
Tooling / MCPOCI Gen-AI gateway (OpenAI-compatible)
OCI Gen-AI gateway is OpenAI-compatible (the same API across every hosted model).
Security postureOCI enterprise cloud controls
OCI enterprise cloud controls — the row's real value (multi-model access 'without compromising security controls'); enterprise/gov regions incl. US classified cloud.
MCPprotocolModel Context Protocol — the open tool-use STANDARDThe open tool-use protocol, not a model vendor — routed to the Agents & Tools desk (R4 IA). Its value is governance + adoption + auth, NOT model/pricing/context/coding.Open its columns + sources →
ModelThe open tool-use standard (not a model)
Open standard introduced by Anthropic (Nov 2024); adopted by OpenAI (Mar 2025), Google DeepMind (Apr 2025).
Tooling / MCPThe standard itself · AAIF (Linux Foundation) · ~97M SDK dl/mo
Governance: donated Dec 2025 to the Agentic AI Foundation (AAIF), a Linux Foundation fund (Anthropic · Block · OpenAI co-founders). Adoption: ~97M monthly SDK downloads (Mar 2026), up from 100K at launch (~970x); ChatGPT · Cursor · Gemini · Copilot · VS Code.
Security postureOAuth 2.1 resource-server auth; prompt-injection / tool-poisoning surface
OAuth 2.1 resource-server authorization model; known prompt-injection / tool-poisoning attack surface that agents must guard. Specific hardening guidance ABSENT this pass.
Compare every vendor
9 model vendors × 6 classes · scroll the grid sideways →The vendor column stays put as you scroll; open a vendor above for each figure's source. Oracle (host) and MCP (protocol) aren't scored on the model classes — see their cards above.
| Vendor | Flagship | API $/1M | Context | SWE-bench | Tooling / MCP | Security |
|---|---|---|---|---|---|---|
| Anthropicclosed | Claude Fable 5 · Mythos 5 | $10 / $50✓ | 1M✓ | 88.6%✓ | Created MCP · native✓ | SOC 2 · ISO 27001/42001✓ |
| OpenAIclosed | GPT-5.5 | $5 / $30✓ | 1.05M✓ | 88.7% (#1)✓ | Native MCP · AAIF✓ | SOC 2 · ISO 27001/27701✓ |
| Googleclosed | Gemini 3.1 Pro | $2 / $12✓ | 1M-2M✓ | 80.6%✓ | MCP in Gemini✓ | FedRAMP High · ISO 42001✓ |
| Metaopen weights | Llama 4 Maverick | ~$0.15 / $0.60 (provider)✓ | 1M✓ | — · proxy: MMLU-Pro 80.5△ | MCP via client✓ | Self-host = own boundary✓ |
| xAIclosed | Grok 4.5 | $2 / $6✓ | 500K✓ | 86.6%✓ | Tool-calling; no native MCP yet△ | GDPR risk · EU investigation✓ |
| Mistralopen weights | Mistral Large 3 (2512) | $0.50 / $1.50✓ | 256K✓ | — · proxy: HumanEval ~92%△ | MCP via client✓ | EU data-residency default✓ |
| DeepSeekopen weights · MIT | DeepSeek V4-Pro | $0.435 / $0.87✓ | 1M✓ | 83.7% (alt 80.6%)✓ | OpenAI + Anthropic compat✓ | Self-host clean; hosted = CN infra✓ |
| Alibabaclosed (API-only) | Qwen3.8-Max | $1.25 / $3.75✓ | 1M✓ | 80.4%✓ | Qwen3.6 MCP-trained✓ | CN-infra caveat · certs — |
| NVIDIAopen weights · OpenMDW-1.1 | Nemotron 3 Ultra | Self-host (open weights) · hosted list —△ | 262K (BF16) / 1M (NVFP4)✓ | 70.7%✓ | OpenAI-compat✓ | Self-host · certs —△ |
Tools by industry
what practitioners actually use · sources on clickTrades & field serviceServiceTitan Voice Agent · Housecall Pro · JobberPhone-intake, scheduling and dispatch automation is the mature use; AI voice agents that book straight to the dispatch board are the fast-rising edge.2 sources · click to open
Law & professional servicesHarvey · Thomson Reuters CoCounsel · Spellbook · Lexis+ AIFirst-pass research, drafting and contract review are dependable and widely adopted (Harvey reports 100k+ lawyers); automated billing is still early.2 sources · click to open
EducationMagicSchool · Khanmigo · Google Gemini in ClassroomTeacher-side planning, assessment and feedback tools are mainstream (MagicSchool crossed 5M educators); student tutoring with real work-checking is the frontier few do well.3 sources · click to open
HealthcareAbridge · Microsoft Dragon Copilot (ex-Nuance DAX) · Suki · NablaAmbient clinical documentation — the AI scribe — is the breakout mature use (on track for ~30% of the market, saving clinicians hours a day); AI coding and orders are rising.2 sources · click to open
Finance & accountingRamp Accounting Agent · Intuit Assist (QuickBooks) · PuzzleTransaction categorization, expense coding and month-end close automation are mature; AI advisory and audit copilots are rising.2 sources · click to open
Real estatePerspective AI / Ylopo (lead qualification) · Listings AI / Rechat (listing copy) · Follow Up Boss (CRM)AI listing copy and conversational lead qualification are common; AI-enhanced CRMs are projected to reach ~89% of top-producing agents.2 sources · click to open
Marketing & agenciesJasper · HubSpot Breeze · Canva Magic Studio · Salesforce AgentforceContent generation, design and campaign optimization are mainstream — ~88% of marketers now use AI daily; autonomous campaign agents are rising.2 sources · click to open
Retail & e-commerceShopify Sidekick & Magic · Klaviyo AI · Amazon listing toolsAI product content, email and recommendations are mature (Shopify now bundles Sidekick free on every plan); agentic store-ops that build automations from a sentence are rising.2 sources · click to open
Manufacturing & logisticsSiemens Industrial Copilot · Blue Yonder · o9 Solutions · Palantir FoundryAI demand forecasting and predictive maintenance are established (15–25% forecast-accuracy gains; ~25% less reactive-maintenance time in Siemens pilots); agentic plant-ops are early.2 sources · click to open
HospitalityIDeaS · Duetto · Atomize (revenue management) · HiJiffy · Canary (guest chat)AI-driven dynamic pricing and revenue management are established (8–15% RevPAR lift reported); AI concierge and guest-messaging agents are rising fast.2 sources · click to open
What's rising
agents · skills · platforms · backup on click1Long-running coding agentsmulti-hour autonomous builds↑ rising
The everyday version: an agent that keeps working through a multi-step build — writing, testing and fixing across many files — instead of one prompt-and-reply at a time.
2Sub-agent orchestrationfan-out review patterns↑ rising
One lead agent splits a job across many specialist sub-agents running in parallel, then checks their work before folding it back — the pattern behind the fastest coding tools.
3MCP tool serversinterop standard settlingnew
MCP is becoming the USB-C of AI tools — one open standard that lets any assistant plug into your apps and data; 500+ public servers already exist.
1Prompt cachingcost lever most teams miss— steady
Reuse a big fixed prompt — instructions, a codebase, a long document — across calls and pay ~10% for the repeated part: up to ~90% cheaper and ~85% faster on long prompts.
2Computer useagents driving real screens↑ rising
The agent sees the screen and moves the mouse and keyboard like a person — filling forms and clicking through apps that have no API. Claude's OSWorld score tripled to 44% in 2026.
3Voice-first agentsfield-work entry pointnew
Talk to the agent and it talks back in real time — answering calls, booking jobs, taking intake. It's how trades and clinics are meeting AI first.
AI Generated - Every source linked The model details, comparisons, and local-model specs are compiled by AI from the cited sources — figures can go stale or be misread. Verify any number against its linked source before you rely on it.