- Anthropic88.5%
- Google7.5%
- OpenAI3.1%
- SpaceXAI0.6%
Permanent board
AI Wars
Who’s ahead in foundation models and coding agents, tracked across several signals that often diverge. Desk ranks for US and international labs sit above Arena preference, OpenRouter volume, Reddit heat for coding tools, lab GitHub stars, and longer-horizon Polymarket odds.
Scroll the charts for history. Open a lab card for the full analysis, or jump to the write-ups collected at the bottom of the page.
Current state
Desk ranking of who holds the field right now, split by US and international labs. Order is positioning + heat (0 to 100). Green dot means primary models are open-weight / self-hostable; red means primary models are closed (side experiments don’t count). Click a company for the full analysis, also listed below under Desk analyses.
United States
UpdatedFrontier labs + the coding products riding them.
Hover for values · click legend to hide a series
Anthropic
Read analysisOwns the preference board and coding-agent conversation. Claude family leads Arena text; Claude Code is the loudest subreddit heat signal.
Positioning96Heat96OpenAI
Read analysisStill the default frontier brand. Codex community heat is real; model preference is contested, but distribution and brand remain unmatched.
Positioning84Heat82SpaceX
Read analysisCursor is the SpaceX AI coding bet after the Anysphere deal — IDE distribution plus Colossus compute, aimed straight at Claude Code and Codex.
Positioning71Heat83Google
Read analysisGemini stays in the Arena top tier with massive vote volume. Product surface is everywhere; research velocity is the question, not reach.
Positioning56Heat74Meta
Read analysisMuse Spark climbing preference tables; Llama remains the open-weight line, but the frontier push is closed. Less API-share theater, more ecosystem influence.
Positioning28Heat64Microsoft
Read analysisCopilot is embedded in work software most people already pay for. Heat is quieter than startups; positioning is distribution.
Positioning38Heat52
International
UpdatedUsage and open-weight pressure from outside the US.
Hover for values · click legend to hide a series
DeepSeek
Read analysisThe OpenRouter volume story. Flash and Pro variants have owned daily token share for weeks — usage leadership outside US labs.
Positioning92Heat95Tencent
Read analysisHY3 spikes keep showing up in routed traffic. A platform company that can turn model capacity into consumer distribution overnight.
Positioning73Heat81MiniMax
Read analysisM3 stays in the global token mix. Competitive on cost and throughput; still building a Western brand story.
Positioning52Heat71Xiaomi
Read analysisMiMo has punched above expectations on OpenRouter. Hardware + software stack gives it a lane most pure labs do not have.
Positioning44Heat70Moonshot
Read analysisKimi keeps a seat in Arena and OpenRouter tops. Long-context reputation; heat comes in waves with each release.
Positioning33Heat63Z.ai
Read analysisGLM family is a consistent presence in routed volume. Strong domestic footprint; international recognition still catching up.
Positioning22Heat61
Coding agents
Subreddit weekly visitors as a community-heat proxy, tracked from July 2026.
Share of weekly visitors
Proportional to latest snapshot · Jul 1, 2026
Coding-agent weekly visitors
Subreddit weekly visitors
Hover for values · click legend to hide a series
Desk-tracked estimates. Heat signal only, not seats or revenue.
Lab repos
Flagship GitHub projects from the tracked labs · Jul 26, 11:19 PM
Prediction markets
Longer-horizon AI odds from Polymarket. Skips markets resolving within ~10 days · as of Jul 26, 2026.
- Anthropic69.5%
- OpenAI16.5%
- Baidu7.2%
- Alibaba4.1%
- MiniMax4.1%
- 1470+85.5%
- 1480+71.5%
- 1490+50.5%
- 1500+29.0%
- 1510+12.2%
- December 31, 202671.0%
- October 31, 202641.5%
- September 30, 20266.5%
- September 15, 20261.1%
- 151053.5%
- 152016.5%
- 153010.0%
- 15408.5%
- 15506.5%
- 156046.0%
- 158035.5%
- 160013.9%
OpenRouter volume
Daily tokens · Apr 27–Jul 25
Hover for values · click legend to hide a series
OpenRouter rankings-daily. Prompt + completion tokens for the public top 50 each day.
Provider share
% of OpenRouter tokens · Apr 27–Jul 25
Hover for values · click legend to hide a series
Share of daily OpenRouter token volume by provider.
Arena Elo
Weekly snapshots · current top models
Hover for values · click legend to hide a series
Weekly Arena text-leaderboard snapshots.
Live boards
Preference and API volume measure different races.
Arena Elo
Human preference · text
- claude-fable-51507
- claude-opus-4-6-thinking1505
- claude-opus-4-7-thinking1502
- claude-opus-4-61498
- muse-spark-1.11495
- claude-opus-4-71494
- muse-spark1488
- gemini-3.1-pro1486
OpenRouter
API tokens · Jul 25
- mimo-v2.51.45T
- deepseek-v4-flash944B
- hy3590B
- nemotron-3-ultra-550b-a55b424B
- deepseek-v4-pro414B
- glm-5.2317B
- minimax-m3262B
- step-3.7-flash205B
Research on Perplexity
Board-shaped queries with cited answers. Opens on Perplexity.
Desk analyses
Full write-ups behind each company score, kept on the page for easy reading. Ranked by positioning + heat within region.
United States
#1 · San Francisco · positioning 96 · heat 96 · closed weight
Anthropic
Owns the preference board and coding-agent conversation. Claude family leads Arena text; Claude Code is the loudest subreddit heat signal.
Signals: Arena text — Claude family holds #1–#4 band (claude-fable-5 ~1507 Elo; opus-thinking variants immediately behind) with large vote bases (tens of thousands). Coding agents — Claude Code ~890K subreddit weekly visitors, ~56% of the desk’s seven-tool weekly-visitor pool (Codex ~440K, Cursor ~104K, Copilot ~92K, OpenCode ~48K). Weights: closed. Stack: preference leader + coding-agent weekly visitors leader in one company — the only US lab with that double.
Positioning 96: frontier preference + developer agent product + enterprise “safe lab” brand + closed-weight lock-in. Weak relative signal: OpenRouter daily tokens often dominated by DeepSeek/Tencent/Xiaomi/MiniMax, not Claude — Anthropic wins quality/agent boards more than cheap routed volume. Concentration risk: Arena + Claude Code move together if either cools.
Heat 96: highest US heat. Drivers — continuous Arena occupancy, Claude Code discourse dominance, every peer forced to answer coding-agent releases. Not driven by OpenRouter share. Net: sets US tempo on preference + agents; does not set global token-price tempo.
#2 · San Francisco · positioning 84 · heat 82 · closed weight
OpenAI
Still the default frontier brand. Codex community heat is real; model preference is contested, but distribution and brand remain unmatched.
Signals: ChatGPT still default consumer/dev surface; Microsoft/Azure distribution intact. Coding — Codex ~440K weekly visitors (#2 on desk board, ~half Claude Code). Arena text — no longer monopoly; Anthropic leads, Gemini/Meta pressing. OpenRouter — not the volume story (DeepSeek/HY3/MiMo/M3 dominate). Weights: closed.
Positioning 84: brand + ChatGPT install + Codex product + partner distribution outweigh missing Arena #1 and missing OpenRouter #1 — still a clear tier below Anthropic’s double (preference + agent weekly visitors). Heat 82: Codex keeps coding narrative hot; preference discourse no longer auto-centers OpenAI on every release.
Cross-pressure: reclaim coding mindshare from Claude Code and SpaceX/Cursor; defend default-frontier status while Chinese labs own cheap tokens. Score logic: #2 US franchise; gap to #1 is structural (Arena ridge + ~2× coding weekly visitors), not a tweak.
#3 · Hawthorne · positioning 71 · heat 83 · closed weight
SpaceX
Cursor is the SpaceX AI coding bet after the Anysphere deal — IDE distribution plus Colossus compute, aimed straight at Claude Code and Codex.
Signals: Corporate — xAI merge + Anysphere/Cursor ~$60B all-stock path; Cursor → SpaceX AI coding wedge. Coding weekly visitors — Cursor ~104K (#3 desk board), behind Claude Code 890K / Codex 440K, ahead of Copilot 92K. Compute narrative — Colossus / owned training capacity. Arena/OpenRouter — SpaceX not a token or Elo leader under its own model names yet. Weights: closed.
Positioning 71: IDE distribution to pro engineers + public-company capital + owned supercompute + Grok-adjacent surface — rare vertical, still below OpenAI’s installed franchise and well below Anthropic’s preference+agent stack. Execution risk: mega-acquisition can blunt product taste.
Heat 83: deal + Cursor culture + Composer/Grok Build shipping talk. Heat > Cursor’s weekly-visitor share because narrative amp is corporate, not just subreddit size. Competitive targets explicit: Claude Code + Codex.
#4 · Mountain View · positioning 56 · heat 74 · closed weight
Gemini stays in the Arena top tier with massive vote volume. Product surface is everywhere; research velocity is the question, not reach.
Signals: Arena — Gemini 3.x-class models recur in top-10 with very large vote counts (often larger sample than flashy new entries). Distribution — Search, Android, Workspace, Cloud. Coding-agent weekly visitors board — no Google-native row in the desk’s seven-tool set. OpenRouter — not a Google volume story vs DeepSeek/Tencent. Weights: primary Gemini closed (Gemma does not flip the dot).
Positioning 56: distribution + Arena presence + cloud is a real floor, not a frontier-war lead. Missing coding-agent seat and OpenRouter volume keep Google a full tier under OpenAI/SpaceX. Heat 74: durable, low-drama — infra/product updates > culture-war launch cycles.
Score logic: Google can lose most exciting lab and still hold mid-50s positioning. Does not need OpenRouter #1; needs Gemini preference sticky and an agent product that shows up on the weekly-visitor board before it can re-enter the 70s.
#5 · Menlo Park · positioning 28 · heat 64 · closed weight
Meta
Muse Spark climbing preference tables; Llama remains the open-weight line, but the frontier push is closed. Less API-share theater, more ecosystem influence.
Signals: Arena — Muse Spark / Muse Spark 1.1 in top preference band (Meta vendor, closed/proprietary primary). Llama — separate open-weight track (ecosystem fine-tunes/hosts); does not make primary frontier open. OpenRouter — Meta rarely owns daily token podium vs DeepSeek/HY3/MiMo. Coding weekly visitors board — no Meta agent row. Dot: closed (Muse Spark primary).
Positioning 28: research + social distribution + Llama ecosystem gravity, but weakest US seat on this board — no coding-agent weekly visitors, weak OpenRouter, frontier monetization secondary to ads. Heat 64: spikes on Muse Spark / Llama drops, then cedes timeline to Claude Code, Codex, Cursor, DeepSeek.
Score logic: open Llama ≠ open primary. Preference signal is Muse Spark (closed). Ecosystem influence ≠ board positioning; Meta can move charts on release week and still sit near the floor until agents or API share show up.
#6 · Redmond · positioning 38 · heat 52 · closed weight
Microsoft
Copilot is embedded in work software most people already pay for. Heat is quieter than startups; positioning is distribution.
Signals: Distribution — Copilot in M365, Windows, GitHub, Azure. Coding weekly visitors — GitHub Copilot ~92K (#4), behind Claude Code / Codex / Cursor. Arena — not Microsoft-branded Elo leadership. OpenRouter — not MS volume story. Weights: primary closed (Phi secondary; dot stays red). Partnership — OpenAI still core to many Copilot paths.
Positioning 38: procurement + seat licenses + Azure is real money, not frontier leadership. ~60 pts behind Anthropic — not a few product tweaks. Heat 52: lowest US heat — weak timeline presence; Copilot weekly visitors real but not discourse-leading.
Score logic: mid-low positioning / low heat on purpose. Microsoft monetizes already installed while builders attention, Arena, and agent weekly visitors live elsewhere. Climbing into the 70s needs preference or agent possession, not another Copilot surface.
International
#1 · Hangzhou · positioning 92 · heat 95 · open weight
DeepSeek
The OpenRouter volume story. Flash and Pro variants have owned daily token share for weeks — usage leadership outside US labs.
Signals: OpenRouter rankings-daily (~90d window) — DeepSeek Flash/Pro lines repeatedly #1 or top-tier on daily tokens; multi-week #1 occupancy in desk samples. Open weight: yes (primary). Arena — present but not the main DeepSeek story vs volume. Coding weekly visitors board — no DeepSeek-native IDE row. US labs win preference/agents; DeepSeek wins routed usage.
Positioning 92: cost/speed leadership + open weights + global builder default for cheap intelligence. Soft spots: Western enterprise trust, consumer brand, regulatory comfort vs Anthropic/OpenAI. Clear international #1 — not a photo-finish with Tencent.
Heat 95: max international heat. OpenRouter moves on ship days; price/speed pressure hits everyone. #1 international on positioning+heat because tokens actually called is the clearest non-US primary metric on this page.
#2 · Shenzhen · positioning 73 · heat 81 · open weight
Tencent
HY3 spikes keep showing up in routed traffic. A platform company that can turn model capacity into consumer distribution overnight.
Signals: OpenRouter — HY3 / HY3-preview / free variants recur in top daily token ranks; multi-day #1 streaks in desk window alongside DeepSeek. Open weight: yes. Distribution — WeChat, games, cloud, ads (domestic). Arena — not the HY3 headline vs volume. Coding weekly visitors — no Tencent row on desk seven-tool board.
Positioning 73: platform install base + ability to manufacture usage via pricing/free tiers + cloud. ~20 pts behind DeepSeek — strong #2, not co-leader. Western brand thinner than domestic machine.
Heat 81: bursty — high when HY3 owns/co-owns OpenRouter podium, quieter in English discourse between spikes. Score pair = heavyweight volume + distribution, not continuous Arena/agent possession.
#3 · Shanghai · positioning 52 · heat 71 · open weight
MiniMax
M3 stays in the global token mix. Competitive on cost and throughput; still building a Western brand story.
Signals: OpenRouter — M3 repeatedly upper-rank / podium-adjacent across days (consistency > single viral peak). Open weight: yes. Arena/coding weekly visitors — secondary. Consumer myth in West — weak vs ChatGPT/Claude/DeepSeek name recognition.
Positioning 52: API utility + cost/throughput for agent stacks; incomplete Western brand. Mid-pack international — well below Tencent’s platform seat, above Xiaomi’s still-forming model identity. Heat 71: token-chart persistence compounds builder mindshare.
Score logic: keep volume → climb; need sharper product narrative for positioning to catch heat. Currently heat-led international mid-tier.
#4 · Beijing · positioning 44 · heat 70 · open weight
Xiaomi
MiMo has punched above expectations on OpenRouter. Hardware + software stack gives it a lane most pure labs do not have.
Signals: OpenRouter — MiMo variants in top token tier over desk window (including multi-day leadership samples). Open weight: yes. Hardware — phones/IoT channel for eventual edge. Arena/coding weekly visitors — not Xiaomi’s primary board signals here.
Positioning 44: usage receipt real; Western AI brand still forming; OpenRouter ≠ full frontier franchise. Hardware loop is upside not yet fully scored. Tier below MiniMax consistency.
Heat 70: > positioning — Xiaomi on the token podium is a high-surprise signal. Sustain vs one-off spike is the watch item for both meters.
#5 · Beijing · positioning 33 · heat 63 · open weight
Moonshot
Kimi keeps a seat in Arena and OpenRouter tops. Long-context reputation; heat comes in waves with each release.
Signals: Product ID — Kimi / long-context (named). Arena — intermittent top-band presence. OpenRouter — seats, rarely multi-week #1 like DeepSeek. Open weight: yes. Coding weekly visitors — none on desk board.
Positioning 33: real-lab tier via Arena + API presence + brand clarity, but far from volume kings. Heat 63: release-wave pattern — spikes then settles. Long-context moat eroded (table stakes in 2026).
Watch: durable usage wedge or coding/agent beachhead required to climb. Intact franchise, not agenda-setter — gap in score terms to DeepSeek/Tencent.
#6 · Beijing · positioning 22 · heat 61 · open weight
Z.ai
GLM family is a consistent presence in routed volume. Strong domestic footprint; international recognition still catching up.
Signals: OpenRouter — GLM family recurring in broader top ranks (steady calls, not usually the #1 plot). Open weight: yes. Brand string — Zhipu / GLM / Z.ai fragmented internationally. Domestic China base — strong. Arena leadership / coding weekly visitors — not primary.
Positioning 22: lowest international set — consistency without owning preference, tokens, or agents narratives. Heat 61: API-visible, not agenda-setting.
Score logic: respected utility lab. Breakout multimodal/agent release that travels could jump both meters; until then, correctly floor-tier on this board’s positioning scale.
About this board
What you’re looking at, and how often it moves.
- What is AI Wars?
- AI Wars is Sonar Mag’s permanent scoreboard for the AI industry. It combines desk assessments of major labs with live signals from LMSYS Chatbot Arena, OpenRouter volume, coding-agent Reddit communities, open-source GitHub stars, and Polymarket prediction markets.
- How are the lab rankings decided?
- US and international labs are scored by the desk on positioning (strategic seat) and heat (near-term momentum), each from 0 to 100. Rank within each region is derived from the sum of those scores, and each company has a full analysis below the board.
- How often does AI Wars update?
- Arena, OpenRouter, Reddit visitor estimates, GitHub stars, and Polymarket odds refresh on page load with short CDN caching. Positioning and heat are a weekly desk pass: every company on the board gets its week’s news read carefully (product drops, funding, distribution deals, usage slips, open-weight releases), then scores, blurbs, and full analyses are updated where the footing actually moved. Each regional column shows when that side was last reviewed.