Permanent board

AI Wars

Who’s ahead in foundation models and coding agents, tracked across several signals that often diverge. Desk ranks for US and international labs sit above Arena preference, OpenRouter volume, Reddit heat for coding tools, lab GitHub stars, and longer-horizon Polymarket odds.

Scroll the charts for history. Open a lab card for the full analysis, or jump to the write-ups collected at the bottom of the page.

United States

Updated

Frontier labs + the coding products riding them.

Positioning over time0 to 100
14365779100Jul 26

Hover for values · click legend to hide a series

  1. Anthropic

    Read analysis

    Owns the preference board and coding-agent conversation. Claude family leads Arena text; Claude Code is the loudest subreddit heat signal.

    Positioning96
    Heat96
  2. OpenAI

    Read analysis

    Still the default frontier brand. Codex community heat is real; model preference is contested, but distribution and brand remain unmatched.

    Positioning84
    Heat82
  3. SpaceX

    Read analysis

    Cursor is the SpaceX AI coding bet after the Anysphere deal — IDE distribution plus Colossus compute, aimed straight at Claude Code and Codex.

    Positioning71
    Heat83
  4. Google

    Read analysis

    Gemini stays in the Arena top tier with massive vote volume. Product surface is everywhere; research velocity is the question, not reach.

    Positioning56
    Heat74
  5. Meta

    Read analysis

    Muse Spark climbing preference tables; Llama remains the open-weight line, but the frontier push is closed. Less API-share theater, more ecosystem influence.

    Positioning28
    Heat64
  6. Microsoft

    Read analysis

    Copilot is embedded in work software most people already pay for. Heat is quieter than startups; positioning is distribution.

    Positioning38
    Heat52

International

Updated

Usage and open-weight pressure from outside the US.

Positioning over time0 to 100
8315477100Jul 26

Hover for values · click legend to hide a series

  1. DeepSeek

    Read analysis

    The OpenRouter volume story. Flash and Pro variants have owned daily token share for weeks — usage leadership outside US labs.

    Positioning92
    Heat95
  2. Tencent

    Read analysis

    HY3 spikes keep showing up in routed traffic. A platform company that can turn model capacity into consumer distribution overnight.

    Positioning73
    Heat81
  3. MiniMax

    Read analysis

    M3 stays in the global token mix. Competitive on cost and throughput; still building a Western brand story.

    Positioning52
    Heat71
  4. Xiaomi

    Read analysis

    MiMo has punched above expectations on OpenRouter. Hardware + software stack gives it a lane most pure labs do not have.

    Positioning44
    Heat70
  5. Moonshot

    Read analysis

    Kimi keeps a seat in Arena and OpenRouter tops. Long-context reputation; heat comes in waves with each release.

    Positioning33
    Heat63
  6. Z.ai

    Read analysis

    GLM family is a consistent presence in routed volume. Strong domestic footprint; international recognition still catching up.

    Positioning22
    Heat61

Share of weekly visitors

Proportional to latest snapshot · Jul 1, 2026

1.6M total
  • Claude Code890K55.9%
  • Codex440K27.6%
  • Cursor104K6.5%
  • GitHub Copilot92K5.8%
  • OpenCode48K3.0%
  • Windsurf9.7K0.6%
  • Cline8K0.5%

Coding-agent weekly visitors

Subreddit weekly visitors

Jul 1, 2026
0236K472K708K943KJul 1

Hover for values · click legend to hide a series

Desk-tracked estimates. Heat signal only, not seats or revenue.

Prediction markets

Longer-horizon AI odds from Polymarket. Skips markets resolving within ~10 days · as of Jul 26, 2026.

polymarket.com
Resolves Jul 1, 2027
Anthropic IPO by __?
  1. December 31, 202671.0%
  2. October 31, 202641.5%
  3. September 30, 20266.5%
  4. September 15, 20261.1%

OpenRouter volume

Daily tokens · Apr 27–Jul 25

Jul 26, 2026
0271B543B814B1.09TAprMayMayJunJunJulJulJul

Hover for values · click legend to hide a series

OpenRouter rankings-daily. Prompt + completion tokens for the public top 50 each day.

Provider share

% of OpenRouter tokens · Apr 27–Jul 25

Jul 26, 2026
0.0%7.0%14%21%28%AprMayMayJunJunJulJulJul

Hover for values · click legend to hide a series

Share of daily OpenRouter token volume by provider.

Arena Elo

Weekly snapshots · current top models

Jul 25, 2026
14841492150015081516May 9May 23Jun 6Jun 20Jul 11Jul 25

Hover for values · click legend to hide a series

Weekly Arena text-leaderboard snapshots.

Arena Elo

Human preference · text

  1. claude-fable-51507
    Anthropic · 15K votes
  2. claude-opus-4-6-thinking1505
    Anthropic · 63K votes
  3. claude-opus-4-7-thinking1502
    Anthropic · 51K votes
  4. claude-opus-4-61498
    Anthropic · 67K votes
  5. muse-spark-1.11495
    Meta · 7.9K votes
  6. claude-opus-4-71494
    Anthropic · 52K votes
  7. muse-spark1488
    Meta · 14K votes
  8. gemini-3.1-pro1486
    Google · 85K votes

OpenRouter

API tokens · Jul 25

  1. mimo-v2.51.45T
    xiaomi
  2. deepseek-v4-flash944B
    deepseek
  3. hy3590B
    tencent
  4. nemotron-3-ultra-550b-a55b424B
    nvidia
  5. deepseek-v4-pro414B
    deepseek
  6. glm-5.2317B
    z-ai
  7. minimax-m3262B
    minimax
  8. step-3.7-flash205B
    stepfun

Research on Perplexity

Board-shaped queries with cited answers. Opens on Perplexity.

perplexity.ai

United States

#1 · San Francisco · positioning 96 · heat 96 · closed weight

Anthropic

Owns the preference board and coding-agent conversation. Claude family leads Arena text; Claude Code is the loudest subreddit heat signal.

Signals: Arena text — Claude family holds #1–#4 band (claude-fable-5 ~1507 Elo; opus-thinking variants immediately behind) with large vote bases (tens of thousands). Coding agents — Claude Code ~890K subreddit weekly visitors, ~56% of the desk’s seven-tool weekly-visitor pool (Codex ~440K, Cursor ~104K, Copilot ~92K, OpenCode ~48K). Weights: closed. Stack: preference leader + coding-agent weekly visitors leader in one company — the only US lab with that double.

Positioning 96: frontier preference + developer agent product + enterprise “safe lab” brand + closed-weight lock-in. Weak relative signal: OpenRouter daily tokens often dominated by DeepSeek/Tencent/Xiaomi/MiniMax, not Claude — Anthropic wins quality/agent boards more than cheap routed volume. Concentration risk: Arena + Claude Code move together if either cools.

Heat 96: highest US heat. Drivers — continuous Arena occupancy, Claude Code discourse dominance, every peer forced to answer coding-agent releases. Not driven by OpenRouter share. Net: sets US tempo on preference + agents; does not set global token-price tempo.

#2 · San Francisco · positioning 84 · heat 82 · closed weight

OpenAI

Still the default frontier brand. Codex community heat is real; model preference is contested, but distribution and brand remain unmatched.

Signals: ChatGPT still default consumer/dev surface; Microsoft/Azure distribution intact. Coding — Codex ~440K weekly visitors (#2 on desk board, ~half Claude Code). Arena text — no longer monopoly; Anthropic leads, Gemini/Meta pressing. OpenRouter — not the volume story (DeepSeek/HY3/MiMo/M3 dominate). Weights: closed.

Positioning 84: brand + ChatGPT install + Codex product + partner distribution outweigh missing Arena #1 and missing OpenRouter #1 — still a clear tier below Anthropic’s double (preference + agent weekly visitors). Heat 82: Codex keeps coding narrative hot; preference discourse no longer auto-centers OpenAI on every release.

Cross-pressure: reclaim coding mindshare from Claude Code and SpaceX/Cursor; defend default-frontier status while Chinese labs own cheap tokens. Score logic: #2 US franchise; gap to #1 is structural (Arena ridge + ~2× coding weekly visitors), not a tweak.

#3 · Hawthorne · positioning 71 · heat 83 · closed weight

SpaceX

Cursor is the SpaceX AI coding bet after the Anysphere deal — IDE distribution plus Colossus compute, aimed straight at Claude Code and Codex.

Signals: Corporate — xAI merge + Anysphere/Cursor ~$60B all-stock path; Cursor → SpaceX AI coding wedge. Coding weekly visitors — Cursor ~104K (#3 desk board), behind Claude Code 890K / Codex 440K, ahead of Copilot 92K. Compute narrative — Colossus / owned training capacity. Arena/OpenRouter — SpaceX not a token or Elo leader under its own model names yet. Weights: closed.

Positioning 71: IDE distribution to pro engineers + public-company capital + owned supercompute + Grok-adjacent surface — rare vertical, still below OpenAI’s installed franchise and well below Anthropic’s preference+agent stack. Execution risk: mega-acquisition can blunt product taste.

Heat 83: deal + Cursor culture + Composer/Grok Build shipping talk. Heat > Cursor’s weekly-visitor share because narrative amp is corporate, not just subreddit size. Competitive targets explicit: Claude Code + Codex.

#4 · Mountain View · positioning 56 · heat 74 · closed weight

Google

Gemini stays in the Arena top tier with massive vote volume. Product surface is everywhere; research velocity is the question, not reach.

Signals: Arena — Gemini 3.x-class models recur in top-10 with very large vote counts (often larger sample than flashy new entries). Distribution — Search, Android, Workspace, Cloud. Coding-agent weekly visitors board — no Google-native row in the desk’s seven-tool set. OpenRouter — not a Google volume story vs DeepSeek/Tencent. Weights: primary Gemini closed (Gemma does not flip the dot).

Positioning 56: distribution + Arena presence + cloud is a real floor, not a frontier-war lead. Missing coding-agent seat and OpenRouter volume keep Google a full tier under OpenAI/SpaceX. Heat 74: durable, low-drama — infra/product updates > culture-war launch cycles.

Score logic: Google can lose most exciting lab and still hold mid-50s positioning. Does not need OpenRouter #1; needs Gemini preference sticky and an agent product that shows up on the weekly-visitor board before it can re-enter the 70s.

#5 · Menlo Park · positioning 28 · heat 64 · closed weight

Meta

Muse Spark climbing preference tables; Llama remains the open-weight line, but the frontier push is closed. Less API-share theater, more ecosystem influence.

Signals: Arena — Muse Spark / Muse Spark 1.1 in top preference band (Meta vendor, closed/proprietary primary). Llama — separate open-weight track (ecosystem fine-tunes/hosts); does not make primary frontier open. OpenRouter — Meta rarely owns daily token podium vs DeepSeek/HY3/MiMo. Coding weekly visitors board — no Meta agent row. Dot: closed (Muse Spark primary).

Positioning 28: research + social distribution + Llama ecosystem gravity, but weakest US seat on this board — no coding-agent weekly visitors, weak OpenRouter, frontier monetization secondary to ads. Heat 64: spikes on Muse Spark / Llama drops, then cedes timeline to Claude Code, Codex, Cursor, DeepSeek.

Score logic: open Llama ≠ open primary. Preference signal is Muse Spark (closed). Ecosystem influence ≠ board positioning; Meta can move charts on release week and still sit near the floor until agents or API share show up.

#6 · Redmond · positioning 38 · heat 52 · closed weight

Microsoft

Copilot is embedded in work software most people already pay for. Heat is quieter than startups; positioning is distribution.

Signals: Distribution — Copilot in M365, Windows, GitHub, Azure. Coding weekly visitors — GitHub Copilot ~92K (#4), behind Claude Code / Codex / Cursor. Arena — not Microsoft-branded Elo leadership. OpenRouter — not MS volume story. Weights: primary closed (Phi secondary; dot stays red). Partnership — OpenAI still core to many Copilot paths.

Positioning 38: procurement + seat licenses + Azure is real money, not frontier leadership. ~60 pts behind Anthropic — not a few product tweaks. Heat 52: lowest US heat — weak timeline presence; Copilot weekly visitors real but not discourse-leading.

Score logic: mid-low positioning / low heat on purpose. Microsoft monetizes already installed while builders attention, Arena, and agent weekly visitors live elsewhere. Climbing into the 70s needs preference or agent possession, not another Copilot surface.

International

#1 · Hangzhou · positioning 92 · heat 95 · open weight

DeepSeek

The OpenRouter volume story. Flash and Pro variants have owned daily token share for weeks — usage leadership outside US labs.

Signals: OpenRouter rankings-daily (~90d window) — DeepSeek Flash/Pro lines repeatedly #1 or top-tier on daily tokens; multi-week #1 occupancy in desk samples. Open weight: yes (primary). Arena — present but not the main DeepSeek story vs volume. Coding weekly visitors board — no DeepSeek-native IDE row. US labs win preference/agents; DeepSeek wins routed usage.

Positioning 92: cost/speed leadership + open weights + global builder default for cheap intelligence. Soft spots: Western enterprise trust, consumer brand, regulatory comfort vs Anthropic/OpenAI. Clear international #1 — not a photo-finish with Tencent.

Heat 95: max international heat. OpenRouter moves on ship days; price/speed pressure hits everyone. #1 international on positioning+heat because tokens actually called is the clearest non-US primary metric on this page.

#2 · Shenzhen · positioning 73 · heat 81 · open weight

Tencent

HY3 spikes keep showing up in routed traffic. A platform company that can turn model capacity into consumer distribution overnight.

Signals: OpenRouter — HY3 / HY3-preview / free variants recur in top daily token ranks; multi-day #1 streaks in desk window alongside DeepSeek. Open weight: yes. Distribution — WeChat, games, cloud, ads (domestic). Arena — not the HY3 headline vs volume. Coding weekly visitors — no Tencent row on desk seven-tool board.

Positioning 73: platform install base + ability to manufacture usage via pricing/free tiers + cloud. ~20 pts behind DeepSeek — strong #2, not co-leader. Western brand thinner than domestic machine.

Heat 81: bursty — high when HY3 owns/co-owns OpenRouter podium, quieter in English discourse between spikes. Score pair = heavyweight volume + distribution, not continuous Arena/agent possession.

#3 · Shanghai · positioning 52 · heat 71 · open weight

MiniMax

M3 stays in the global token mix. Competitive on cost and throughput; still building a Western brand story.

Signals: OpenRouter — M3 repeatedly upper-rank / podium-adjacent across days (consistency > single viral peak). Open weight: yes. Arena/coding weekly visitors — secondary. Consumer myth in West — weak vs ChatGPT/Claude/DeepSeek name recognition.

Positioning 52: API utility + cost/throughput for agent stacks; incomplete Western brand. Mid-pack international — well below Tencent’s platform seat, above Xiaomi’s still-forming model identity. Heat 71: token-chart persistence compounds builder mindshare.

Score logic: keep volume → climb; need sharper product narrative for positioning to catch heat. Currently heat-led international mid-tier.

#4 · Beijing · positioning 44 · heat 70 · open weight

Xiaomi

MiMo has punched above expectations on OpenRouter. Hardware + software stack gives it a lane most pure labs do not have.

Signals: OpenRouter — MiMo variants in top token tier over desk window (including multi-day leadership samples). Open weight: yes. Hardware — phones/IoT channel for eventual edge. Arena/coding weekly visitors — not Xiaomi’s primary board signals here.

Positioning 44: usage receipt real; Western AI brand still forming; OpenRouter ≠ full frontier franchise. Hardware loop is upside not yet fully scored. Tier below MiniMax consistency.

Heat 70: > positioning — Xiaomi on the token podium is a high-surprise signal. Sustain vs one-off spike is the watch item for both meters.

#5 · Beijing · positioning 33 · heat 63 · open weight

Moonshot

Kimi keeps a seat in Arena and OpenRouter tops. Long-context reputation; heat comes in waves with each release.

Signals: Product ID — Kimi / long-context (named). Arena — intermittent top-band presence. OpenRouter — seats, rarely multi-week #1 like DeepSeek. Open weight: yes. Coding weekly visitors — none on desk board.

Positioning 33: real-lab tier via Arena + API presence + brand clarity, but far from volume kings. Heat 63: release-wave pattern — spikes then settles. Long-context moat eroded (table stakes in 2026).

Watch: durable usage wedge or coding/agent beachhead required to climb. Intact franchise, not agenda-setter — gap in score terms to DeepSeek/Tencent.

#6 · Beijing · positioning 22 · heat 61 · open weight

Z.ai

GLM family is a consistent presence in routed volume. Strong domestic footprint; international recognition still catching up.

Signals: OpenRouter — GLM family recurring in broader top ranks (steady calls, not usually the #1 plot). Open weight: yes. Brand string — Zhipu / GLM / Z.ai fragmented internationally. Domestic China base — strong. Arena leadership / coding weekly visitors — not primary.

Positioning 22: lowest international set — consistency without owning preference, tokens, or agents narratives. Heat 61: API-visible, not agenda-setting.

Score logic: respected utility lab. Breakout multimodal/agent release that travels could jump both meters; until then, correctly floor-tier on this board’s positioning scale.

What is AI Wars?
AI Wars is Sonar Mag’s permanent scoreboard for the AI industry. It combines desk assessments of major labs with live signals from LMSYS Chatbot Arena, OpenRouter volume, coding-agent Reddit communities, open-source GitHub stars, and Polymarket prediction markets.
How are the lab rankings decided?
US and international labs are scored by the desk on positioning (strategic seat) and heat (near-term momentum), each from 0 to 100. Rank within each region is derived from the sum of those scores, and each company has a full analysis below the board.
How often does AI Wars update?
Arena, OpenRouter, Reddit visitor estimates, GitHub stars, and Polymarket odds refresh on page load with short CDN caching. Positioning and heat are a weekly desk pass: every company on the board gets its week’s news read carefully (product drops, funding, distribution deals, usage slips, open-weight releases), then scores, blurbs, and full analyses are updated where the footing actually moved. Each regional column shows when that side was last reviewed.

← Back to homepage