Same benchmark, separate publisher runs

Model rankings, with receipts

Pick a benchmark. Compare publisher-reported scores, open every source, and see missing results as missing.

18 models18 routable6 boardsRevised 2026-09-02

Coding

Artificial Analysis Coding Index · Higher is better

Price

USD per 1M tokens, blended 3:1 input to output · Lower is better

Speed

Output tokens per second · Higher is better

Latency

Seconds to first answer token, including reasoning · Lower is better

Source: Artificial Analysis

10 of 18 models report this result

One benchmark · publisher figures · 18 routable now

Score spread

Publisher-reported DeepSWE v1.1, on one zero-to-100 scale.

  1. Rank 1
    GPT-5.6 Sol

    OpenAI

    72.7%

    DeepSWE v1.1

    RoutableOpenAI
  2. Rank 2
    GLM-5.3

    Z.ai

    66.9%

    DeepSWE v1.1

    RoutableZ.ai
  3. Rank 3
    Grok 4.6

    xAI

    65.9%

    DeepSWE v1.1

    RoutablexAI
DeepSWE price performance

Directional only: publisher harnesses and effort settings may differ.

  1. #1533.9
    GLM-5.3-Flash

    Launch discount

    $0.12 / 1M · 63.4

  2. #2266.9
    GLM-5.3-Flash

    List

    $0.24 / 1M · 63.4

  3. #387.4
    MiniMax M3

    Standard

    $0.70 / 1M · 61.2%

  4. #478.9
    GPT-5.4 mini

    Standard

    $0.79 / 1M · 62.1%

  5. #560.8
    Kimi K3

    Standard

    $1.05 / 1M · 63.8%

  6. #643.5
    Gemini 3.7 Flash

    Introductory

    $1.50 / 1M · 65.3%

  7. #722.9
    Qwen3.8-2.4T-A95B

    Model Studio API

    $2.48 / 1M · 56.6

  8. #822.0
    Grok 4.6

    Standard

    $3.00 / 1M · 65.9%

  9. #921.8
    Gemini 3.7 Flash

    Standard

    $3.00 / 1M · 65.3%

  10. #106.5
    GPT-5.6 Sol

    Current list price

    $11.25 / 1M · 72.7%

Unpriced

Revised 2026-09-02

Method, sources, and raw data
What is ranked
Publisher scores inside one named benchmark. No composite score.
What a gap means
No publisher result. MicroRouter does not estimate one.
How to compare
Check harness, scaffold, and effort settings before comparing.