Model intelligence
Releases, with the sources attached
Reviewed publisher claims, beside a live feed of newly routable models.
Editorial reference
18 sourced · 18 routable · revised 2026-09-02
Live catalog signal
8 arrivals · through 2 Sept 2026
Reviewed release dossiers
Pricing, specs, and reported benchmarks from publisher announcements.
Newest first · revised 2026-09-02
- Google
Gemini 3.8 Flash
Google's third Flash release in six weeks, announced alongside a Cyber variant, with three thinking-effort levels, a 1M context window, and introductory pricing held to Dec 31, 2026.
Routable on MicroRouter
- Alibaba
Qwen3.8-Max-0902
An upgraded snapshot of Alibaba's 2.4-trillion-parameter Qwen3.8-Max, further post-trained on coding and collaborative-agent work while keeping the 1M context window and thinking mode.
Routable on MicroRouter
- Anthropic
Claude Fable 5.1
Anthropic's most capable generally available model, sharing weights with the gated Claude Mythos 5.1, with unchanged list prices, cache reads cut 75%, and a nine-benchmark launch table against Fable 5, Opus 5, and GPT-5.6 Sol.
Routable on MicroRouter
- Tencent
Tencent Hy4 preview
Tencent's open-weight 770B sparse MoE with 49B active parameters, a 1M context window, Apache 2.0 weights, and publisher-reported coding and reasoning results Tencent places at the open-source frontier.
Routable on MicroRouter
- Z.ai
GLM-5.3-Flash
Z.ai's natively multimodal coding model, tested anonymously as ox-alpha, pairing 320B total with 18B active parameters, MIT-licensed weights, and a 50% launch discount through Sep 9, 2026.
Routable on MicroRouter
- Alibaba
Qwen3.8-Flash
Alibaba's hosted production version of Qwen3.8-Flash-Next, adding a 1M-token default context and built-in tools on top of the open-weight base for coding, agent, and visual work.
Routable on MicroRouter
- Alibaba
Qwen3.8-Flash-Next
Alibaba's open-weight 125B MoE with 6B active parameters, positioned as an early preview of the Qwen4 architecture, with vision input, selectable reasoning effort, and publisher-reported coding and reasoning results.
Routable on MicroRouter
- DeepSeek
DeepSeek-V4-Flash-Vision-Exp
DeepSeek's first experimental multimodal V4 model, adding visual modules and continued training to V4-Flash at V4-Flash pricing, with publisher-reported agent results DeepSeek places close to Claude Opus 4.8.
Routable on MicroRouter
- Z.ai
GLM-5.3
Z.ai's coding-model update, announced with post-training gains, three reasoning-effort levels, and publisher-reported coding and cyber benchmarks.
Routable on MicroRouter
- Google
Gemini 3.7 Flash
Google's latest Flash-tier model, announced with introductory API pricing and three publisher-reported coding benchmarks.
Routable on MicroRouter
- DeepSeek
DeepSeek-V4-Pro
DeepSeek moved V4-Pro from preview to general availability across app, web, and API, citing enhanced agent capabilities, reasoning-effort levels, and a peak/off-peak pricing change from Aug 16.
Routable on MicroRouter
- xAI
Grok 4.6
xAI's frontier model for coding, agentic tasks, and knowledge work, positioned around long-running agents with enhanced self-testing, launched in Cursor, Grok Build, and the xAI API.
Routable on MicroRouter
- OpenAI
GPT-5.4 mini
OpenAI's small-tier refresh for high-volume agents, positioned around tool-calling reliability at mini-tier prices across coding, science, and long-context work.
Routable on MicroRouter
- Meta
Muse Glimmer 30B
Meta's 30B open-weight dense multimodal model for local, long-running agentic workflows, distilled from the larger Muse Spark teacher and sized to run on a single consumer GPU.
Routable on MicroRouter
- MiniMax
MiniMax M3
MiniMax's sparse MoE release positioned around cheap agentic throughput, with interleaved thinking on by default and a listed API price well under a dollar per million input tokens.
Routable on MicroRouter
- Moonshot AI
Kimi K3
Moonshot's agentic coding update to Kimi K2, positioned around long-horizon tool use, with open weights and publisher-reported coding and science benchmarks.
Routable on MicroRouter
- Alibaba
Qwen3.8-2.4T-A95B
The open-weight release of Qwen3.8-Max: a 2.4-trillion-parameter sparse MoE with 95B active parameters, positioned around autonomous coding, workplace tasks, and long-horizon execution.
Routable on MicroRouter
- OpenAI
GPT-5.6 Sol
OpenAI's flagship GPT-5.6 tier for long-running coding, research, and professional workflows, with a revised API list price and publisher-reported coding benchmarks.
Routable on MicroRouter
Newly routable
Production-catalog arrivals from the last 14 days.
8 current arrivals · latest 2 Sept 2026
- google
Gemini 3.8 Flash
google/gemini-3.8-flash
Gemini 3.8 Flash is the high-efficiency, cost-effective powerhouse of the Gemini 3 family. It delivers Pro-level agentic capabilities, major leaps in code generation and terminal execution. 3.8 Flash serves as the primary agentic workhorse in the Gemini 3 family, bridging the gap between deep-reasoning Pro models and high-throughput Flash-Lite models while delivering high token efficiency and multi-step multimodal processing.
$0.788 input / $3.938 output per 1M
- alibaba
Qwen3.8 Max 0902
alibaba/qwen3.8-max-0902
Qwen3.8-Max-0902 is an upgraded snapshot of qwen3.8-max.
$2.1 input / $6.3 output per 1M
- anthropic
Claude Fable 5.1
anthropic/claude-fable-5.1
Claude Fable 5.1 improves on Fable 5 across long-running agentic coding, knowledge work, and research, following instructions precisely over sessions that run unattended for hours.
$10.5 input / $52.5 output per 1M
- tencent
Tencent Hy4 Preview
tencent/hy4-preview
Tencent Hy4 preview is Tencent Hy’s open-source large language model, featuring 770B total parameters, 49B active parameters, and a context window exceeding 1 million tokens. Built for real-world productivity, it supports long-horizon coding, cross-document analysis, office content creation, game development, and scientific reasoning, making it well-suited to complex agents and multi-step professional workflows.
$0.876 input / $2.626 output per 1M
- alibaba
Qwen 3.8 Flash
alibaba/qwen3.8-flash
Qwen3.8-Flash is Qwen’s fast, cost-efficient multimodal model, combining advanced reasoning and generation with a native 1M-token context window. Built for coding, agentic workflows, and visual understanding, it handles large codebases, long documents, charts, videos, and desktop applications.
$0.168 input / $0.494 output per 1M
- alibaba
Qwen 3.8 Flash Next
alibaba/qwen3.8-flash-next
Qwen3.8-Flash-Next is Qwen’s experimental open-weight multimodal language model, pairing 125B parameters with just 6B activated for efficient reasoning and generation. With native 262K-token context extensible to 1M, vision support, and an architecture optimized for lower long-context latency, it is built for demanding agentic coding, tool use, visual reasoning, and multimodal automation.
$0.126 input / $0.42 output per 1M
- zai
GLM 5.3 Flash
zai/glm-5.3-flash
GLM-5.3-Flash is Z.ai’s native multimodal coding model, featuring 320B total parameters, 18B activated parameters, and a 1M-token context window. Its efficient hybrid attention architecture supports visual coding, tool use, and end-to-end professional workflows across code, browsers, documents, and graphical interfaces.
$0.158 input / $0.525 output per 1M
- deepseek
DeepSeek V4 Flash Vision Exp
deepseek/deepseek-v4-flash-vision-exp
DeepSeek-V4-Flash-Vision-Exp is the first experimental multimodal model in the DeepSeek-V4 family. It builds on DeepSeek-V4-Flash by adding visual modules and continued training for visual understanding, with substantially improved multimodal agent capabilities while remaining comparable on text-only agent tasks.
$0.231 input / $0.693 output per 1M
Source: MicroRouter's live model catalog ↗