Skip to content
  • Models
  • Rankings
  • Ori
Sign Up
Sign Up
OpenRouterOpenRouter
© 2026 OpenRouter, Inc

Product

  • Chat
  • Rankings
  • Benchmarks
  • Apps
  • Discover
  • Models
  • Collections
  • Providers
  • Tools
  • Pricing
  • Business
  • Enterprise
  • Labs

Company

  • About
  • Blog
  • Careers
    Hiring
  • Privacy
  • Terms of Service
  • Trust Center
  • Support
  • Works With OR
  • Data
  • Brand

Developer

  • Documentation
  • API Reference
  • Developer Platform
  • Status
  • AI Site Map

Connect

  • Discord
  • GitHub
  • LinkedIn
  • X
  • YouTube
Collections/Coding

Best AI Models for Coding, Ranked by Real Usage

Model rankings updated September 2026 based on real usage data.

Models in this coding collection are ranked by total prompt and completion tokens processed through the OpenRouter API over the trailing 7 days. Showing the top 10 models. Rankings measure usage on OpenRouter and reflect adoption, not model quality or benchmark performance.

The best AI models for coding are ranked below by how much developers actually use them on OpenRouter over the past week, not by benchmark scores. The current top models are DeepSeek V4.1 Flash, GLM 5.3 Flash, and Hy4 preview. Each listing shows the model's context length and pricing where OpenRouter has them, so you can weigh cost against usage before you pick one. Every model is available through a single OpenRouter API key, so you can test them on your own codebase and switch without rewriting your integration.

Browse All ModelsCompare Models

LLM Leaderboard for Programming Models

1.
Favicon for z-ai
GLM 5.3 Flash
by z-ai
7.58T
14.8%
2.
Favicon for deepseek
Deepseek V4.1 Flash
by deepseek
6.96T
13.6%
3.
Favicon for stealth
Space Bunny Alpha
by stealth
5.31T
10.4%
4.
Favicon for nvidia
Nemotron 3 Ultra 550B A55B (free)
by nvidia
4.42T
8.6%
5.
Favicon for xiaomi
Mimo V2.6 Flash
by xiaomi
3.53T
6.9%
6.
Favicon for xiaomi
Mimo V2.5
by xiaomi
2.74T
5.4%
7.
Favicon for deepseek
Deepseek V4 Flash
by deepseek
1.71T
3.3%
8.
Favicon for z-ai
GLM 5.3
by z-ai
1.49T
2.9%
9.
Favicon for tencent
Hy4 Preview
by tencent
1.46T
2.9%
10.
Favicon for unknown
Others
16T
31.2%

Top Coding Models on OpenRouter

Favicon for deepseek

DeepSeek: DeepSeek V4.1 Flash

19.8T tokens
Academia (#1)
Finance (#1)
Health (#2)
Legal (#3)
Marketing (#1)

DeepSeek V4.1 Flash is a sparse mixture-of-experts model from DeepSeek, and the first built on the company's Causal Encoder-Decoder (CED) architecture. It activates 8B parameters on input and 16B on output from a 552B-parameter backbone, an asymmetric split that keeps per-token compute low relative to the model's total size. Image understanding is native to the architecture, with visual and text embeddings trained jointly from the start of pre-training rather than added afterward as in the earlier experimental V4 Flash Vision Exp.

It is suited for coding, terminal, and computer-use agents, along with long-horizon tasks that must run to completion across many steps and long-context analysis. Compressed KV caching cuts cache memory to roughly a quarter of the previous Flash generation, significantly reducing costs on agentic workloads. DeepSeek positions it as the cost-efficient tier of the V4.1 family and reports that it exceeds V4 Pro on performance, speed, and task completion time.

by deepseek1.05M context$0.035/M input tokens$0.29/M output tokens
Favicon for z-ai

Z.ai: GLM 5.3 Flash

19.6T tokens
Academia (#3)
Finance (#3)
Health (#1)
Legal (#5)
Marketing (#2)

GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while reducing compute overhead.

by z-ai1.31M context$0.04/M input tokens$0.50/M output tokens
Favicon for tencent

Tencent: Hy4 preview

11.4T tokens
Academia (#8)
Finance (#10)
Health (#48)
Legal (#11)
Marketing (#8)

Tencent: Hy4 preview is a mixture-of-experts model from Tencent, with 49B active parameters out of 770B total. It is designed for coding agents, complex tool-use workflows, and productivity tasks that require planning, context continuity, and sustained multi-step execution.

by tencent1.05M context$0.834/M input tokens$2.501/M output tokens
Favicon for openai

OpenAI: GPT-5.6 Luna

8.81T tokens
Academia (#5)
Finance (#5)
Health (#5)
Legal (#2)
Marketing (#5)

GPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series. It is suited for high-volume, latency-sensitive tasks such as chat, classification, and lightweight agentic workflows, providing capable reasoning for its price tier.

by openai1.05M context$0.20/M input tokens$1.20/M output tokens
Favicon for deepseek

DeepSeek: DeepSeek V4 Flash 0731

8.48T tokens
Academia (#2)
Finance (#2)
Health (#4)
Legal (#4)
Marketing (#3)

DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows. This is the GA release of DeepSeek V4 Flash.

by deepseek1.31M context$0.021/M input tokens$0.32/M output tokens
Favicon for stealth

Space Bunny Alpha

6.96T tokens
Finance (#50)
Programming (#20)
Science (#16)
Technology (#43)
Translation (#21)

Space Bunny Alpha is an anonymous large model with blazing-fast inference, strong coding capabilities and native multimodal input support. It delivers adjustable reasoning effort, and a 1M-token context window.

Space Bunny Alpha is a stealth model. It is developed and operated by a third-party provider who has chosen to remain anonymous during this preview. OpenRouter routes requests to it and is not its developer, owner, or provider. Prompts and completions may be retained by the provider but are not used for training; all other use is governed by the Stealth Model Terms.

by stealth1M context$0/M input tokens$0/M output tokens
Favicon for nvidia

NVIDIA: Nemotron 3 Ultra (free)

5.36T tokens
Programming (#16)

NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it supports text input and output with a context window of up to 1M tokens. It is suited for long-running agentic workflows, including agent orchestration, coding agents, deep research, and complex enterprise tasks.

It is particularly strong at multi-step reasoning and planning, with high-throughput inference designed for high-volume agent pipelines. It is part of the NVIDIA Nemotron family of open models for agentic AI.

by nvidia1M context$0/M input tokens$0/M output tokens
Favicon for xiaomi

Xiaomi: MiMo-V2.5

4.21T tokens
Academia (#10)
Finance (#13)
Health (#7)
Marketing (#30)
SEO (#13)

MiMo-V2.5 is a native omnimodal model by Xiaomi. It delivers Pro-level agentic performance at roughly half the inference cost, while surpassing MiMo-V2-Omni in multimodal perception across image and video understanding tasks. Its 1M context window supports complete documents, extended conversations, and complex task contexts in a single pass, making it ideal for integration with agent frameworks where strong reasoning, rich perception, and cost efficiency all matter.

by xiaomi1.05M context$0.119/M input tokens$0.238/M output tokens15% off
Favicon for deepseek

DeepSeek: DeepSeek V4 Flash 0423

3.47T tokens
Academia (#4)
Finance (#7)
Health (#6)
Legal (#6)
Marketing (#4)

DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated parameters, supporting a 1M-token context window. It is designed for fast inference and high-throughput workloads, while maintaining strong reasoning and coding performance.

The model includes hybrid attention for efficient long-context processing. Reasoning efforts high and xhigh are supported; xhigh maps to max reasoning. It is well suited for applications such as coding assistants, chat systems, and agent workflows where responsiveness and cost efficiency are important.

by deepseek1.05M context$0.03/M input tokens$1.28/M output tokens
Favicon for xiaomi

Xiaomi: MiMo-V2.6-Flash

3.38T tokens
Academia (#35)
Finance (#41)
Marketing (#48)
Programming (#4)
Roleplay (#13)

MiMo-V2.6-Flash is an open-source foundation model developed by Xiaomi. Built on a Mixture-of-Experts architecture with 309B total parameters and 15B activated per token, it employs a hybrid attention mechanism for greater computational efficiency. The model features a 1M-token context window and native multimodal capabilities. Optimized for agentic workflows, it delivers strong performance across coding, visual, general, and research scenarios, excelling at complex, long-horizon tasks with robust generalization across a diverse range of agent harnesses.

by xiaomi1.05M context$0.14/M input tokens$0.28/M output tokens

Explore more collections

  • Free Models
  • Discounted Models
  • Roleplay
  • Vision Models
  • Tool Calling
  • OpenClaw
  • Image Models
  • Video Models
  • Audio Models
  • Text-to-Speech
  • Speech-to-Text
  • Embedding Models
  • Rerank Models
  • Distillable Models
  • All collections