Best code AI models

45 code models compared on specs, pricing and community reviews.

  • o4-mini — openai: o4-mini delivers strong reasoning performance at lower cost and latency than o3. It supports tool use and is optimized for math and coding. It's designed for hi
  • Gemini 2.5 Pro — google: Gemini 2.5 Pro is Google's flagship model with native multimodality and a 1M token context window, featuring built-in 'thinking' for complex reasoning. It leads
  • Gemini 2.5 Flash-Lite — google: Gemini 2.5 Flash-Lite is optimized for high-throughput, latency-sensitive tasks at the lowest cost in the Gemini 2.5 family. It retains a 1M token context windo
  • Claude Opus 4.1 — anthropic: Claude Opus 4.1 is Anthropic's top-tier model, excelling at agentic coding, complex reasoning and long-horizon tasks. It supports a 200k context window and visi
  • Mistral Large 2 — mistral: Mistral Large 2 is a 123B parameter dense model with strong multilingual, coding and reasoning capabilities and a 128k context window. It's available under a re
  • Codestral 2508 — mistral: Codestral is fine-tuned specifically for code generation, completion and understanding across 80+ programming languages, with a 256k context window. It supports
  • Gemini 2.5 Flash — google: Gemini 2.5 Flash balances speed, cost and reasoning quality with a configurable thinking budget. It supports a 1M token context and full multimodal input. It's
  • Claude Sonnet 4.5 — anthropic: Claude Sonnet 4.5 offers strong coding and agentic capabilities at a mid-tier price point, with a 200k context window (1M in beta). It's positioned as Anthropic
  • Mistral Small 3.2 — mistral: Mistral Small 3.2 is a 24B parameter open-weight model with multimodal support, tuned for instruction following and reduced repetition errors. It offers a 128k
  • DeepSeek-V3.1 — deepseek: DeepSeek-V3.1 is a 671B parameter (37B active) mixture-of-experts model combining fast and thinking modes in one model. It offers strong coding and reasoning at
  • DeepSeek-R1 — deepseek: DeepSeek-R1 is a reasoning-focused model trained with large-scale reinforcement learning, achieving performance comparable to OpenAI's o1 on math and coding ben
  • Qwen3-235B-A22B — alibaba: Qwen3-235B-A22B is a mixture-of-experts model with 235B total/22B active parameters, supporting seamless switching between thinking and non-thinking modes. It s
  • Qwen 3 235B — alibaba: Alibaba's Qwen 3 235B is an exceptionally powerful open-weight model optimized for elite multilingual understanding and advanced programming tasks. Built on a m
  • Claude Sonnet 5 — anthropic: Claude Sonnet 5 by Anthropic delivers an exceptional balance of speed and high-tier intelligence, specifically optimized for advanced chat and complex coding ta
  • Claude Sonnet 4.5 — anthropic: Claude 4.5 Sonnet is Anthropic's premier mid-tier model designed to deliver elite-level reasoning, coding, and comprehension capabilities. It excels at deeply u
  • GPT-5.6 Terra — openai: GPT-5.6 Terra is a balanced powerhouse within OpenAI's latest model family, engineered to deliver a cost-effective blend of advanced intelligence and efficiency
  • Claude Fable 5 — anthropic: Claude Fable 5 is Anthropic's premier model designed to power highly sophisticated, autonomous, long-running agents. It features a massive 1M-token context wind
  • Claude Opus 5 — anthropic: Claude Opus 5 is Anthropic's flagship model designed specifically for complex agentic coding and heavy enterprise workloads. Featuring an expansive 1-million-to
  • GPT-5.6 Sol — openai: GPT-5.6 Sol represents OpenAI's absolute pinnacle of machine intelligence, specifically engineered to tackle ultra-complex reasoning challenges and autonomous a
  • Gemini 3.1 Pro (Preview) — google: Gemini 3.1 Pro is a highly advanced multimodal reasoning model from Google, specifically optimized for complex chat interactions and deep analytical coding. Equ
  • Grok 4.5 — xai: xAI's frontier model for coding and agentic work. 500k context, $2/$6 per 1M tokens under 200k prompt tokens ($4/$12 above).
  • Mistral Large 3 — mistral: Mistral Large 3 is Mistral AI's premier European frontier model, engineered to deliver top-tier reasoning, advanced coding capabilities, and highly sophisticate
  • DeepSeek V4 Pro — deepseek: Mixture-of-experts model with 1.6T total and 49B active parameters, hybrid compressed attention, and a 1M-token context window.
  • Qwen3.5-397B-A17B — alibaba: Alibaba's first Qwen3.5 release: a 397B-parameter MoE with 17B active parameters, Apache 2.0 licensed, hosted as Qwen3.5-Plus on Model Studio.
  • Llama 4 Maverick — meta: Meta's natively multimodal MoE with 400B total and 17B active parameters, under the Llama 4 Community License.
  • Gemini 3.6 Flash — google: Gemini 3.6 Flash is Google's premier speed-optimized model, engineered to deliver rapid responses without sacrificing deep intelligence or coding capabilities.
  • Grok 3 — xai: Grok 3 introduced 'Think' mode for extended reasoning and DeepSearch for web-integrated answers, with a 131k context window. It was trained on xAI's Colossus su
  • Command A — cohere: Command A is Cohere's most capable model, optimized for enterprise use cases like RAG, tool use and agents, with a 256k context window. It runs efficiently on j
  • Jamba 1.6 — ai21: Jamba 1.6 combines Mamba and Transformer architectures in a mixture-of-experts design, offering a 256k context window with efficient long-context inference. It
  • Phi-4 — microsoft: Phi-4 is a 14B parameter dense model trained with a focus on data quality, achieving strong performance on reasoning and math benchmarks relative to its size. I
  • Nemotron-4 340B — nvidia: Nemotron-4 340B is a dense 340B parameter model optimized to generate high-quality synthetic training data for other LLMs. It's released under the NVIDIA Open M
  • Llama-3.1-Nemotron-Ultra-253B — nvidia: This model is a pruned and post-trained derivative of Llama 3.1 405B, optimized by Nvidia for reasoning tasks with toggleable reasoning mode. It supports a 128k
  • Kimi K2 — moonshot ai: Kimi K2 is a 1T parameter (32B active) mixture-of-experts model trained for agentic tool-use and coding, released with open weights under a modified MIT license
  • GLM-4.5 — zhipu ai: GLM-4.5 is a 355B parameter (32B active) mixture-of-experts model designed for agentic, reasoning and coding tasks with a hybrid thinking mode. It's open-weight
  • MiniMax M1 — minimax: MiniMax M1 is a 456B parameter (45.9B active) hybrid-attention MoE reasoning model supporting up to 1M tokens of context, trained with efficient large-scale rei
  • ERNIE 4.5 — baidu: ERNIE 4.5 is a family of mixture-of-experts models (up to 424B parameters) supporting text and multimodal understanding, released with open weights for several
  • Granite 3.3 8B — ibm: Granite 3.3 8B is IBM's open-weight model tuned for enterprise use cases like RAG, function calling and fill-in-the-middle code, with a 128k context window. It'
  • Claude Opus 4.5 — anthropic: Claude 4.5 Opus is Anthropic's flagship model designed to tackle the most demanding cognitive tasks, offering unmatched depth in reasoning and precision. It is
  • Llama 4 405B — meta: Llama 4 405B represents Meta's frontier-class open-weights model, offering state-of-the-art general reasoning, coding, and chat capabilities. Designed to compet
  • Claude 3.5 Sonnet — anthropic: Claude 3.5 Sonnet introduced computer use (agentic screen control) and improved coding performance over Claude 3 Opus at a lower price. It supports a 200k conte
  • DeepSeek R2 — deepseek: DeepSeek R2 is a highly efficient, reasoning-focused open-weights model designed specifically to excel in complex mathematical synthesis and advanced coding tas
  • GPT-5.5 — openai: GPT-5.5 represents OpenAI's next-generation frontier model, specifically optimized for highly complex reasoning, advanced coding tasks, and multi-step analytica
  • GPT-4.1 mini — openai: GPT-4.1 mini balances cost and capability, offering a 1M token context window at a fraction of GPT-4.1's price. It maintains strong coding and reasoning perform
  • GPT-4.1 — openai: GPT-4.1 offers a 1M token context window with major improvements in coding and instruction following over GPT-4o. It's available in standard, mini and nano size
  • o3 — openai: o3 is a reasoning-focused model that uses extended chain-of-thought to solve complex math, science and coding problems. It supports tool use during reasoning. I