Best chat AI models
70 chat models compared on specs, pricing and community reviews.
- Qwen3.8-Omni-Flash — alibaba: Qwen3.8-Omni-Flash is Alibaba's cutting-edge omni-modal AI model designed to seamlessly process text, image, and audio inputs simultaneously. Built specifically
- Gemini 3.8 Flash — google: Released in September 2026, Gemini 3.8 Flash serves as Google's highly efficient workhorse model optimized for rapid coding tasks and autonomous agent orchestra
- GPT-6.1 Sol — openai: OpenAI's updated Sol model with near-flagship coding, computer-use and professional performance at a lower price.
- Qwen 3 235B — alibaba: Alibaba's Qwen 3 235B is an exceptionally powerful open-weight model optimized for elite multilingual understanding and advanced programming tasks. Built on a m
- Claude Sonnet 5.5 — anthropic: Anthropic's faster, more efficient Sonnet upgrade in the Claude 5.5 family for everyday coding and document work.
- GPT-5.5 — openai: GPT-5.5 represents OpenAI's next-generation frontier model, specifically optimized for highly complex reasoning, advanced coding tasks, and multi-step analytica
- Ember-1 — fireworks ai: Fireworks Research's first self-trained model, post-trained from Kimi K3 to use about 40% fewer reasoning tokens.
- Holo4-27B — h company: H Company's 27B open-weight vision-language model for computer-use agents across GUIs, code, MCP and APIs.
- Gemini 3.7 Flash — google: Gemini 3.7 Flash represents Google's premier high-efficiency model as of August 2026, engineered specifically for fast-paced coding assistance and agentic workf
- Claude Fable 5.1 — anthropic: Claude Fable 5.1 represents Anthropic's September 2026 frontier refresh, delivering highly specialized performance across chat and coding modalities. It sets a
- GPT-5.6 Terra — openai: GPT-5.6 Terra is a balanced powerhouse within OpenAI's latest model family, engineered to deliver a cost-effective blend of advanced intelligence and efficiency
- GPT-5 — openai: OpenAI's flagship unified model that combines fast responses with built-in deep reasoning, replacing the separate GPT-4o and o-series split. It set new state-of
- Claude Opus 4.5 — anthropic: Claude 4.5 Opus is Anthropic's flagship model designed to tackle the most demanding cognitive tasks, offering unmatched depth in reasoning and precision. It is
- Claude Opus 5.5 — anthropic: Claude Opus 5.5 is Anthropic's premier model built specifically for advanced agentic coding and complex knowledge-work workflows. Featuring a massive 1-million-
- Claude Fable 5 — anthropic: Claude Fable 5 is Anthropic's premier model designed to power highly sophisticated, autonomous, long-running agents. It features a massive 1M-token context wind
- Claude Sonnet 4.5 — anthropic: Claude 4.5 Sonnet is Anthropic's premier mid-tier model designed to deliver elite-level reasoning, coding, and comprehension capabilities. It excels at deeply u
- Grok 4.7 — xai: Released in September 2026, Grok 4.7 is xAI's premier frontier model designed for complex coding, agentic workflows, and deep knowledge work. It features a larg
- GPT-5.6 Sol — openai: GPT-5.6 Sol represents OpenAI's absolute pinnacle of machine intelligence, specifically engineered to tackle ultra-complex reasoning challenges and autonomous a
- GPT-5.4 mini — openai: GPT-5.4 mini is OpenAI's highly efficient and cost-effective model optimized for high-volume automated tasks. It excels in driving specialized coding sub-agents
- GPT-6 Sol — openai: GPT-6 Sol is OpenAI's powerful mid-tier model engineered specifically to tackle complex coding challenges and advanced agentic workflows. With an expansive 1.05
- GPT-5.4 nano — openai: GPT-5.4 nano is OpenAI's highly optimized, lightweight model designed for high-throughput, low-latency text processing tasks. Positioned as the most cost-effect
- Gemini 3.1 Pro (Preview) — google: Gemini 3.1 Pro is a highly advanced multimodal reasoning model from Google, specifically optimized for complex chat interactions and deep analytical coding. Equ
- Qwen3.5-397B-A17B — alibaba: Alibaba's Qwen3.5-397B-A17B is the debut model of the Qwen3.5 family, leveraging a massive Mixture of Experts architecture with 397 billion total and 17 billion
- Qwen3.8-27B — alibaba: Dense 27B open-weight Qwen3.8 release for self-hosting and fine-tuning.
- Qwen3.8-Flash-Next — alibaba: Multimodal 125B MoE with just 6B active parameters per token — an early preview of the Qwen4 architecture, open-weighted with FP8 checkpoint.
- GLM-5.3 — zhipu ai: Z.ai's 744B flagship; open weights released after a two-week cyber-capability safety review.
- GPT-6 Luna — openai: OpenAI's GPT-6 Luna is a highly efficient chat model engineered specifically for high-volume, focused tasks at a fraction of the cost of other frontier models.
- Claude Haiku 4.5 — anthropic: Claude Haiku 4.5 is Anthropic's fastest and most cost-effective model, optimized for near-instantaneous response times. It delivers exceptional performance on h
- MiMo-V2.6-Pro — xiaomi: Xiaomi's MiMo-V2.6-Pro is a flagship, MIT-licensed Mixture-of-Experts (MoE) model boasting 1.02 trillion total parameters with 42 billion active. This open-weig
- Gemini 3.6 Flash — google: Gemini 3.6 Flash is Google's premier speed-optimized model, engineered to deliver rapid responses without sacrificing deep intelligence or coding capabilities.
- Llama 4 Scout — meta: Llama 4 Scout by Meta is an efficient Mixture of Experts (MoE) chat model featuring 109 billion total and 17 billion active parameters. It stands out in the lan
- Gemini 3 Pro — google: Google's most capable model at launch, with state-of-the-art reasoning and deep multimodal understanding across text, image, audio and video. It powers the Gemi
- Gemini 3.1 Flash Lite — google: Gemini 3.1 Flash Lite is Google’s highly optimized, ultra-low-cost model designed for high-throughput text processing and conversational tasks at massive scale.
- Gemini 3.8 Flash Cyber — google: Gemini 3.8 Flash Cyber is a highly specialized, restricted-access variant of Google's lightweight model engineered specifically for advanced cybersecurity opera
- Llama 4 405B — meta: Llama 4 405B represents Meta's frontier-class open-weights model, offering state-of-the-art general reasoning, coding, and chat capabilities. Designed to compet
- Llama 4 70B — meta: Meta's Llama 4 70B is the premier sweet spot for open-weights self-hosting, masterfully balancing state-of-the-art conversational quality with a highly manageab
- DeepSeek R2 — deepseek: DeepSeek R2 is a highly efficient, reasoning-focused open-weights model designed specifically to excel in complex mathematical synthesis and advanced coding tas
- DeepSeek-V3.2 — deepseek: Open-weight successor to V3.2-Exp that introduces DeepSeek Sparse Attention for efficient long-context reasoning and agentic work. A high-compute Speciale varia
- Qwen3-Max — alibaba: Alibaba's largest Qwen release, a trillion-parameter-scale mixture-of-experts model with a production thinking mode for coding, reasoning and agentic tasks. It
- MiMo-V2.6-Flash — xiaomi: Xiaomi's MiMo-V2.6-Flash is an open-source, highly efficient Mixture-of-Experts (MoE) model released under the permissive MIT license. Operating on 15B active p
- Step 5 Preview — stepfun: Step 5 Preview is the early-access iteration of StepFun's fifth-generation multimodal foundation model, showcasing advanced capabilities in conversational chat
- GLM-4.6 — zhipu ai: Zhipu's upgrade to GLM-4.5, expanding context from 128k to 200k tokens with better agentic, reasoning and coding behaviour. It has become a widely adopted open
- Gemini 3.5 Flash-Lite — google: Gemini 3.5 Flash-Lite is Google's most economical general-availability model, specifically optimized for high-volume agentic workflows and data processing. Offe
- Claude Opus 5 — anthropic: Claude Opus 5 is Anthropic's flagship model designed specifically for complex agentic coding and heavy enterprise workloads. Featuring an expansive 1-million-to
- Claude Sonnet 5 — anthropic: Claude Sonnet 5 by Anthropic delivers an exceptional balance of speed and high-tier intelligence, specifically optimized for advanced chat and complex coding ta
- GPT-5.6 Luna — openai: GPT-5.6 Luna by OpenAI is a highly efficient chat-based model designed specifically to handle high-volume, cost-sensitive workloads without compromising on mode
- Gemini 3.5 Flash — google: Google's Gemini 3.5 Flash is a highly optimized, cost-effective model designed specifically for high-speed, high-volume conversational tasks and multi-turn agen
- Gemini 3.1 Pro (Preview) — google: Preview Gemini Pro for multimodal understanding, agents and coding. $4/$18 above 200k-token prompts.
- Grok 4 — xai: Grok 4 is xAI's flagship conversational AI, engineered for advanced analytical reasoning and deep conceptual synthesis. Building on its predecessor's strengths,
- Qwen3.5-397B-A17B — alibaba: Alibaba's first Qwen3.5 release: 397B-parameter MoE with 17B active parameters, Apache 2.0.
- Llama 4 Maverick — meta: Llama 4 Maverick is a state-of-the-art natively multimodal Mixture-of-Experts (MoE) model developed by Meta, engineered to excel in high-performance chat and co
- GPT-6 Astra — openai: OpenAI's first GPT-6-generation model. SOTA on computer use, browsing, software engineering and science; saturates ARC-AGI-3 (99.9%) and FrontierMath Tier 4 (98
- Claude Mythos 5.1 — anthropic: Trusted-access counterpart to Fable 5.1 with elevated cyber capability, gated under Anthropic's access program.
- Kimi K2.8 Preview — moonshot ai: Moonshot AI's Kimi K2.8 Preview is an advanced agentic long-context model designed specifically for complex tool utilization and extended programming tasks. Bui
- GLM-5.3-Flash — zhipu ai: 320B-A18B hybrid-attention MoE (MIT), first natively multimodal GLM-5 model; ran anonymously on OpenRouter as 'Ox Alpha' before launch.
- Kimi K3 — moonshot ai: Moonshot AI's August 2026 release — the largest and best-performing open-weights model at launch.
- Hunyuan Hy4 Preview — tencent: Tencent's 770B-A49B open-weight flagship (Apache 2.0) with 1M context, sparse attention and native speculative decoding.
- Granite 4.2 — ibm: IBM's open reasoning family (3B/8B/30B, Apache 2.0) with native chain-of-thought, a thinking switch, and agentic RL post-training.
- Fugu Ultra v2.0 — sakana ai: Sakana AI's September 2026 flagship, released alongside Fugu Max.
- Grok 4.6 — xai: xAI reasoning model for coding and agentic workloads, with selectable reasoning effort and a 500k-token context window.