# Reviuws — full AI model dataset (plain text) Source: https://reviuws.com | Generated: 2026-08-21 | Models: 101 Each entry below is a single AI model tracked by Reviuws. Fields: provider, modality, license, parameters, context window, input/output price per million tokens, community rating, review count, summary, typical use case and canonical URL. ## Amazon Nova Micro URL: https://reviuws.com/models/amazon-nova-micro Provider: amazon Modality: text License: closed Parameters: n/a Context window: 128k tokens Input price: $0.035/M tokens Output price: $0.14/M tokens Community rating: not yet rated Released: 2024-12-03 Summary: Amazon Nova Micro is a text-only model optimized for the lowest latency and cost within the Nova family, with a 128k context window. It's available via AWS Bedrock. It's designed for simple, high-volume text tasks. Typical use case: Used for lightweight chatbots, text classification, and summarization at massive scale. Ideal for cost-constrained AWS workloads. ## Amazon Nova Pro URL: https://reviuws.com/models/amazon-nova-pro Provider: amazon Modality: text, vision, video License: closed Parameters: n/a Context window: 300k tokens Input price: $0.8/M tokens Output price: $3.2/M tokens Community rating: not yet rated Released: 2024-12-03 Summary: Amazon Nova Pro is a multimodal model supporting text, image and video understanding with a 300k token context window, available exclusively via AWS Bedrock. It's optimized for accuracy, speed and cost balance. It's part of the broader Nova family (Micro, Lite, Pro, Premier). Typical use case: Used for enterprise document/video analysis and agentic workflows on AWS. Suited for AWS-native applications needing multimodal reasoning. ## Aya Expanse 32B URL: https://reviuws.com/models/aya-expanse-32b Provider: cohere Modality: text License: open Parameters: 32B Context window: 128k tokens Input price: n/a Output price: n/a Community rating: not yet rated Released: 2024-10-24 Summary: Aya Expanse 32B is an open-weight model covering 23 languages, developed by Cohere's research lab to close multilingual performance gaps in open models. It supports a 128k context window. It's released for research use with broad language coverage. Typical use case: Used for multilingual research, translation-adjacent tasks, and underserved-language chat applications. Suited for global applications needing broad language support. ## Claude 3.5 Sonnet URL: https://reviuws.com/models/claude-3-5-sonnet Provider: anthropic Modality: text, vision, code License: closed Parameters: n/a Context window: 200k tokens Input price: $3/M tokens Output price: $15/M tokens Community rating: not yet rated Released: 2024-10-22 Summary: Claude 3.5 Sonnet introduced computer use (agentic screen control) and improved coding performance over Claude 3 Opus at a lower price. It supports a 200k context window and vision input. It was widely adopted for coding assistant products. Typical use case: Used for coding assistants, agentic browser/computer control tasks, and general chat. Was a leading choice for developer tools before Claude 4. ## Claude Fable 5 URL: https://reviuws.com/models/claude-fable-5 Provider: anthropic Modality: chat, code License: closed Parameters: n/a Context window: 1000k tokens Input price: $10/M tokens Output price: $50/M tokens Community rating: not yet rated Released: 2026-06-09 Summary: Claude Fable 5 is Anthropic's premier model designed to power highly sophisticated, autonomous, long-running agents. It features a massive 1M-token context window, an unprecedented 128k maximum output, and always-on adaptive thinking for deep problem-solving. This makes it a formidable tool for complex programming, analysis, and conversational workflows. Typical use case: An enterprise development team can use Claude Fable 5 to automate the refactoring of a massive legacy codebase. By ingesting the entire repository within its 1M-token context window, the model can autonomously identify structural inefficiencies, map out dependencies, and generate thousands of lines of updated, production-ready code in a single 128k output run. ## Claude Haiku 4.5 URL: https://reviuws.com/models/claude-haiku-4-5 Provider: anthropic Modality: text, vision License: closed Parameters: n/a Context window: 200k tokens Input price: $1/M tokens Output price: $5/M tokens Community rating: not yet rated Released: 2025-10-15 Summary: Claude Haiku 4.5 is optimized for speed and affordability while retaining solid reasoning ability inherited from the Claude 4 family. It supports a 200k context window. It's suited for latency-sensitive, high-volume tasks. Typical use case: Used for chat support, content moderation, and lightweight agent tasks. Ideal where response time and cost matter more than peak capability. ## Claude Haiku 4.5 URL: https://reviuws.com/models/claude-haiku-45 Provider: anthropic Modality: chat License: closed Parameters: n/a Context window: 200k tokens Input price: $1/M tokens Output price: $5/M tokens Community rating: not yet rated Released: 2025-10-01 Summary: Claude Haiku 4.5 is Anthropic's fastest and most cost-effective model, optimized for near-instantaneous response times. It delivers exceptional performance on high-volume, repetitive tasks like classification, structured data extraction, and entry-level conversational agents. This model offers enterprises a highly efficient way to scale AI workloads without incurring massive token expenses. Typical use case: A global e-commerce brand deploys Claude Haiku 4.5 to manage its high-volume customer service routing. The model processes tens of thousands of incoming customer emails per minute in real time, instantly classifying the intent, extracting order IDs, and assigning them to the correct departments while generating instant, helpful auto-responses for simple FAQs. ## Claude Opus 4.1 URL: https://reviuws.com/models/claude-opus-4-1 Provider: anthropic Modality: text, vision, code License: closed Parameters: n/a Context window: 200k tokens Input price: $15/M tokens Output price: $75/M tokens Community rating: not yet rated Released: 2025-08-05 Summary: Claude Opus 4.1 is Anthropic's top-tier model, excelling at agentic coding, complex reasoning and long-horizon tasks. It supports a 200k context window and vision input. It's designed for the most demanding enterprise workloads. Typical use case: Used for advanced coding agents, research analysis, and enterprise-grade reasoning tasks. Suited for tasks requiring maximum quality regardless of cost. ## Claude Opus 4.5 URL: https://reviuws.com/models/claude-opus-45 Provider: anthropic Modality: chat, code License: closed Parameters: n/a Context window: 200k tokens Input price: $15/M tokens Output price: $75/M tokens Community rating: not yet rated Summary: Claude 4.5 Opus is Anthropic's flagship model designed to tackle the most demanding cognitive tasks, offering unmatched depth in reasoning and precision. It is the premier choice for complex software engineering, multi-step agentic workflows, and highly nuanced creative or technical writing. While it sets a new benchmark for intelligence, its elite performance is accompanied by premium pricing and higher latency. Typical use case: A multinational financial services firm can utilize Claude 4.5 Opus to power autonomous developer agents that safely migrate monolithic legacy codebases to cloud-native microservices. The model can ingest thousands of lines of undocumented legacy code, map complex dependency trees, systematically refactor the logic into clean Python, and automatically write comprehensive unit and integration tests. ## Claude Opus 5 URL: https://reviuws.com/models/claude-opus-5 Provider: anthropic Modality: chat, code License: closed Parameters: n/a Context window: 1000k tokens Input price: $5/M tokens Output price: $25/M tokens Community rating: not yet rated Released: 2026-06-09 Summary: Claude Opus 5 is Anthropic's flagship model designed specifically for complex agentic coding and heavy enterprise workloads. Featuring an expansive 1-million-token context window and a groundbreaking 128k maximum output capacity, it excels at managing massive, multi-step analytical tasks. It stands as the premier choice for organizations requiring deep reasoning, reliable execution, and extensive code generation. Typical use case: A financial services institution uses Claude Opus 5 to automate the compliance auditing of thousands of pages of regulatory documents alongside their entire transaction processing codebase. By loading both the massive compliance PDF libraries and the application code into the 1M context window, the model autonomously analyzes the logic, identifies compliance gaps, and generates hundreds of pages of refactored, compliant code and audit reports in a single run. ## Claude Sonnet 4.5 URL: https://reviuws.com/models/claude-sonnet-4-5 Provider: anthropic Modality: text, vision, code License: closed Parameters: n/a Context window: 200k tokens Input price: $3/M tokens Output price: $15/M tokens Community rating: not yet rated Released: 2025-09-29 Summary: Claude Sonnet 4.5 offers strong coding and agentic capabilities at a mid-tier price point, with a 200k context window (1M in beta). It's positioned as Anthropic's best model for real-world agentic tasks. It supports vision and extended thinking. Typical use case: Used for coding assistants, autonomous agents, and everyday enterprise chat applications. Balances cost and capability for production use. ## Claude Sonnet 4.5 URL: https://reviuws.com/models/claude-sonnet-45 Provider: anthropic Modality: chat, code License: closed Parameters: n/a Context window: 200k tokens Input price: $3/M tokens Output price: $15/M tokens Community rating: not yet rated Summary: Claude 4.5 Sonnet is Anthropic's premier mid-tier model designed to deliver elite-level reasoning, coding, and comprehension capabilities. It excels at deeply understanding complex prompts, maintaining context over large datasets, and producing highly articulate responses. This makes it an ideal balance of state-of-the-art intelligence and cost efficiency for demanding enterprise workflows. Typical use case: A software development team can employ Claude 4.5 Sonnet to automate the refactoring of a legacy codebase by feeding it entire directories of old code. The model can accurately map dependencies, identify security vulnerabilities, and generate clean, updated code alongside comprehensive unit tests, dramatically accelerating the modernization process. ## Claude Sonnet 5 URL: https://reviuws.com/models/claude-sonnet-5 Provider: anthropic Modality: chat, code License: closed Parameters: n/a Context window: 1000k tokens Input price: $3/M tokens Output price: $15/M tokens Community rating: not yet rated Summary: Claude Sonnet 5 by Anthropic delivers an exceptional balance of speed and high-tier intelligence, specifically optimized for advanced chat and complex coding tasks. With a massive 1-million-token context window, it effortlessly processes expansive documents and entire code repositories in a single query. It represents a highly cost-effective yet powerful solution for developers requiring deep reasoning without sacrificing performance. Typical use case: A software development team can utilize Claude Sonnet 5 to perform comprehensive code reviews and automated refactoring across a massive enterprise codebase. By ingestion of hundreds of source files into the 1M token context window, the model can trace complex dependencies, identify security vulnerabilities, and generate optimized, production-ready code suggestions in seconds. ## Codestral 2508 URL: https://reviuws.com/models/codestral-2508 Provider: mistral Modality: code, text License: open Parameters: 22B Context window: 256k tokens Input price: $0.3/M tokens Output price: $0.9/M tokens Community rating: not yet rated Released: 2025-08-01 Summary: Codestral is fine-tuned specifically for code generation, completion and understanding across 80+ programming languages, with a 256k context window. It supports fill-in-the-middle for IDE integration. It's optimized for low-latency developer tooling. Typical use case: Used for IDE code completion, code review automation, and codebase Q&A. Suited for developer productivity tools. ## Cohere Rerank 3.5 URL: https://reviuws.com/models/cohere-rerank-3-5 Provider: cohere Modality: text, embedding License: closed Parameters: n/a Context window: 4k tokens Input price: n/a Output price: n/a Community rating: not yet rated Released: 2024-12-02 Summary: Rerank 3.5 improves search relevance by reordering candidate documents using reasoning and better multilingual understanding. It supports 100+ languages and a 4k context per document. It's used as a post-retrieval step in RAG pipelines. Typical use case: Used to improve search and RAG pipeline accuracy by reranking initial retrieval results. Suited for enterprise search applications. ## Command A URL: https://reviuws.com/models/command-a Provider: cohere Modality: text, code License: open Parameters: 111B Context window: 256k tokens Input price: $2.5/M tokens Output price: $10/M tokens Community rating: not yet rated Released: 2025-03-13 Summary: Command A is Cohere's most capable model, optimized for enterprise use cases like RAG, tool use and agents, with a 256k context window. It runs efficiently on just two GPUs. Weights are available for research use. Typical use case: Used for enterprise chatbots, RAG pipelines, and multilingual business applications. Suited for regulated industries needing on-prem deployment options. ## DeepSeek R2 URL: https://reviuws.com/models/deepseek-r2 Provider: deepseek Modality: chat, code License: open Parameters: 236B Context window: 128k tokens Input price: n/a Output price: n/a Community rating: not yet rated Summary: DeepSeek R2 is a highly efficient, reasoning-focused open-weights model designed specifically to excel in complex mathematical synthesis and advanced coding tasks. Built by DeepSeek, it delivers state-of-the-art logical reasoning capabilities at a fraction of the cost of larger proprietary alternatives. This makes it an exceptionally accessible option for developers and researchers who require heavy-duty computational logic. Typical use case: A software engineering team can deploy DeepSeek R2 locally to automate the debugging and refactoring of legacy codebases. By leveraging its reasoning-focused architecture, the model can analyze complex, multi-file code dependencies, identify silent logical errors, and automatically generate optimized, secure patch suggestions, accelerating the development cycle while maintaining data privacy. ## DeepSeek V4 Pro URL: https://reviuws.com/models/deepseek-v4-pro Provider: deepseek Modality: chat, code License: open Parameters: 1600B Context window: 1000k tokens Input price: n/a Output price: n/a Community rating: not yet rated ## DeepSeek-R1 URL: https://reviuws.com/models/deepseek-r1 Provider: deepseek Modality: text, code License: open Parameters: 671B Context window: 128k tokens Input price: $0.55/M tokens Output price: $2.19/M tokens Community rating: not yet rated Released: 2025-01-20 Summary: DeepSeek-R1 is a reasoning-focused model trained with large-scale reinforcement learning, achieving performance comparable to OpenAI's o1 on math and coding benchmarks. It's released fully open-weight under MIT license. It sparked wide adoption due to low cost and openness. Typical use case: Used for research into reasoning models, math/coding benchmarks, and cost-efficient reasoning applications. Widely used for distillation into smaller models. ## DeepSeek-V3.1 URL: https://reviuws.com/models/deepseek-v3-1 Provider: deepseek Modality: text, code License: open Parameters: 671B Context window: 128k tokens Input price: $0.56/M tokens Output price: $1.68/M tokens Community rating: not yet rated Released: 2025-08-21 Summary: DeepSeek-V3.1 is a 671B parameter (37B active) mixture-of-experts model combining fast and thinking modes in one model. It offers strong coding and reasoning at very low API cost. It's released with open weights for research and commercial use. Typical use case: Used for cost-sensitive high-volume inference, coding agents, and research on MoE architectures. Popular for self-hosting on large GPU clusters. ## Eleven Multilingual v2 URL: https://reviuws.com/models/eleven-multilingual-v2 Provider: elevenlabs Modality: audio License: closed Parameters: n/a Context window: n/a Input price: n/a Output price: n/a Community rating: not yet rated Released: 2024-01-01 Summary: Eleven Multilingual v2 generates lifelike speech across 29+ languages with emotional expressiveness and voice cloning support. It's used broadly for narration, dubbing and voice agents. It's accessed via ElevenLabs' API and platform. Typical use case: Used for audiobook narration, video dubbing, and conversational voice agents. Popular for content localization workflows. ## ElevenLabs v3 URL: https://reviuws.com/models/elevenlabs-v3 Provider: elevenlabs Modality: audio License: closed Parameters: n/a Context window: n/a Input price: n/a Output price: n/a Community rating: not yet rated Summary: ElevenLabs v3 is a state-of-the-art audio generation model specializing in high-fidelity text-to-speech and voice cloning. It offers unprecedented control over emotional expression, tone, and pacing, making synthetic voices virtually indistinguishable from human speakers. The model also excels at multilingual generation, seamlessly maintaining voice characteristics across dozens of languages. Typical use case: An audiobook publisher utilizes ElevenLabs v3 to automate the production of their expansive backlist catalog. By cloning the original author's voice with permission, they generate expressive, multi-character narrations that capture subtle emotional nuances without the prohibitive time and cost of traditional studio recording. ## Embed v4 URL: https://reviuws.com/models/embed-v4 Provider: cohere Modality: embedding, vision License: closed Parameters: n/a Context window: 128k tokens Input price: $0.12/M tokens Output price: n/a Community rating: not yet rated Released: 2025-04-15 Summary: Embed v4 generates unified embeddings for text, images and mixed documents (like PDFs with charts), supporting a 128k token context. It's designed for enterprise search and RAG over multimodal content. It supports multiple languages. Typical use case: Used for multimodal semantic search, document retrieval, and enterprise knowledge bases. Good for RAG systems ingesting mixed media content. ## ERNIE 4.5 URL: https://reviuws.com/models/ernie-4-5 Provider: baidu Modality: text, vision, code License: open Parameters: 424B Context window: 128k tokens Input price: $0.55/M tokens Output price: $2.2/M tokens Community rating: not yet rated Released: 2025-06-30 Summary: ERNIE 4.5 is a family of mixture-of-experts models (up to 424B parameters) supporting text and multimodal understanding, released with open weights for several variants. It's designed for both cloud API and open deployment. It emphasizes strong Chinese and English performance. Typical use case: Used for enterprise chat, multimodal search, and Chinese-language applications. Suited for organizations needing strong bilingual (Chinese/English) capability. ## FLUX.1 [schnell] URL: https://reviuws.com/models/flux-1-schnell Provider: black forest labs Modality: image License: open Parameters: 12B Context window: n/a Input price: n/a Output price: n/a Community rating: not yet rated Released: 2024-08-01 Summary: FLUX.1 [schnell] is a distilled 12B parameter text-to-image model optimized for very fast generation (1-4 steps), released under Apache 2.0. It offers strong quality despite its speed focus. It's popular for local and open-source deployment. Typical use case: Used for rapid prototyping of images and self-hosted image generation apps. Popular in the open-source diffusion community. ## FLUX.1.1 [pro] URL: https://reviuws.com/models/flux-1-1-pro Provider: black forest labs Modality: image License: closed Parameters: 12B Context window: n/a Input price: n/a Output price: n/a Community rating: not yet rated Released: 2024-10-03 Summary: FLUX1.1 [pro] is a 12B parameter diffusion transformer producing high-fidelity images from text prompts with improved speed over the original FLUX.1. It's available via API and partners like Replicate. It offers ultra mode for high-resolution output. Typical use case: Used for creative image generation, marketing assets, and design prototyping. Popular among developers building image-generation apps. ## FLUX.2 [dev] URL: https://reviuws.com/models/flux-2-dev Provider: black-forest-labs Modality: image License: source-available Parameters: 32B Context window: n/a Input price: n/a Output price: n/a Community rating: not yet rated Summary: FLUX 2 Dev is an advanced open-weights image generation model developed by Black Forest Labs, designed to deliver state-of-the-art visual quality and prompt adherence. Building upon its predecessor's strengths, this model offers developers and enterprises a highly customizable and self-hostable solution for large-scale image synthesis. It successfully bridges the gap between enterprise-grade performance and local deployment flexibility. Typical use case: A marketing agency self-hosts FLUX 2 Dev on their own cloud infrastructure to generate highly specific, high-resolution product placement and advertising imagery at scale. By hosting the model themselves, they can dynamically handle thousands of generation requests daily while keeping client data entirely private and avoiding costly per-image API subscription fees. ## FLUX.2 [pro] URL: https://reviuws.com/models/flux-2-pro Provider: black-forest-labs Modality: image License: closed Parameters: n/a Context window: n/a Input price: n/a Output price: n/a Community rating: not yet rated Summary: FLUX 2 Pro by Black Forest Labs is a top-tier text-to-image model engineered engineered specifically for elite-level photorealism and precise visual rendering. It excels at producing complex textures, realistic human anatomy, and highly legible embedded text with unmatched fidelity. This model represents a massive leap forward for professional creators who demand exceptional asset quality and prompt adherence. Typical use case: A commercial advertising agency utilizes FLUX 2 Pro to rapidly generate high-fidelity visual assets for an upcoming print campaign. By inputting detailed prompts that specify exact copy, lighting, and product placement, designers can produce production-ready promotional imagery with perfectly rendered typography, cutting down creative iteration cycles from weeks to hours. ## Gemini 2.5 Flash URL: https://reviuws.com/models/gemini-2-5-flash Provider: google Modality: text, vision, audio, code License: closed Parameters: n/a Context window: 1049k tokens Input price: $0.3/M tokens Output price: $2.5/M tokens Community rating: not yet rated Released: 2025-06-17 Summary: Gemini 2.5 Flash balances speed, cost and reasoning quality with a configurable thinking budget. It supports a 1M token context and full multimodal input. It's designed for high-volume production workloads. Typical use case: Used for chatbots, summarization, and multimodal apps requiring low latency at scale. Good for cost-sensitive production deployments. ## Gemini 2.5 Flash-Lite URL: https://reviuws.com/models/gemini-2-5-flash-lite Provider: google Modality: text, vision, code License: closed Parameters: n/a Context window: 1049k tokens Input price: $0.1/M tokens Output price: $0.4/M tokens Community rating: not yet rated Released: 2025-07-22 Summary: Gemini 2.5 Flash-Lite is optimized for high-throughput, latency-sensitive tasks at the lowest cost in the Gemini 2.5 family. It retains a 1M token context window. It targets classification and simple generation tasks at scale. Typical use case: Used for high-volume classification, tagging, and simple chat tasks. Suited for cost-constrained applications. ## Gemini 2.5 Pro URL: https://reviuws.com/models/gemini-2-5-pro Provider: google Modality: text, vision, audio, code License: closed Parameters: n/a Context window: 1049k tokens Input price: $1.25/M tokens Output price: $10/M tokens Community rating: not yet rated Released: 2025-03-25 Summary: Gemini 2.5 Pro is Google's flagship model with native multimodality and a 1M token context window, featuring built-in 'thinking' for complex reasoning. It leads on many coding and math benchmarks. It's available via the Gemini API and Vertex AI. Typical use case: Used for complex reasoning, coding agents, and large-document/video analysis. Suited for enterprise applications needing long context. ## Gemini 3.1 Flash Lite URL: https://reviuws.com/models/gemini-31-flash-lite Provider: google Modality: chat License: closed Parameters: n/a Context window: 1000k tokens Input price: $0.05/M tokens Output price: $0.2/M tokens Community rating: not yet rated Summary: Gemini 3.1 Flash Lite is Google’s highly optimized, ultra-low-cost model designed for high-throughput text processing and conversational tasks at massive scale. Built to deliver exceptionally low latency, it provides developers with a budget-friendly option for high-frequency chat interfaces and basic data manipulation. While it sacrifices some deep reasoning capabilities, it excels at keeping operational costs minimal for volume-heavy enterprise workflows. Typical use case: A global e-commerce platform integrates Gemini 3.1 Flash Lite to power its customer support triage system, processing hundreds of thousands of incoming chat queries every day. The model instantly classifies the intent of each user message, extracts critical metadata like order IDs, and generates a structured summary for human agents, resolving simple inquiries automatically while keeping API overhead virtually negligible. ## Gemini 3.1 Pro (Preview) URL: https://reviuws.com/models/gemini-31-pro Provider: google Modality: chat, code License: closed Parameters: n/a Context window: 1000k tokens Input price: $2/M tokens Output price: $12/M tokens Community rating: not yet rated Summary: Gemini 3.1 Pro is a highly advanced multimodal reasoning model from Google, specifically optimized for complex chat interactions and deep analytical coding. Equipped with an industry-leading 2-million-token context window, it effortlessly digests massive datasets and entire software architectures in a single prompt. This model serves as an elite assistant for developers and researchers who require comprehensive understanding over massive volumes of information. Typical use case: A software development firm can utilize Gemini 3.1 Pro to execute a seamless, automated migration of a legacy enterprise application to a modern microservices architecture. By feeding the entire legacy codebase, system logs, and architectural documentation into the model's massive 2M-token context window, engineers can obtain a fully refactored, production-ready codebase with all dependency conflicts resolved and comprehensive unit tests generated. ## Gemini 3.5 Flash URL: https://reviuws.com/models/gemini-35-flash Provider: google Modality: chat License: closed Parameters: n/a Context window: 1000k tokens Input price: $1.5/M tokens Output price: $9/M tokens Community rating: not yet rated Summary: Google's Gemini 3.5 Flash is a highly optimized, cost-effective model designed specifically for high-speed, high-volume conversational tasks and multi-turn agentic workflows. Built to deliver near-instantaneous responses, it excels at real-time data extraction and structured processing. It serves as the ultimate developer workhorse for applications requiring low latency and ultra-low operational costs. Typical use case: A global e-commerce enterprise deploys Gemini 3.5 Flash to power a network of customer service agents handling millions of inquiries daily. The model instantly analyzes incoming support tickets, extracts order IDs or tracking numbers, and autonomously executes API calls to update delivery statuses or issue refunds in real-time. This high-throughput capability drastically reduces customer wait times while keeping operational costs at a minimum. ## Gemini 3.5 Flash-Lite URL: https://reviuws.com/models/gemini-35-flash-lite Provider: google Modality: chat License: closed Parameters: n/a Context window: 1000k tokens Input price: $0.3/M tokens Output price: $2.5/M tokens Community rating: not yet rated Summary: Gemini 3.5 Flash-Lite is Google's most economical general-availability model, specifically optimized for high-volume agentic workflows and data processing. Offered at an ultra-low price point of $0.30 per million input tokens, it brings reliable performance to massive scalability demands. It serves as an ideal solution for organizations looking to deploy chat agents and translation services at a fraction of the cost of larger models. Typical use case: A global e-commerce platform leverages Gemini 3.5 Flash-Lite to run its customer service chatbots and translate millions of user reviews in real time. Because of its low cost, the platform can handle millions of customer interactions daily, routing queries to agents and instantly translating product feedback across dozens of languages without breaking the budget. ## Gemini 3.6 Flash URL: https://reviuws.com/models/gemini-36-flash Provider: google Modality: chat, code License: closed Parameters: n/a Context window: 1000k tokens Input price: $1.5/M tokens Output price: $7.5/M tokens Community rating: not yet rated Summary: Gemini 3.6 Flash is Google's premier speed-optimized model, engineered to deliver rapid responses without sacrificing deep intelligence or coding capabilities. Leveraging Google's robust search grounding, it excels at providing highly accurate, real-time information for dynamic conversational and programming workflows. It offers an accessible entry point for developers with a generous free tier and competitive pay-as-you-go pricing. Typical use case: A software development team can integrate Gemini 3.6 Flash into their continuous integration pipeline as a real-time debugging assistant. When a build fails, the model can instantly analyze the error logs, query the web via search grounding to find similar issues, and write the corrected code block directly in the team's chat channel, allowing developers to resolve deployment blockers in seconds. ## Gemini Embedding 2 URL: https://reviuws.com/models/gemini-embedding-2 Provider: google Modality: embedding License: closed Parameters: n/a Context window: n/a Input price: $0.2/M tokens Output price: n/a Community rating: not yet rated ## Gemma 2 9B URL: https://reviuws.com/models/gemma-2-9b Provider: google Modality: text License: open Parameters: 9B Context window: 8k tokens Input price: n/a Output price: n/a Community rating: not yet rated Released: 2024-06-27 Summary: Gemma 2 9B is a dense open-weight model trained with knowledge distillation from larger models, offering strong performance for its size. It supports an 8k context window. It's released under the Gemma license for broad commercial use. Typical use case: Used for lightweight chatbots, edge deployment, and fine-tuning experiments. Popular in the open-source community for research and small-scale production. ## Gemma 3 URL: https://reviuws.com/models/gemma-3 Provider: google Modality: text, vision License: open Parameters: 27B Context window: 128k tokens Input price: n/a Output price: n/a Community rating: not yet rated Released: 2025-03-12 Summary: Gemma 3 is a family of open-weight models (1B-27B) with multimodal and multilingual support, derived from Gemini research. It supports a 128k context window. It's released under a permissive open license for commercial use. Typical use case: Used for on-device and self-hosted deployments, research, and fine-tuning. Suited for developers wanting open weights with strong performance. ## GLM-4.5 URL: https://reviuws.com/models/glm-4-5 Provider: zhipu ai Modality: text, code License: open Parameters: 355B Context window: 128k tokens Input price: $0.6/M tokens Output price: $2.2/M tokens Community rating: not yet rated Released: 2025-07-28 Summary: GLM-4.5 is a 355B parameter (32B active) mixture-of-experts model designed for agentic, reasoning and coding tasks with a hybrid thinking mode. It's open-weight and competitive with leading proprietary models on agentic benchmarks. It supports a 128k context window. Typical use case: Used for coding agents, tool-use pipelines and research applications. Suited for developers wanting an open frontier-class Chinese LLM. ## GPT Image 2 URL: https://reviuws.com/models/gpt-image-2 Provider: openai Modality: image License: closed Parameters: n/a Context window: n/a Input price: n/a Output price: n/a Community rating: not yet rated Summary: GPT Image 2 is OpenAI's state-of-the-art visual model designed for advanced image generation and editing directly through an API. It allows developers and creators to synthesize high-quality graphics from textual prompts and perform complex modifications on existing images. Integrated seamlessly into the OpenAI ecosystem, it offers a scalable solution for modern digital design pipelines. Typical use case: An e-commerce platform can integrate GPT Image 2 to dynamically generate and customize product backgrounds for their online storefront. By automatically replacing generic studio backdrops with contextual scenes (like placing a hiking boot on a mountain trail) based on user search queries, the system enhances visual appeal and personalization, leading to increased customer engagement and conversion rates. ## GPT-4.1 URL: https://reviuws.com/models/gpt-4-1 Provider: openai Modality: text, vision, code License: closed Parameters: n/a Context window: 1048k tokens Input price: $2/M tokens Output price: $8/M tokens Community rating: not yet rated Released: 2025-04-14 Summary: GPT-4.1 offers a 1M token context window with major improvements in coding and instruction following over GPT-4o. It's available in standard, mini and nano sizes. It targets developers building agentic and coding workflows. Typical use case: Well suited for large codebase analysis, long-document QA, and agentic coding assistants. Also used for enterprise workflows requiring long context. ## GPT-4.1 mini URL: https://reviuws.com/models/gpt-4-1-mini Provider: openai Modality: text, vision, code License: closed Parameters: n/a Context window: 1048k tokens Input price: $0.4/M tokens Output price: $1.6/M tokens Community rating: not yet rated Released: 2025-04-14 Summary: GPT-4.1 mini balances cost and capability, offering a 1M token context window at a fraction of GPT-4.1's price. It maintains strong coding and reasoning performance. It's targeted at cost-sensitive production workloads. Typical use case: Used for chatbots, document processing and coding assistants where cost matters. Good fit for high-throughput agentic pipelines. ## GPT-4o URL: https://reviuws.com/models/gpt-4o Provider: openai Modality: text, vision, audio License: closed Parameters: n/a Context window: 128k tokens Input price: $2.5/M tokens Output price: $10/M tokens Community rating: not yet rated Released: 2024-05-13 Summary: GPT-4o is OpenAI's natively multimodal model supporting text, image and audio inputs with fast, low-latency responses. It powers ChatGPT and the API with a 128k context window. It offers strong reasoning, coding, and multilingual capabilities at a lower cost than GPT-4 Turbo. Typical use case: Used for chat assistants, multimodal document/image understanding, and voice-enabled applications. Suited for production apps needing balance of cost, latency and quality. ## GPT-4o mini URL: https://reviuws.com/models/gpt-4o-mini Provider: openai Modality: text, vision License: closed Parameters: n/a Context window: 128k tokens Input price: $0.15/M tokens Output price: $0.6/M tokens Community rating: not yet rated Released: 2024-07-18 Summary: GPT-4o mini is a cost-efficient version of GPT-4o designed for high-volume tasks. It supports text and vision inputs with a 128k context window. It offers strong performance for its price point, replacing GPT-3.5 Turbo in many use cases. Typical use case: Ideal for high-volume, low-latency applications like customer support bots and simple classification tasks. Also used for cost-sensitive multimodal apps. ## GPT-5.4 mini URL: https://reviuws.com/models/gpt-54-mini Provider: openai Modality: chat License: closed Parameters: n/a Context window: 400k tokens Input price: $0.75/M tokens Output price: $4.5/M tokens Community rating: not yet rated Summary: GPT-5.4 mini is OpenAI's highly efficient and cost-effective model optimized for high-volume automated tasks. It excels in driving specialized coding sub-agents and managing repetitive, large-scale chat workflows without compromising on speed. This model represents an ideal balance of lightweight agility and robust developer-focused capabilities, making it a powerful tool for scaling enterprise pipelines. Typical use case: A software engineering department can deploy GPT-5.4 mini to power a fleet of autonomous development sub-agents that continuously monitor code repositories. These agents can automatically analyze pull requests, generate boilerplate unit tests, and flag immediate syntax errors across hundreds of files simultaneously. This allows the engineering team to maintain high code quality and continuous integration speeds at a fraction of the cost of larger models. ## GPT-5.4 nano URL: https://reviuws.com/models/gpt-54-nano Provider: openai Modality: chat License: closed Parameters: n/a Context window: 400k tokens Input price: $0.2/M tokens Output price: $1.25/M tokens Community rating: not yet rated Summary: GPT-5.4 nano is OpenAI's highly optimized, lightweight model designed for high-throughput, low-latency text processing tasks. Positioned as the most cost-effective option in the GPT-5.4 lineup, it excels at streamlined text classification, data extraction, and content ranking. It is ideal for developers requiring rapid-fire conversational responses and high-volume utility operations without the overhead of larger models. Typical use case: A high-volume e-commerce platform can integrate GPT-5.4 nano to screen millions of incoming user reviews and customer support tickets in real-time. The model automatically classifies the sentiment of reviews, extracts specific product defects mentioned in the text, and ranks support tickets by urgency to route them to human agents instantly, dramatically reducing response times while maintaining minimal operational costs. ## GPT-5.5 URL: https://reviuws.com/models/gpt-55 Provider: openai Modality: chat, code License: closed Parameters: n/a Context window: 400k tokens Input price: $5/M tokens Output price: $30/M tokens Community rating: 5.0/5 from 1 reviews Summary: GPT-5.5 represents OpenAI's next-generation frontier model, specifically optimized for highly complex reasoning, advanced coding tasks, and multi-step analytical workflows. Operating primarily across chat and code modalities, it delivers unprecedented depth in problem-solving and logical synthesis. It serves as the go-to engine for developers and researchers requiring rigorous logical validation and sophisticated instruction-following capabilities. Typical use case: A software engineering team at a financial institution can deploy GPT-5.5 to automate the refactoring and security auditing of legacy COBOL databases into modern, secure Python microservices. The model acts as an expert co-developer, translating highly complex business logic, identifying hidden edge-case vulnerabilities, and draft-testing the migration in real-time, drastically reducing transit times and human oversight errors. ## GPT-5.6 Luna URL: https://reviuws.com/models/gpt-56-luna Provider: openai Modality: chat License: closed Parameters: n/a Context window: 1050k tokens Input price: $1/M tokens Output price: $6/M tokens Community rating: not yet rated Summary: GPT-5.6 Luna by OpenAI is a highly efficient chat-based model designed specifically to handle high-volume, cost-sensitive workloads without compromising on modern intelligence standards. Featuring a massive 1.05-million-token context window, it allows organizations to process vast datasets and complex conversational histories at a highly competitive price point of $1 per million input tokens and $6 per million output tokens. This model effectively bridges the gap between deep contextual understanding and budget-conscious enterprise scale. Typical use case: A global enterprise customer support system can deploy GPT-5.6 Luna to manage millions of daily customer inquiries, feeding entire user accounts' multi-year interaction histories directly into the prompt. Because of the model's huge 1.05M context window, virtual agents can reference previous technical tickets, purchase histories, and long-form service agreements in a single chat session to resolve complex issues instantly. The aggressive pricing of $1/$6 per million tokens ensures the enterprise can maintain this hyper-personalized, high-throughput automation at a fraction of the cost of previous model generations. ## GPT-5.6 Sol URL: https://reviuws.com/models/gpt-56-sol Provider: openai Modality: chat, code License: closed Parameters: n/a Context window: 1050k tokens Input price: $5/M tokens Output price: $30/M tokens Community rating: not yet rated Summary: GPT-5.6 Sol represents OpenAI's absolute pinnacle of machine intelligence, specifically engineered to tackle ultra-complex reasoning challenges and autonomous agentic workflows. By prioritizing deep cognitive processing over raw speed, it excels at decomposing multi-step problems that leave lesser models stalled. While it carries a premium price tag and operates with higher latency, its unmatched analytical capabilities make it the definitive choice for enterprise-level decision making and advanced software engineering. Typical use case: A pharmaceutical research enterprise deploys GPT-5.6 Sol to orchestrate autonomous multi-agent pipelines for drug discovery. The model acts as a master coordinator, analyzing vast biomedical datasets to identify target proteins, writing and self-correcting python code to run molecular simulations, and autonomously generating comprehensive regulatory compliance documentation. ## GPT-5.6 Terra URL: https://reviuws.com/models/gpt-56-terra Provider: openai Modality: chat, code License: closed Parameters: n/a Context window: 1050k tokens Input price: $2.5/M tokens Output price: $15/M tokens Community rating: not yet rated Summary: GPT-5.6 Terra is a balanced powerhouse within OpenAI's latest model family, engineered to deliver a cost-effective blend of advanced intelligence and efficiency for text and code tasks. Armed with an impressive 1.05-million token context window, it allows developers to process vast amounts of information simultaneously without incurring prohibitive costs. It represents a highly competitive middle-ground option for enterprise-grade automation and deep analytical processing. Typical use case: A software development firm can leverage GPT-5.6 Terra to automate the auditing and refactoring of legacy codebases. By feeding entire repositories and years of documentation into its massive 1.05M context window, engineering teams can instantly identify architectural flaws, generate comprehensive unit tests, and receive precise modernization recommendations. This significantly accelerates the development lifecycle while maintaining a highly predictable budget due to the model's cost-efficient pricing structure. ## GPT-Realtime 2.1 URL: https://reviuws.com/models/gpt-realtime-21 Provider: openai Modality: audio, speech License: closed Parameters: n/a Context window: n/a Input price: n/a Output price: n/a Community rating: not yet rated Summary: GPT-Realtime 2.1 by OpenAI represents a major leap in conversational AI, offering native speech-to-speech interaction with extremely low latency. Combining advanced reasoning with real-time tool execution, it enables fluid, context-aware verbal interactions. This model sets a new standard for natural and highly interactive voice-based applications. Typical use case: A global logistics company integrates GPT-Realtime 2.1 into their telephone dispatch system. When a delivery driver encounters an unexpected roadblock, they can speak naturally to the AI assistant, which instantly analyzes the traffic data, uses a routing API tool to calculate an alternative path, and verbally guides the driver through the new directions without any disruptive delay. ## Granite 3.3 8B URL: https://reviuws.com/models/granite-3-3-8b Provider: ibm Modality: text, code License: open Parameters: 8B Context window: 128k tokens Input price: $0.2/M tokens Output price: $0.2/M tokens Community rating: not yet rated Released: 2025-04-16 Summary: Granite 3.3 8B is IBM's open-weight model tuned for enterprise use cases like RAG, function calling and fill-in-the-middle code, with a 128k context window. It's released under Apache 2.0 with a focus on transparency and data provenance. It's available on watsonx and Hugging Face. Typical use case: Used for enterprise RAG applications, code assistance, and regulated-industry deployments needing transparent training data. Good for on-prem enterprise AI. ## Grok 2 URL: https://reviuws.com/models/grok-2 Provider: xai Modality: text, image License: closed Parameters: n/a Context window: 131k tokens Input price: $2/M tokens Output price: $10/M tokens Community rating: not yet rated Released: 2024-08-13 Summary: Grok 2 improved reasoning, coding and multilingual capabilities over Grok 1.5, and introduced image generation through a FLUX.1 partnership. It's available on X Premium and via API. It has a 131k context window. Typical use case: Used within X for chat and image generation, and via API for developer applications. Was xAI's flagship before Grok 3/4. ## Grok 3 URL: https://reviuws.com/models/grok-3 Provider: xai Modality: text, code License: closed Parameters: n/a Context window: 131k tokens Input price: $3/M tokens Output price: $15/M tokens Community rating: not yet rated Released: 2025-02-19 Summary: Grok 3 introduced 'Think' mode for extended reasoning and DeepSearch for web-integrated answers, with a 131k context window. It was trained on xAI's Colossus supercomputer. It remains available via API alongside Grok 4. Typical use case: Used for chat, reasoning tasks, and search-augmented queries within the X ecosystem. Also used in third-party apps via API. ## Grok 4 URL: https://reviuws.com/models/grok-4 Provider: xai Modality: chat License: closed Parameters: n/a Context window: 256k tokens Input price: $3/M tokens Output price: $12/M tokens Community rating: not yet rated Summary: Grok 4 is xAI's flagship conversational AI, engineered for advanced analytical reasoning and deep conceptual synthesis. Building on its predecessor's strengths, it utilizes real-time web access and direct data streams to deliver exceptionally current responses. The model represents a major step forward in handling complex, multi-step logical tasks through a streamlined chat interface. Typical use case: A quantitative researcher can utilize Grok 4 to conduct live sentiment analysis and gather immediate context on sudden market anomalies. By querying the model, they can cross-reference real-time social media trends on X with breaking financial news to generate an instant, comprehensive market report before traditional data sources have even updated. ## Grok 4.3 URL: https://reviuws.com/models/grok-43 Provider: xai Modality: chat License: closed Parameters: n/a Context window: 1000k tokens Input price: $1.25/M tokens Output price: $2.5/M tokens Community rating: not yet rated Summary: Grok 4.3 by xAI is a highly efficient conversational model designed to handle massive datasets with its expansive 1-million-token context window. It offers a highly competitive pricing structure, particularly for prompts under 200k tokens, making deep-text analysis far more accessible. This model combines xAI's signature real-time information access with the capacity to process entire libraries of information in a single query. Typical use case: A financial research firm uses Grok 4.3 to upload decades of annual reports, earnings call transcripts, and market regulations for a specific industry all at once. The model quickly synthesizes the historical data, identifies long-term market trends, and drafts a comprehensive investment thesis, saving analysts days of manual cross-referencing and document chunking. ## Grok 4.5 URL: https://reviuws.com/models/grok-45 Provider: xai Modality: chat, code License: closed Parameters: n/a Context window: 500k tokens Input price: $2/M tokens Output price: $6/M tokens Community rating: not yet rated ## Ideogram 3.0 URL: https://reviuws.com/models/ideogram-3-0 Provider: ideogram Modality: image License: closed Parameters: n/a Context window: n/a Input price: n/a Output price: n/a Community rating: not yet rated Released: 2025-03-26 Summary: Ideogram 3.0 generates high-quality images with especially accurate in-image text rendering and realistic styles. It offers style reference and consistent character features. It's accessible via web app and API. Typical use case: Used for marketing graphics, logos with text, and social media content creation. Popular among designers needing reliable text-in-image generation. ## Imagen 4 URL: https://reviuws.com/models/imagen-4 Provider: google Modality: image License: closed Parameters: n/a Context window: n/a Input price: n/a Output price: n/a Community rating: not yet rated Summary: Imagen 4 is Google's premier flagship text-to-image model, engineered to deliver class-leading photorealism and precise prompt adherence. It excels at rendering complex details, including intricate text overlays and diverse artistic styles, with exceptional clarity. Fully integrated into Google’s enterprise ecosystem, it offers businesses a highly scalable and secure solution for high-fidelity visual generation. Typical use case: A global marketing agency can leverage Imagen 4 to rapidly prototype and generate high-resolution advertising assets for a multi-channel campaign. By inputting detailed prompts, designers can instantly produce photorealistic product mockups situated in various lifestyle settings, complete with crisp, embedded brand slogans, significantly accelerating the creative brainstorming phase and reducing reliance on expensive physical photo shoots. ## Jamba 1.6 URL: https://reviuws.com/models/jamba-1-6 Provider: ai21 Modality: text, code License: open Parameters: 398B Context window: 256k tokens Input price: $0.2/M tokens Output price: $0.4/M tokens Community rating: not yet rated Released: 2025-03-05 Summary: Jamba 1.6 combines Mamba and Transformer architectures in a mixture-of-experts design, offering a 256k context window with efficient long-context inference. It comes in Large and Mini sizes. It's open-weight under the Jamba Open Model License. Typical use case: Used for long-document processing, RAG, and enterprise applications needing efficient long-context handling. Suited for cost-sensitive large-context workloads. ## Kimi K2 URL: https://reviuws.com/models/kimi-k2 Provider: moonshot ai Modality: text, code License: open Parameters: 1000B Context window: 128k tokens Input price: $0.6/M tokens Output price: $2.5/M tokens Community rating: not yet rated Released: 2025-07-11 Summary: Kimi K2 is a 1T parameter (32B active) mixture-of-experts model trained for agentic tool-use and coding, released with open weights under a modified MIT license. It achieves state-of-the-art results among open models on agentic benchmarks. It supports a 128k context window. Typical use case: Used for autonomous coding agents, tool-calling workflows, and research into large-scale MoE training. Popular among developers wanting an open GPT-4-class agentic model. ## Llama 3.3 70B URL: https://reviuws.com/models/llama-3-3-70b Provider: meta Modality: text License: open Parameters: 70B Context window: 128k tokens Input price: n/a Output price: n/a Community rating: not yet rated Released: 2024-12-06 Summary: Llama 3.3 70B delivers performance comparable to Llama 3.1 405B at a fraction of the size, focused purely on text tasks. It supports a 128k context window and multiple languages. It's tuned for instruction following and dialogue. Typical use case: Used for cost-efficient self-hosted chatbots and text generation pipelines. Popular for fine-tuning on domain-specific tasks. ## Llama 4 405B URL: https://reviuws.com/models/llama-4-405b Provider: meta Modality: chat, code License: open Parameters: 405B Context window: 128k tokens Input price: n/a Output price: n/a Community rating: not yet rated Summary: Llama 4 405B represents Meta's frontier-class open-weights model, offering state-of-the-art general reasoning, coding, and chat capabilities. Designed to compete directly with leading proprietary models, it provides enterprises with the flexibility of local deployment or cloud-based API access. Its massive scale makes it a premier choice for complex, high-security AI workloads that require full model ownership. Typical use case: An enterprise software development firm integrates Llama 4 405B into their secure on-premise infrastructure to automate code generation, conduct deep vulnerability audits, and generate internal technical documentation. Because they handle highly sensitive proprietary client codebases, using an open-weights model of this scale allows them to keep all data within their private cloud while achieving performance comparable to closed proprietary APIs. ## Llama 4 70B URL: https://reviuws.com/models/llama-4-70b Provider: meta Modality: chat License: open Parameters: 70B Context window: 128k tokens Input price: n/a Output price: n/a Community rating: not yet rated Summary: Meta's Llama 4 70B is the premier sweet spot for open-weights self-hosting, masterfully balancing state-of-the-art conversational quality with a highly manageable computational footprint. Optimized for chat and complex reasoning, it delivers near-proprietary level performance while running efficiently on accessible enterprise hardware. It serves as the go-to option for organizations demanding complete data sovereignty without sacrificing advanced model intelligence. Typical use case: A major financial institution deploys Llama 4 70B on-premises to power an internal virtual assistant that helps analysts synthesize complex market reports, query legacy databases, and generate compliant client communications. By hosting the 70B parameter model locally on a dual-GPU node, the company ensures absolute data privacy and compliance with strict financial regulations. The model's advanced chat capabilities allow it to act as a highly competent, context-aware writing partner that understands nuanced financial terminology without leaking sensitive proprietary data to third-party APIs. ## Llama 4 Maverick URL: https://reviuws.com/models/llama-4-maverick Provider: meta Modality: chat, code License: source-available Parameters: 400B Context window: 1000k tokens Input price: n/a Output price: n/a Community rating: not yet rated Released: 2025-04-05 ## Llama 4 Scout URL: https://reviuws.com/models/llama-4-scout Provider: meta Modality: chat License: source-available Parameters: 109B Context window: 10000k tokens Input price: n/a Output price: n/a Community rating: not yet rated Released: 2025-04-05 ## Llama-3.1-Nemotron-Ultra-253B URL: https://reviuws.com/models/llama-3-1-nemotron-ultra-253b Provider: nvidia Modality: text, code License: open Parameters: 253B Context window: 128k tokens Input price: n/a Output price: n/a Community rating: not yet rated Released: 2025-04-07 Summary: This model is a pruned and post-trained derivative of Llama 3.1 405B, optimized by Nvidia for reasoning tasks with toggleable reasoning mode. It supports a 128k context window. It's designed for high accuracy with reduced inference cost versus the original. Typical use case: Used for enterprise reasoning tasks and as a base for further fine-tuning. Suited for teams wanting Llama-derived open weights with efficiency gains. ## Luma Ray3 URL: https://reviuws.com/models/luma-ray3 Provider: luma ai Modality: video License: closed Parameters: n/a Context window: n/a Input price: n/a Output price: n/a Community rating: not yet rated Released: 2025-06-17 Summary: Ray3 is Luma's reasoning-capable video model, the first to generate HDR video with improved physics and motion coherence. It supports text and image-to-video generation. It's available via Luma's Dream Machine platform and API. Typical use case: Used for professional video content creation, VFX previsualization, and marketing videos. Suited for creators needing HDR-quality output. ## MiniMax M1 URL: https://reviuws.com/models/minimax-m1 Provider: minimax Modality: text, code License: open Parameters: 456B Context window: 1000k tokens Input price: $0.4/M tokens Output price: $2.2/M tokens Community rating: not yet rated Released: 2025-06-16 Summary: MiniMax M1 is a 456B parameter (45.9B active) hybrid-attention MoE reasoning model supporting up to 1M tokens of context, trained with efficient large-scale reinforcement learning. It's released under Apache 2.0. It targets long-context agentic and reasoning tasks. Typical use case: Used for long-document analysis, extended agentic workflows, and research on efficient long-context RL training. Suited for tasks needing million-token context. ## Mistral Large 2 URL: https://reviuws.com/models/mistral-large-2 Provider: mistral Modality: text, code License: open Parameters: 123B Context window: 128k tokens Input price: $2/M tokens Output price: $6/M tokens Community rating: not yet rated Released: 2024-07-24 Summary: Mistral Large 2 is a 123B parameter dense model with strong multilingual, coding and reasoning capabilities and a 128k context window. It's available under a research license with commercial API access. It focuses on efficiency and low hallucination rates. Typical use case: Used for enterprise chat, coding assistance, and multilingual applications. Suited for teams wanting an open-weight alternative to GPT-4 class models. ## Mistral Large 3 URL: https://reviuws.com/models/mistral-large-3 Provider: mistral Modality: chat, code License: open Parameters: 675B Context window: n/a Input price: n/a Output price: n/a Community rating: not yet rated Released: 2025-12-02 Summary: Mistral Large 3 is Mistral AI's premier European frontier model, engineered to deliver top-tier reasoning, advanced coding capabilities, and highly sophisticated multilingual performance. It serves as a formidable, sovereign alternative to US-centric models, excelling in complex analytical tasks and enterprise-grade chat interactions. Its design prioritizes strong logical deduction and deep comprehension across major global languages. Typical use case: A European retail group can deploy Mistral Large 3 to run a centralized, multilingual customer service and IT automation hub. The model can simultaneously interpret complex user queries in German, French, and Spanish, debug internal APIs on the fly, and draft precise, localized responses or code fixes without needing translating middleware. ## Mistral Small 3.2 URL: https://reviuws.com/models/mistral-small-3-2 Provider: mistral Modality: text, vision, code License: open Parameters: 24B Context window: 128k tokens Input price: $0.1/M tokens Output price: $0.3/M tokens Community rating: not yet rated Released: 2025-06-20 Summary: Mistral Small 3.2 is a 24B parameter open-weight model with multimodal support, tuned for instruction following and reduced repetition errors. It offers a 128k context window under Apache 2.0 license. It's designed for efficient local and cloud deployment. Typical use case: Used for on-device assistants, cost-sensitive chatbots, and vision-language tasks. Good for developers wanting fully open commercial-use weights. ## Nano Banana 2 (Gemini 3.1 Flash Image) URL: https://reviuws.com/models/nano-banana-2 Provider: google Modality: image License: closed Parameters: n/a Context window: n/a Input price: $0.5/M tokens Output price: $60/M tokens Community rating: not yet rated Summary: Nano Banana 2 (Gemini 3.1 Flash Image) is Google's highly efficient model designed for rapid image generation and editing. It offers an exceptionally cost-effective solution, costing a fraction of a cent per image for both standard and high-resolution 4K outputs. This model is optimized for developers and businesses requiring high-throughput visual asset creation without high latency or premium costs. Typical use case: An e-commerce platform can integrate Nano Banana 2 into its seller dashboard to automate product photography editing at scale. When a merchant uploads a raw product photo, the model can instantly remove the background, generate a realistic lifestyle setting based on the product type, and upscale the final asset to a crisp 4K resolution for a polished, ready-to-publish storefront listing. ## Nano Banana 2 (Gemini Image) URL: https://reviuws.com/models/nano-banana-2-gemini-image Provider: google Modality: image License: closed Parameters: n/a Context window: n/a Input price: n/a Output price: n/a Community rating: not yet rated Summary: Nano Banana 2 (Gemini Image) is an advanced image generation model by Google that excels at converting complex textual prompts into high-fidelity visuals. It sets a high benchmark in the industry with its standout ability to render crisp, accurate text within images and maintain identical character features across diverse scenes. This makes the model particularly suited for narrative-driven and commercial design workflows. Typical use case: An independent comic book creator can utilize Nano Banana 2 to storyboard and generate complete scenes featuring a recurring protagonist. By exploiting the model's excellent character consistency, the creator can place their hero in various environments and action poses without losing their distinct facial features. Furthermore, they can generate panels that naturally incorporate readable background text, such as neon signs and newspapers, directly in the original render. ## Nemotron-4 340B URL: https://reviuws.com/models/nemotron-4-340b Provider: nvidia Modality: text, code License: open Parameters: 340B Context window: 4k tokens Input price: n/a Output price: n/a Community rating: not yet rated Released: 2024-06-14 Summary: Nemotron-4 340B is a dense 340B parameter model optimized to generate high-quality synthetic training data for other LLMs. It's released under the NVIDIA Open Model License. It also functions as a capable general chat/instruct model. Typical use case: Used for generating synthetic fine-tuning datasets and as a reward model in RLHF pipelines. Also usable as a general instruction-following model. ## o3 URL: https://reviuws.com/models/o3 Provider: openai Modality: text, code License: closed Parameters: n/a Context window: 200k tokens Input price: $2/M tokens Output price: $8/M tokens Community rating: not yet rated Released: 2025-04-16 Summary: o3 is a reasoning-focused model that uses extended chain-of-thought to solve complex math, science and coding problems. It supports tool use during reasoning. It's aimed at the most demanding analytical tasks. Typical use case: Used for advanced research, complex coding tasks, and multi-step agentic reasoning. Suited for scientific and mathematical problem solving. ## o4-mini URL: https://reviuws.com/models/o4-mini Provider: openai Modality: text, code License: closed Parameters: n/a Context window: 200k tokens Input price: $1.1/M tokens Output price: $4.4/M tokens Community rating: not yet rated Released: 2025-04-16 Summary: o4-mini delivers strong reasoning performance at lower cost and latency than o3. It supports tool use and is optimized for math and coding. It's designed for high-volume reasoning tasks. Typical use case: Used for coding agents and math tutoring where cost-efficient reasoning is needed. Good for high-throughput agentic applications. ## OLMo 2 32B URL: https://reviuws.com/models/olmo-2-32b Provider: allen institute for ai Modality: text License: open Parameters: 32B Context window: 4k tokens Input price: n/a Output price: n/a Community rating: not yet rated Released: 2025-03-13 Summary: OLMo 2 32B is the largest model in Allen Institute's fully open OLMo family, released with training data, code, and checkpoints for full reproducibility. It's competitive with similarly sized open-weight models like Qwen2.5. It's designed to advance open science in LLM research. Typical use case: Used for academic research into LLM training dynamics and as a transparent base for fine-tuning. Suited for researchers needing full training reproducibility. ## Phi-4 URL: https://reviuws.com/models/phi-4 Provider: microsoft Modality: text, code License: open Parameters: 14B Context window: 16k tokens Input price: n/a Output price: n/a Community rating: not yet rated Released: 2024-12-12 Summary: Phi-4 is a 14B parameter dense model trained with a focus on data quality, achieving strong performance on reasoning and math benchmarks relative to its size. It's released under MIT license. It's designed for efficient deployment on limited hardware. Typical use case: Used for on-device reasoning assistants and cost-efficient enterprise apps. Popular for research into small-model efficiency. ## Phi-4-multimodal URL: https://reviuws.com/models/phi-4-multimodal Provider: microsoft Modality: text, vision, audio License: open Parameters: 5.6B Context window: 128k tokens Input price: n/a Output price: n/a Community rating: not yet rated Released: 2025-02-26 Summary: Phi-4-multimodal is a 5.6B parameter model that unifies text, image and audio understanding in a single small model with a 128k context window. It's released under MIT license. It targets efficient multimodal inference on edge devices. Typical use case: Used for on-device multimodal assistants and speech-to-text-integrated apps. Suited for resource-constrained multimodal deployments. ## Pixtral Large URL: https://reviuws.com/models/pixtral-large Provider: mistral Modality: text, vision License: open Parameters: 124B Context window: 128k tokens Input price: $2/M tokens Output price: $6/M tokens Community rating: not yet rated Released: 2024-11-18 Summary: Pixtral Large is a 124B parameter model combining a 1B vision encoder with Mistral Large 2's text backbone, supporting a 128k context window. It excels at document, chart and diagram understanding. It's released under the Mistral Research License. Typical use case: Used for document AI, chart/diagram QA, and multimodal enterprise applications. Suited for teams needing strong visual reasoning with open weights. ## Qwen 3 235B URL: https://reviuws.com/models/qwen-3-235b Provider: alibaba Modality: chat, code License: open Parameters: 235B Context window: 128k tokens Input price: n/a Output price: n/a Community rating: not yet rated Summary: Alibaba's Qwen 3 235B is an exceptionally powerful open-weight model optimized for elite multilingual understanding and advanced programming tasks. Built on a massive 235-billion parameter architecture, it offers a highly capable open alternative to proprietary giants. It excels at complex reasoning, making it ideal for sophisticated chat and technical workflows. Typical use case: An international software firm can leverage Qwen 3 235B as a centralized developer assistant. The model can simultaneously interpret complex technical requirements from global clients in multiple languages and automatically generate high-quality, production-ready code to accelerate dev pipelines. ## Qwen2.5-VL-72B URL: https://reviuws.com/models/qwen2-5-vl-72b Provider: alibaba Modality: text, vision License: open Parameters: 72B Context window: 128k tokens Input price: n/a Output price: n/a Community rating: not yet rated Released: 2025-01-28 Summary: Qwen2.5-VL-72B provides advanced visual understanding including document parsing, video comprehension and object grounding. It supports a 128k context window and is Apache 2.0 licensed for the smaller variants. It excels in agentic visual tasks like screen understanding. Typical use case: Used for document AI, video analysis, and GUI agent applications. Suited for developers needing strong open vision-language capability. ## Qwen3-235B-A22B URL: https://reviuws.com/models/qwen3-235b-a22b Provider: alibaba Modality: text, code License: open Parameters: 235B Context window: 128k tokens Input price: n/a Output price: n/a Community rating: not yet rated Released: 2025-04-29 Summary: Qwen3-235B-A22B is a mixture-of-experts model with 235B total/22B active parameters, supporting seamless switching between thinking and non-thinking modes. It supports 100+ languages and a 128k context window. It's released under Apache 2.0. Typical use case: Used for multilingual chat, agentic tool use, and coding tasks. Popular for self-hosted deployments requiring open licensing. ## Qwen3.5-397B-A17B URL: https://reviuws.com/models/qwen-35-397b-a17b Provider: alibaba Modality: chat, code License: open Parameters: 397B Context window: n/a Input price: n/a Output price: n/a Community rating: not yet rated Released: 2026-02-17 ## Reka Flash 3 URL: https://reviuws.com/models/reka-flash-3 Provider: reka ai Modality: text, vision License: open Parameters: 21B Context window: 128k tokens Input price: n/a Output price: n/a Community rating: not yet rated Released: 2025-03-11 Summary: Reka Flash 3 is a 21B parameter multimodal model with reasoning capabilities, trained via reinforcement learning and released under Apache 2.0. It supports a 128k context window. It's designed to be a strong compact alternative to larger closed models. Typical use case: Used for on-device or self-hosted multimodal reasoning tasks. Suited for developers wanting a smaller open reasoning model. ## Runway Gen-4 URL: https://reviuws.com/models/runway-gen-4 Provider: runway Modality: video License: closed Parameters: n/a Context window: n/a Input price: n/a Output price: n/a Community rating: not yet rated Summary: Runway Gen-4 represents the pinnacle of production-grade AI video generation, offering filmmakers and creators unprecedented control over physics, camera movement, and character consistency. Building on its predecessors, this model delivers photorealistic high-fidelity outputs suitable for commercial workflows. With its advanced steering capabilities, it bridges the gap between generative AI and traditional cinematic pipelines. Typical use case: An advertising agency utilizes Runway Gen-4 to rapidly prototype and generate high-end visual effects for a car commercial. By leveraging the model's precise camera controls and motion brushes, the team generates hyper-realistic shots of a vehicle navigating a rain-slicked mountain pass, eliminating the need for expensive on-location scouting and physical stunt drivers during the pitch phase. ## SDXL URL: https://reviuws.com/models/sdxl Provider: stability Modality: image License: open Parameters: n/a Context window: n/a Input price: n/a Output price: n/a Community rating: not yet rated Summary: Stable Diffusion XL (SDXL) is the definitive open-source standard for text-to-image generation, renowned for its massive community-driven ecosystem. Utilizing a dual-model architecture, it produces high-resolution 1024x1024 images with improved prompt adherence and photorealism compared to its predecessors. It remains the premier choice for creators seeking complete creative control, local deployment, and extensive fine-tuning capabilities. Typical use case: An indie game development studio leverages SDXL to rapidly prototype concept art and generate in-game assets such as environment textures and item icons. By running the model locally, they maintain absolute data privacy and eliminate recurring API costs. They train custom LoRAs on their unique art style, enabling them to generate thousands of cohesive assets that match their game's specific aesthetic. ## Sonar Pro URL: https://reviuws.com/models/sonar-pro Provider: perplexity Modality: text License: closed Parameters: n/a Context window: 200k tokens Input price: $3/M tokens Output price: $15/M tokens Community rating: not yet rated Released: 2025-01-21 Summary: Sonar Pro is a real-time web-search-grounded model built on top of open-weight LLMs, providing cited answers with a 200k context window. It's optimized for complex, multi-step search queries. It's available via Perplexity's API. Typical use case: Used for building search-augmented chat products, research assistants, and applications needing up-to-date cited information. Suited for real-time Q&A over the web. ## Sora 2 URL: https://reviuws.com/models/sora-2 Provider: openai Modality: video, audio License: closed Parameters: n/a Context window: n/a Input price: n/a Output price: n/a Community rating: not yet rated Summary: Sora 2 is OpenAI's cutting-edge text-to-video generator that seamlessly integrates high-fidelity synchronized audio directly into its video outputs. Building upon its predecessor's impressive visual capabilities, this model streamlines multimedia creation by generating cohesive sights and sounds from a single text prompt. With a flexible billing model priced per second, it offers a scalable solution for creators looking to produce cinematic-quality content. Typical use case: An independent advertising agency can leverage Sora 2 to rapidly prototype and produce high-quality video ads for social media campaigns. Instead of spending days coordinating separate video editing, sound design, and voiceover tracks, a designer can input a descriptive prompt to generate a fully realized 15-second commercial complete with realistic environmental sounds. This significantly slashes production timelines and budget, allowing the team to test multiple creative directions in real-time. ## Stable Diffusion 3.5 Large URL: https://reviuws.com/models/stable-diffusion-3-5-large Provider: stability ai Modality: image License: open Parameters: 8.1B Context window: n/a Input price: n/a Output price: n/a Community rating: not yet rated Released: 2024-10-22 Summary: Stable Diffusion 3.5 Large is an 8.1B parameter multimodal diffusion transformer generating high-quality images with improved prompt adherence and diversity. It's released under the Stability Community License. It includes a Turbo variant for faster generation. Typical use case: Used for creative image generation, self-hosted generation pipelines, and fine-tuning for custom styles. Popular in the open-source generative art community. ## text-embedding-004 URL: https://reviuws.com/models/text-embedding-004 Provider: google Modality: embedding License: closed Parameters: n/a Context window: 2k tokens Input price: n/a Output price: n/a Community rating: not yet rated Released: 2024-05-14 Summary: text-embedding-004 generates vector representations for text used in search, clustering and classification. It's accessible via the Gemini API. It supports task-specific embedding types. Typical use case: Used for RAG systems and semantic search within Google Cloud/Gemini ecosystem. Also used for recommendation systems. ## text-embedding-3-large URL: https://reviuws.com/models/text-embedding-3-large Provider: openai Modality: embedding License: closed Parameters: n/a Context window: n/a Input price: n/a Output price: n/a Community rating: not yet rated Summary: Text-embedding-3-large is OpenAI's flagship embedding model, engineered to provide highly accurate vector representations for complex semantic search and retrieval tasks. It supports up to 3,072 dimensions and features native support for shortening embeddings, allowing developers to balance storage costs and search latency without losing significant performance. This model offers substantial improvements in multilingual capabilities and conceptual understanding compared to its predecessors. Typical use case: An enterprise e-commerce platform can implement text-embedding-3-large to power its global semantic search engine. When a customer types a conversational query like 'lightweight waterproof jacket for summer hiking', the platform converts both the query and the product catalog into dense vector representations. Because the model supports flexible dimensions, the engineering team can truncate the vectors to 1,024 dimensions to optimize search speeds and reduce vector database costs while still delivering highly accurate, localized results. ## text-embedding-3-small URL: https://reviuws.com/models/text-embedding-3-small Provider: openai Modality: embedding License: closed Parameters: n/a Context window: n/a Input price: n/a Output price: n/a Community rating: not yet rated Summary: Text-embedding-3-small is OpenAI's highly efficient and cost-effective embedding model designed to convert textual data into numerical vectors. It offers a significant performance upgrade and drastically lower pricing compared to its predecessor, text-embedding-ada-002. Additionally, it supports native dimension reduction, allowing developers to balance storage costs and retrieval accuracy for high-volume applications. Typical use case: A global e-commerce enterprise can leverage text-embedding-3-small to build a highly scalable semantic search and recommendation engine. By embedding millions of product listings and customer search queries, the system can instantly match user intent with relevant products in real-time. This architecture supports cost-efficient Retrieval-Augmented Generation (RAG) for virtual shopping assistants while keeping vector database storage costs to a minimum. ## Veo 3 URL: https://reviuws.com/models/veo-3 Provider: google Modality: video, audio License: closed Parameters: n/a Context window: n/a Input price: n/a Output price: n/a Community rating: not yet rated Summary: Google's Veo 3 is a flagship generative model that seamlessly translates text and image prompts into high-definition video. By incorporating native audio generation, it ensures that visual elements and accompanying soundtracks are perfectly synchronized from the ground up. This unified approach makes it a highly capable tool for modern digital storytelling and rapid content creation. Typical use case: An independent filmmaker can use Veo 3 to generate high-fidelity B-roll footage and matching soundscapes directly from their screenplay's descriptive text. By typing in a scene description, the director can instantly produce cinematic cutaways—such as a rainy neon-lit city street with the matching sound of pattering rain and distant sirens—saving thousands of dollars on location scouting and Foley recording. ## Veo 3.1 URL: https://reviuws.com/models/veo-31 Provider: google Modality: video, audio License: closed Parameters: n/a Context window: n/a Input price: n/a Output price: n/a Community rating: not yet rated Summary: Google's Veo 3.1 is a cutting-edge generative AI model that produces high-quality video with native, synchronized audio. Supporting resolutions up to 1080p, it offers creators a flexible pricing model featuring a standard tier at $0.40/sec and a budget-friendly fast tier starting at $0.10/sec. This integration of video and native audio generation significantly streamlines multimedia production pipelines. Typical use case: A boutique advertising agency can leverage Veo 3.1 to rapidly generate social media ad concepts and mood boards for clients. By inputting descriptive prompts, designers can quickly produce a variety of 720p or 1080p video ads complete with realistic, perfectly synchronized background scores and sound effects. Utilizing the $0.10/second fast tier allows them to cost-effectively iterate on dozens of creative directions in minutes before committing to final production. ## Voyage-3-large URL: https://reviuws.com/models/voyage-3-large Provider: voyage ai Modality: embedding License: closed Parameters: n/a Context window: 32k tokens Input price: $0.18/M tokens Output price: n/a Community rating: not yet rated Released: 2025-01-07 Summary: voyage-3-large delivers state-of-the-art retrieval quality across domains including code, legal and finance, with flexible output dimensions and quantization options. It supports a 32k token context window. It's used by Anthropic and others for RAG pipelines. Typical use case: Used for high-accuracy semantic search and retrieval-augmented generation systems. Suited for enterprise RAG needing top retrieval benchmarks. ## Whisper Large v3 URL: https://reviuws.com/models/whisper-large-v3 Provider: openai Modality: speech License: open Parameters: n/a Context window: n/a Input price: n/a Output price: n/a Community rating: not yet rated Summary: Whisper Large v3 is OpenAI's state-of-the-art open-source speech recognition model, trained on millions of hours of diverse audio data to deliver industry-leading multilingual transcription and translation. Building upon its predecessors, this iteration offers enhanced performance in low-resource languages and exhibits superior robustness against background noise and varied accents. Because it is open-source, developers can run it locally or deploy it in secure environments to maintain full data ownership. Typical use case: A global media organization can integrate Whisper Large v3 into their post-production pipeline to automatically transcribe and translate multi-speaker video footage in over 90 languages. The model generates highly accurate, time-synced subtitles even when processing field recordings with heavy background noise, significantly reducing the manual labor and turnaround time required for international content distribution. ## Whisper large-v3-turbo URL: https://reviuws.com/models/whisper-large-v3-turbo Provider: openai Modality: audio License: open Parameters: 0.809B Context window: n/a Input price: n/a Output price: n/a Community rating: not yet rated Released: 2024-10-01 Summary: Whisper large-v3-turbo is a pruned decoder version of large-v3 offering significantly faster transcription with minimal accuracy loss. It's released open-source under MIT license. It supports multilingual transcription and translation. Typical use case: Used for real-time transcription applications and self-hosted speech-to-text pipelines. Popular for latency-sensitive voice apps. ## Whisper-1 URL: https://reviuws.com/models/whisper-1 Provider: openai Modality: audio License: closed Parameters: n/a Context window: n/a Input price: n/a Output price: n/a Community rating: not yet rated Released: 2023-03-01 Summary: Whisper is a general-purpose speech recognition model trained on diverse multilingual audio. It's available open-source and via OpenAI's API. It handles transcription and translation across many languages. Typical use case: Used for transcription services, subtitles, and voice interfaces. Also used in accessibility tools.