Gemini 3.5 Flash is now generally available (GA), with stable performance suitable for large-scale production use. As our most intelligent Flash model, it delivers sustained leading performance at scale for agent execution, coding, and long-running tasks.
Google's most cost-effective model, optimized for high-volume agent tasks, translation, and simple data processing.
GPT-5.5 is OpenAI's flagship large language model released on April 24, 2026. Positioned as a new form of intelligence for real-world work and agents, its core breakthrough is autonomous planning and execution of multi-step complex tasks. It excels in programming, computer operation, scientific research analysis, and other areas, delivering higher efficiency with fewer token costs. The suffixes high/low/medium/xhigh indicate reasoning effort level.
DeepSeek-V4-Flash is a lightweight version of the DeepSeek V4 series, focused on high cost-effectiveness and high throughput. It is suitable for general conversation and basic text tasks, while supporting million-token long context and efficient reasoning.
DeepSeek-V4-Pro is a high-performance open-source large model from DeepSeek, offering top-tier reasoning and agent capabilities, support for ultra-long context, adaptation to domestic Ascend chips, and excellent cost-effectiveness.
GPT-5.5 Pro is now available for processing Responses API requests, including operations through the BBatch API, supporting multi-turn model interactions before responding to API requests, with additional advanced API features planned for the future. Because GPT-5.5 Pro is designed to tackle complex problems, some requests may take several minutes to complete. To avoid timeouts, try using background mode.
MiniMax-M2.7 reaches or refreshes industry SOTA performance in productivity scenarios such as programming, tool calling and search, and office work.
GPT-5.4 nano is the lightest and fastest version of GPT-5.4, designed for tasks with extremely high speed and cost requirements.
GPT-5.4 mini brings the strengths of GPT-5.4 into a faster, more efficient model designed for high-volume workloads.
GPT-5.4 Pro uses more compute to think more deeply and provide consistently better answers. It is accessible only through the Responses API, supporting multi-turn model interactions before responding to Responses API requests, as well as future advanced API features.
GPT-5.4 is our frontier model for complex professional work. The suffixes high/low/medium/xhigh indicate reasoning effort level.
Gemini 3.1 Flash Lite Preview is Google's high-efficiency model optimized for high-volume use cases. It outperforms Gemini 2.5 Flash Lite on overall quality and approaches Gemini 2.5 Flash performance across key capabilities. Improvements span audio input/ASR, RAG snippet ranking, translation, data extraction, and code completion. Supports full thinking levels (minimal, low, medium, high) for fine-grained cost/performance trade-offs. Priced at half the cost of Gemini 3 Flash.
GPT-5.3-Codex redefines AI's role in programming and broader productivity through performance improvements, capability generalization, and safety upgrades.
Gemini 3.1 is Google's most intelligent model family to date, built on advanced reasoning capabilities. It is designed to turn any idea into reality by mastering agent workflows, autonomous coding, and complex multimodal tasks. gemini-3.1-pro-preview is best suited for complex tasks that require extensive world knowledge and advanced cross-modal reasoning.
gemini-3-flash-preview is our most intelligent model, combining speed with frontier intelligence and offering excellent search and grounding capabilities.
GPT-5.2 pro is available only in the Responses API to support multi-turn model interactions and other advanced API features before responding to API requests. Because GPT-5.2 pro is designed to solve difficult problems, some requests may take several minutes to complete.
GPT-5.2 is the best model for coding and intelligent tasks across industries.
Vidu Q3 Turbo: The high-speed, cost-effective tier of the Q3 series. Operating at 720p (muted, no native audio) with foundational reference locking, this API is heavily optimized for ultra-fast rendering speeds and massive-scale video draft generation.
Vidu Q3 Pro: The absolute pinnacle of Shengshu's video API lineup. Outputting at 1080p with native audio-sync, it supports multi-image locking and directional frame transition. Capable of generating highly stable 16-second clips with cinematic camera moves and intricate facial micro-expressions.