Claude Sonnet 4Claude Sonnet 4, part of the Claude 4 family, is a significant upgrade to Claude Sonnet 3.7. It excels in coding (72.7% on SWE-bench) and reasoning, responding more precisely to instructions. Sonnet 4 offers an optimal mix of capability and practicality, with enhanced steerability, and supports extended thinking with tool use. | | May 22, 2025 | | 72.7% | - | - | - | - | |
Claude Opus 4Claude Opus 4 is Anthropic's most powerful model and the world's best coding model, part of the Claude 4 family. It delivers sustained performance on complex, long-running tasks and agent workflows. Opus 4 excels at coding, advanced reasoning, and can use tools (like web search) during extended thinking. It supports parallel tool execution and has improved memory capabilities. | | May 22, 2025 | | 72.5% | - | - | - | - | |
Claude 3.7 SonnetThe most intelligent Claude model and the first hybrid reasoning model on the market. Claude 3.7 Sonnet can produce near-instant responses or extended, step-by-step thinking that is made visible to the user. Shows particularly strong improvements in coding and front-end web development. | | Feb 24, 2025 | | 70.3% | - | - | - | - | |
Gemini 2.5 Pro Preview 06-05The latest preview version of Google's most advanced reasoning Gemini model, capable of solving complex problems. Built for the agentic era with enhanced reasoning capabilities, multimodal understanding (text, image, video, audio), and a 1M token context window. Features thinking preview, code execution, grounding with Google Search, system instructions, function calling, and controlled generation. Supports up to 3,000 images per prompt, 45-60 minutes of video, and 8.4 hours of audio. | | Jun 5, 2025 | | 67.2% | 82.2% | - | 69.0% | - | |
Gemini 2.5 ProOur most intelligent AI model, built for the agentic era. Gemini 2.5 Pro leads on common benchmarks with enhanced reasoning, multimodal capabilities (text, image, video, audio input), and a 1M token context window. | | May 20, 2025 | | 63.2% | 76.5% | - | - | - | |
Gemini 2.5 FlashA thinking model designed for a balance between price and performance. It builds upon Gemini 2.0 Flash with upgraded reasoning, hybrid thinking control, multimodal capabilities (text, image, video, audio input), and a 1M token input context window. | | May 20, 2025 | | 60.4% | 61.9% | - | - | - | |
Claude 3.5 SonnetClaude 3.5 Sonnet is a powerful AI model with industry-leading software engineering skills. It excels in coding, planning, and problem-solving, with significant improvements in agentic coding and tool use tasks. The model includes computer use capabilities in public beta, allowing it to interact with computer interfaces like a human user. | | Oct 22, 2024 | | 49.0% | - | 93.7% | - | - | |
Claude 3.5 HaikuClaude 3.5 Haiku is Anthropic's fastest model, delivering advanced coding, tool use, and reasoning capabilities at an accessible price. It excels at user-facing products, specialized sub-agent tasks, and generating personalized experiences from large data volumes. The model is particularly well-suited for code completions, interactive chatbots, data extraction, and real-time content moderation. | | Oct 22, 2024 | | 40.6% | - | 88.1% | - | - | |
Gemini 2.5 Flash-LiteGemini 2.5 Flash-Lite is a model developed by Google DeepMind, designed to handle various tasks including reasoning, science, mathematics, code generation, and more. It features advanced capabilities in multilingual performance and long context understanding. It is optimized for low latency use cases, supporting multimodal input with a 1 million-token context length. | | Jun 17, 2025 | Creative Commons Attribution 4.0 License | 31.6% | 26.7% | - | 33.7% | - | |
| | Feb 5, 2025 | | - | - | - | - | - | |