Scenario guide
Best AI models for Coding Assistant
A pair-programming assistant for IDE / agent loops. Heavy on coding benchmarks, with a real-world agentic component (SWE-bench) and some weight on cost since coding loops burn tokens. A small slice goes to the recovery-rate reliability axis because real-world agent loops live or die on whether the model self-corrects.
Rankings use the same scenario weights and cost blending as the interactive leaderboard on AI Model Analyzer. Data is min-max normalised per benchmark; missing scores are skipped without penalty.
- 1Gemini 3 FlashGoogleScore 85.3Q 90.3In $0.30/M
- 2Gemini 3 ProGoogleScore 83.9Q 95.0In $1.25/M
- 3GPT-5.5OpenAIScore 83.5Q 95.3In $1.50/M
- 4GPT-5.4OpenAIScore 81.5Q 92.8In $1.50/M
- 5DeepSeek R1DeepSeekScore 80.7Q 85.0In $0.55/M
- 6GPT-5OpenAIScore 77.8Q 87.3In $1.25/M
- 7DeepSeek V3 (Thinking)DeepSeekScore 76.4Q 76.5In $0.27/M
- 8Claude Opus 4.6AnthropicScore 76.2Q 95.3In $15.00/M
- 9GPT-5.2OpenAIScore 76.0Q 85.1In $1.25/M
- 10Claude Opus 4.7AnthropicScore 75.4Q 94.2In $15.00/M
- 11o4-miniOpenAIScore 74.9Q 80.9In $1.10/M
- 12Claude Sonnet 4.6AnthropicScore 74.7Q 86.0In $3.00/M
- 13Kimi K2Moonshot (Kimi)Score 74.2Q 77.4In $0.60/M
- 14Qwen3 235BAlibaba (Qwen)Score 74.1Q 71.4In $0.20/M
- 15Claude Sonnet 4.5AnthropicScore 74.0Q 85.1In $3.00/M