Scenario guide
Best AI models for Local / Self-Hosted Coding
A coding assistant you run on your own hardware — no API bills, full privacy, works offline. Restricted to open-weight, self-hostable models and weighted toward practical coding and agentic ability, with a slice for the self-correction that smaller local models lean on inside a good harness. There is no cost axis because inference runs on hardware you already own.
Rankings use the same scenario weights and cost blending as the interactive leaderboard on AI Model Analyzer. Data is min-max normalised per benchmark; missing scores are skipped without penalty.
- 1DeepSeek R1DeepSeekScore 86.2Q 86.2In $0.55/M
- 2DeepSeek V3 (Thinking)DeepSeekScore 77.9Q 77.9In $0.27/M
- 3Kimi K2Moonshot (Kimi)Score 75.0Q 75.0In $0.60/M
- 4Qwen3 235BAlibaba (Qwen)Score 73.1Q 73.1In $0.20/M
- 5Qwen3 235B (Thinking)Alibaba (Qwen)Score 63.1Q 63.1In $0.20/M
- 6DeepSeek V3DeepSeekScore 61.0Q 61.0In $0.27/M
- 7GLM-4.7Zhipu AI (GLM)Score 60.4Q 60.4In $0.50/M
- 8GLM-4.6Zhipu AI (GLM)Score 60.2Q 60.2In $0.50/M
- 9Llama 3.3 70B InstructMetaScore 58.3Q 58.3In $0.88/M
- 10Qwen2.5 72B InstructAlibaba (Qwen)Score 50.7Q 50.7In $0.90/M
- 11Llama 3.1 70B InstructMetaScore 46.6Q 46.6In $0.88/M
- 12Llama 4 ScoutMetaScore 24.9Q 24.9In $0.18/M
- 13Llama 3.1 405B InstructMetaScore 20.8Q 20.8In $3.50/M
- 14Mistral Large 2MistralScore 20.3Q 20.3In $2.00/M
- 15Llama 4 MaverickMetaScore 19.2Q 19.2In $0.27/M