

Open-source models closed the gap with closed frontier labs faster than almost anyone expected this year — reshaping the same coding-agent competition playing out elsewhere, and several are now genuinely competitive on real benchmarks rather than just “impressive for open-source.”
DeepSeek V4-Pro currently leads several open leaderboards outright — 80.6% on SWE-Bench Verified and 90.1% on GPQA Diamond, with a 1-million-token context window. It uses a mixture-of-experts design with 1.6 trillion total parameters but only 49 billion active per forward pass, and ships under an MIT license.
Kimi K2.6 is built specifically for agentic work — multi-step tool use and long-running workflows — and posts a 58.6% SWE-Bench Pro score that reportedly beats some closed frontier models on that specific benchmark. Licensed under a modified MIT license.
GLM-4.6 is the value pick for coding work, priced around $0.43 per million input tokens and $1.74 per million output tokens through hosted providers, while reportedly landing “near Claude Sonnet 4” quality — and it’s light enough to self-host after quantization.
Qwen3-235B-A22B is the cheapest genuinely capable option on this list — roughly $0.09 input / $0.10 output per million tokens — licensed under the permissive Apache 2.0 and fluent across 100+ languages.
Llama 4 Maverick remains the only multimodal model in this group (text and images), with a 1M-token context window, though its August 2024 knowledge cutoff and slower release cadence mean it’s fallen behind the newer Chinese models on most current benchmarks.
The bottom line: the “open-source is always behind” assumption doesn’t really hold anymore — DeepSeek and Kimi are competing on merit, not just price, and Qwen3 makes a genuinely capable model available at a fraction of frontier-API cost.
Sources: tech-insider.org — Best Open Source LLM 2026
—