The top LLMs of August 2026 are no longer separated by capability alone. They’re separated by price, token efficiency, and whether you can trust the benchmark numbers. Our pick: Claude Opus 5, which lands within 0.5% of Anthropic’s own flagship on CursorBench 3.2 at half the price. Below you’ll find all five ranked, with real pricing (dated), the benchmark caveats vendors won’t volunteer, and a straight answer on when the open-weight models are the smarter buy.
How we picked these top LLMs
We’re an AI Agent Development Company, an engineering studio that ships agents and RAG systems into production, so we rank models the way we buy them: cost per completed task, not leaderboard vanity. We weighed vendor-published benchmarks against independent leaderboards (they disagree, often badly), checked pricing pages directly as of August 13, 2026, and dropped anything we couldn’t confirm. Meta didn’t make the cut. Llama 4 dates to April 2025 and the Behemoth variant never shipped publicly, which leaves the open-weight seat to newer contenders.
| Model | Price per M tokens (Aug 2026) | Standout number | Best for |
|---|---|---|---|
| Claude Opus 5 | $5 in / $25 out | Within 0.5% of Fable 5 on CursorBench 3.2 | Near-frontier work without frontier prices |
| GPT-5.6 Terra | $2 in / $12 out | 1.05M-token context, 128K output | Long-context production pipelines |
| Gemini 3.1 Pro | $2 in / $12 out (≤200K prompt) | 77.1% on ARC-AGI-2 (self-reported) | Reasoning-heavy tasks on a budget |
| Grok 4.5 | $2 in / $6 out | ~4.2x fewer tokens per SWE-Bench Pro task | High-volume agentic coding |
| Kimi K3 | Open weights, free to self-host | 2.8T params, 1M context | Teams that need to own the model |
1. Claude Opus 5 — Best value at near-frontier capability
Here’s the trade nobody frames honestly. Anthropic sells two flagships, and the cheaper one is the right buy for almost everyone. Claude Opus 5 shipped July 24, 2026 at $5/M input and $25/M output, unchanged from Opus 4.8, per Anthropic’s pricing documentation. Claude Fable 5, the “Mythos-class” model released June 9, costs exactly double at $10/$50. Anthropic’s own numbers say Opus 5 gets within 0.5% of Fable 5 on CursorBench 3.2. Half the price for 99.5% of the coding capability is the most concrete same-lab value comparison any vendor published this year.
“Surpasses all other models.” That’s Anthropic’s claim for Opus 5 on Frontier-Bench v0.1, its own internal coding benchmark, from the Opus 5 announcement. Internal benchmark, internal harness. Weigh it accordingly.
What we like in practice:

What could a custom AI agent take off your plate?
We build production-grade AI systems that quietly handle the busywork, so your team can focus on the work that actually matters.
- More than double Opus 4.8’s Frontier-Bench score, a generational jump rather than a point release.
- Roughly 3x the next-best model on ARC-AGI-3, though that benchmark has no independent public dataset yet.
- An optional fast mode at 2x price for about 2.5x speed. Useful for latency-sensitive agents, and easy to forget you left on.
- Cache-hit pricing at 90% off and 50%-off batch rates on the Fable tier make heavy pipelines cheaper than sticker price suggests.
The honest cons: output tokens at $25/M still run 4x Grok 4.5’s rate, and the headline benchmarks are Anthropic’s own. If your workload is hard reasoning where the last half-percent matters, Fable 5 exists. For everyone else, Opus 5 is the one we keep deploying. Anthropic’s mid-tier Sonnet 5, made permanent at $2/$10 on August 10, 2026, covers the boring high-volume work.
2. GPT-5.6 (Sol, Terra, Luna) — Best long-context spread
Skip Sol unless you’ve measured a reason not to. That’s the counterintuitive read on OpenAI’s July 9, 2026 flagship family. All three tiers share the same 1.05M-token context window and 128K max output, so the spec sheet doesn’t change as you move down the price ladder. Sol runs $5/M input and $30/M output. Terra costs $2/$12. Luna, after OpenAI’s late-July cut of up to 80%, sits at $0.20/$1.20.
Terra is the sweet spot. Same million-token window as Sol at 40% of the input price, which matters when you’re stuffing entire codebases or contract sets into context. Fair warning on the prior flagship: GPT-5.5 remains available at $5/$30, but its model card documents a 2x input and 1.5x output surcharge past 272K tokens. We’ve watched that surcharge quietly double a long-context bill. Budget for it or stay under the line.
One more reason OpenAI stays on this list: GDPval, its real-world economic-value benchmark spanning 44 occupations, built from tasks written by professionals averaging 14 years of experience (research paper). OpenAI open-sourced a 220-task gold subset with public grading, and it’s become the closest thing 2026 has to a cross-lab yardstick.
Pick Terra for production pipelines that live past 200K tokens. Pick Luna when volume dwarfs difficulty.
3. Gemini 3.1 Pro — Best reasoning per dollar, with an asterisk
The number is startling. Google says Gemini 3.1 Pro scores 77.1% on ARC-AGI-2 and delivers more than double the reasoning performance of Gemini 3 Pro, per the DeepMind model card. The asterisk: that’s Google’s figure, not one verified against the ARC Prize’s public leaderboard. Vendor-reported and independently scored numbers diverged all year, so treat it as a claim, not a fact.
The pricing is easier to trust. Google’s Gemini API pricing page lists the 3.1 Pro preview at $2.00/M input and $12.00/M output for prompts up to 200K tokens, rising to $4.00/$18.00 beyond that. Batch runs at half rate. That’s Terra money for a model Google positions against Sol and Fable 5.
Two things stood out when we put it to work. The four configurable thinking levels (Minimal through High) give you a real cost dial per request, which agent builders should love. And the output ceiling is 65,000 tokens against a 1M input window, so plan multi-step generation instead of asking for one giant artifact. It’s also still labeled a preview as of August 2026, released to developers in February. Google’s cheaper Gemini 3.6 Flash ($1.50/$7.50 after a 16.7% output cut) handles the overflow work.
Right for teams already on Google Cloud, and for anyone whose workload is reasoning-shaped rather than agentic-coding-shaped.
4. Grok 4.5 — Best cost per completed coding task
Don’t buy Grok 4.5 for peak accuracy. xAI barely pretends otherwise. On its own published comparisons, the July 8, 2026 flagship scores 62.0% on DeepSWE 1.0 against Fable 5’s 66.1%, and 64.7% on SWE-Bench Pro against Fable 5’s 80.4%. A step behind, plainly.
Buy it for the economics. At $2/M input and $6/M output, per xAI’s launch announcement, it’s the cheapest closed-weight model on this list, and xAI’s central claim is that it uses roughly 4.2x fewer tokens than Opus 4.8 on equivalent SWE-Bench Pro tasks. Token efficiency compounds brutally in agentic loops. A model that’s 15 points worse per attempt but a fifth the cost per completed task wins a lot of real workloads. Co-training with Cursor shows: it’s built for editor-driven coding, runs about 80 tokens/second, and ships a 500K context window with function calling, web and X search, and code execution.
The gap to know about. Scale AI’s public SWE-Bench Pro leaderboard, covering 1,865 long-horizon tasks across 41 repositories, shows its August 10, 2026 leader at just 59.1%. Every vendor’s blog numbers run hot against independent harnesses. Grok’s included.
Best for high-volume agentic coding where cost per merged fix beats cost per benchmark point.
5. Kimi K3 — Best open-weight model, full stop
This is the seat that didn’t exist a year ago. Moonshot AI published full weights for Kimi K3 on July 27, 2026: a 2.8-trillion-parameter MoE with a Stable LatentMoE architecture (16 of 896 experts active), native text, image, and video understanding, and a 1M-token context window. It’s the largest open-weight model ever released, and it reportedly ranks third on GDPval-AA v2 behind only Fable 5 and GPT-5.6 Sol while topping Arena’s Frontend Code board. Those placements come from aggregators rather than Moonshot itself, so hold them loosely. Moonshot also hasn’t posted a confirmable API rate, which is why our value pick below differs.
For healthcare and financial-services teams we work with, the pitch isn’t the leaderboard. It’s that the weights live on your hardware, in your compliance boundary, with no per-token meter running.
If cost is the whole question, look one shelf down:
- DeepSeek V4-Pro: 1.6T parameters (49B active), 32T+ training tokens, and official API pricing of $0.435/M input (cache miss) and $0.87/M output. Orders of magnitude under the closed flagships.
- DeepSeek V4-Flash: roughly one-sixth of even that price, at 284B total and 13B active parameters.
- Qwen3.8-Max: Alibaba’s 2.4T MoE, announced August 2 to 3, 2026, around $2.0/$6.0 per M tokens, strongest on long-horizon agentic coding endurance.
Kimi K3 for capability you can own. DeepSeek V4-Pro for the lowest credible cost per token in 2026.
Which of these top LLMs should you actually use?
Match the model to the constraint that’s actually binding, not the leaderboard.
- Hardest reasoning, budget secondary: Claude Fable 5, or Gemini 3.1 Pro if you trust Google’s benchmark suite. The two vendors genuinely disagree about who leads.
- Best all-around default: Claude Opus 5. That 0.5% CursorBench gap at half of Fable’s price is hard to argue with.
- Contexts past 200K tokens in production: GPT-5.6 Terra, $2/$12 with the full 1.05M window.
- Agent fleets running thousands of coding tasks daily: Grok 4.5, on token efficiency.
- Data can’t leave your infrastructure: Kimi K3 self-hosted, or DeepSeek V4-Pro when the meter matters most.
The common mistake is picking one model for everything. Every production system we’ve shipped this year routes across at least two tiers, a frontier model for the hard 10% and a cheap one for the rest.
FAQ
What is the best LLM in August 2026?
Claude Opus 5 is the best overall pick as of August 2026. It costs $5/M input and $25/M output, more than doubled its predecessor’s score on Anthropic’s Frontier-Bench v0.1, and lands within 0.5% of the pricier Claude Fable 5 on CursorBench 3.2 at half the price.
What is the cheapest capable LLM right now?
DeepSeek V4-Pro, at $0.435/M input (cache miss) and $0.87/M output on DeepSeek’s official API, is the cheapest credible frontier-adjacent option. Among closed models, GPT-5.6 Luna costs $0.20/$1.20 after OpenAI’s late-July 2026 price cut.
Are open-weight LLMs as good as closed models in 2026?
Close enough to matter. Kimi K3 reportedly ranks third on GDPval-AA v2, behind only Fable 5 and GPT-5.6 Sol, and wins Arena’s Frontend Code category. The remaining gap is at the very top end, and you pay a large premium to cross it.
Can you trust vendor benchmark numbers?
Not at face value. Vendor blogs report SWE-Bench Pro scores in the 60s and 80s, while Scale AI’s independent public leaderboard tops out at 59.1% as of August 10, 2026. Different task subsets and harnesses, so compare models within one source, never across sources.
How to run your own bake-off
Curious what AI could do for your business?
No jargon and no hard sell. Just a friendly look at where AI fits, and where it doesn't.
Benchmarks got you a shortlist. Your workload picks the winner. Take 20 to 30 real tasks from your own backlog, run them through Opus 5, Terra, and Grok 4.5, and score cost per acceptable completion rather than cost per token. The rankings shift more than you’d expect. Grok’s token efficiency wins some workloads that Fable 5 wins on paper. Then check the fine print that moves bills: GPT-5.5’s 272K surcharge, Gemini’s 200K price step, Anthropic’s cache discounts. If you’d rather not run that evaluation alone, our AI integration audit does exactly this against your production workload. Start with Opus 5 as the control. Beat it or keep it.





