Another DeepSeek moment
Cheap and efficient models are changing the token economy. Is it still worth it to pay for more intelligent models, or is harness engineering the safer bet?
If you haven’t been following the world of AI this summer, I don’t blame you. Things are moving at breakneck speed, and every week there seems to be something new on the horizon.
This weekend it was DeepSeek V4 Flash 0731, a model released in yet another “DeepSeek moment”.
I spent the weekend stress-testing it. And while Anthropic’s models are still in a league of their own when it comes to quality and intelligence, I simply cannot ignore the price difference. In some cases, DeepSeek is more than 100x cheaper than Claude. Yes, 100x.
The unfair comparison
Comparing DeepSeek V4 Flash to Fable or Opus 5 really isn’t comparing apples to apples. V4 Flash is a small ~300B model that was never intended to compete with the larger models, and in some arenas it shows.
Artificial Analysis provides independent benchmarks of models, which is really important when comparing them. Their numbers paint a picture of a small model that performs on par with much larger ones.
Intelligence
Artificial Analysis provides three generalized intelligence scores: their own Intelligence Index, Coding Index and Agentic Index.
What is impressive is that the model trails Opus by ~18% on the Intelligence Index while trading blows with GLM 5.2, a point behind it on intelligence but ahead of it on both coding and agentic work. And on the Coding Index it comes even closer to Opus 5, a gap of only ~11%. This confirms my initial impression: for bread-and-butter coding tasks, DeepSeek has given us a cost-effective model.
Costs
When it comes to cost, there are some important considerations to make. A model might be cheap, but not very token-efficient, resulting in a higher total cost per task. DeepSeek does burn a lot of tokens when solving the benchmark, considerably more than the other models.
However, looking at the cost-to-run, which tells us how much it costs to run the Artificial Analysis benchmarks, it is clear that this really has no impact on the total cost.
It costs only $72 to run the benchmark, versus $3,836 on Opus 5. That is the real gap right there, and this is what makes the business case for this model so enticing to me.
Looking at the chart above (bear in mind, the x-axis is logarithmic) we can see that DeepSeek and Opus sit at opposite ends of the spectrum. But the vertical gap between the two is much smaller than we might think. Remember Sonnet 4.6? It had an Intelligence Index of 34.
The new token economy
This is what I call the new token economy. Frontier models bring frontier-level reasoning and intelligence, but that comes at a cost. My take, and my advice, is not to use only smaller and more cost-efficient models, but to use the right model for the right task. I do this myself when coding with agents: Opus for specifications, DeepSeek for well-defined and scoped coding tasks. If your AI strategy is simply to buy tokens for a single model, your agentic workloads will be cost-inefficient. Instead, engineers must assess which models are the right ones for the job, and have access to those models when they need them.
The emerging field of AI engineers
Harness engineering is the field that matters, now. Building effective and cost-efficient AI solutions is just as much about the harness you build around the model as it is about the model’s intelligence. But you can’t improve what you can’t measure. So as you build, evaluate your agentic workloads with quantifiable metrics, gather feedback from subject matter experts, and invest your resources into that loop rather than skimping on good evaluations.
Building truly agentic applications means acknowledging that the days when a simple application making a few API calls to an LLM endpoint was good enough are over. Now the focus is on orchestration, context management, evals, agentic loops, structured outputs and ontologies. The field is emerging.
Outlook
Yesterday we got Qwen 3.8, which is much more of a frontier-level model and should go head-to-head with Anthropic and Kimi K3. Early benchmarks look promising, but I’m personally more excited about next week, when the Qwen 3.8 27B-parameter version is expected to drop. Will we see another small model that punches well above its weight?
Here’s the question I will leave you with:
Are you spending your money on the latest, most expensive model — or on building a harness around your agent?