No Datacenter

Hello all, I will host an online update next Tuesday 10am AEST, please register here.

The past week saw some of the biggest advances in AI since DeepSeek rocked the world with an open-weight model well over a year ago. 

Again, it came from China, only this time it was Alibaba's family of models, Qwen.

Specifically the 27 billion parameter version of Qwen 3.8, which is small enough to work on a Macbook Pro with near state-of-the-art performance. I'm getting a solid 14 tokens/s on my laptop with minimal configuration, but others are reporting far higher speeds.

It's competitive with Chinese competitors that are 100x the size (as measured by total parameters, 2.7 trillion parameters for Kimi-K3 Max vs 27 billion).

And is ahead of Opus 4.8 Max on some benchmarks (which was itself excellent).

So you can now run an Opus 4.8-level model on your own computer. No need for a Claude/ChatGPT subscription, internet access, sharing of data, or … a datacenter.

Right on queue Apple released an upgraded series of M5 computers which match or exceed Nvidia's consumer hardware for this kind of thing (caveat caveat).

And of course by the end of the year we will see much better local models. At the very least, it will force the leaders to keep prices low and responses fast.

It's striking how at odds this is with the latest podcast circuit commentary.

Dwarkesh and Dylan Patel - AI-commentating royalty - just aired claims that
a) Claude and ChatGPT would run most compute, as they are best able to monetize it, and
b) Less and less compute will go to inference

These developments do not support that.

Firstly, open-weight and now local models are serious market-share contenders.

Secondly, it looks like OpenAI is allocating more to inference, closing compute intensive initiatives like their web browser, Sora. On August 19 Sam Altman said they had temporarily paused some frontier RL training, ostensibly for safety reasons, but perhaps there’s no coincidence they’ve been able to dramatically cut prices and just released a fantastic family of models.

This looks like the right move, as their cheap and fast Luna model is now easily competitive with Chinese models. Which to be honest is a bit of a relief, as OpenAI's codex harness and enterprise security is better.

Meanwhile, Claude's Fable growth has stalled, with the Financial Times reporting that it is only accounting for 11% of Anthropic usage after two months. Which is strange, given all the pomp and fanfare on its release.

Usually new Claude models have rapid uptake, not least because everyone needs to clean up the slop generated by prior versions.

But now it seems customers are unwilling to pay up for something smarter, even within the Claude userbase, which is already ignoring better and cheaper alternatives. The cost/intelligence is just not adding up, and it’s hard to see why that would change for a better, even more expensive model.

Most of us are just grinding out white collar jobs, not doing frontier scientific work that might justify paying vastly more.

The price war continued elsewhere. OpenAI also cut the cost of their flagship Sol model by 33%/20% for input/output tokens.

All of which is suggesting the Claude Code regime of the last 10 months might be nearing an end, and as usual around turning points it’s unclear what replaces it.

Advances like these move surprisingly slowly through markets. It took over three months for November's Claude Code release to really sink in, and we may see a similar timeline now.

We now have:

  • a serious price war at the frontier,

  • evidence that users are happy with non-frontier level of intelligence (which is still exceptionally powerful and more than capable of handling most tasks), and

  • the first serious models that can run on a laptop with no datacenter or subscription required at all.

And if you don't know how to get this set up, just ask AI to do it for you.

There are easily available unrestricted versions of the model that will, for example, give recipes for crystal meth, guides to building weapons, and honest political histories of certain prickly nations.

As a further indication of the direction of travel, two days ago Nvidia announced the $6 billion acquisition of Poolside to push forward their own open-weight models, no doubt optimized for their own consumer hardware.

The pace is not letting up… as I write a new Qwen model has already been announced, apparently better than the one above, and trained at 1/9th the cost.

And if there was any doubt that everything is back in play, OpenAI announced a new chip that significantly outperforms Nvidia's Blackwell stack for inference.

Amidst all this, Nvidia is about to report quarterly earnings. Definitely one to watch live.

Mike

As mentioned above I will host an online update next Tuesday 10am AEST, please register here.