All opinions are my own.
Back in early May, we held a launch event showcasing the latest progress on our systems and software stack. The thing that stuck with me most — and got me most excited — was a single slide with two words on it: Run Anything.
Not long after, I happened to catch Debbie Marr’s ISCA keynote (Computing at the Crossroads: Architecture, Economics, and the Next Era). She made three points: it’s a great time to build a high-performance CPU, a great time to build a RISC-V CPU, and a great time to build general purpose computing.
In an era where everyone is talking about LLM acceleration, those three statements cut against the current. But I don’t want to just restate her argument. I want to lay out my own reasons for why, right now, we still need — and should be building — general purpose compute.
Let me be precise about what I mean. General purpose here isn’t the narrow “go back to building one big CPU.” What I mean is: don’t hard-bind your architecture and microarchitecture to one model structure or one operator shape. Programmable, reschedulable, able to absorb a new workload without a new tape-out — those are the operative words. A CPU is one form this takes, but not the only one. As for exactly where the line should be drawn — how general is general enough — I’ll save that for the end, because the answer depends on the three points below.
I. You can’t know what the next generation of models looks like
LLM capability growth is slowing down. This is probably my most contentious claim, and it started as a subjective impression. Opus 4.8 is only three months old, but Anthropic immediately followed it with Fable 5 and Opus 5, OpenAI shipped GPT-5.6 Sol, and on the open-source side GLM-5.2, GLM-5.3, K2.7, and K3 all landed in quick succession. Yet as a power user, I haven’t felt a qualitative jump. The problems that AI couldn’t crack for me on Opus 4.8 were, for the most part, still uncracked after I switched to Fable 5 the day it shipped. Rumor puts Opus 5 and Fable 5 around 5T parameters — and if that much scale still yields linear or sublinear improvement in practice, then scaling law has stopped paying off at the margin.
An impression is just an impression, though, so I went and looked at the Artificial Analysis Intelligence Index. On the August 2026 board, Claude Opus 5 sits at 63.0 and the tenth-place model at 56.8 — the entire top ten is packed into 6.2 points, spread across six different teams. Closer to my own experience: Opus 4.8, the model I had been using, ranks ninth at 57.3. Which means my move from 4.8 to Fable 5 (62.1) bought me under 5 points on that ruler.
That kind of crowding at the top is itself a signal, I think. When teams with different approaches, different data, different scale, and different compute budgets all converge into the same narrow band, it looks less like someone found a better recipe and more like everyone hit the same wall. A real capability jump should look like a discontinuity, not like a horizontal line with a crowd standing on it.
The other place things are stuck is context length. The industry has sat around the 1M mark for a while now, and long context is exactly what a lot of genuinely hard problems require. Even if it gets pushed further, where’s the ceiling? And how do we absorb the ever-rising memory and compute pressure that comes with it? I would love for the model companies to prove me wrong, but as things stand, I think the transformer path is starting to show its limits. Which brings me to the second point.
There will be more model architectures. People have been trying new architectures for a long time — JEPA, Mamba, Titans (Beyond LLMs: A Post-Transformer World Emerges) — and the attempts keep surfacing. Different modalities will use different models or model variants; diffusion transformers are one example. I believe different architectures will turn out to have distinct advantages in different application domains.
And this goes well beyond the chat box. Protein structure prediction, genomic analysis, weather and materials simulation — each of these fields is growing its own lineage of models, and the backbone often isn’t a transformer at all: equivariant networks, graph networks, state space models, even hybrids of classical numerical methods and neural networks. Their hardware demands look different too. Some are dominated by irregular memory access, some by sparsity, and for some the real bottleneck is preprocessing rather than matmul. If you design for the shape of an LLM and nothing else, you’ll find you aren’t meaningfully faster than anyone else in those domains.
And LLMs themselves will accelerate how fast these attempts arrive. There’s an analogy to the arrival of EDA here: people used the compute they had to design and explore the next generation of computing platforms. In the same way, we’ll use the AI we have to explore and iterate on the next generation of AI.
Put those two together: model capability is slowing, architectures are diverging, and silicon design, manufacturing, ramp, and system stability naturally trail software by two to three years. The architecture you commit to today is a bet on a workload you haven’t seen yet, two or three years out. In that combination, welding yourself to one particular model is an enormous risk.
II. Even if you did know, what you’re building is a computer
AI eats a lot of compute, but not all compute. Nearly everyone in this field has heard of Amdahl’s law. The essence of it is that if you only optimize part of a computation, your total gain has a hard ceiling. You can take 20% of the work down to zero and still lose to a system that made the whole thing 30% faster.
This year might fairly be called year one of agents: AI got tool use, and can now write code and execute it. So what does my AI agent actually spend most of its time doing, as an architect? Searching for a pattern across an enormous body of code and data (to localize a problem), compiling, running architectural simulation, running RTL simulation, running CI. I once left an agent running overnight and it landed a grand total of three edits — because one compile takes 20 minutes and one full simulation run takes 2 hours. And that isn’t even particularly long by our standards.
So once code generation gets to thousands of tokens per second, the bottleneck shifts wholesale from writing to verifying. And verification, to this day, still lands mostly on general purpose compute.
The point I want to make is not “therefore, bolt on a big CPU.” Any decent SoC today already carries a pile of small accelerators — codecs, crypto, compression, DSP, image signal processing — each doing its own job, none of them aspiring to become a matmul unit. Real workloads will only push further in that direction: gene sequencing has its own alignment and assembly engines, scientific computing has its own sparse and irregular-access demands, and unglamorous stages like data preprocessing, retrieval, and serialization can each grow dedicated hardware. So what this section is really about is the whole machine: matmul is one part of it, and the rest of the parts determine how fast it actually runs.
This is an extension of something I wrote in It’s not rocket science — a lot of people forget that they aren’t just designing a number cruncher, they’re designing a computer. Put more concretely for today: what we should be building is an AI computer, not a faster multiplier. The multiplier only covers that 20%. The computer has to cover the other 80%.
And this point differs from the previous section in one important way: it doesn’t depend on any judgment of mine about where models are heading. Even if transformers turn out to be the end state tomorrow and model architectures never change again, as long as AI still has to be verified, fed with data, and wired into real systems, general purpose compute stays on the critical path.
There are really two different axes of “general” here, worth separating. The previous section was about generality across model architectures: can the same machine follow along as model shapes change? This section is about generality across workload types: compiling, searching, preprocessing, sequencing, simulation, CI — none of these are matmul, and none of them plan to become matmul. The first axis tests whether you left slack in your dataflow and scheduling. The second tests whether your machine has the units to absorb this work at all.
III. The more expensive something is, the more it needs to be kept fed
After Moore’s law “died” around 2018 and 2019, the field broadly assumed that the next step in computing would come from specialization. That idea is reasonable on its face, but it hides an assumption: if we truly lived in a world of chip abundance — where anyone could tape out, package, test, and program a chip cheaply and quickly — then of course you would customize for every person and every scenario. The cost of specialization would be near zero, and the worries in the previous two sections would evaporate.
The problem is that this assumption doesn’t hold, and it’s becoming less true over time: manufacturing is advancing, but getting more expensive at the same time.
I joke with friends that whoever couldn’t afford HBM ten years ago still can’t afford it today. Moore’s law did fail — but what failed was its economic meaning. Cost per transistor no longer falls as you move down nodes, and at leading-edge nodes both design cost and tape-out cost per reticle keep climbing. Process is still advancing slowly, but what actually picked up the baton is memory, IO, and packaging.
Consider that GDDR5 (2008) to GDDR6 (2018) took a full decade. DDR and LPDDR went through similarly long waits after 2010, to the point that everyone carries around that classic chart in their head: memory progress falling far behind compute. Yet HBM has reached its fourth generation since 2013, with per-stack capacity and bandwidth both up by more than an order of magnitude — unthinkable by the old cadence. And with 3D DRAM arriving, memory may have a larger leap still ahead. Interconnect tells the same story: when AMD first announced chiplet-based CPUs in 2017, die-to-die bandwidth was on the order of 50GB/s; on the latest generation of accelerator products, unidirectional die-to-die is in the TB/s range — two orders of magnitude in under a decade. There are product-category and architectural differences baked into that comparison, so it isn’t strictly apples to apples, but what can’t be dismissed is the explosive progress in die-to-die technology over these ten years.
Sports fans like to say that in the face of talent, technique counts for nothing. Chips are similar: against a large enough gap in raw numbers, architecture can narrow the distance but rarely erase it. And if you want to deliver best-in-class performance, you’re forced onto the most advanced process, the most advanced packaging, and the most advanced memory — which means more risk, lower yield, higher manufacturing premiums, tighter supply-chain coupling, and longer lead times. All of which turns chips into a game of burning money.
So, back to the assumption this section opened with: chips did not become abundant. Seven years on, the path of trading customization for performance didn’t get easier with Moore’s law gone. It got more expensive.
And the conclusion this drives is different from the previous two sections. Those were about hedging: because you can’t know, don’t weld yourself shut. This one is about amortization: when the most expensive thing in the system is no longer logic but HBM, advanced packaging, and those few reticles, you have no choice but to keep them busy.
A chip that’s only fed while running matmul is a chip whose HBM and packaging sit idle during compilation, scheduling, data preprocessing, and verification. And if you dodge that by building two separate systems, you end up paying twice for the expensive parts. So my expectation actually runs the other way: CPU and matrix engine may end up more tightly coupled, sharing the same memory and the same system, not less.
To be clear about what that does not mean: sharing a system is not a return to specialization. The units can be specialized; the system is general. Generality here doesn’t live in any one engine — it lives in how many different kinds of work can call on that expensive pool of resources. This also answers a natural objection: flexible control flow and scheduling cost area and power, so is it worth it? If that flexibility buys you one set of HBM and packaging kept fed by many workloads in turn, then the money isn’t wasted. What it saves you is the second system you would otherwise have had to buy.
So
The three sections are really the same thing viewed from three distances. Up close it’s “you can’t know”: models will change, architectures will diverge, and your chip has to place its bet two or three years early. At middle distance it’s “even if you did know”: computing was never only matmul, and the stronger agents get, the more the bottleneck moves to the least glamorous stages — compilation, verification, preprocessing, and whatever specialized work each domain brings with it. From far away it’s “the more expensive something is, the more it needs to be kept fed”: when the expensive part is no longer the compute itself but the system carrying it, generality stops being insurance and becomes efficiency.
Then why is everyone specializing
“Everyone out there is moving toward specialization” — that’s true, but it doesn’t contradict any of the above, because we aren’t doing the same arithmetic. The optimal degree of specialization is set jointly by your market position and your payback window. If you already hold meaningful share of a large market that exists today, binding your architecture tightly to today’s workload is rational: short payback, volume to amortize the cost, and the bet stays in your own hands.
But that calculation inverts in two situations. One is when your payback window is longer than the model iteration cycle. The other is when the part you specialized for has already been squeezed dry: once you’ve driven it down to near zero, more specialization buys no further speedup or throughput, and everything left in the runtime is somewhere else — which is Amdahl’s law from the second section biting back from the other end. An incumbent already at the table with real share is thinking about how to carve out a bigger slice of the market that already exists. I’m thinking about what the next generation of AI computer looks like. Those are two different questions.
Generality isn’t a hedge, it’s the norm
I used the word “hedging” earlier, but that word smuggles in an assumption: that there exists some stable steady state where specialization is safe, and generality is just insurance you buy while waiting for it to arrive.
The problem is that computing has never concentrated — it has never narrowed onto a single workload, and it has never stood still in any one shape. Mainframes to PCs to mobile to cloud to today’s AI: in every wave, someone declared the new workload had settled and could be frozen into silicon. Some of those bets did pay off — video codecs, baseband DSPs, network forwarding, each backed by a standard that stayed frozen for years, and each an excellent trade. But note what they were betting on: a stable level of abstraction, not one specific implementation. What has repeatedly come up empty is welding one generation of model, or one particular dataflow, directly into silicon.
GPUs are the most instructive case. Born as a fixed-function graphics pipeline, they moved steadily toward programmability while steadily absorbing specialized units — tensor cores, RT cores, codec engines. They never chose between getting more general and getting more specialized. They did exactly what the third section describes: the units can be specialized; the system is general. And what let them go on to eat graphics, HPC, mining, and deep learning — workloads nobody planned for — was the second half of that sentence.
So to my mind, generality isn’t a hedge against uncertainty. It’s what computing has always looked like. The thing that actually needs justifying is never generality — it’s specialization. You have to explain why this time is different, and why this workload will stay still long enough to cover your design, tape-out, ramp, and depreciation cycles.
That isn’t to say generality is free. Flexible dataflow, scheduling with slack in it, the units that absorb all that other work — all of it takes area, power, and design effort, and a chip that stripped every bit of slack out can genuinely embarrass you on this generation’s workload. Generality can lose a generation. But the two ways of losing aren’t the same shape: generality loses a round and still gets to play the next one, while specialization bound at the wrong level usually doesn’t lose one notch of performance. It loses everything.
So how general is general enough
You can bind to a broad class of computation, but not to a specific dataflow and control flow.
An architecture can absolutely be built for large-scale parallel numerical work — that category is quite stable, and across the several paradigm shifts of the last decade-plus, “lots of parallel math” never stopped being true. What determines how long a chip stays alive isn’t its peak throughput; it’s whether it can be rearranged once the operator shapes, data movement, and scheduling above it change. So being aggressive in the arithmetic units is fine. Leaving slack in control flow and scheduling is not optional.
That also gives a fairly practical self-check: imagine your target model’s dataflow changes next year. Do you ship a new compiler and scheduler, or do you need a new tape-out? If it’s the latter, you bound yourself at too low a level.
So general purpose isn’t merely elegant design. It’s an unavoidable consideration for every architect in this era. An architect shouldn’t trade the momentary thrill of topping a benchmark for a whole product generation — and their own judgment — staked on an uncertain spin of the wheel.
Back to those two words. Run Anything reads most easily as a claim about coverage — how many models you can run today. But the weight of it is in time: the things that haven’t been invented yet, next year and the year after, run too.