All opinions are my own.
My manager, when we’re discussing about architecture, has a habit of dropping the phrase “it’s not rocket science” without much ceremony. I’ve come to agree with it more and more, so I stole it for the title.
The biggest difference between building products and doing research is that products also have to be novel — but the risk has to stay bounded. There’s a pile of constraints, and you have to think through every corner case a user might hit. An architect isn’t just designing a chip; they’re designing the single most consequential part of a computer, the part that everything else hangs off of. So a good architect can’t just know semiconductors and applications — they need to know how every one of their decisions ripples through the rest of the machine.
I’ve been thinking lately about what shape and architecture the next generation of AI chips should take, and where the real opening is. After turning it over enough times, the question compresses into one sentence: how do you design a computer that can actually be used at scale — not a machine that only works if you have the most advanced process node, the deepest pockets, and the most specialized team on earth.
I. An Era Sliding Back from Consensus into Chaos
Computer architecture is currently sliding back from consensus into chaos.
It’s not just the conclusions that are chaotic — which architecture suits ever-larger models, which products the big companies will actually pay for. The conclusions are chaotic because the research methods themselves, and the future technology tree, are both uncertain. In the past, workloads and their characteristics were relatively fixed, the tech tree tracked Moore’s Law closely enough to be predictable, and performance could still be simulated, quantified, forecast, and analyzed with reasonable accuracy.
The last few years have been nothing like that. You can barely predict what scale next year’s frontier model will be. You can’t predict what new application patterns users will invent. You can’t simulate the performance of a thousand-GPU or ten-thousand-GPU cluster with much confidence. You can’t tell which technology path will mature first, pull yield up, and push price down — 3D DRAM, 3D SRAM, HBF, CPO — none of it is settled yet. Worse, problems that most people wouldn’t even think to worry about keep showing up one after another: reliability, software maturity — things that used to be considered “solved” are back, and uglier than before.
But here’s the thing worth noticing: almost none of these new problems are soley “deeper science problems.” They’re not solely some physical limit sitting there waiting to be broken through — they’re engineering problems created by scale. More chips, higher power, a wider and messier set of users — so things that a handful of experts could previously just carry on their backs now have to be solved by the design itself. In other words: it used to be enough to stare at raw performance. Now you can’t just chase performance — you also have to make sure this machine can actually be put to work by thousands upon thousands of chips and thousands upon thousands of users.
That shift is exactly what this piece is about.
II. What to Bet On, and How Much
With this many open problems and a market this frenzied, a lot of companies choose aggressive technology bets to win capital’s favor and produce impressive-looking performance projections.
I don’t think there’s anything inherently wrong with being aggressive. In an industry where the technology path hasn’t converged, not betting is impossible — every process node, every memory type, every packaging technology you pick is a bet. The real question is whether the bet is placed with a clear head.
Two things are worth working out carefully here.
First: multiple bets multiply, they don’t add. A chip takes three years from definition to mass production, and you’re never really betting on “will this technology succeed” — you’re betting on “will it mature, at the yield and price I need, in exactly the quarter I tape out.” That alone isn’t a high-probability event. Stack three or four such bets together and your final success rate is their product — and they’re usually not independent either. An aggressive technology choice tends to simultaneously reshape floorplan, power delivery, packaging, and the software stack, so when one bet fails, the others tend to collapse with it.
Second: you have to look at what’s left after you lose. Some bets, if they fail, mean the product ships six months late and a tier below spec — but it still ships. Others, if they fail, mean the entire project goes to zero. These sound equally “aggressive,” but the nature of the risk is completely different.
So the criterion shouldn’t just be “how likely is this to succeed” — it should be two questions asked together: how likely is it to succeed, and if it fails, what do I still have left? A plan with a 50% success rate that can still ship a degraded product on failure is very often a better bet than a plan with an 80% success rate that requires starting over from scratch if it fails.
This is where rockets and computers diverge. A rocket can bet everything on a single mission, because it only needs to succeed once, on that one day. If the bet is wrong, the cost is one launch. A computer has to be manufactured in the thousands and handed to thousands of people you’ll never meet — if the bet is wrong, the cost is the entire product line disappearing, along with the software ecosystem that was growing up around it. Anything meant for use at scale should, by its nature, place bets that leave more room to maneuver.
Under uncertainty, the most valuable thing isn’t the bet itself — it’s the optionality: the other possibilities you keep for yourself while placing the bet. Keeping a fallback isn’t a lack of nerve. It’s precisely what lets you afford to bet at all.
A lot of people overlook this, or forget it: they’re not just designing a number cruncher, not a rocket. They’re designing a “computer.”
III. Chip Architecture Is Not Rocket Science
Saying chip architecture isn’t rocket science doesn’t mean it isn’t hard — it means it’s hard in a different way than rockets are hard. Put a few things side by side in a table and the difference becomes clear:
| Dimension | Ideal AI Computer | GPU | AI Chip Startup | Rocket |
|---|---|---|---|---|
| TCO predictability | Predictable, budgetable | Expensive, but priceable | Expensive and hard to estimate | Extremely expensive, priced per mission |
| Volume production | Mass produced | Mass produced, capacity-constrained | Hard to scale to volume | Produced in very small numbers |
| Reliability | Reliable long-term at room temperature and pressure | High but manageable yield and failure rates | Persistently high yield loss and failure rates | Reliable briefly, under extreme conditions |
| User | Almost anyone | Mid-to-large enterprises | Mid-to-large enterprises | Trained aerospace professionals |
| Purpose | General-purpose | General-purpose | Single-purpose | Get payload into space |
| Extensibility | High | Medium | Low | Low |
The gap between the leftmost and rightmost columns falls almost entirely on the two ends of one axis: can this thing actually be used at scale? Can the price be pinned down, can production capacity keep up, does the user need specialized training, does it need a special facility to operate in. The two middle columns are where this industry actually stands today.
Having written this far, though, I want to give rockets their due.
The literal meaning of “it’s not rocket science” is “this isn’t that hard,” and what I’m borrowing it to say is “what we’re building isn’t a rocket.” But here’s the interesting part: the company that actually turned rockets into a business isn’t really doing rocket science either. Falcon 9’s value doesn’t come from how high it flies or how much thrust it produces — it comes from the fact that it can be recovered, reused, scheduled, and its cost can be written into a price quote. What SpaceX did was essentially turn the rocket from a “one-off scientific spectacle” into a “mass-producible, predictable transportation product.” They won on engineering, not science. Look back at that table — the rows SpaceX actually rewrote in the “rocket” column are “volume production” and “TCO predictability,” not “thrust.”
So the real opposition was never “computer vs. rocket” — it’s engineering vs. show/demo. An accelerator that only produces impressive numbers under ideal conditions, and needs the vendor to send an engineer to sit on-site tuning it, is still in the show/demo stage no matter how astonishing its peak FLOPS look. It might win you a funding round. It won’t win a customer’s second order.
IV. The Architect as Head Chef
An architect isn’t a craftsman, and isn’t a mad scientist — an architect is closer to the head chef running a kitchen: a limited budget, limited time, limited staff, and a menu to put together that the guest can actually finish and actually afford. A chef who refuses to cook unless handed the finest ingredients and a Michelin-caliber brigade isn’t much of a chef.
This metaphor is really just another way of saying “it’s not rocket science”: the hard part was never mastering some unfathomable secret technique — it’s deciding what to do and what not to do given the conditions you actually have. And now that AI is this advanced, “drafting the menu” itself is becoming more and more replaceable. What’s genuinely irreplaceable is the judgment to make trade-offs under real constraints.
As for the computer itself — it’s meant to be bought and used however the buyer wants. A good computer should let its users extend it, hack it, debug it themselves, rather than requiring the vendor to personally show up for every single problem.
“Usable at scale” is really three things stacked on top of each other: it can be built, it can be deployed widely, and it can be afforded. None of these three are performance metrics — but without any one of them, performance is beside the point. The three goals in this section look at these three things from the product and company side; the six things in the next section look at the same three things from the architecture side.
Supply-chain cycles and risk — first, it has to be buildable at all. Over the past couple of years advanced packaging lines have been essentially fully booked, with a handful of top customers eating most of the capacity and lead times from order to delivery measured in years. HBM is even tighter, because it consumes three to four times more wafer area per gigabyte than ordinary DRAM, and AI demand keeps pulling capacity away from commodity memory. That means you have to commit to a volume before the architecture is even frozen, and packaging capacity is often scheduled in lockstep with wafer capacity.
The same logic applies to process nodes. Leading-edge capacity is similarly locked up by top customers — expensive, slow to ramp yield, and you’re not guaranteed a spot in line at all. Architects tend to carry a default assumption: of course I’ll use the most advanced process, the newest generation of memory, the most aggressive packaging. But that assumption is itself a bet, and the most expensive kind — it raises cost, stretches the timeline, and puts you in someone else’s queue. The real skill is building something still competitive on a process node one generation behind, using more ordinary materials. Only if you can do that does the product have a shot at being cheap, predictable, and actually available — which is the precondition for being usable at scale at all.
So when settling on a design, the question isn’t just “how fast does this run” — there’s a more fundamental question: is our product’s fate held in someone else’s hands? In concrete terms, that breaks down into: who does this choice tie me to, what does it cost to untie myself, and if their capacity, yield, or priorities shift, what do I have left?
Development cycle and cost — buildable in principle isn’t the same as buildable in time. If the architecture is too complex and development too slow, by the time you ship, the underlying landscape may have already turned over once — memory has moved to a new generation, the process has shifted a node, the interconnect standard has a new revision — and on the market side, the trend, the subsidy window, the demand may all have moved on too. You end up delivering the perfect answer to a three-year-old question.
And the cost of being late isn’t linear — it’s a cliff. A product that’s two quarters late usually doesn’t just earn 20% less — it often has no slot at all. The customer’s rack is already full of someone else’s cards, and the software stack has already grown up around a competitor. So architectural complexity is itself a risk, not just a cost. A simpler design ships earlier; shipping earlier means getting real workload feedback earlier and entering the next iteration sooner. The penalty for a complex design compounds.
Volume production and delivery potential — buildable, on time, and still able to scale out. Yield, cost of ramping capacity, binning, fallback options when something goes wrong — I’ll get into those later. But one thing I think is the most underrated: the architecture’s own flexibility in ratios. The optimal balance among IO, bandwidth, capacity, and compute keeps shifting — training and inference are different, prefill and decode are different — and it shifts faster than a single chip generation’s development cycle. The problem is that many technology choices tightly couple these resources together: 3D DRAM stacks DRAM directly on the compute die, so adding capacity forces you to add compute alongside it. HBM is the same — bandwidth and capacity are bound into the same stack. The cost of that isn’t some mild “less flexible” — it’s that you’re forced to sell the customer several resources they don’t need, and those happen to be the most expensive ones.
To be clear, this coupling isn’t a design flaw — it’s exactly where this class of solution gets its performance advantage. Decouple it, and bandwidth drops right back down. So this is a genuine trade-off, and simultaneously the solution’s greatest strength and its greatest risk. There’s no universal answer to whether the risk is worth taking — it can only be decided by the architect’s own read on where the market and the technology curve are headed. And that’s exactly where this role is genuinely irreplaceable: not in calculating which option is faster, but in judging whether, over the next three years, that coupling turns into a moat or a millstone.
V. Six Things That Matter Most in Architecture, Given All This
Before diving in, one thread is worth naming up front: under how demanding a set of conditions does this machine have to hold up?
A rocket can demand extreme conditions: a dedicated launch site, a specific weather window, a specially trained team, cost-no-object investment, and a single well-defined mission. The more demanding the conditions, the less possible it is to make it widespread — but that’s not a problem for a rocket, because a rocket was never meant to be widespread. Of course, as noted earlier, SpaceX is changing exactly this, and the way they’re changing it is by pushing the rocket toward the “computer” side of the table: raise the volume, drive down the cost per launch, standardize the support process.
A computer is the opposite. It has to be sold to someone you’ve never met, running in an ordinary server room, executing a workload you never anticipated. You cannot require every customer to have access to the most advanced process, to have deep enough pockets, to only run the handful of models you’ve personally tuned for, to staff a team that understands your architecture. Scale itself is the source of this difference: the complexity of one machine can be absorbed by people. The complexity of ten thousand machines can only be absorbed by design. Every unit of complexity you build in gets copied thousands of times over, spreading into places you will never see.
The complexity of one feature is absorbed by the verification team. An aggressive packaging choice is absorbed by the backend team. An unobservable design is absorbed by the field engineer. An unattributable performance gap is absorbed by the user. A bug with no fallback is absorbed by the entire production line. An interface that can’t be reused is absorbed by you.
So the six things below are, at bottom, the same thing said six ways: bringing down, one by one, the demanding conditions this machine requires in order to work at all. This is also the real weight “it’s not rocket science” carries for me — it’s not saying this is easy, it’s saying this shouldn’t be built to only work for a handful of companies under a pile of demanding preconditions. Go back to that earlier table — the difference in the “user” row is exactly what this is about.
The six things are:
- Design for verifiability — proving it’s correct and clearly defining the boundaries takes more time than the design itself.
- Don’t dump every hard problem on backend and packaging — they’re the ones who actually have to face physical reality last.
- The system must be quickly testable and debuggable — in a cluster of a thousand-plus GPUs, being unable to find a problem is the same as being unable to deploy at all.
- Think for the person writing the software — whether a user can run the newest model on day one matters as much as peak performance.
- The chip needs a Plan B — it has to still be usable even when it isn’t perfect.
- Design for reuse and extensibility — this determines how much market the product can cover and how long it can live.
1 | Verification-Friendly
Verification is the process of proving that this chip won’t compute the wrong answer under any possible circumstance. It’s the single most time-consuming part of the whole project, usually eating over half the headcount — every feature you add is like adding another dial to a combination lock; the number of combinations you have to check multiplies. And the real difficulty isn’t even running the tests — it’s first figuring out what to test, which means clearly defining the boundaries of the design.
So when is a design actually done? Finishing the RTL doesn’t count. PPA closure and verification sign-off are both hard gates, and either one can become the final bottleneck. The difference is that PPA has numbers on the table — area, frequency, power — scrutinized from day one. Verification’s cost has no equivalent number; it only shows up on the calendar. That’s exactly why it’s the most underestimated line item at the architecture stage.
Two things matter most in practice. First, boundaries need to be sharp: the interface between modules should be fully specified by the protocol itself, not by two engineers who happen to know how the other one thinks. The moment a boundary blurs, verification can no longer divide and conquer — it’s forced into full-system random testing, and convergence speed drops by an order of magnitude.
Second, correctness and performance should be two separable claims. Correctness should be built to be provable; performance is a different problem altogether. The value of separating them is that failure modes converge: if a performance mechanism breaks, the consequence should be “runs too slow,” not “computes the wrong answer.” The worst case is missed performance, but the thing still works. Conversely, if an optimization added purely for performance also happens to carry correctness semantics, then its failure is fatal — and during verification you can no longer converge on the two dimensions separately.
Performance verification has gotten much heavier and much harder to deal with in recent years, precisely because it has no natural golden model. But I don’t think it should be left entirely to after-the-fact profiling — the performance design space should equally be constrained by assertions and rules: queue depth has an upper bound, arbitration latency has an upper bound, a given state isn’t allowed to stall for more than N cycles without progress. Without that, performance degrades into a black box you can’t reason about, and a black box can’t be made to converge.
I want to dwell on one point here. I’ve seen plenty of designs blur a boundary to save one or two cycles — a direct cross-module handshake, two pipeline stages fused together, a signal punched through a stage early. Those one or two cycles are a real, measurable win in the performance model, but they simultaneously leave both PD and verification with a harder problem: one more critical timing path that’s hard to close, one more state whose ownership is ambiguous during verification.
My view is that system performance should be improved through better architecture and design, not traded for verification and PD complexity. The former’s payoff is structural and carries forward into the next generation; the latter’s payoff is one-time, and the cost has to be paid over and over within this generation.
This, in turn, demands that the design be more deterministic. Mechanisms that rely on unpredictable runtime behavior to gain average-case performance are essentially trading bounded performance for unbounded performance — you might win a few percentage points on a benchmark, at the cost of a system whose performance can no longer be proven, explained, or reused analytically in the next generation. In one line: correctness should be provable, performance should be bounded.
The last point might be the most counterintuitive. Once the design space is locked in, AI can already automate a substantial part of the design itself. Verification, though, can’t be automated the same way — because what verification does is precisely define that design space. AI can certainly generate UVM code, but the space UVM defines — which behaviors are legal, which combinations must be covered, which corner cases are worth spending two weeks chasing — still comes from a person. Generating a testbench is execution. Drawing the boundary is judgment. So even as more people try to apply AI to verification, verification will likely remain, for a very long time, the part of the flow that needs the most human effort.
2 | Packaging-Friendly
On an architecture diagram, a lot of things are free. Drawing a line takes one second, but that line eventually has to become actual metal. You can easily sketch an eight-chiplet topology, tens of TB/s of die-to-die bandwidth, 3D SRAM stacked on top of logic — all of it holds together in a block diagram, and it looks especially good on a slide. But down at the implementation level, it’s a different group of people who have to face bump pitch that won’t fit, IR drop that can’t be sustained, hotspots that can’t be dissipated, warpage that tanks yield, timing across die boundaries that refuses to converge. None of these problems show up in the architecture spec — but any one of them can make the chip unbuildable.
There’s a dangerous asymmetry here: the architect’s cost of “proposing a requirement” is nearly zero, while backend’s cost of “implementing that requirement” is enormous. Left uncorrected, this asymmetry causes the architecture to systematically drift toward aggressiveness, because the cost of that drift is never borne by the person drawing the diagram.
So packaging-friendliness isn’t really a technical property — it’s a process requirement. Any decision that pushes difficulty downstream must be laid out, risk and payoff both, on the table with backend, packaging, thermal, and power delivery teams before it’s finalized. This isn’t asking the architect to do backend’s job — it’s asking them to acknowledge that every line they draw is eventually implemented by someone else.
I’m not saying you can never be aggressive. Sometimes the market genuinely requires 3D, requires pushing bandwidth to the limit. The difference is only this: that should be a decision the whole team makes together, with everyone understanding what’s being bet, rather than a consequence backend has to absorb after the architect has already finished the diagram. The same aggressive plan, in a team that aligned on the risk beforehand versus a team where “the diagram’s done, you figure it out” — those are two completely different things.
An architect’s drawing is, in the end, always implemented by someone else. An architect who doesn’t understand the cost of implementation isn’t drawing an architecture — they’re drawing a wish list.
3 | System-Level Testing, Debugging, and Profiling
Once a server room has a thousand-plus GPUs wired together, “something breaks” stops being an accident and becomes a daily occurrence. A thousand-to-ten-thousand-GPU cluster stacks up almost every known risk source at once: the thinner voltage margins of advanced process nodes, the thermal cycling and power-delivery noise from extreme power density, multi-layer stacks like HBM, and optical interconnect and CPO now entering the system — lasers have finite lifetimes, coupling is sensitive to temperature and mechanical deformation, and failure modes are completely different from electrical interconnect. Any one of these looks “low probability” in isolation, but multiplied by the GPU count, it isn’t. And synchronous training makes it worse with a bucket-brigade effect: one bad card out of ten thousand, and the entire job stalls.
And when stability itself is the problem, the ability to quickly localize a fault stops being an operational convenience and becomes the precondition for shipping at scale at all. If a batch of cards goes into a customer’s server room and every failure takes days to trace back to a specific card — is it swappable, can the job keep running — then deployment speed is locked to debug speed. You can build it, but you can’t roll it out — the bottleneck of large-scale deployment shifts from the fab to the field.
Performance problems are even harder to chase than stability problems. A crash at least leaves you a clear scene. A performance problem in a cluster usually just shows up as “20% slower than expected” — no error, just an ugly curve. Which card is dragging things down? Is it actually slow, or is it waiting on someone else? Do cross-machine traces even share a common time base you can align them against? None of these are hard on a single machine. Every one of them is hard at ten-thousand-GPU scale. So the chip needs a diagnostic path independent of the main data path, one that still functions in a faulted state, and cross-chip timestamp synchronization with enough precision to matter. Scan chains and JTAG are useful in a lab and nearly useless in a rack — observability has to be an architectural feature, not a patch bolted on afterward.
There’s also something engineers commonly overlook: profiling is actually a sales tool. A customer preparing to place a large order will always run a POC first. If the numbers they get come in below spec and neither side can explain why, that deal is probably dead — not because the chip isn’t fast enough, but because the customer can’t prove to their own boss that the money was well spent. Conversely, a toolchain that can break performance down to specific bottlenecks and point at how to fix them sells more than performance — it sells certainty. In large-scale procurement, certainty is worth more than peak numbers. The standard here is simple: a user who sees below-expected performance the first time won’t give up on you right away. But if they spend two weeks and still can’t get an explanation for why it’s slow, there won’t be a second chance. Performance that can be explained is performance users dare to depend on; performance that can be optimized is performance users are willing to buy more of.
4 | Software-Friendly
The chip is hardware, but what users deal with every day is software. A model architecture changes every few weeks; a chip takes three years. So when a new model comes out and the customer wants to deploy it next week but your software stack can’t support it, what they lose isn’t a few percent of performance — it’s everything they could have been doing in that gap and weren’t. Their cards are still spinning in the rack, running last generation’s model. That’s an enormous opportunity cost, and it never shows up on any benchmark table. A chip with immature software is, in practice, selling the customer a dependency that requires them to keep waiting on you — it takes away the user’s own agency.
So during the architecture phase, you should keep running the same thought experiment over and over: when my chip is plugged into an eight-GPU server, and those servers are wired into a cluster of a thousand-plus, what does the person writing parallel software actually run into? What can I guarantee for them?
The keyword is guarantee. The difficulty of software is, at bottom, the sum total of guarantees the hardware refuses to provide. Is memory ordering clearly defined? Are cross-card atomics hardware-supported? Does collective communication have hardware semantics backing it? Is network delivery ordered, and who owns retransmission? Is there an upper bound on the cost of a synchronization primitive? Every guarantee hardware doesn’t provide becomes a case software has to handle — and in a multi-GPU setting, it becomes a distributed case. The difficulty isn’t additive, it’s multiplicative by an order of magnitude. An interface that’s “occasionally out of order” on a single machine gets fixed with a barrier. The same problem at ten-thousand-GPU scale might take three months just to reproduce.
One more thing worth flagging explicitly: don’t assume AI will simply solve all of this. Parallel software is hard even for AI. Its difficulty isn’t in writing the code — it’s in reasoning about concurrent behavior, guaranteeing correctness in a system with no well-defined semantics, and localizing root causes from phenomena that can’t be reproduced. This is the same principle as the verification section: AI is good at executing within a well-defined space, and the difficulty of parallel software is precisely that the space itself was never well-defined to begin with.
So this ends up being a trade-off that’s especially easy to get wrong. Architects naturally love clean, elegant, lean designs — strip out coherence, strip out hardware synchronization, strip out anything “software can handle anyway,” and the block diagram instantly looks more elegant, and the area and power numbers look better too. But whatever got stripped out didn’t vanish — it just moved to the software side, and usually to the hardest part of it: the multi-GPU, asynchronous, unreproducible part. Trading architectural cleanliness for software complexity is a deal that always looks good on paper and is always expensive once it actually ships at scale — the architecture’s elegance is a one-time gain, while the software’s complexity is a cost every user and every deployment has to pay all over again.
5 | The Chip Needs a Plan B
Silicon is never perfect. Some imperfections are random manufacturing defects, some are design bugs — and by the time either is discovered, the chip has already taped out. That’s not pessimism, it’s statistics. So the question isn’t “will something go wrong” — it’s “when something goes wrong, what do I have left.” A chip shouldn’t have only two states, “perfect” and “scrap.” There should be a lot of room in between.
First: can it survive binning? Imperfect dies are the norm, and the job of the architecture is to make sure they’re still sellable — which means reserving trimmable dimensions at design time: tile count, frequency tier, memory channel count. The key is that these dimensions have to be something software can handle transparently. If disabling two tiles breaks the compiler’s partitioning strategy or makes certain operators unrunnable, then that dimension is useless for binning, and a chunk of your yield curve is gone. Ultimately, whether you can bin at all depends on how adaptive the software stack is — and how adaptive the software stack can be depends, in large part, on how much flexibility the architecture left it in the first place.
Next: can it survive different compositions? In the chiplet era, a product is a combination of several kinds of dies. So when one particular die’s supply hits a problem, its yield falls short, or you’re forced to swap in a new revision, can the rest of the system keep working as-is? That requires the freedom to recombine to actually exist in the design, not just on paper — those combinations need to be verified and software-supported configurations, not “theoretically possible but nobody’s ever tried it.” Otherwise so-called modularity is just a block diagram.
Third: don’t leave anything that’s “hard to verify but critical” — this is really the production-stage bill for point 1. If a feature sits on a critical functional path and is also practically impossible to verify exhaustively, it’s a time bomb: it will most likely pass every regression test, and then blow up after tapeout, on some customer’s workload, in a way nobody saw coming. So every time you add a feature, it’s worth asking specifically: if this is wrong, is there a way to route around it? If the answer is “no, the whole chip is dead” — then its verification priority needs to go to the top, or it needs to be split into a fast path that can be disabled, backed by a slow-but-provably-correct fallback path.
Fourth ties the previous points together: there should be options to degrade gracefully, not just on/off. When some aggressive timing path is flaky, or some high-bandwidth mode has a bug, you should still have a mid-tier performance configuration to fall back to. Some performance is lost, but the product is still alive. And in this industry, the gap between a chip that’s 20% slower and a chip that can’t ship at all isn’t 20% — it’s everything.
These fallbacks have to be designed in at the architecture stage. Thinking about them the day a bug appears is already too late — by then the only thing left to change is software, and software can’t route around a path the hardware never left open.
At bottom, Plan B and supply chain are two sides of the same question: one asks “is our fate held in someone else’s hands,” the other asks “is it held in luck’s hands.” And the binning dimensions, the chiplet recombination freedom, the routable-around features, the degradable configurations — those are the answer to the second question. The bit of area and elegance you spend at the architecture stage buys you cards still left to play the day something goes wrong.
6 | Modularity and Extensibility
How many times the same design can be reused determines a chip company’s economics. And this isn’t just “preparing for next generation” — it should start paying off within the current generation.
Start with vertical reuse. A well-parameterized IP block can be instantiated repeatedly, in different configurations, within the same SKU. What that saves isn’t just design hours — more importantly, it saves verification hours. Something already verified costs approximately nothing extra the next time it’s used, and that’s exactly the most expensive resource there is. One step further is hardened IP: can it be stamped down directly, without re-running synthesis and PD? If so, backend closure risk, timing predictability, and even tapeout schedule all improve — and this tapeout benefits, not just the next one. So the boundary of parameterization needs to be settled at the architecture stage: an IP block that can be configured to do anything can never be fully verified; one that can’t be configured at all is dead after a single use.
Now look at horizontal scale. Whether a chip’s performance scales roughly linearly with the number of interconnected dies is a question this generation of product has to answer. A customer who buys eight cards expects something close to 8x. If they get 5x, that other 3x is silicon they paid for, power they paid for, rack space they paid for, for nothing. Linearity doesn’t emerge on its own — it’s engineered: the ratio between cross-die bandwidth and on-die compute, topology diameter and bisection bandwidth, whether collective communication has hardware support, whether synchronization overhead stays bounded as scale grows. Get these ratios wrong at the architecture stage and software can’t buy them back later — you just get to watch the scaling curve bend over at some size, and that inflection point tends to become the ceiling on your product’s addressable market.
More fundamentally, extensibility should be treated as a basic requirement, not something you check for after the design is done. That means deliberately leaving margin when defining interfaces, setting ratios, and partitioning modules — and margin isn’t free: an extra hop of latency, a bit more area, a somewhat higher per-card cost.
I think that money is worth spending. Per-card cost is a loss you can calculate precisely, paid once. The loss from insufficient flexibility is a loss you can’t calculate precisely at all: it shows up as a customer configuration you can’t meet, a scale point where your scaling curve bends, a market window that opens and closes before you can spin a new SKU. The former is written on the BOM. The latter is written into the orders you never got.
By the same logic, different resources need to be able to scale independently — otherwise all you can offer a customer is a handful of points on a single line, while their needs are scattered across an entire plane. Finally, the interface should outlive the implementation: the inside of a tile can be redone every generation, but the protocol between tiles should stay stable, and the interconnect should lean toward standards wherever possible. Every proprietary protocol you introduce is a debt against future compatibility and ecosystem — and that debt usually starts coming due within the very generation that created it.
At bottom, extensibility was never a technical metric — it’s an economic one. It measures how many times what you’ve verified gets reused, how many markets what you’ve taped out can cover, how many generations the compiler you’ve written can serve. Add all those numbers up, and what they determine is exactly how many people this machine can ultimately reach.
VI. Closing
None of the six things above are flashy. They’re all about cost, fallback, constraint, and other people’s schedules — not some dazzling new idea.
But that’s exactly the point. In an era where the technology tree is deeply uncertain, the biggest value an architect can contribute isn’t picking the single most aggressive path — it’s making sure the whole system can still keep moving no matter which path it ends up on. Keeping a fallback isn’t a failure of nerve — it’s the opposite: it’s what lets you afford to bet at all, what lets you sit back down at the table after losing a hand.
And making that judgment call might be the one truly irreplaceable part of being an architect. Testbenches can be generated. RTL can be generated. Once the design space is locked in, even the design itself can be substantially automated. But deciding what to verify, what to bet on, which piece of complexity to keep for yourself instead of pushing downstream — none of that can be outsourced, because none of it is an execution problem. It’s a judgment problem.
And what “usable at scale” ultimately points to is something plainer than all of this: can more people actually get to use this machine. Today’s AI compute is still concentrated in a handful of companies, partly because of capital and manufacturing capacity — but partly because of people like us doing the architecture. If a chip can only be built with the most expensive materials and the most advanced process, if a machine only runs well when the people who built it are standing next to it, if it only works given a pile of demanding preconditions — then it was destined to belong to a handful of people from the start.
Chip design really is hard, but it’s hard because of how many constraints there are, not because the underlying ideas run deep. Manufacturing is hard — maneuvering at nanometer scale, controlling electrons, reasoning through electromagnetic environments. Architecture itself, though, is more like a craft of trade-offs: converging a pile of conflicting constraints into something someone else can pick up, run under ordinary conditions, and keep running indefinitely. Buildable. Deployable at scale. Affordable to use. In the end, that’s where a computer’s value actually lives.
The best architecture usually isn’t the one that leaves people in awe. It’s the one the most people can pick up and trust.
So — chip architecture is not rocket science. And as the path SpaceX proved shows, even rockets, these days, aren’t really rocket science anymore.
— Shibo Chen