Why AI is a hardware story now
Modern AI is mostly enormous linear algebra under brutal constraints: power, memory bandwidth, latency, interconnect, and cost. GPUs won because they are programmable and massively parallel. Custom chips exist because hyperscalers can save huge money when a workload is predictable enough to harden into silicon.
The key trade-off is specialization versus flexibility. A custom ASIC can be far more efficient than a general GPU, but only if the models, operators, compiler, and deployment scale justify the loss of flexibility.
Choose constraints before a chip
These are qualitative questions, not benchmark scores. Chip families overlap: a TPU is an ASIC, while NPU can describe several accelerator designs. A specific device and supported runtime determine feasibility.
- 1. Does the model and its KV cache fit in accessible memory?
- 2. Does the compiler/runtime support every required operator and precision?
- 3. What latency and throughput does the exact workload need?
- 4. What power and cooling budget is available?
- 5. What are the hardware, hosting and engineering costs?
GPU
GPU: broad software support and flexible kernels. Check actual memory, interconnect and supported precision.
TPU
TPU: tensor accelerator with a specific compiler and distributed-system stack. Verify model support and access to the required generation.
NPU
NPU: often intended for local inference within power limits. Memory, supported operators and vendor tooling can restrict the model.
ASIC
Custom ASIC: specializes hardware around a workload. Design cost, volume and software support determine whether that specialization pays.
FPGA
FPGA: reconfigurable logic for a custom pipeline. Mapping operators and sustaining memory throughput require engineering.
The strategic reason companies build their own chips
Custom silicon is not just about peak speed. It is about supply, margin, roadmap control, and avoiding total dependence on one vendor. Google TPUs, AWS Trainium and Inferentia, Microsoft Maia, and Meta MTIA all reflect the same pressure: AI demand is so large that hardware becomes product strategy.
But the risk is real: if the software stack is weak or model architecture shifts, a specialized chip can become an expensive dead end. The winning platform is usually chip plus compiler plus kernels plus networking plus developer experience.
JAX Scaling Book: Rooflines