Custom Chips for AI

Intermediate

Why GPUs, TPUs, NPUs, FPGAs, and custom ASICs matter for modern AI — and how to pick the right accelerator for a workload.

Last updated: Sep 13, 2026

Why AI is a hardware story now

Modern AI is mostly enormous linear algebra under brutal constraints: power, memory bandwidth, latency, interconnect, and cost. GPUs won because they are programmable and massively parallel. Custom chips exist because hyperscalers can save huge money when a workload is predictable enough to harden into silicon.

The key trade-off is specialization versus flexibility. A custom ASIC can be far more efficient than a general GPU, but only if the models, operators, compiler, and deployment scale justify the loss of flexibility.

Choose constraints before a chip

These are qualitative questions, not benchmark scores. Chip families overlap: a TPU is an ASIC, while NPU can describe several accelerator designs. A specific device and supported runtime determine feasibility.

  1. 1. Does the model and its KV cache fit in accessible memory?
  2. 2. Does the compiler/runtime support every required operator and precision?
  3. 3. What latency and throughput does the exact workload need?
  4. 4. What power and cooling budget is available?
  5. 5. What are the hardware, hosting and engineering costs?
GPU

GPU: broad software support and flexible kernels. Check actual memory, interconnect and supported precision.

TPU

TPU: tensor accelerator with a specific compiler and distributed-system stack. Verify model support and access to the required generation.

NPU

NPU: often intended for local inference within power limits. Memory, supported operators and vendor tooling can restrict the model.

ASIC

Custom ASIC: specializes hardware around a workload. Design cost, volume and software support determine whether that specialization pays.

FPGA

FPGA: reconfigurable logic for a custom pipeline. Mapping operators and sustaining memory throughput require engineering.

The strategic reason companies build their own chips

Custom silicon is not just about peak speed. It is about supply, margin, roadmap control, and avoiding total dependence on one vendor. Google TPUs, AWS Trainium and Inferentia, Microsoft Maia, and Meta MTIA all reflect the same pressure: AI demand is so large that hardware becomes product strategy.

But the risk is real: if the software stack is weak or model architecture shifts, a specialized chip can become an expensive dead end. The winning platform is usually chip plus compiler plus kernels plus networking plus developer experience.

JAX Scaling Book: Rooflines