Timeline

AWS unveils Trainium3, its first 3nm AI chip, at re:Invent

Its first accelerator on a 3-nanometre process, AWS claimed roughly 4.4x the compute and 4x the performance-per-watt of Trainium2, scaling to hundreds of thousands of chips per cluster.

  • Compute & infrastructure
  • Notable

AWS announced Trainium3, its fourth-generation custom AI accelerator and the first built on a 3-nanometre process, alongside general availability of EC2 Trn3 UltraServers built around it, at its re:Invent conference.

AWS said Trainium3 delivered roughly 4.4 times the compute performance, about 3.9 times the memory bandwidth and around 4 times the performance-per-watt of its predecessor, Trainium2, with each chip carrying 144GB of HBM3e memory and 4.9TB/s of memory bandwidth — increases the company attributed to the smaller process node and architectural changes. A fully configured Trn3 UltraServer combines up to 144 chips into a single system, delivering a claimed 362 petaflops of FP8 compute and 706TB/s of aggregate memory bandwidth; AWS said the servers link into EC2 UltraClusters capable of scaling to hundreds of thousands of chips. On Amazon Bedrock, AWS claimed up to three times faster performance and five times more output tokens per megawatt using the new hardware.

All figures were AWS’s own; no independent third-party benchmarking accompanied the announcement. Trainium3 continued Amazon’s push to offer an in-house alternative to Nvidia GPUs for both training and inference, part of a broader industry pattern in which hyperscalers built custom silicon to reduce dependence on a single chip supplier and to control the unit economics of running increasingly large models at scale.