You can edit almost every page by Creating an account and confirming your email.

DPU Dojo Processing Unit

From EverybodyWiki Bios & Wiki









tesla front

The DOJO project

Dojo Processing Unit (DPU) is described by the AI Tesla team as a « virtual device that can be sized according to the application needs ». As the CPU or GPU the DPU is a computing hardware technology.

The Dojo chip is a technology designed by Tesla initially to improve the computing power of their Autopilot AI system available in their cars.[1] Their goal is to scale computing power based on a distributed compute architecture in maximizing the bandwidth and minimizing the latency, exploiting temporal and spatial localities.

The goals of project Dojo are to :

  • Achieve better AI training performance especially in order to train the Neural Network serving Tesla Autopilot,
  • Enable larger and more complex neural networks,
  • Be power efficient and cost effective.

Training node

This is a 64-bit superscalar CPU optimized around matrix multiplication units and vector SIMD. It supports floating point 32 (FP32), binary float point 16 (BFP16) and a new format configurable FP8 (CPF8). All this is backed by 1.25 MB of fast ECC protected SRAM and the low-latency high bandwidth fabric that they have designed:[2]

  • 1024 GFLOPS (BF16/CFP8) or 1 TFLOPS
  • CPU 64 GFLOPS (FP32)
  • 512 GB/s in each cardinal direction of the network
  • 4 Way Multithreaded, allowing compute and data transfer simultaneously
  • Custom ISA optimized for machine learning kernels

Compute array (D1)

The compute array consists of 354 training nodes linked together:[3]

  • 354 Training nodes network[4]
  • Delivering 362 TFLOPS with BF16 and CFP8 of machine learning compute
  • And 22.6 TFLOPS with FP32
  • 10 TBps/dir of transfer for the bisection bandwidth
  • 576 High speed low latency lanes at 122GB
  • 4 TBps/edge off-chip bandwidth (i/o)
  • 400W TDP (Thermal Design Power)
  • 645mm² of surface
  • 7 nm engraved technology
  • 50 Billion transistors
  • 11+ Miles of wires (~17km)

It is described by the AI team lead engineer as having:

  1. GPU level of compute,
  2. CPU level flexibility,
  3. Twice the network chip level I/O

Training tile

A training tile is 25 Dojo chips (D1) put together with electronic components that aims to assume the power of compute such as:[5]

  • Twenty-five of Dojo chip (x25 D1) dies onto a wafer (a thin slice of semiconductor) process seamlessly attached together.
  • High bandwidth with high density connectors to preserve the I/O of 9TB/s on each side of the « training tile » which brings it to 36 TB/s out tile bandwidth.
  • Power is supplied by a custom voltage regulator module that could be reflowed directly onto this fan out wafer.[6]
  • Finally they integrated the entire electrical thermal and mechanical pieces with a 52 volt DC input.
  • Achieving 9 Peta Flops of computing.

Training matrix

The training matrix is a network of 2x3 training tiles linked together:[7]

  • 2x3 Tiles x 2 trays in cabinet
  • 100+ PFLOPS/Cabinet
  • 12 TBps Bisection Bandwidth

ExaPOD

The Exapod is the supercomputer assembled with several training matrix to initially train the model that is used by the Tesla Autopilot. The specifications unveiled at the conference are:[8]

  • 1.1 EFLOP (Exaflop) with BF16 and CFP8
  • 120 Training Tiles
  • 3000 D1 Chips
  • > 1M Training Nodes
  • Still with uniform high bandwidth and low-latency fabric

References

  1. "https://www.businessinsider.in/tech/news/tesla-ai-day-elon-musk-is-working-on-humanoid-robots-a-new-custom-chip-and-ways-to-make-cars-safer/articleshow/85479325.cms". External link in |title= (help)
  2. "Mind blowing specs of Tesla Dojo announced on Tesla AI event". August 20, 2021.
  3. Bos, Chanan (August 22, 2021). "Tesla's Dojo Supercomputer Breaks All Established Industry Standards — CleanTechnica Deep Dive, Part 3". CleanTechnica.
  4. "Enter Dojo: Tesla Reveals Design for Modular Supercomputer & D1 Chip". HPCwire. August 20, 2021.
  5. Novet, Jordan (August 20, 2021). "Tesla unveils chip to train A.I. models inside its data centers". CNBC.
  6. Patel, Dylan. "Tesla Dojo, Unique Packaging and Chip Design Allow An Order Magnitude Advantage Over Competing AI Hardware". semianalysis.substack.com.
  7. "Tesla Packs 50 Billion Transistors Onto D1 Dojo Chip Designed to Conquer Artificial Intelligence Training". Tom's Hardware. August 20, 2021.
  8. "Tesla's AI Day Reveals a New AI-Learning Server, Dojo — And Tesla Bot?". interestingengineering.com. August 20, 2021.


This article "DPU Dojo Processing Unit" is from Wikipedia. The list of its authors can be seen in its historical and/or the page Edithistory:DPU Dojo Processing Unit. Articles copied from Draft Namespace on Wikipedia could be seen on the Draft Namespace of Wikipedia and not main one.