Alphabet Inc (NASDAQ:GOOG)’s search giant Google has claimed that the processors powering its supercomputers used to train its PaLM AI model beats market leader Nvidia Corp’s comparable processors in both speed and energy efficiency.
A new scientific paper on Google’s Tensor Processing Unit (TPU), which is a custom-designed chip that the company uses for over 90% of its AI training work, detailed how it has strung more than 4,000 of the chips together into a supercomputer, using custom-developed optical switches to connect individual machines.
Google said that for comparably sized systems, its chips are up to 1.7 times faster and 1.9 times more power-efficient than a system based on Nvidia's current-generation A100 chip.
Improving connections between chips has become a key point of competition among companies that build AI supercomputers used to power technologies like Google's Bard or OpenAI's ChatGPT.
"Circuit switching makes it easy to route around failed components," said Google’s Norm Jouppi and David Patterson in a blog post. "This flexibility even allows us to change the topology of the supercomputer interconnect to accelerate the performance of an ML (machine learning) model."
ChatGPT’s model was trained using up 1,000 Nvidia processors, while Google's PaLM model was trained by splitting it across two of the 4,000-chip supercomputers over 50 days.
Google has been using the supercomputer since 2020 in a data centre in Mayes County, Oklahoma, though has only recently begun releasing details of the system.
Google hinted that it might be working on a new TPU that would compete with Nvidia's next-gen H100 processor but provided no details.