The cheaper and more efficient artificial intelligence model unveiled by China's DeepSeek has left technology firms around the world at a crossroads over how to press on, according to Jefferies analysts.
Citing “negative implications” for data centre builders, Jefferies said pressure was now on for firms to justify ever-increasing capital expenditure plans.
“The key question for the data centre builders is whether it continues to be a ‘build at all costs’ strategy with accelerated model improvements, or whether focus now shifts towards higher capital efficiency,” the analysts said.
DeepSeek, having last week rolled out its free bot, sparked a global technology stock sell-off on Monday as fears built around so-far dominant names, such as Nvidia Corp.
Jefferies assured that any effects on the likes of data centre demand would likely take over a year to seep through and impact earnings.
“We see limited risk of alterations or cancellations to existing orders and expect at this stage a shift in expectations to higher [returns] on existing investments driven by more efficient models.
“Overall, we remain bullish on the sector where scale leaders benefit from a widening moat and higher pricing power.”
What is DeepSeek?
Other tech analysts at the brokerage noted that DeepSeek has developed an open-source large language model that matches the performance of OpenAI's GPT-4o using a fraction of computing power.
The startup is 100% owned by a highly successful AI-driven quant fund in China, High-Flyer, which created DeepSeek in April 2023 to focus on artificial general intelligence and LLMs, with its second version (V2) launched in May 2024 that achieved a number-7 ranking at the University of Waterloo's LLM leaderboard.
Last month, it launched V4, trained on a data set of 14.8 trillion tokens (compared to 13 trillion for GPT4o), at a training cost of US$5.6 million (assuming US$2/H800 hour rental cost), which the analysts said was less than 10% of the cost of Meta's Llama LLM.
DeepSeek also indicated V3's performance exceeded that of Llama 3.1 and Alibaba's Qwen 2.5, while matching GPT4o and Claude 3.5 Sonnet.
The analysts said the computer architecture is based on Mixture of Experts (MoE) and Multi-head Latent Attention (MLA).
"Each MoE model has ~200bn data parameters, and each query would activate only ~20bn parameters, which lowers the inferencing cost and shortens the response time."
As an open-source model, other AI developers could use it.
"We believe V3 will allow AI developers to develop applications at a much lower cost," the Jefferies tech analysts said, but noting that DeepSeek is "not focused on commercialization, and has not accelerated any AI commercialization".