Skip to main content
The Markets by Proactive
Go to Proactive UK
Proactive UK has moved. Proactive’s coverage of London’s small caps continues on proactiveinvestors.com Go there →
Advertisement
The Markets
by Proactive
Proactive UK has moved.
Coverage of London’s small caps continues on proactiveinvestors.com
Go to Proactive UK
The Markets
by Proactive
Proactive UK has moved.
Small-cap coverage continues on .com
Go to Proactive UK
Advertisement
The Markets
by Proactive
Proactive UK has moved.
Small-cap coverage continues on .com
Go to Proactive UK

Hardware & electrical equipment

Tech Bytes: Alibaba claims 82% cut in Nvidia GPU demand with new AI pooling system

Alibaba Group (NYSE:BABA) has unveiled a home-grown system it says can slash its need for Nvidia Corp (NASDAQ:NVDA, ETR:NVD) chips by more than 80%, marking one of the clearest signs yet that Chinese cloud giants are finding workarounds to US export controls.

The system — called Aegaeon — was developed by Alibaba Cloud engineers and academic partners at Peking University, and it was detailed this month in a paper presented at the ACM Symposium on Operating Systems Principles (SOSP 2025) in Seoul.

According to the research, Aegaeon allows Alibaba to serve multiple large-language models (LLMs) on the same graphics card rather than dedicating one graphics processing unit (GPU) to each model — effectively multiplying hardware efficiency.

Token-level efficiency

Traditional cloud inference systems often waste vast GPU capacity because each model is hosted separately. Even popular models like Qwen or Llama can see fluctuating demand, while niche or “cold” models tie up expensive silicon for only occasional requests, according to the researchers.

Aegaeon changes that through a method the authors call “token-level auto-scaling.” Instead of switching workloads only when an entire model finishes responding, it dynamically reallocates GPU memory and computation at the level of individual output tokens — effectively sharing resources in milliseconds.

That fine-grained approach is what lets Alibaba pool up to seven models per GPU, compared with two or three under existing multi-model serving systems.

Benchmarks in the paper show Aegaeon sustaining 2–2.5 times higher request rates and 1.5–9 times greater “goodput” (useful work performed) than conventional serverless AI platforms. In live deployment at Alibaba Cloud Model Studio, the company said it cut the number of GPUs needed to serve tens of models from 1,192 to 213 — an 82% resource saving.

Built for a new hardware reality

The timing is not coincidental. Since Washington tightened export rules on high-end AI chips, Chinese firms such as Alibaba, Baidu and Tencent have been rationing access to Nvidia’s H800 and A100 GPUs while racing to optimise domestic alternatives like Biren and Huawei Ascend.

By getting more out of each GPU, Alibaba can soften the supply crunch and maintain AI services without relying on constant imports of cutting-edge chips.

The paper emphasises that Aegaeon’s efficiency isn’t theoretical: it has been beta-deployed for more than three months, serving models ranging from 1.8 billion to 72 billion parameters. The team claims the framework’s full-stack optimisations — from memory management to cache synchronisation — cut auto-scaling overhead by 97%.

Strategic implications

If Alibaba can reliably reproduce those savings at scale, the impact extends beyond its cloud division. Lower GPU demand reduces capital expenditure and operating costs, but also signals that software innovation is becoming a key front in the AI hardware race.

In practice, Aegaeon could let Alibaba offer cheaper AI inference services to third-party developers while staying competitive with US hyperscalers. It could also embolden other Chinese tech groups to publish similar breakthroughs as part of a broader effort to demonstrate technological independence from US chip suppliers.

At the same time, GPU pooling introduces new trade-offs — such as higher engineering complexity and potential latency spikes during heavy loads — and will likely complement rather than replace physical GPU expansion.

The bigger picture

For investors, the story underlines how cloud and AI infrastructure are evolving in response to geopolitics as much as demand. Nvidia’s market dominance is being challenged not only by export controls and rival chipmakers but also by algorithmic efficiency — the ability to do more with less.

If Aegaeon’s reported 82% efficiency gain holds up in production, it could reshape cloud economics for AI in China and, eventually, globally.

Advertisement
The Markets
by Proactive
Proactive UK has moved.
Small-cap coverage continues on .com
Go to Proactive UK