September 24, 2026, (Inside AI) — The AI infrastructure industry is pivoting from a singular focus on raw GPU count to a more nuanced metric: token efficiency. At the 2026 Apsara Conference, Chinese technology firm H3C unveiled a comprehensive portfolio aimed at optimizing every layer of the AI stack, from compute and networking to storage and software, to maximize the number of useful tokens produced per GPU. The move signals a maturing market where operational costs and real-world throughput trump theoretical peak performance.
The shift comes as agentic AI systems, which continuously call foundation models, drive demand for trillion-token-scale services. Simply adding more GPUs no longer guarantees proportional performance gains. Idle compute, network congestion, insufficient data supply, and the operational challenges of large-scale clusters erode the value of additional hardware. H3C's strategy addresses these bottlenecks through system-level optimization rather than hardware stacking.
Beyond GPUs: The System-Level Imperative
Zhu Shiyin, general manager of H3C's Advanced Technology Research Department, described AI clusters as highly integrated systems requiring close hardware-software coordination. As clusters grow from hundreds to tens of thousands of GPUs, chip-to-chip communication, data transfer, task scheduling, power supply, and cooling directly affect overall compute utilization.
H3C's UniPoD S80000 Series SuperPod, showcased at the event, embodies this approach. It supports configurations from 32 to 1,024 GPUs and scales to 16,384 GPUs, while accommodating heterogeneous resources including CPUs, GPUs, NPUs, and DPUs. The focus is on coordinating compute, networking, storage, cloud, security, and operations as a unified system. In other words, the competition is shifting from how many GPUs a system has to how many useful tokens each GPU can produce.
Read: Huawei Ascend 960 SuperPoD Challenges NVIDIA AI Dominance
This emphasis on efficiency is not unique to H3C. Rival firms like Nvidia and AMD have also been vocal about the importance of system-level optimization. Nvidia's DGX SuperPOD and InfiniBand networking solutions, for instance, aim to reduce communication overhead. However, H3C's approach integrates storage and software more tightly, potentially offering a more holistic solution for enterprises deploying agentic AI at scale.
Interconnects: The New Bottleneck
As GPU performance improves, the volume of data exchanged between GPUs grows, making the network a critical bottleneck. H3C showcased three interconnect scenarios: Scale-Up, Scale-Out, and Scale-Across.
Within a node, the S9828-128EO, a 102.4T NPO silicon-photonics intelligent computing switch, uses co-packaged optics to reduce end-to-end latency by 15%. Between nodes, the 1.6T intelligent computing switch S9828-64FP adopts 224G SerDes technology, cutting power consumption by over 20% in high-density environments. Across data centers, the 800G intelligent computing DCI switch S12500R-64EP supports intra-city cross-data-center compute scheduling and gradient synchronization.
These innovations address a core question: as AI clusters grow, how can interconnects avoid becoming brakes on compute performance? The answer lies not just in bandwidth but in protocols, congestion control, link reliability, and fault diagnosis. Closed technical ecosystems could create new silos, making open protocols and industry standards increasingly important. H3C's support for mainstream frameworks and open standards positions it well in this regard.
Data is another frequently underestimated challenge. GPUs cannot run without data. In large-model training, data preparation, model training, and parameter exchange involve massive data movement. During inference, KV Cache adds further pressure on storage and caching resources. H3C's UniStor X20000 series X20836 delivers up to 200GB/s bandwidth and 3 million IOPS per node, supporting block, file, object, and HDFS protocols. According to H3C, its full-speed engine reduces GPU waiting time by 30%, while XCache inference acceleration cuts time-to-first-token latency by up to 90%.
These figures, if independently verified, represent significant improvements. Inside AI could not independently verify these claims. However, the broader trend is clear: AI infrastructure is evolving from a compute center into a comprehensive data processing system. Every period of waiting translates into cost, and storage is no longer a passive repository but an active participant in token generation.
Software may be the next battleground. Zhu noted that everything from the operating system and communication libraries to resource management and scheduling platforms must work closely with the underlying hardware. Communication optimization, overlapping computation with communication, task scheduling, dynamic load balancing, and unified resource management all help reduce GPU idle time and communication conflicts. Modern AI infrastructure increasingly resembles a super-system: chips provide compute, networks connect resources, storage supplies data, and power and cooling keep the system running. Software orchestrates these resources into a functioning whole.
If any component becomes a bottleneck, the impact is no longer limited to a single hardware metric. It ultimately affects how many useful tokens the system can generate per second and how much each token costs to produce. This holistic view is gaining traction across the industry. For instance, Google has emphasized the role of its TPU and Vertex AI in optimizing the entire ML pipeline, while Microsoft has integrated its Azure AI stack with Nvidia GPUs and InfiniBand networking. H3C's approach, however, is notable for its inclusion of storage and software as first-class citizens in the token production pipeline.
The next phase of AI infrastructure competition may not be determined simply by who has more GPUs. As model sizes reach the trillion-parameter scale and AI agents increasingly call foundation models, metrics such as compute utilization, token throughput, latency, energy consumption, and cost per token are becoming increasingly important. That is the deeper meaning behind the pursuit of extreme token cost efficiency: it is not about putting more GPUs together, but about making every GPU, every network link, and every unit of storage spend less time waiting and more time working.
As AI enters the stage of large-scale production, compute is no longer just a chip problem. It is becoming a competition over the efficiency of the entire infrastructure system. H3C's showcase at the Apsara Conference underscores this shift, and the industry would do well to heed the message. The company's focus on token efficiency, rather than raw GPU count, may well define the next wave of AI infrastructure leadership.