Announced at the INSPIRE 2026 conference in Shanghai in June, the stack addresses the compute, memory, and orchestration demands of persistent, multi-day autonomous agent workloads. Rather than treating agent operations as isolated inference calls, Huawei Cloud packages foundational hardware layers with higher-level model training and agent frameworks.
Scaled compute, context memory, and sandboxed runtimes
The physical and resource-management layer introduces four targeted components:
AI Cluster Service (AICS): Built on Huawei'sUnifiedBus (UB)interconnect, the service is engineered to scale across clusters exceeding 100,000 accelerator cards, providing up to 200 EFLOPS of total compute. Huawei claims the system achieves per-token generation latency under 10 milliseconds, throughput of 5 million tokens per second per 1,000 cards, and 99.95% online service availability.Agentic Memory Storage (AMS): Connects neural processing units directly to Context Memory Storage hardware via NPU passthrough to establish petabyte-scale memory spaces. The architecture supports tiered key-value (KV) cache pooling, designed to lower inference expenses while preserving state across multi-day tasks.CCE VolcanoNext: A scheduling engine that unifies general-purpose and AI computing pools. Huawei states that combining shared training-inference capacity with fragmentation consolidation improves hardware resource utilization by more than 30%.AgentSphere: A secure agent execution environment powered by lightweight sandboxing. Huawei claims the runtime starts instances within 100 milliseconds and can batch-provision hundreds of thousands of isolated environments per minute.
Model routing and agent platform availability
Above the raw infrastructure, Huawei introduced lifecycle tools spanning foundation model routing and production agent construction:
ModelArtsNext: A training and inference engine featuring Reinforcement Learning as a Service (RLaaS), confidential inference, dynamic model routing across three policies (experience-first, efficiency-first, and balanced mode), and a model matrix. Huawei reports it serves over 15 state-of-the-art model services, claiming dynamic routing achieves over 95% scheduling accuracy and reduces invocation costs by an average of 20%.AgentArts: An enterprise-grade agent development platform covering long-running task workflows, enterprise security, and end-to-end observability. Huawei said it was available for open beta testing (OBT) at the announcement.openJiuwen: An open-source edition ofAgentArtslaunched alongside the enterprise platform, sharing more than 90% of its core kernel.AgentArts Orchard: A unified portal automating workflows from intent understanding to cloud resource provisioning using skill- and CLI-based tools.
For confidential execution, Huawei paired the release with Hold Your Own Key (HYOK) encryption, isolated data capsules, and confidential computing virtual machines with PCIe-based NPU passthrough. The company also announced an AI Model Partner Program with more than 20 model providers, including DeepSeek, Zhipu AI, MiniMax, Kimi, and Baidu, while slating its healthcare AI platform and CloudRobo robotics platform for open beta testing on June 30, 2026.
