Mon-Sun 9:00-21:00
Powered by the "PingHu" Architecture
Defining Flagship AI Compute for the Generative AI Era
MTT S5000 is a Universal GPU engineered for the generative AI era, purpose-built for large model training, inference, and high-performance computing. With Moore Threads' next-generation PH100 chip at its core, and built on the advanced "PingHu" architecture, MTT S5000 delivers full-precision compute support from FP8 to FP64. The PH100 chip is among the first to receive security and reliability certification from the China Information Technology Security Evaluation Center (2026 No. 2 Certification).
Powered by the fourth-generation MUSA full-stack platform, MTT S5000 transcends ecosystem barriers with native support for PyTorch, Megatron-LM, vLLM, and SGLang, enabling near-zero-effort code migration. Whether building a 10K-GPU training cluster or deploying high-concurrency, low-latency inference services, MTT S5000 delivers capabilities benchmarked against the world's leading flagship GPUs. Backed by the PH100 chip's performance and high-level security, it provides a solid, accessible, and fully sovereign foundation for compute.
Native FP8 Computing
The New Standard for Trillion-Parameter Training and Inference
As model scale continues to grow, native FP8 has become a critical precision standard for training and inference of frontier models. MTT S5000 is among the first GPUs to implement hardware-native FP8 compute, delivering significant memory efficiency gains while preserving model convergence. To address bandwidth and compute bottlenecks in large-scale clusters, its FP8 engine provides full support for frontier model architectures, including DeepSeek and Qwen, improving training performance by over 30% and significantly reducing the total cost of ownership.
The Benchmark for Trillion-Parameter Model Training on 10K-GPU Clusters
Addressing the extreme demands of trillion-parameter model training in both compute scale and system stability, MTT S5000 is designed to deliver near-linear performance scaling from a single GPU to a 10K-GPU cluster. In large-scale pre-training workloads at the 10K-GPU scale, MTT S5000 demonstrates strong numerical consistency, with training outcomes closely aligned with leading international flagship GPUs. The breakthrough Model FLOPs Utilization (MFU) enables highly efficient compute utilization, ensuring stable execution and near-peak performance across large-scale training workloads.

Comprehensive Training Performance
Multimodal fine-tuning advantage: In fine-tuning tasks for multimodal large models, MTT S5000 surpasses leading international flagship products①.

Large-Scale Cluster Efficiency
DeepSeek-236B training: Model FLOPs Utilization (MFU) exceeds 40%.

10K-GPU Cluster Training Precision
Superior downstream performance: Under equivalent data volumes, downstream task evaluation scores exceed those of international flagship products①, validating the precision capabilities of the 10K-GPU cluster.

Near-Linear Cluster Scaling

Reliability, Availability, and Serviceability
MTT S5000 builds a comprehensive RAS framework from chip to system level. It supports fault detection, reporting, and error isolation to rapidly identify and replace failed, slow, or silent data corruption (SDC) nodes. Proactive monitoring and self-healing mechanisms safeguard long-term cluster health, ensuring consistent performance and result integrity. Leveraging chip-to-system RAS capabilities, MTT S5000 improves training success rates in 10K-GPU clusters by over 30%.
Instant Response and Real-Time Interaction
Delivering High-Performance Inference for Large Models
MTT S5000 is designed to eliminate the barrier between compute and interaction. Through deep hardware-software co-optimization, it delivers near-zero-latency responses for AI applications. For large language model inference, MTT S5000 applies full-stack optimization, integrating high-performance inference stacks including SGLang and vLLM to fully exploit its architectural advantages in long-sequence processing and high-frequency decoding.
For long-context understanding tasks, MTT S5000 dramatically accelerates the Prefill stage, reducing time-to-first-token (TTFT) and enabling immediate responses to complex instructions. For content generation, deep optimization of Decode throughput delivers exceptional token output speeds. This inference performance not only meets user demand for real-time interaction but also delivers superior compute efficiency and reliable service quality under high-concurrency workloads for enterprises.

Prefill Breakthrough
Instant First-Token Response for Long Contexts
In 16K long-sequence input benchmarks, MTT S5000 achieves single-GPU Prefill throughput 2.5× that of leading international flagship products②. This means faster context understanding when processing long-text prompts.
Empowering Agentic AI and AI Coding with Real-Time Inference
In latency-sensitive scenarios such as agentic AI and AI coding assistance, MTT S5000 demonstrates exceptional inference performance. Optimized for the rapid agent-to-agent communication and instant code generation demanded by these workloads, MTT S5000 achieves token generation rates far exceeding industry benchmarks on frontier models, including DeepSeek.
This near-imperceptible output latency ensures real-time instruction flow and feedback in multi-agent collaboration (Agent-to-Agent), dramatically raising productivity ceilings and establishing a robust foundation for the performance of next-generation autonomous, collaborative AI applications.
≥4000Tokens/s
Single-GPU Prefill Throughput
ISL: 4K; mean TTFT ≤ 4s; MTP on. PD-disaggregated inference with 2P4D cluster configuration. Hardware: 6 × MTT S5000 8-GPU servers. Software: SiliconFlow high-performance inference engine with Moore Threads underlying software stack. Tested at native FP8 precision under conditions consistent with high-load production environments.
≥1000Tokens/s
Single-GPU Decode Throughput
ISL/OSL: 1K/2K; mean TPOT ≤ 100ms; MTP on. PD-disaggregated inference with 2P4D cluster configuration. Hardware: 6 × MTT S5000 8-GPU servers. Software: SiliconFlow high-performance inference engine with Moore Threads underlying software stack. Tested at native FP8 precision under conditions consistent with high-load production environments.
Accelerating Multimodal Foundation Model Training and Inference
Efficient Multimodal Training
Accelerating Video Generation Innovation
Leveraging exceptional tensor compute and a highly optimized parallel framework, MTT S5000 delivers top-tier engineering performance in video and image generation. Built on the FSDP2 framework, MTT S5000 has completed full-model training validation for Wan2.1 video generation.
In a 2-node, 16-GPU configuration, MTT S5000 achieves a training throughput of 61.83 samples/s with a Model FLOPs Utilization (MFU) of 51%. Generated outputs align precisely with industry benchmarks in video logic, visual fidelity, and motion consistency. From efficient pre-training to high-concurrency inference, MTT S5000 provides full-stack compute support for the ongoing evolution of video generation models.
Drone view of waves crashing against the rugged cliffs along Big Sur's garay point beach.The crashing blue waters create white-tipped waves,while the golden light of the setting sun illuminates the rocky shore. A small island with a lighthouse sits in the distance, and green shrubbery covers the cliffs edge. The steep drop from the road down to the beach is adramatic feat, with the cliff's edges jutting out over the sea. This is a view that captures the raw beauty of the coast and the rugged landscape of the Pacific Coast Highway.
Text-to-Video Inference at Scale
MTT S5000 extensively optimizes text-to-video models. Leveraging native FP8 hardware acceleration, it dramatically improves inference speed while maintaining lossless output quality, delivering a highly competitive video generation inference solution for enterprise deployments.
Strong Performance
Single-node performance reaches 64%–79% of leading international flagship products①, balancing high throughput with compelling ROI.
Native FP8 Lossless Inference
Hardware-native FP8 precision dramatically improves inference throughput while preserving video generation quality, achieving the optimal balance of speed and fidelity.
Flexible, Scalable Deployment
Fully compatible with mainstream video generation models, with seamless scaling from single-node to large-scale cluster deployments to meet diverse business requirements.
Native FP64 Double-Precision Compute
Advancing Scientific Computing and AI for Science Across Generations
MTT S5000's powerful native FP64 double-precision compute makes it a high-performance engine for scientific computing and AI for Science (AI4S). Through deep collaboration and co-optimization with national laboratories, it delivers significant performance gains across key scientific computing domains.
Full-Precision Hardware Support
Full-precision compute from FP8 to FP64 ensures the extreme data precision required in scientific simulation, providing a reliable foundation for high-fidelity computation and the convergence of AI and science.
Molecular Dynamics Breakthrough
In the SPONGE simulation engine, MTT S5000's highly efficient parallel compute architecture achieves 1.7× the performance of leading international flagship products③.
Biopharma Acceleration
In benchmarks using the molecular docking tool DSDP, MTT S5000 demonstrates an overwhelming performance advantage, achieving 8.1× the performance of leading international flagship products③.
Broad Domestic Ecosystem Compatibility
Rapid integration with mainstream scientific computing software stacks, providing an efficient and sovereign compute foundation for materials science, bioengineering, weather forecasting, and more.
High-Performance Multimedia Processing
As a Universal GPU, MTT S5000 integrates a high-performance multimedia codec engine, delivering powerful media processing capabilities for cloud multimedia and AI convergence workloads. Hardware-native video decode support: H.264, H.265, VP9, AV1, AVS2, AVS+, VP8. Hardware-native video encode support: H.264, H.265, AV1.
160Streams
High-Concurrency Decode
Up to 160 concurrent 1080P30 video decode streams, Meeting the demands of high-density video analytics workloads
40Streams
High-Definition Decode
Up to 40 concurrent 4K30 video decode streams, Enabling high-fidelity video analysis at scale
Hardware-Level Security
MTT S5000 implements a hardware-level Trusted Execution Environment (TEE), establishing a full-stack security perimeter from the underlying hardware to the application layer, providing robust protection for sensitive data and model assets.
Hardware Root of Trust (HRoT)
An integrated hardware root of trust supports secure boot, secure firmware update, and firmware protection, ensuring security and control across the full device lifecycle from power-on to runtime.
Confidential Computing Support
Hardware isolation technology provides encryption protection during data processing, preventing unauthorized access and enabling secure data sharing in high-sensitivity industries such as finance and healthcare.
Native Support for SM Cryptographic Algorithms
Comprehensive support for Chinese national cryptographic standards (SM algorithms), meeting the encryption requirements of trusted and innovative IT (Xinchuang) compliance and high-security-level business workloads.
Near-Zero-Cost Software Migration Path
Thanks to MUSA's high compatibility with CUDA syntax, developers can directly reuse existing code and expertise, enabling seamless ecosystem migration at near-zero cost.
MUSIFY Automatic Porting Tool
Moore Threads' proprietary MUSIFY tool automatically identifies and converts code, compressing what was once weeks of adaptation work into hours, enabling true plug-and-play migration.
No Core Code Refactoring Required
The MUSA SDK provides complete support for Runtime APIs, Driver APIs, and mainstream math libraries. Existing GPU applications run efficiently on MTT S5000 without rewriting core logic.
Native Support for Mainstream AI Frameworks
Deep integration with PyTorch, PaddlePaddle, and other leading frameworks, with full support via plugins such as Torch-MUSA, ensures rapid production deployment.
All-Around Compute Engine
Powering Innovation in Frontier Research and Cross-Disciplinary Science
With full-precision compute spanning FP8 to FP64, MTT S5000 leverages a unified, general-purpose architecture to handle high-precision and mixed-precision AI+ workloads flexibly. It is applicable across a wide range of frontier application domains, including quantum technology, embodied intelligence, world models, medical imaging, genomics, industrial large models, physics simulation, and molecular dynamics.
AI + Quantum Technology Quantum Technology
AI + Embodied Intelligence Embodied Intelligence
AI + Medical Imaging Medical Imaging
AI + Genomics Genomics
AI + Industrial Large Models Industrial Models
AI + Physics Simulation Physics Simulation
AI + Molecular Dynamics Molecular Dynamics
A Comprehensive MUSA Software Ecosystem for Every Workload, Accelerating Innovation
The full-stack MUSA software platform encompasses drivers, the MUSA SDK, KUAE Training Suite, and KUAE Inference Suite, providing end-to-end coverage for AI training, AI inference, and scientific computing workloads.
The MUSA SDK includes core components such as deep learning acceleration libraries and math compute libraries that efficiently optimize matrix operations, common deep learning operators, collective communication, linear algebra, and general-purpose math. Built on the MUSA software stack, MTT S5000 delivers outstanding compute efficiency and a streamlined migration experience, helping research institutions and enterprises rapidly achieve performance gains and production deployment in large model training, scientific computing, image processing, and beyond.
MTT S5000 Specifications
Flexible Hardware Form Factors to Meet Diverse Deployment Needs

OAM Compute Module
The liquid-cooled variant is purpose-built for high-density green data centers, maximizing compute density while significantly reducing Power Usage Effectiveness (PUE) and energy consumption.
The air-cooled variant is compatible with standard off-the-shelf servers, offering flexible, straightforward deployment that minimizes operational overhead and the long-term cost of ownership.

MTT MGX 8-GPU Modular Platform

MTT SGX5000 Server
Fully optimized for thermal management, power delivery, and I/O expandability, the SGX5000 is available in both air-cooled and liquid-cooled configurations to meet the requirements of diverse data center environments.
Available with pre-installed, optimized training and inference software stacks for integrated hardware-software delivery — ready out of the box.
① Baseline 1: Dense FP16 compute 989 TFLOPs, memory bandwidth 3.35 TB/s; ② Baseline 2: Dense FP16 compute 148 TFLOPs, memory bandwidth 4.0 TB/s; ③ Baseline 3: Dense FP16 compute 312 TFLOPs, memory bandwidth 2.0 TB/s.

中文

