1. Products
    • Hardware
    • AI Data Center

    Entertainment and Creation

  2. Solutions
  3. MetaPark
  4. Support
  5. Company
  6. language 中文
MTT S5000

MTT S5000

Powered by the "PingHu" Architecture

Defining Flagship AI Compute for the Generative AI Era


MTT S5000 is a Universal GPU engineered for the generative AI era, purpose-built for large model training, inference, and high-performance computing. With Moore Threads' next-generation PH100 chip at its core, and built on the advanced "PingHu" architecture, MTT S5000 delivers full-precision compute support from FP8 to FP64. The PH100 chip is among the first to receive security and reliability certification from the China Information Technology Security Evaluation Center (2026 No. 2 Certification).

Powered by the fourth-generation MUSA full-stack platform, MTT S5000 transcends ecosystem barriers with native support for PyTorch, Megatron-LM, vLLM, and SGLang, enabling near-zero-effort code migration. Whether building a 10K-GPU training cluster or deploying high-concurrency, low-latency inference services, MTT S5000 delivers capabilities benchmarked against the world's leading flagship GPUs. Backed by the PH100 chip's performance and high-level security, it provides a solid, accessible, and fully sovereign foundation for compute.

Native FP8 Computing

The New Standard for Trillion-Parameter Training and Inference


As model scale continues to grow, native FP8 has become a critical precision standard for training and inference of frontier models. MTT S5000 is among the first GPUs to implement hardware-native FP8 compute, delivering significant memory efficiency gains while preserving model convergence. To address bandwidth and compute bottlenecks in large-scale clusters, its FP8 engine provides full support for frontier model architectures, including DeepSeek and Qwen, improving training performance by over 30% and significantly reducing the total cost of ownership.

The Benchmark for Trillion-Parameter Model Training on 10K-GPU Clusters


Addressing the extreme demands of trillion-parameter model training in both compute scale and system stability, MTT S5000 is designed to deliver near-linear performance scaling from a single GPU to a 10K-GPU cluster. In large-scale pre-training workloads at the 10K-GPU scale, MTT S5000 demonstrates strong numerical consistency, with training outcomes closely aligned with leading international flagship GPUs. The breakthrough Model FLOPs Utilization (MFU) enables highly efficient compute utilization, ensuring stable execution and near-peak performance across large-scale training workloads.

Comprehensive Training Performance

Comprehensive Training Performance

Mainstream model benchmarks: In core training scenarios across computer vision (CV) and large language models (LLMs), MTT S5000 achieves over 74% of the performance of leading international flagship products.

Multimodal fine-tuning advantage: In fine-tuning tasks for multimodal large models, MTT S5000 surpasses leading international flagship products.
Large-Scale Cluster Efficiency

Large-Scale Cluster Efficiency

Llama3-70B training: Model FLOPs Utilization (MFU) exceeds 60%.
DeepSeek-236B training: Model FLOPs Utilization (MFU) exceeds 40%.
10K-GPU Cluster Training Precision

10K-GPU Cluster Training Precision

Full precision alignment: During DeepSeek-236B training on a 10K-GPU cluster, the Loss curve maintains a relative precision deviation of just 0.6% compared to clusters running international flagship products, measured over the first 30,000 training steps.

Superior downstream performance: Under equivalent data volumes, downstream task evaluation scores exceed those of international flagship products, validating the precision capabilities of the 10K-GPU cluster.
Near-Linear Cluster Scaling

Near-Linear Cluster Scaling

In large-scale AI model training, MTT S5000 achieves breakthrough cluster linearity of up to 95% at the 10K-GPU scale. As compute scales from thousands to tens of thousands of GPUs, overall training throughput continues to grow at near-linear rates, enabling larger models and more complex workloads while significantly compressing pre-training timelines.
Reliability, Availability, and Serviceability

Reliability, Availability, and Serviceability

Reliability, availability, and serviceability (RAS) are the foundational infrastructure capabilities that keep large-scale AI training running continuously and stably.

MTT S5000 builds a comprehensive RAS framework from chip to system level. It supports fault detection, reporting, and error isolation to rapidly identify and replace failed, slow, or silent data corruption (SDC) nodes. Proactive monitoring and self-healing mechanisms safeguard long-term cluster health, ensuring consistent performance and result integrity. Leveraging chip-to-system RAS capabilities, MTT S5000 improves training success rates in 10K-GPU clusters by over 30%.
Instant Response and Real-Time Interaction

Delivering High-Performance Inference for Large Models


MTT S5000 is designed to eliminate the barrier between compute and interaction. Through deep hardware-software co-optimization, it delivers near-zero-latency responses for AI applications. For large language model inference, MTT S5000 applies full-stack optimization, integrating high-performance inference stacks including SGLang and vLLM to fully exploit its architectural advantages in long-sequence processing and high-frequency decoding.

For long-context understanding tasks, MTT S5000 dramatically accelerates the Prefill stage, reducing time-to-first-token (TTFT) and enabling immediate responses to complex instructions. For content generation, deep optimization of Decode throughput delivers exceptional token output speeds. This inference performance not only meets user demand for real-time interaction but also delivers superior compute efficiency and reliable service quality under high-concurrency workloads for enterprises.

Prefill Breakthrough

Instant First-Token Response for Long Contexts

Prefill Breakthrough
Instant First-Token Response for Long Contexts

MTT S5000 delivers deep optimization of Prefill-stage processing efficiency. Under ultra-long sequence inputs, it substantially accelerates prompt pre-processing, enabling faster context comprehension and first-token response, directly addressing latency bottlenecks in large-scale knowledge retrieval and long-document analysis.

In 16K long-sequence input benchmarks, MTT S5000 achieves single-GPU Prefill throughput 2.5× that of leading international flagship products. This means faster context understanding when processing long-text prompts.

Empowering Agentic AI and AI Coding with Real-Time Inference


In latency-sensitive scenarios such as agentic AI and AI coding assistance, MTT S5000 demonstrates exceptional inference performance. Optimized for the rapid agent-to-agent communication and instant code generation demanded by these workloads, MTT S5000 achieves token generation rates far exceeding industry benchmarks on frontier models, including DeepSeek.

This near-imperceptible output latency ensures real-time instruction flow and feedback in multi-agent collaboration (Agent-to-Agent), dramatically raising productivity ceilings and establishing a robust foundation for the performance of next-generation autonomous, collaborative AI applications.


≥4000Tokens/s

Single-GPU Prefill Throughput

ISL: 4K; mean TTFT ≤ 4s; MTP on. PD-disaggregated inference with 2P4D cluster configuration. Hardware: 6 × MTT S5000 8-GPU servers. Software: SiliconFlow high-performance inference engine with Moore Threads underlying software stack. Tested at native FP8 precision under conditions consistent with high-load production environments.

≥1000Tokens/s

Single-GPU Decode Throughput

ISL/OSL: 1K/2K; mean TPOT ≤ 100ms; MTP on. PD-disaggregated inference with 2P4D cluster configuration. Hardware: 6 × MTT S5000 8-GPU servers. Software: SiliconFlow high-performance inference engine with Moore Threads underlying software stack. Tested at native FP8 precision under conditions consistent with high-load production environments.


Accelerating Multimodal Foundation Model Training and Inference

Efficient Multimodal Training
Accelerating Video Generation Innovation

Leveraging exceptional tensor compute and a highly optimized parallel framework, MTT S5000 delivers top-tier engineering performance in video and image generation. Built on the FSDP2 framework, MTT S5000 has completed full-model training validation for Wan2.1 video generation.

In a 2-node, 16-GPU configuration, MTT S5000 achieves a training throughput of 61.83 samples/s with a Model FLOPs Utilization (MFU) of 51%. Generated outputs align precisely with industry benchmarks in video logic, visual fidelity, and motion consistency. From efficient pre-training to high-concurrency inference, MTT S5000 provides full-stack compute support for the ongoing evolution of video generation models.

Drone view of waves crashing against the rugged cliffs along Big Sur's garay point beach.The crashing blue waters create white-tipped waves,while the golden light of the setting sun illuminates the rocky shore. A small island with a lighthouse sits in the distance, and green shrubbery covers the cliffs edge. The steep drop from the road down to the beach is adramatic feat, with the cliff's edges jutting out over the sea. This is a view that captures the raw beauty of the coast and the rugged landscape of the Pacific Coast Highway.

Text-to-Video Inference at Scale


MTT S5000 extensively optimizes text-to-video models. Leveraging native FP8 hardware acceleration, it dramatically improves inference speed while maintaining lossless output quality, delivering a highly competitive video generation inference solution for enterprise deployments.

Strong Performance

Single-node performance reaches 64%–79% of leading international flagship products, balancing high throughput with compelling ROI.

Native FP8 Lossless Inference

Hardware-native FP8 precision dramatically improves inference throughput while preserving video generation quality, achieving the optimal balance of speed and fidelity.

Flexible, Scalable Deployment

Fully compatible with mainstream video generation models, with seamless scaling from single-node to large-scale cluster deployments to meet diverse business requirements.

Native FP64 Double-Precision Compute

Advancing Scientific Computing and AI for Science Across Generations


MTT S5000's powerful native FP64 double-precision compute makes it a high-performance engine for scientific computing and AI for Science (AI4S). Through deep collaboration and co-optimization with national laboratories, it delivers significant performance gains across key scientific computing domains.

Full-Precision Hardware Support

Full-precision compute from FP8 to FP64 ensures the extreme data precision required in scientific simulation, providing a reliable foundation for high-fidelity computation and the convergence of AI and science.

Molecular Dynamics Breakthrough

In the SPONGE simulation engine, MTT S5000's highly efficient parallel compute architecture achieves 1.7× the performance of leading international flagship products.

Biopharma Acceleration

In benchmarks using the molecular docking tool DSDP, MTT S5000 demonstrates an overwhelming performance advantage, achieving 8.1× the performance of leading international flagship products.

Broad Domestic Ecosystem Compatibility

Rapid integration with mainstream scientific computing software stacks, providing an efficient and sovereign compute foundation for materials science, bioengineering, weather forecasting, and more.

High-Performance Multimedia Processing


As a Universal GPU, MTT S5000 integrates a high-performance multimedia codec engine, delivering powerful media processing capabilities for cloud multimedia and AI convergence workloads. Hardware-native video decode support: H.264, H.265, VP9, AV1, AVS2, AVS+, VP8. Hardware-native video encode support: H.264, H.265, AV1.


160Streams

High-Concurrency Decode

Up to 160 concurrent 1080P30 video decode streams, Meeting the demands of high-density video analytics workloads

40Streams

High-Definition Decode

Up to 40 concurrent 4K30 video decode streams, Enabling high-fidelity video analysis at scale

Hardware-Level Security


MTT S5000 implements a hardware-level Trusted Execution Environment (TEE), establishing a full-stack security perimeter from the underlying hardware to the application layer, providing robust protection for sensitive data and model assets.

Hardware Root of Trust (HRoT)

An integrated hardware root of trust supports secure boot, secure firmware update, and firmware protection, ensuring security and control across the full device lifecycle from power-on to runtime.

Confidential Computing Support

Hardware isolation technology provides encryption protection during data processing, preventing unauthorized access and enabling secure data sharing in high-sensitivity industries such as finance and healthcare.

Native Support for SM Cryptographic Algorithms

Comprehensive support for Chinese national cryptographic standards (SM algorithms), meeting the encryption requirements of trusted and innovative IT (Xinchuang) compliance and high-security-level business workloads.

Near-Zero-Cost Software Migration Path


Thanks to MUSA's high compatibility with CUDA syntax, developers can directly reuse existing code and expertise, enabling seamless ecosystem migration at near-zero cost.

MUSIFY Automatic Porting Tool

Moore Threads' proprietary MUSIFY tool automatically identifies and converts code, compressing what was once weeks of adaptation work into hours, enabling true plug-and-play migration.

No Core Code Refactoring Required

The MUSA SDK provides complete support for Runtime APIs, Driver APIs, and mainstream math libraries. Existing GPU applications run efficiently on MTT S5000 without rewriting core logic.

Native Support for Mainstream AI Frameworks

Deep integration with PyTorch, PaddlePaddle, and other leading frameworks, with full support via plugins such as Torch-MUSA, ensures rapid production deployment.

All-Around Compute Engine

Powering Innovation in Frontier Research and Cross-Disciplinary Science


With full-precision compute spanning FP8 to FP64, MTT S5000 leverages a unified, general-purpose architecture to handle high-precision and mixed-precision AI+ workloads flexibly. It is applicable across a wide range of frontier application domains, including quantum technology, embodied intelligence, world models, medical imaging, genomics, industrial large models, physics simulation, and molecular dynamics.

AI + Quantum Technology

AI + Quantum Technology Quantum Technology

AI + Embodied Intelligence

AI + Embodied Intelligence Embodied Intelligence

AI + Medical Imaging

AI + Medical Imaging Medical Imaging

AI + Genomics

AI + Genomics Genomics

AI + Industrial Large Models

AI + Industrial Large Models Industrial Models

AI + Physics Simulation

AI + Physics Simulation Physics Simulation

AI + Molecular Dynamics

AI + Molecular Dynamics Molecular Dynamics

A Comprehensive MUSA Software Ecosystem for Every Workload, Accelerating Innovation


The full-stack MUSA software platform encompasses drivers, the MUSA SDK, KUAE Training Suite, and KUAE Inference Suite, providing end-to-end coverage for AI training, AI inference, and scientific computing workloads.

The MUSA SDK includes core components such as deep learning acceleration libraries and math compute libraries that efficiently optimize matrix operations, common deep learning operators, collective communication, linear algebra, and general-purpose math. Built on the MUSA software stack, MTT S5000 delivers outstanding compute efficiency and a streamlined migration experience, helping research institutions and enterprises rapidly achieve performance gains and production deployment in large model training, scientific computing, image processing, and beyond.


MTT S5000 Specifications

Flexible Hardware Form Factors to Meet Diverse Deployment Needs

OAM Compute Module

OAM Compute Module

Designed to the OAM standard, MTT S5000 is available in two compute module form factors to meet high compute density requirements in a single node.

The liquid-cooled variant is purpose-built for high-density green data centers, maximizing compute density while significantly reducing Power Usage Effectiveness (PUE) and energy consumption.

The air-cooled variant is compatible with standard off-the-shelf servers, offering flexible, straightforward deployment that minimizes operational overhead and the long-term cost of ownership.
MTT MGX 8-GPU Modular Platform

MTT MGX 8-GPU Modular Platform

A modular platform purpose-built for AI and high-performance computing. Eight MTT S5000 OAM compute modules are interconnected via MTLink high-speed fabric, delivering massive compute scale for large model training, inference, and scientific computing workloads.
MTT SGX5000 Server

MTT SGX5000 Server

An all-in-one AI server equipped with eight MTT S5000 OAM compute modules, featuring premium compute, storage, and networking configurations to power hundred-billion- and trillion-parameter large model workloads at scale.

Fully optimized for thermal management, power delivery, and I/O expandability, the SGX5000 is available in both air-cooled and liquid-cooled configurations to meet the requirements of diverse data center environments.

Available with pre-installed, optimized training and inference software stacks for integrated hardware-software delivery — ready out of the box.

① Baseline 1: Dense FP16 compute 989 TFLOPs, memory bandwidth 3.35 TB/s; ② Baseline 2: Dense FP16 compute 148 TFLOPs, memory bandwidth 4.0 TB/s; ③ Baseline 3: Dense FP16 compute 312 TFLOPs, memory bandwidth 2.0 TB/s.

phone phone
Live
Agent
400-667-5666

Mon-Sun 9:00-21:00