Mon-Sun 9:00-21:00
MTT KUAE is Moore Threads Full-Stack Solution for AI Data Centers. It based on MTT S5000 GPU and the dual-processor 8-GPU server MTT SGX5000. This integrated solution tackles the challenges inherent in deploying large-scale GPU computing power with efficiency and effectiveness.
Quick Deployment
Cluster construction in 30 days
Best Practices
Optimization of computing power, storage, and networking
Ready to Use
Complete suite of intuitive tools and software stack
Unmatched Performance
Distributed training for trillion-parameter models
Ready-to-Use Integrated Solution
MTT KUAE full-stack solution is built on Moore Threads' Universal GPUs and seamlessly integrates hardware and software. MTT KUAE Platform for cluster management and MTT KUAE Model Studio for accessing model services fully support MTT KUAE to achieve maximum optimization. This end-to-end solution greatly simplifies the deployment and operation of large-scale GPU computational infrastructure.
Core Features
MTT KUAE full-stack solution fully leverages the advantages of Moore Threads GPUs.
Product Portfolio
MTT KUAE Core Components

MTT KUAE Platform
- Deeply integrates full-stack GPU computing, networking, and storage capabilities, with centralized GPU driver management to reduce deployment complexity and operational costs
- Provides multi-dimensional isolation for different organizations and users through enterprise workspaces and projects
- Supports GPU sharing and incorporates built-in best practices for multi-GPU-aware scheduling to improve resource utilization and maximize performance
- Delivers a unified observability platform for physical nodes, storage, networking, cluster components, and workloads, accelerating issue diagnosis while reducing troubleshooting costs
- Deeply integrates application and infrastructure telemetry, enabling proactive issue detection through diagnostic management and fine-grained monitoring and alerting

MTT KUAE ModelStudio
Model Development
- Launch ready-to-use development environments (VS Code & Jupyter) with a single click, preconfigured with required dependencies and mounted datasets to accelerate development
- Support multiple development workspaces with persistent storage, providing isolated development environments for different projects
- Support for mainstream distributed training frameworks, with rapid fault detection and checkpoint recovery in under 10 minutes
- Innovative training insights, 3D parallelism visualization for rapid identification of slow-performing nodes, and built-in operator profiling for large-scale training optimization
Key Features of MTT KUAE
Modular design of large-scale GPU computing power with flexible deployment
Optimization of the linear speedup ratio of GPU computing power
Deployment of a high-speed parameter transmission network
Deployment and scheduling of heterogeneous computing clusters
Design and deployment of a computational power service support system
Scheduling of elastic computing power for cloud-native GPU clusters
Reliability and security of computing and storage
Highly reliable automatic problem diagnosis and recovery

中文

