Mon-Sun 9:00-21:00
LLM Training and Finetuning

Training Platform Architecture

Training and Finetuning Examples

Cluster Scaling Efficiency
Large Model Inference Service Platform
MTT S4000, equipped with 128 Tensor Cores and 48 GB of memory, effectively supports inference for mainstream LLMs, such as LLaMa, ChatGLM, Qwen, and Baichuan.
KUAE ModelStudio
A training, finetuning, and inference platform for developers of LLM apps. It is based on Moore Threads' GPUs and models.
MUSA Serving
A software for high-performance and distributed inference services, supporting backend models such as LLMs, image and video generation, and traditional AI.
MT Transformer
A distributed inference acceleration framework for Moore Threads' GPUs. It achieves inference acceleration for LLMs based on transformer architectures.
TensorX
An inference acceleration framework for Moore Threads’ GPUs. It is ideal for inference acceleration in image and video generation and traditional AI.
From Chips to Clusters
Accelerating the Scale-Up of Chinese Computing Power
Supports the KUAE Series
Optimized server for clusters to train large models, providing excellent support for Moore Threads Full-Stack Solution for AI Data Centers.
Supports the KUAE AIDC Software Stack
Supports More than Just Large Models
MCCX D800
Specifications
FP32: 200 TFLOPS
FP16: 800 TFLOPS
Data drives: 4 × 3.84TB PCIe Gen 4 NVMe SSDs
2 × 2 Port 25G Fiber Network Cards

中文

