Mon-Sun 9:00-21:00
LLM Training and Finetuning

Training Platform Architecture

Training and Finetuning Examples

Cluster Scaling Efficiency
Large Model Inference Service Platform
MTT S4000, equipped with 128 Tensor Cores and 48 GB of memory, effectively supports inference for mainstream LLMs, such as LLaMa, ChatGLM, Qwen, and Baichuan.
KUAE ModelStudio
A training, finetuning, and inference platform for developers of LLM apps. It is based on Moore Threads' GPUs and models.
MUSA Serving
A software for high-performance and distributed inference services, supporting backend models such as LLMs, image and video generation, and traditional AI.
MT Transformer
A distributed inference acceleration framework for Moore Threads' GPUs. It achieves inference acceleration for LLMs based on transformer architectures.
TensorX
An inference acceleration framework for Moore Threads’ GPUs. It is ideal for inference acceleration in image and video generation and traditional AI.
Supporting KUAE Cluster Products
MTT KUAE is Moore Threads' full-stack solution for artificial intelligence data centers. It is based on the S4000 GPU and MCCX D800 all-in-one cluster computing unit, which is equipped with eight dual-processor S4000 GPUs. This integrated solution tackles the challenges inherent in deploying large-scale GPU computing power with efficiency and effectiveness.
Advanced-Generation Tensor
Moore Threads' advanced-generation Tensor Cores assist in the training, finetuning, and inference of LLMs. MTT S4000 includes 8,192 vector cores and 128 Tensor Cores. It supports mainstream precision computing formats such as FP64, FP32, TF32, FP16, BF16, and INT8.
Third-Generation MUSA Software Stack
MUSA is Moore Threads' self-developed metacomputing unified architecture, which includes an instruction set architecture, MUSA programming model, driver, runtime library, operator library, communication library, and mathematical library. CUDA programs can be smoothly migrated to MUSA through Moore Threads' self-developed MUSIFY tool.
Fully Supports Mainstream Graphics APIs
MTT S4000 supports mainstream graphics APIs such as DirectX, Vulkan, OpenGL, and OpenGL ES. It provides versatile graphics rendering for a variety of scenarios, including digital twins, cloud gaming, cloud rendering, and content creation. It also supports large model inference capabilities, functioning as a one-stop solution for multi-modal scenarios like AIGC.

中文

