Retrieved article excerpt
Open article ยท Retrieved 2026-09-15T15:22:54.743302+00:00
**English** | [็ฎไฝไธญๆ](https://github.com/baidu-baige/LoongForge/blob/master/README_zh.md)
LoongForge
### Train LLMs, VLMs, diffusion and embodied models, faster.
[GitHub stars](https://github.com/baidu-baige/LoongForge/stargazers)
[License: Apache-2.0](https://github.com/baidu-baige/LoongForge/blob/master/LICENSE)
[Docker images on Docker Hub](https://hub.docker.com/u/loongforge)
[PRs welcome](https://github.com/baidu-baige/LoongForge/blob/master/CONTRIBUTING.md)
[Training throughput speedup up to 5.04x over open-source baselines](https://github.com/baidu-baige/LoongForge#performance)
[40+ ready-to-run model examples](https://github.com/baidu-baige/LoongForge#models)
[Runs on NVIDIA GPUs and Kunlun XPUs](https://camo.githubusercontent.com/e3904c0c1db653db13c438803938257ec11b0ea0d47a36730792f79a93b3b994/68747470733a2f2f696d672e736869656c64732e696f2f62616467652ff09f96a55f48617264776172652d4e56494449412532424b756e6c756e2d454334383939)
[Proven in production with runs up to 5,000+ XPUs](https://camo.githubusercontent.com/5100ab04a7ddced47a1ed69e91cc0bc6417c8a58bfb391d7a8bf63d28779a77c/68747470733a2f2f696d672e736869656c64732e696f2f62616467652ff09f8fad5f50726f64756374696f6e2d353030302532425f585055732d444232373737)
[**๐ Website**](https://baidu-baige.github.io/LoongForge/)
ย ยทย
[**๐ Docs**](https://loongforge.readthedocs.io/en/latest/index.html)
ย ยทย
[**โ๏ธ Blog**](https://baidu-baige.github.io/LoongForge/blog/)
ย ยทย
[**โก Quick Start**](https://github.com/baidu-baige/LoongForge#quickstart)
ย ยทย
[**๐ Performance**](https://github.com/baidu-baige/LoongForge#performance)
ย ยทย
[**๐๏ธ Supported Models**](https://github.com/baidu-baige/LoongForge#models)
ย ยทย
[**๐ฌ Contact Us**](https://github.com/baidu-baige/LoongForge#contact)
---
###### Example: embodied-model training on LoongForge โ DreamZero at 4.38ร baseline throughput, loss curves aligned
[DreamZero training run compared side by side: LoongForge reaches 4.38x the baseline throughput while the training loss curves stay aligned](https://baidu-baige.github.io/LoongForge/assets/video/dreamzero-comparison.mp4)
## ๐ก Why LoongForge?
**LoongForge** is an open-source training framework developed by the [Baidu AI Cloud Baige team](https://cloud.baidu.com/product/aihc.html), built to deliver [faster training](https://github.com/baidu-baige/LoongForge#performance) for mainstream LLMs, VLMs, diffusion, and embodied models, thereby significantly reducing costs.
- **Easy to Use** โ [Ready-to-run configs](https://github.com/baidu-baige/LoongForge/blob/master/configs/models) and [launch examples](https://github.com/baidu-baige/LoongForge/blob/master/examples) for supported models, covering pre-training, continued pre-training, SFT, and LoRA.
- **High Performance** โ Built on multiple training backends (Megatron-LM and torch-native), with **deep optimizations** for each model family across parallelism strategy, memory footprint, communication overlap, and kernel efficiency, while keeping **training loss curves aligned with the baseline**.
- **Proven at Scale** โ Open-sourced from [AIAK-Training-LLM](https://cloud.baidu.com/doc/AIHC/s/Alyo476jr), a training acceleration suite serving enterprise customers in Education, Computer Vision, and Embodied AI, **with the largest production runs reaching 5,000+ XPUs**.
> ๐ LoongForge is named after the traditional Chinese **loong boat (้พ่)**, a symbol of coordinated power and forward momentum.
## ๐๏ธ Architecture
Since optimal training strategies differ across model families and scales, LoongForge adopts a multi-backend architecture.
LoongForge architecture: a patched-Megatron stack for LLMs, VLMs and diffusion models alongside a torch-native stack for embodied models
- **Megatron Stack** โ For LLMs, VLMs, and diffusion models. Powered by a [patched Megatron-LM](https://github.com/baidu-baige/Loong-Megatron) and extended with MoE parallelism, per-component heterogeneous parallelism, long-sequence optimizations, etc.
- **Torch-Native Stack** โ For embodied models (VLA and WAM). A standalone [torch-native subsystem](https://github.com/baidu-baige/LoongForge/blob/master/loongforge/embodied) featuring **DDP / ZeRO-1 / FSDP / HSDP**, with deep optimizations for representative models across I/O, communication strategy, kernel efficiency, etc.
## ๐ฅ Latest News
- **[2026/09]** โจ Added **[Kimi-K3](https://github.com/baidu-baige/LoongForge/blob/master/examples/kimi_k3)** BF16 training support for both LLMs and VLMs.
- **[2026/09]** โก Added an optimized **[DreamZero Wan2.2-5B FSDP recipe](https://github.com/baidu-baige/LoongForge/blob/master/examples/embodied/dreamzero/run_dreamzero_wan22_5b_full_fsdp_finetune.sh)** with cache-aware data loading, compiled attention blocks, frozen-module handling, and Delta-FP8 AllGather.
- **[2026/08]** ๐ค Added VLA training support for **[Wall-OSS-0.5](https://github.com/baidu-baige/LoongForge/blob/master/examples/embodied/wall_oss_0_5)**, with custom fused operators for higher training throughput.
- **[2026/08]** ๐ Released the **[TAOT paper](https://arxiv.org/abs/2608.03676)** โ topology-aware dynamic expert replica placement that tackles expert-parallel (**EP**) load imbalance in **MoE** training, cutting overhead by up to **74%** over industry solutions, with **1.43ร speedup** measured on a real training case. [[blog](https://baidu-baige.github.io/LoongForge/blog/2026-08-taot-topology-aware-expert-placement.html)]
- **[2026/08]** โจ Added training support for **GLM-5.2**, along with a **[GLM-5.2 + MoonViT](https://github.com/baidu-baige/LoongForge/blob/master/configs/models/glm5.2_vit)** custom-composition [example](https://github.com/baidu-baige/LoongForge/blob/master/examples/glm5.2_vit) for extending GLM with multimodal capabilities.
- **[2026/08]** โจ Added training support for **MiniCPM-V-4.6** and **Qwen3.8-27B**.
- **[2026/08]** ๐งช Introduced a unified [**evaluation module**](https://github.com/baidu-baige/LoongForge/blob/master/loongforge/embodied/eval) for the embodied stack, currently covering **Pi0.5 / xVLA / GR00T**, with more models on the way.
- **[2026/07]** ๐ณ Unified the **prebuilt Docker images** โ all model families (LLM / VLM / VLA / Diffusion) now share a single image.
- **[2026/07]** ๐ค Released **[LoongForge-Embodied](https://github.com/baidu-baige/LoongForge/blob/master/loongforge/embodied)**, a torch-native DDP/FSDP training subsystem for embodied models (Pi0.5, GR00T-N1.6/N1.7, xVLA, LingBot-VA, FastWAM, DreamZero, and Cosmos3), with up to **4.38ร speedup**. [[blog](https://baidu-baige.github.io/LoongForge/blog/2026-07-announcing-loongforge-embodied.html)]
- **[2026/07]** โจ Added training support for **Qwen-Image-Edit-2511**.
- **[2026/07]** โจ Added training support for **DeepSeek-V4-Flash / DeepSeek-V4-Pro**.
**๐
More**
- **[2026/06]** ๐ค Expanded VLA coverage with **GR00T N1.6**; **2.3ร speedup** on GR00T training. [[blog](https://baidu-baige.github.io/LoongForge/blog/2026-06-loongforge-groot-n16-acceleration.html)]
- **[2026/05]** โก Accelerated **Wan 2.2** training by **116%**, and added CP and data packing support.
- **[2026/05]** โจ Added training support for **Kimi K2.5 / K2.6**, and introduced **INT4 / NVFP4** PTQ.
- **[2026/05]** ๐ **v0.1.0** โ first official tagged release of LoongForge.
- **[2026/05]** ๐ Powered the training and public release of **LLaVA-OneVision-2.0**.
- **[2026/04]** ๐งฉ Added training support for **MiniMax-M2.7** on both NVIDIA GPU and Kunlun XPU.
- **[2026/04]** ๐ LoongForge source code publicly available on GitHub. [[blog](https://baidu-baige.github.io/LoongForge/blog/2026-04-announcing-loongforge.html)]
- **[2025/10]** ๐ Powered the training and public release of **LLaVA-OneVision-1.5** under **AIAK-Training-LLM**, the predecessor of LoongForge. [[blog](https://baidu-baige.github.io/LoongForge/blog/2025-10-llava-onevision-case-study.html)]
## โจ Key Features
**๐ Foundation Models**
- **MoE EP Communication Optimization** โ Overlapped All2All / activation offload / compute, with **further memory reduction** beyond upstream Megatron-LM on DeepSeek-V3, Qwen3-MoE, etc.
- **MoE Expert Load Balancing** โ Topology-aware dynamic replication of hot experts to balance EP workloads, with up to **74%** lower overhead than industry solutions. [[TAOT Paper](https://arxiv.org/pdf/2608.03676)]
- **Adaptive FP8 Training** โ End-to-end FP8 for LLMs and VLMs with standard **blockwise FP8**; optional **adaptive** mode picks per-operator precision by GEMM shape and efficiency.
- **Custom Fused Operators** โ Fused kernels like **FusedDSA** for DSA-style models โ TileLang version open-sourced, high-performance CUDA version available on Baidu Baige platform.
- **Long-Sequence Training** โ **Context Parallel (CP)** with **chunked-pipeline scheduling** scales LLM training to long sequence lengths.
**๐งฉ Multi-Modal Models**
- **Flexible Multi-Modal Composition** โ Assemble VLMs from interchangeable ViT and LLM components (e.g. **GLM-5.2 + MoonViT**) straight from config โ no custom model code.
- **Heterogeneous Parallelism** โ Independent TP / DP / recompute / freeze per model component (e.g., ViT vs. LLM) for optimal throughput and memory. [[blog](https://baidu-baige.github.io/LoongForge/blog/2026-05-loongforge-heterogeneous-parallel-training.html)]
- **Decoupled Encoder-Decoder Training** โ Separates ViT and LLM into independent tasks, eliminating encoder-induced pipeline bubbles.
- **DP Load Balancing** โ Load-aware data redistribution mitigates sequence-packing imbalance, improving multi-node scaling efficiency. [[blog](https://baidu-baige.github.io/LoongForge/blog/2026-05-loongforge-dp-load-balancing.html)]
**๐ค Embodied Models**
- **VLA & WAM Training** โ A dedicated **torch-native DDP/FSDP** subsystem for **VLA and world-action (WAM)** models, decoupled from the Megatron core, with flexible **DDP / ZeRO-1 / FSDP / HSDP** strategies. [[README](https://github.com/baidu-baige/LoongForge/blob/master/loongforge/embodied)]
- **Delta-FP8 FSDP Communication** โ Optionally compresses BF16 FSDP2 AllGather deltas into blockwise FP8 on supported NVIDIA GPUs while keeping model computation in BF16. [[Usage](https://github.com/baidu-baige/LoongForge/blob/master/docs/source/features/delta_fp8_allgather.md)]
- **Per-Model Deep Optimization** โ Training code deeply customized for each supported model across I/O, communication strategy, and kernel efficiency โ **1.79รโ4.38ร** over official baselines in our [benchmarks](https://github.com/baidu-baige/LoongForge#performance).
- **Unified Evaluation** โ Evaluate trained policies on **LIBERO / CALVIN / SimplerEnv / RoboTwin**, with coverage expanding continuously.
**๐งฐ Workflow & Compatibility**
- **Versatile Pipelines & Data Tools** โ Out-of-the-box **Pretrain / MidTrain / SFT / LoRA**, with built-in dataset format conversion and sequence packing.
- **Flexible Checkpointing** โ Offline bidirectional **Megatron โ HuggingFace** conversion plus native online HF load/save โ no format barriers across your workflow.
- **Heterogeneous Hardware** โ Native support for **NVIDIA GPUs** and **Kunlun XPUs** via a minimally-intrusive plugin design.
> ๐ Deep-dive: [LLM](https://loongforge.readthedocs.io/en/latest/llm_tutorial/features_index.html) ยท [VLM](https://loongforge.readthedocs.io/en/latest/vlm_tutorial/features_index.html) ยท [Embodied Model](https://loongforge.readthedocs.io/en/latest/embodied_tutorial/overview.html)
## ๐ Performance
Training throughput speedups over mainstream open-source baselines โ each model and its baseline were benchmarked on the same machine type with the same training hyperparameters:
[LoongForge benchmark speedups over open-source baselines โ from 1.45x on Qwen3-VL up to 5.04x on DeepSeek-V3.2 Lite](https://github.com/baidu-baige/LoongForge/blob/master/docs/assets/images/benchmark_speedup.png)
> DeepSeek-V3.2 Lite reflects DSA operator-level optimizations and was validated on a reduced-layer configuration due to test-bed scale limits.
> Numbers were measured at a point in time a