2026-10-11 16:37 UTC

Baidu AI Cloud's Baige team claims its released LoongForge framework accelerates supported model-training workloads by up to 5.04 times over specified open-source baselines while aligning training loss curves, potentially reducing training costs and configuration work across multiple model families.

state: seedheat: mediumuncertainty: mediumnovelscott: lowmodel-training ai-infrastructure open-source-mlBaidu AI Cloud Baige team

What is this?

LoongForge is an Apache-2.0 open-source training framework from Baidu AI Cloudโ€™s Baige team for LLMs, vision-language, diffusion, and embodied models on NVIDIA GPUs and Kunlun XPUs. Its project pages describe Megatron-LM and torch-native backends, ready-to-run configurations for pre-training, continued pre-training, SFT, and LoRA, and optimizations to parallelism, memory, communication, and kernels. The developers advertise up to roughly 5ร— training speedups and say comparisons use the same machine type and training hyperparameters while keeping loss curves aligned with baselines. The supplied snippets do not show the benchmark rows establishing the precise 5.04ร— figure, independent validation, or measured cost and configuration-work savings.

Why it matters to Scott

The closest practical connection is Scottโ€™s Salesforce fine-tuning-data factory, but the hits do not establish a supported training workload he currently runs or that LoongForge would improve it; his archived crypto training used TensorFlow/Keras rather than the advertised backends. No supplied radar page tracks LoongForge itself, and its unvalidated speedup claims neither challenge a load-bearing Scott position nor establish a consequential convergence with one.
dev:project.redditradar:concept.training-efficiencyradar:concept.model-trainingradar:concept.lora
queries asked of Scott's wikis
  • fine-tuning SFT LoRA training pipelines projects
  • training economics compute efficiency optimization
  • hardware portability accelerator sovereignty vendor lock-in
  • reproducible benchmarks loss curves model quality equivalence
  • multi-backend ML tooling configuration automation

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 625h
points/hour across evidence ยท reading as of 2026-10-12 02:59:37.977291+11:00 ยท deterministic, not a model opinion

How the heat travelled

09-15 15:29 (minted)โญ origin echo-reconstructedLoongForge provides an Apache-2.0 multi-backend training framework for NVIDIA GPUs and Kunlun XPUs with 40-plus model examples and claimed s
Baidu AI Cloud Baige team on github (echo) ยท attributed from hn.story.49713630 ยท published time unknown
โ€”
09-15 15:03first on hacker news ยท published ยท lag ?Show HN: LoongForge-Train LLMs, VLMs, diffusion and embodied models, faster
nullnonenilNULL
โ€”
09-15 15:03amplified on hacker news ๐Ÿ‘‘hn.story.49713630
nullnonenilNULL
peak 2 ยท 0 comments ยท 98% of case engagement
09-15 15:20our radar first saw it ยท lag ?discovery anchor: hn.story.49713630โ€”
pace: p23 vs 1032 stories at the 336h mark (now 625h old) โ€” ahead of aafp-commons-signed-agent-notebook (2.0x), behind agentgate-signed-agent-receipts (0.7x)

Evidence (2) โ€” โญ canonical anchor

sourceobjectauthorscorecomments
๐ŸŸง hnShow HN: LoongForge-Train LLMs, VLMs, diffusion and embodied models, faster
Retrieved article excerpt

Open article ยท Retrieved 2026-09-15T15:22:54.743302+00:00

**English** | [็ฎ€ไฝ“ไธญๆ–‡](https://github.com/baidu-baige/LoongForge/blob/master/README_zh.md)

LoongForge

### Train LLMs, VLMs, diffusion and embodied models, faster.

[GitHub stars](https://github.com/baidu-baige/LoongForge/stargazers)
[License: Apache-2.0](https://github.com/baidu-baige/LoongForge/blob/master/LICENSE)
[Docker images on Docker Hub](https://hub.docker.com/u/loongforge)
[PRs welcome](https://github.com/baidu-baige/LoongForge/blob/master/CONTRIBUTING.md)

[Training throughput speedup up to 5.04x over open-source baselines](https://github.com/baidu-baige/LoongForge#performance)
[40+ ready-to-run model examples](https://github.com/baidu-baige/LoongForge#models)
[Runs on NVIDIA GPUs and Kunlun XPUs](https://camo.githubusercontent.com/e3904c0c1db653db13c438803938257ec11b0ea0d47a36730792f79a93b3b994/68747470733a2f2f696d672e736869656c64732e696f2f62616467652ff09f96a55f48617264776172652d4e56494449412532424b756e6c756e2d454334383939)
[Proven in production with runs up to 5,000+ XPUs](https://camo.githubusercontent.com/5100ab04a7ddced47a1ed69e91cc0bc6417c8a58bfb391d7a8bf63d28779a77c/68747470733a2f2f696d672e736869656c64732e696f2f62616467652ff09f8fad5f50726f64756374696f6e2d353030302532425f585055732d444232373737)

[**๐ŸŒ Website**](https://baidu-baige.github.io/LoongForge/)
ย ยทย 
[**๐Ÿ“– Docs**](https://loongforge.readthedocs.io/en/latest/index.html)
ย ยทย 
[**โœ๏ธ Blog**](https://baidu-baige.github.io/LoongForge/blog/)
ย ยทย 
[**โšก Quick Start**](https://github.com/baidu-baige/LoongForge#quickstart)
ย ยทย 
[**๐Ÿ“Š Performance**](https://github.com/baidu-baige/LoongForge#performance)
ย ยทย 
[**๐Ÿ›๏ธ Supported Models**](https://github.com/baidu-baige/LoongForge#models)
ย ยทย 
[**๐Ÿ’ฌ Contact Us**](https://github.com/baidu-baige/LoongForge#contact)

---

###### Example: embodied-model training on LoongForge โ€” DreamZero at 4.38ร— baseline throughput, loss curves aligned

[DreamZero training run compared side by side: LoongForge reaches 4.38x the baseline throughput while the training loss curves stay aligned](https://baidu-baige.github.io/LoongForge/assets/video/dreamzero-comparison.mp4)

## ๐Ÿ’ก Why LoongForge?

**LoongForge** is an open-source training framework developed by the [Baidu AI Cloud Baige team](https://cloud.baidu.com/product/aihc.html), built to deliver [faster training](https://github.com/baidu-baige/LoongForge#performance) for mainstream LLMs, VLMs, diffusion, and embodied models, thereby significantly reducing costs.

- **Easy to Use** โ€” [Ready-to-run configs](https://github.com/baidu-baige/LoongForge/blob/master/configs/models) and [launch examples](https://github.com/baidu-baige/LoongForge/blob/master/examples) for supported models, covering pre-training, continued pre-training, SFT, and LoRA.
- **High Performance** โ€” Built on multiple training backends (Megatron-LM and torch-native), with **deep optimizations** for each model family across parallelism strategy, memory footprint, communication overlap, and kernel efficiency, while keeping **training loss curves aligned with the baseline**.
- **Proven at Scale** โ€” Open-sourced from [AIAK-Training-LLM](https://cloud.baidu.com/doc/AIHC/s/Alyo476jr), a training acceleration suite serving enterprise customers in Education, Computer Vision, and Embodied AI, **with the largest production runs reaching 5,000+ XPUs**.

> ๐Ÿ‰ LoongForge is named after the traditional Chinese **loong boat (้พ™่ˆŸ)**, a symbol of coordinated power and forward momentum.

## ๐Ÿ—๏ธ Architecture

Since optimal training strategies differ across model families and scales, LoongForge adopts a multi-backend architecture.

LoongForge architecture: a patched-Megatron stack for LLMs, VLMs and diffusion models alongside a torch-native stack for embodied models

- **Megatron Stack** โ€” For LLMs, VLMs, and diffusion models. Powered by a [patched Megatron-LM](https://github.com/baidu-baige/Loong-Megatron) and extended with MoE parallelism, per-component heterogeneous parallelism, long-sequence optimizations, etc.
- **Torch-Native Stack** โ€” For embodied models (VLA and WAM). A standalone [torch-native subsystem](https://github.com/baidu-baige/LoongForge/blob/master/loongforge/embodied) featuring **DDP / ZeRO-1 / FSDP / HSDP**, with deep optimizations for representative models across I/O, communication strategy, kernel efficiency, etc.

## ๐Ÿ”ฅ Latest News

- **[2026/09]** โœจ Added **[Kimi-K3](https://github.com/baidu-baige/LoongForge/blob/master/examples/kimi_k3)** BF16 training support for both LLMs and VLMs.
- **[2026/09]** โšก Added an optimized **[DreamZero Wan2.2-5B FSDP recipe](https://github.com/baidu-baige/LoongForge/blob/master/examples/embodied/dreamzero/run_dreamzero_wan22_5b_full_fsdp_finetune.sh)** with cache-aware data loading, compiled attention blocks, frozen-module handling, and Delta-FP8 AllGather.
- **[2026/08]** ๐Ÿค– Added VLA training support for **[Wall-OSS-0.5](https://github.com/baidu-baige/LoongForge/blob/master/examples/embodied/wall_oss_0_5)**, with custom fused operators for higher training throughput.
- **[2026/08]** ๐Ÿ“„ Released the **[TAOT paper](https://arxiv.org/abs/2608.03676)** โ€” topology-aware dynamic expert replica placement that tackles expert-parallel (**EP**) load imbalance in **MoE** training, cutting overhead by up to **74%** over industry solutions, with **1.43ร— speedup** measured on a real training case. [[blog](https://baidu-baige.github.io/LoongForge/blog/2026-08-taot-topology-aware-expert-placement.html)]
- **[2026/08]** โœจ Added training support for **GLM-5.2**, along with a **[GLM-5.2 + MoonViT](https://github.com/baidu-baige/LoongForge/blob/master/configs/models/glm5.2_vit)** custom-composition [example](https://github.com/baidu-baige/LoongForge/blob/master/examples/glm5.2_vit) for extending GLM with multimodal capabilities.
- **[2026/08]** โœจ Added training support for **MiniCPM-V-4.6** and **Qwen3.8-27B**.
- **[2026/08]** ๐Ÿงช Introduced a unified [**evaluation module**](https://github.com/baidu-baige/LoongForge/blob/master/loongforge/embodied/eval) for the embodied stack, currently covering **Pi0.5 / xVLA / GR00T**, with more models on the way.
- **[2026/07]** ๐Ÿณ Unified the **prebuilt Docker images** โ€” all model families (LLM / VLM / VLA / Diffusion) now share a single image.
- **[2026/07]** ๐Ÿค– Released **[LoongForge-Embodied](https://github.com/baidu-baige/LoongForge/blob/master/loongforge/embodied)**, a torch-native DDP/FSDP training subsystem for embodied models (Pi0.5, GR00T-N1.6/N1.7, xVLA, LingBot-VA, FastWAM, DreamZero, and Cosmos3), with up to **4.38ร— speedup**. [[blog](https://baidu-baige.github.io/LoongForge/blog/2026-07-announcing-loongforge-embodied.html)]
- **[2026/07]** โœจ Added training support for **Qwen-Image-Edit-2511**.
- **[2026/07]** โœจ Added training support for **DeepSeek-V4-Flash / DeepSeek-V4-Pro**.

**๐Ÿ“… More**

- **[2026/06]** ๐Ÿค– Expanded VLA coverage with **GR00T N1.6**; **2.3ร— speedup** on GR00T training. [[blog](https://baidu-baige.github.io/LoongForge/blog/2026-06-loongforge-groot-n16-acceleration.html)]
- **[2026/05]** โšก Accelerated **Wan 2.2** training by **116%**, and added CP and data packing support.
- **[2026/05]** โœจ Added training support for **Kimi K2.5 / K2.6**, and introduced **INT4 / NVFP4** PTQ.
- **[2026/05]** ๐ŸŽ‰ **v0.1.0** โ€” first official tagged release of LoongForge.
- **[2026/05]** ๐ŸŒŸ Powered the training and public release of **LLaVA-OneVision-2.0**.
- **[2026/04]** ๐Ÿงฉ Added training support for **MiniMax-M2.7** on both NVIDIA GPU and Kunlun XPU.
- **[2026/04]** ๐Ÿš€ LoongForge source code publicly available on GitHub. [[blog](https://baidu-baige.github.io/LoongForge/blog/2026-04-announcing-loongforge.html)]
- **[2025/10]** ๐ŸŒŸ Powered the training and public release of **LLaVA-OneVision-1.5** under **AIAK-Training-LLM**, the predecessor of LoongForge. [[blog](https://baidu-baige.github.io/LoongForge/blog/2025-10-llava-onevision-case-study.html)]

## โœจ Key Features

**๐Ÿš€ Foundation Models**

- **MoE EP Communication Optimization** โ€” Overlapped All2All / activation offload / compute, with **further memory reduction** beyond upstream Megatron-LM on DeepSeek-V3, Qwen3-MoE, etc.
- **MoE Expert Load Balancing** โ€” Topology-aware dynamic replication of hot experts to balance EP workloads, with up to **74%** lower overhead than industry solutions. [[TAOT Paper](https://arxiv.org/pdf/2608.03676)]
- **Adaptive FP8 Training** โ€” End-to-end FP8 for LLMs and VLMs with standard **blockwise FP8**; optional **adaptive** mode picks per-operator precision by GEMM shape and efficiency.
- **Custom Fused Operators** โ€” Fused kernels like **FusedDSA** for DSA-style models โ€” TileLang version open-sourced, high-performance CUDA version available on Baidu Baige platform.
- **Long-Sequence Training** โ€” **Context Parallel (CP)** with **chunked-pipeline scheduling** scales LLM training to long sequence lengths.

**๐Ÿงฉ Multi-Modal Models**

- **Flexible Multi-Modal Composition** โ€” Assemble VLMs from interchangeable ViT and LLM components (e.g. **GLM-5.2 + MoonViT**) straight from config โ€” no custom model code.
- **Heterogeneous Parallelism** โ€” Independent TP / DP / recompute / freeze per model component (e.g., ViT vs. LLM) for optimal throughput and memory. [[blog](https://baidu-baige.github.io/LoongForge/blog/2026-05-loongforge-heterogeneous-parallel-training.html)]
- **Decoupled Encoder-Decoder Training** โ€” Separates ViT and LLM into independent tasks, eliminating encoder-induced pipeline bubbles.
- **DP Load Balancing** โ€” Load-aware data redistribution mitigates sequence-packing imbalance, improving multi-node scaling efficiency. [[blog](https://baidu-baige.github.io/LoongForge/blog/2026-05-loongforge-dp-load-balancing.html)]

**๐Ÿค– Embodied Models**

- **VLA & WAM Training** โ€” A dedicated **torch-native DDP/FSDP** subsystem for **VLA and world-action (WAM)** models, decoupled from the Megatron core, with flexible **DDP / ZeRO-1 / FSDP / HSDP** strategies. [[README](https://github.com/baidu-baige/LoongForge/blob/master/loongforge/embodied)]
- **Delta-FP8 FSDP Communication** โ€” Optionally compresses BF16 FSDP2 AllGather deltas into blockwise FP8 on supported NVIDIA GPUs while keeping model computation in BF16. [[Usage](https://github.com/baidu-baige/LoongForge/blob/master/docs/source/features/delta_fp8_allgather.md)]
- **Per-Model Deep Optimization** โ€” Training code deeply customized for each supported model across I/O, communication strategy, and kernel efficiency โ€” **1.79ร—โ€“4.38ร—** over official baselines in our [benchmarks](https://github.com/baidu-baige/LoongForge#performance).
- **Unified Evaluation** โ€” Evaluate trained policies on **LIBERO / CALVIN / SimplerEnv / RoboTwin**, with coverage expanding continuously.

**๐Ÿงฐ Workflow & Compatibility**

- **Versatile Pipelines & Data Tools** โ€” Out-of-the-box **Pretrain / MidTrain / SFT / LoRA**, with built-in dataset format conversion and sequence packing.
- **Flexible Checkpointing** โ€” Offline bidirectional **Megatron โ†” HuggingFace** conversion plus native online HF load/save โ€” no format barriers across your workflow.
- **Heterogeneous Hardware** โ€” Native support for **NVIDIA GPUs** and **Kunlun XPUs** via a minimally-intrusive plugin design.

> ๐Ÿ“– Deep-dive: [LLM](https://loongforge.readthedocs.io/en/latest/llm_tutorial/features_index.html) ยท [VLM](https://loongforge.readthedocs.io/en/latest/vlm_tutorial/features_index.html) ยท [Embodied Model](https://loongforge.readthedocs.io/en/latest/embodied_tutorial/overview.html)

## ๐Ÿ“Š Performance

Training throughput speedups over mainstream open-source baselines โ€” each model and its baseline were benchmarked on the same machine type with the same training hyperparameters:

[LoongForge benchmark speedups over open-source baselines โ€” from 1.45x on Qwen3-VL up to 5.04x on DeepSeek-V3.2 Lite](https://github.com/baidu-baige/LoongForge/blob/master/docs/assets/images/benchmark_speedup.png)

> DeepSeek-V3.2 Lite reflects DSA operator-level optimizations and was validated on a reduced-layer configuration due to test-bed scale limits.  
> Numbers were measured at a point in time a
nullnonenilNULL20
๐ŸŸง echo.github โญLoongForge provides an Apache-2.0 multi-backend training framework for NVIDIA GPUs and Kunlun XPUs with 40-plus model examples and claimed sBaidu AI Cloud Baige teamโ€”โ€”

Interpretation history

Decision trace