Black Forest Labs โ the German lab behind the FLUX image-model family, co-founded by Robin Rombach (original Stable Diffusion work) โ released FLUX 3 Action as an open-weights 7B 'World Action Model' for robotics, claiming first place on NVIDIA's RoboLab-120 benchmark with a 42.92% overall success rate. The claim is that it beats the prior best open model by 6.1 points with 44% of the parameters (7B vs NVIDIA's 16B Cosmos3-Nano-Policy) and runs ~1.43x faster. It is the robotics offshoot of the FLUX 3 multimodal family announced July 2026, which jointly trains image/video/audio/action in one architecture; a partner variant (FLUX-mimic, with mimic robotics) is reportedly running on production tasks at Audi. Caveat for a stranger: the leaderboard numbers are vendor-reported via BFL's own announcement and press coverage echoes those claims โ no independent benchmark results appear in the supplied material.
The radar already tracks this exact development in an open case: radar:flux-3-omnimodal-backbone predicts BFL demonstrating FLUX.3 unifying image, video, audio, and action in one flow-model backbone โ this release is that prediction materializing, now with a concrete twist (open-weights 7B embodied policy beating NVIDIA's Cosmos3-Nano-Policy on RoboLab). That benchmark claim is vendor-reported with no independent replication in the supplied material, so it sits at the announcement end of Scott's Evidence Class Ladder; the notable residual fact for him is that a diffusion-image lab, not a robotics-native one, now ships the top open embodied policy โ connecting his FLUX-based local tooling (gamepc, Replicate) to the radar's open-robotics thread.
ip:concept.evidence-class-ladderdev:project.gamepcdev:technology.replicateradar:flux-3-omnimodal-backboneradar:concept.open-modelsradar:concept.world-modelsradar:concept.roboticsradar:concept.benchmark-integrity
queries asked of Scott's wikis
- open-weights release strategy fine-tuning workflow local deployment economics
- diffusion world models video backbone as dynamics foundation
- unified multimodal single architecture vs specialized per-modality models
- vendor-reported benchmark claims leaderboard skepticism independent evals
- robotics foundation models embodied policy open model options
- Black Forest Labs FLUX dev projects local inference
now 0 pts/hpeak 3 pts/hcomments 0/hpeers p14momentum: steady3 platformsage 482h
points/hour across evidence ยท reading as of 2026-10-12 02:59:37.977291+11:00 ยท deterministic, not a model opinion
2026-09-28T03:36:44Z
The HN attachment is the same first-party BFL release post reaching a third platform (10 pts, 0 comments) โ duplicate coverage of an artifact already echoed into the case, not a new fact or independent check. Meaning is unchanged: a vendor-reported RoboLab-120 claim still awaiting replication; measured rates (1.67 pts/h, 0 comments/h) say residual drift despite a high peer percentile, so low heat holds.
2026-09-28T03:25:10Z
evidence attached: hn.story.49871965 โ First-party Black Forest Labs release post for the already-open Flux 3 Action case โ the primary artifact behind the RoboLab benchmark claims.
2026-09-25T17:56:10Z
The one-day velocity spike decayed to zero by 09-26 with no new evidence; the leaderboard-link comment remains the only quasi-independent check and is unverified. Meaning is unchanged โ a vendor-reported benchmark claim still awaiting replication โ so the case cools to low heat and stays open for independent evals or hands-on results with the released weights.
2026-09-23T23:14:28Z
origin walked (opencode/cheap-glm, conf 0.95): anchor reddit.post.1wof9q2 -> echo.blog.1d983c5f38 by Black Forest Labs
2026-09-23T22:06:16Z
grounded: known/medium โ The radar already tracks this exact development in an open case: radar:flux-3-omnimodal-backbone predicts BFL demonstrating FLUX.3 unifying image, video, audio,
2026-09-23T22:00:27Z
case created โ Two proposals are the same BFL release with concrete first-party benchmark claims; consolidated with the more detailed post as anchor.