2026-10-11 17:14 UTC

Formal-Swordfish-228 reports that released Cosmos3 INT4 weights and MLX/CUDA code enable local text-to-image and image-to-video generation, including a roughly five-minute clip generation on a 128GB M4 Max, potentially making the 64B model usable on high-memory personal hardware.

state: seedheat: lowuncertainty: highconvergesscott: mediummultimodal-models local-inference image-generationgtrg55JuliaMLFormal-Swordfish-228

What is this?

NVIDIA’s Cosmos 3 is a family of omnimodal world models; its official repository lists the Super tier at 64B parameters and recommends data-center hardware, while supplied model and report snippets establish image/video generation and structured-prompt agentic refinement. The case attributes to Formal-Swordfish-228 a report of INT4 weights and MLX/CUDA code enabling local generation, including a clip generated in roughly five minutes on a 128GB M4 Max. The supplied search results do not independently establish that community release, its authorship, or its hardware performance; the web answer’s claim of “five-minute clips” also confuses generation time with clip duration.

Why it matters to Scott

The reported INT4 MLX/CUDA implementation converges with Scott’s hardware-aware local-inference practice and offers a concrete candidate to evaluate alongside his gamepc image/video stack and RTX-3090-tuned BRIA generator. The community release and roughly five-minute generation time remain unverified, and suitability for Scott’s hardware is not established; the radar tracks related local-generation efforts but the supplied hits do not show this Cosmos3 development already covered.
dev:concept.hardware-aware-local-inferencedev:project.gamepcdev:project.briaradar:concept.local-inferenceradar:concept.quantizationradar:concept.mlxradar:concept.image-generationradar:concept.video-generationradar:minimax-h3-comfyui-local-validationradar:sana-cpp-local-inference-speedup
queries asked of Scott's wikis
  • local inference quantization memory limits Apple Silicon MLX CUDA
  • open-weight models personal hardware deployment economics
  • local image video generation workflows projects
  • multimodal agent harnesses structured prompts iterative refinement
  • model benchmark reproducibility harness effects hardware performance

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p0momentum: steady2 platformsage 770h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-09 15:25 (minted)⭐ origin echo-reconstructedThe Reddit post identifies this repository as the code for Cosmos3 INT4 text-to-image and image-to-video generation on MLX and CUDA.
gtrg55 on github (echo) · attributed from reddit.post.1wbmz1y · published time unknown
—
09-09 14:21first on r/LocalLLaMA · published · lag ?SOTA ImageGen Locally NVIDIA Cosmos3(64B) INT4 quants CUDA/MLX
Formal-Swordfish-228
—
09-09 14:21amplified on r/LocalLLaMA 👑reddit.post.1wbmz1y
Formal-Swordfish-228
peak 65 · 23 comments · 100% of case engagement
09-09 15:20our radar first saw it · lag ?discovery anchor: reddit.post.1wbmz1y—
pace: p66 vs 519 stories at the 720h mark (now 770h old) — ahead of geiger-local-agent-access-inventory (1.1x), behind antfly-v02-zig-rewrite (1.0x)

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditSOTA ImageGen Locally NVIDIA Cosmos3(64B) INT4 quants CUDA/MLX
LocalLLaMA
Formal-Swordfish-2285823
🟧 echo.github ⭐The Reddit post identifies this repository as the code for Cosmos3 INT4 text-to-image and image-to-video generation on MLX and CUDA.gtrg55——

Interpretation history

Decision trace