2026-10-11 16:38 UTC

QRUN reports that provisioning, loading, failures, and teardown raise MiniMax H3 per-clip costs to 2.9–14 times steady-state estimates in its seven three-clip GPU runs, making session overhead material to short video-generation rental decisions.

state: seedheat: mediumuncertainty: mediumknownscott: lowinference-economics video-generation gpu-cloudQRUN

What is this?

Provider snippets describe MiniMax H3 as MiniMax’s multimodal video-generation model, supporting text-, image-, and reference-driven generation with synchronized audio. The case attributes to QRUN a seven-configuration, four-provider GPU rental comparison in which session overhead raises per-clip costs to 2.9–14 times steady-state estimates for three-clip runs. None of the supplied web results contains QRUN’s report or establishes who QRUN is, so those measurements and their methodology remain unverified here. WaveSpeed’s cost discussion supports including idle time, maintenance, failures, and retries when comparing rented or local GPUs with APIs, but does not corroborate QRUN’s specific figures.

Why it matters to Scott

Scott’s AI Unit Economics page already requires infrastructure and remediation costs in per-transaction accounting; QRUN’s reported session overhead illustrates that position and touches the startup/loading boundary in his Beam.cloud evaluation, but the hits establish neither an active H3 rental decision nor a consequential new adopter of his position. The radar already follows H3 deployment practicality in minimax-h3-comfyui-local-validation, though not this rental comparison; QRUN’s unverified figures are insufficient here to change Scott’s build choices or substantiate a publishing opportunity.
ip:concept.ai-unit-economicsdev:project.beamradar:concept.inference-economicsradar:minimax-h3-comfyui-local-validationradar:nvidia-sol-h3-spark-video
queries asked of Scott's wikis
  • inference economics end-to-end cost versus steady-state benchmarks
  • GPU rental cold starts model loading teardown overhead
  • local inference versus hosted API utilization break-even
  • video generation pipelines short-batch experimentation
  • benchmark harness failed runs retries cost per usable output

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 698h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-12 14:00⭐ origin echo-reconstructedQRUN publishes its own MiniMax H3 runs across seven GPU configurations at four providers, reporting steady-state costs of $0.0125–0.0375 per
QRUN on blog (echo) · attributed from hn.story.49711906
—
09-15 13:01first on hacker news · published · +71.0hMiniMax H3: measured cost per clip on seven rented GPUs at four providers
susorokin
—
09-15 13:01amplified on hacker news 👑hn.story.49711906
susorokin
peak 2 · 0 comments · 98% of case engagement
09-15 13:20our radar first saw it · +71.3hdiscovery anchor: hn.story.49711906—
pace: p9 vs 1032 stories at the 336h mark (now 698h old) — behind addom-local-coding-harness (0.5x)

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnMiniMax H3: measured cost per clip on seven rented GPUs at four providers
Retrieved article excerpt

Open article · Retrieved 2026-09-15T13:26:31.629544+00:00

Measured GPU runs

# What rented GPUs actually cost per unit of work.

Measured runs on rented GPUs: what each provider actually charged, and the cost per finished unit of work — a video clip, an image, a thousand training steps, a million texts or tokens.

Last updated 13 September 2026

Every number is from a run we paid for; the provider’s own billing is the source of the cost column.

All runs on this page are our own: test workloads we chose, run on our own provider accounts at our own expense. Measurements we make for clients are confidential — we have no right to publish them, and none are here.

11–13 September 2026

## MiniMax H3 video: one 5-second clip on seven cards at four providers.

Four providers, seven cards, the same clip. In the steady state a clip costs $0.0125–0.0375; the same clip costs $0.049–0.185 once the whole session — start-up, loading the weights and the tail until the machine is deleted — is divided over the three clips of the run.

| Run | Provider · card | Mode · rate | Host vCPU / RAM | s per clipsteady, clips 2–3 | $ per clipsteady | Chargedfor the run | $ per clipsession | Ratiosession ÷ steady |
| --- | --- | --- | --- | --- | --- | --- | --- | --- |
| W5-01 | Vast.aiRTX 4090 24 GB | spot, $0.483/hbid $0.40 + 150 GB disk $0.083 | 32 vCPU / 108 GB | 93.193.2 / 93.0 | $0.0125 | $0.52two attempts, 31 min; the first failed before the download started | $0.175 | 14× |
| W5-02 | HyperstackH100 PCIe 80 GB | spot, $2.007/h$2.00 + public IP | 28 vCPU / 177 GB | 67.266.9 / 67.5 | $0.0375 | $0.38686 s billed of a 14.8 min VM lifetime | $0.127 | 3.4× |
| W5-03 | RunPod SecureRTX 4090 24 GB | on-demand, $0.756/h$0.74 + 120 GB disk $0.016 | 15 vCPU / 86 GB | 91.691.6 / 91.6 | $0.0193 | $0.2017.5 min | $0.068 | 3.5× |
| W5-04 | Nebius eu-north1L40S 48 GB | preemptible, $0.947/h$0.92 + 200 GB disk $0.027 | 24 vCPU / 94 GB | 89.388.9 / 89.7 | $0.0235 | ≈ $0.5535.1 min; preset price × lifetime — Nebius has no balance API | ≈ $0.185 | 7.9× |
| W5-05 | RunPod CommunityRTX 5090 32 GB | on-demand, $0.707/h$0.69 + 120 GB disk $0.017 | 16 vCPU / 94 GB | 68.768.4 / 68.9 | $0.0135 | $0.1512.6 min, 12 September | $0.049 | 3.6× |
| W5-06 | RunPod SecureRTX 5090 32 GB | on-demand, $1.007/h$0.99 + 120 GB disk $0.017 | 15 vCPU / 117 GB | 100.2100.4 / 100.0 — of which 37.5 s is the mp4 encode on this pod's CPU (the host, not the card) | $0.0280 | $0.2515.5 min, 12 September | $0.083 | 2.9× |
| W5-07 | RunPod SecureRTX PRO 6000 Blackwell 96 GB | on-demand, $2.107/h$2.09 + 120 GB disk $0.017 | 16 vCPU / 283 GB | 46.947.0 / 46.9 — the whole 67 GB of weights fits in the card, so nothing is streamed | $0.0275 | $0.4914.4 min, 13 September | $0.164 | 6.0× |

MiniMax H3, the official Comfy-Org repack: int8 DiT (34.0 GB), int8 Qwen3-VL-32B text encoder (27.1 GB), fp16 video VAE (5.2 GB), fp32 audio VAE (0.6 GB) — 67.0 GB of weights; stock ComfyUI v0.35.0 text-to-video template without LoRA (res\_multistep / simple, BasicGuider, VAE decode + audio decode to mp4); 864×480, 124 frames = 5.17 s at 24 fps, 20 steps; the same prompt and seeds (1000 / 1001 / 1002) on every card; torch 2.9.0 + CUDA 13.0 everywhere; three clips per card: the first includes loading the weights into memory, the second and third are the steady state and differ by at most 1 %. $3.13 in provider charges for the seven runs, including three failed attempts; the two RTX 5090 rows were added on 12 September ($0.39) and the RTX PRO 6000 row on 13 September ($1.07).

2–3 September 2026

## Phase 3: 15 runs, four workloads, three providers.

The same four scripts on different cards and hosts at RunPod Secure, Vast.ai and Nebius, each run to completion and paid for. Within a workload, rows go from the cheapest unit to the most expensive.

| Run | Workload | Provider · card | Mode · rate | Hoursbilled | Charged | GPU utilmean | Cost per unit | Notes |
| --- | --- | --- | --- | --- | --- | --- | --- | --- |
| SDXL image generation — cost per imageStable Diffusion XL 1.0, 1024×1024, 30 steps, guidance 7.5, diffusers with SDPA attention, no refiner; 2,000 images per run (500 in the fp32 run). | | | | | | | | |
| --- | --- | --- | --- | --- | --- | --- | --- | --- |
| W1-02 | fp16, batch 4 | Vast.aiRTX 4090 24 GB | spot$0.30/h bid | 2.13 | $0.72 | 97 % | $0.00036 | no interruptions |
| W1-03 | fp16, batch 4 | Nebius eu-north1H100 SXM 80 GB | preemptible$2.15/h | 0.97 | $2.11 | 91 % | $0.0011 | 6.2 min VM provisioning; charge estimated from the preset price |
| W1-01 | fp16, batch 4 | RunPod SecureL40S 48 GB | on-demand$0.99/h | 2.13 | $2.21 | 95 % | $0.0011 |  |
| H5-W1-N | fp32, batch 1 (diffusers defaults), 500 images | RunPod SecureH100 PCIe 80 GB | on-demand$2.89/h | 1.94 | $5.61 | 99 % | $0.0112 | 13.8 s per image |
| LoRA fine-tuning of SDXL — cost per 1,000 optimizer stepsThe diffusers LoRA training script for SDXL: 2,000 optimizer steps, batch 1 × gradient accumulation 4, 1024 px, rank 8, fp16, gradient checkpointing, 8-bit Adam. Same script on every host — only the machine differs. | | | | | | | | |
| --- | --- | --- | --- | --- | --- | --- | --- | --- |
| W2-03 | host with 24 vCPU | Vast.aiRTX 4090 24 GB | spot$0.30/h bid | 1.07 | $0.38 | 75 % | $0.19 | 1.84 s per step; no interruptions |
| W2-01 | host with 5 vCPU | RunPod SecureRTX 4090 24 GB | on-demand$0.74/h | 1.95 | $1.52 | 40 % | $0.76 | 3.44–3.72 s per step: the dataloader on 5 vCPU, not the GPU, sets the pace |
| W2-02 | host with 16 vCPU | Nebius eu-north1H100 SXM 80 GB | on-demand$3.85/h | 1.70 | $6.58 | 37 % | $3.29 | 2.68 s per step — slower than the RTX 4090 above: single-thread CPU work per step; charge estimated from the preset price |
| Text embeddings — cost per 1,000,000 textsbge-large-en-v1.5 (335M parameters) over 2,000,000 texts with sentence-transformers: fp16, max sequence length 512, batch 256, normalised vectors. | | | | | | | | |
| --- | --- | --- | --- | --- | --- | --- | --- | --- |
| W3-03 | host with 2 vCPU | RunPod SecureRTX A6000 48 GB | on-demand$0.53/h | 2.06 | $1.10 | 78 % | $0.55 | dataset streaming stalls on 2 vCPU: GPU at 0 % for at least a tenth of the samples |
| W3-02 | host with 16 vCPU | Nebius eu-north1L40S 48 GB | preemptible$0.83/h | 1.78 | $1.54 | 75 % | $0.77 | no interruptions; charge estimated from the preset price |
| W3-01 | host with 16 vCPU | RunPod SecureA100 SXM 80 GB | on-demand$1.59/h | 1.22 | $1.97 | 72 % | $0.99 | 15 % of samples at 0 % GPU: streaming-download stalls |
| Fine-tuning a 7B model — cost per 1,000,000 tokensQwen2.5-7B on the first 30,000 alpaca-cleaned examples, one epoch, batch 4 × gradient accumulation 4, max sequence length 1,024 — 4.97M dataset tokens counted with the Qwen2.5 tokenizer. QLoRA rows: nf4 base model, gradient checkpointing, paged 8-bit AdamW. bf16 rows: no quantisation, no gradient checkpointing, AdamW. | | | | | | | | |
| --- | --- | --- | --- | --- | --- | --- | --- | --- |
| W4-02 | QLoRA nf4 | Vast.aiRTX 4090 24 GB | spot$0.30/h bid | 2.76 | $0.99 | 86 % | $0.20 | 3 host-side container restarts, 2 resumes from checkpoint |
| W4-05 | bf16 LoRA | Nebius us-central1RTX PRO 6000 96 GB | on-demand$1.80/h | 0.70 | $1.33 | 79 % | $0.27 | 0.78× the H100's speed at 55 % of its price; two preemptible attempts failed first (driver stack, then preempted after ~10 min); charge estimated from the preset price |
| H5-W4-N | bf16 LoRA | RunPod SecureH100 SXM 80 GB | on-demand$3.29/h | 0.44 | $1.50 | 80 % | $0.30 | 71 GB of VRAM in use — only fits an 80 GB card |
| W4-03 | QLoRA nf4 | Vast.aiA100 SXM 80 GB | on-demand$1.056/h | 1.73 | $1.90 | 93 % | $0.38 | 1.2× the RTX 4090's speed at 3.5× its price |
| W4-01 | QLoRA nf4 | RunPod SecureH100 SXM 80 GB | on-demand$3.29/h | 0.94 | $3.14 | 90 % | $0.63 | 1.71 s per step against 0.74 s for bf16 on the same card |

$33.55 in provider charges, including the failed attempts. Hours are the billed lifetime of the instance, from creation to verified deletion. Charged is the provider’s billing plus a few cents of checkpoint storage; for Nebius it is an estimate — see the method.

Method

## Method, in short.

What counts as the actual charge
:   The provider’s own billing, read through its API: the account balance before and after the run (RunPod, Vast.ai) or the billing history (Hyperstack: 686 s × $2.00672/h = $0.3824, matching the balance delta to the cent), plus the checkpoint storage in R2 — a few cents per run. Nebius has no balance API and publishes billing with a delay, so its runs are estimated as the preset price × the VM’s lifetime plus disk; those rows say so.

Hours
:   The billed lifetime of the instance from creation to deletion confirmed through the API. Provisioning, dependency install, model download, checkpoint sync and the tail before deletion are all in. Three phase 3 runs on Nebius include 2–6 minutes of manual deletion after a self-termination bug in our orchestrator.

Cost per unit
:   The actual charge divided by the units that run produced: images, optimizer steps, texts, tokens. Tokens are the 4.97M dataset tokens counted with the Qwen2.5 tokenizer, not the padded count.

Steady versus session (H3 table)
:   Steady is the hourly rate × the mean time of clips 2 and 3, on a machine that already has the weights in memory. Session is everything the provider charged for the run, divided by its three clips. The first clip is not in the steady number because it includes loading 67 GB of weights into memory — the whole first clip takes 106.7 s on the H100 and 692.3 s on the L40S, where the boot disk is a network SSD. In clips 2 and 3 the text encoder does not run at all: the prompt is unchanged and ComfyUI serves it from cache. A new prompt adds 15–23 s on a machine with local NVMe. The mp4 encode at the end of each clip runs on the pod’s CPU: on one RunPod pod it took 38 s of a 100 s clip (W5-06), 2–6 s on the other five; the CPU you get inside a pod is not in the offer.

GPU utilisation
:   The mean of periodic `nvidia-smi` samples over the workload; in the H3 table over clips 2–3.

Why measure instead of quoting from a price list
:   Our first 9 phase 3 runs were quoted before anything was measured, from provider price lists and throughput assumptions: the quotes missed by 33 % on average, and every one of them was an overestimate. On the 8 runs used to calibrate model v2, its cost estimates were within ±8 % of the measured costs. This is an in-sample fit, not an independent test of accuracy on future runs. Applied to configurations it had not measured, the calibrated model missed by +85 % (fp32 on a card we had not measured), +151 % (a CPU-bound task on a new host type) and +46 % (spot with three host restarts).

Caveats

## Read these before reusing a number.

- SDXL 1.0 is a 2023 model and a small one. The share of host-CPU work per step is higher than on current models, so the 40 % GPU utilisation on a 5-vCPU host (W2-01) is a fact about that script on that host, not something to generalise.
- “QLoRA is 2.3× slower than bf16 on the same H100” is per step (1.71 s against 0.74 s). By wall clock it is 2.1× (56 min against 26.5 min). The two runs also differ in gradient checkpointing — on for QLoRA, off for bf16 — and in the optimizer, so part of the gap is activation recomputation rather than dequantisation. It is not a one-variable comparison.
- The H100 has no native FP4: nf4 weights are dequantised in software on every step. On Blackwell (B100, B200) the slowdown may not reproduce. Our Blackwell cards so far — the RTX PRO 6000 (bf16) and the RTX 5090 (int8) — have not run nf4.
- The RTX 5090 has 32 GB and the int8 DiT of MiniMax H3 is 32.4 GB: on both 5090 pods VRAM sat at 31.8 GB through the whole sampler, so ComfyUI still streams roughly 1–3 GB of weights per step (9–10 GB on a 24 GB card). Per step it is still the fastest card in the table; the streaming is smaller, not gone.
- RunPod Community pods of the same card come with different RAM. For 45 minutes every Community RTX 5090 on offer had 46–54 GB; this set 
susorokin20
🟧 echo.blog ⭐QRUN publishes its own MiniMax H3 runs across seven GPU configurations at four providers, reporting steady-state costs of $0.0125–0.0375 perQRUN——

Interpretation history

Decision trace