2026-10-11 16:37 UTC

Linum claims its released JiT-DDT pixel-space encoder-decoder trains a text-to-image model with 3.6 times fewer GPU-hours than its Linum v2 baseline while generating four times as many pixels, potentially lowering diffusion-training costs.

state: seedheat: lowuncertainty: highnovelscott: lowdiffusion-training training-optimization ai-infrastructureLinumSahil ChopraManu Chopra

What is this?

The case describes Linum’s claimed release of JiT-DDT, a pixel-space encoder-decoder for text-to-image diffusion training, with code and weights reportedly under an Apache license. Linum claims it uses 3.6 times fewer GPU-hours than its Linum v2 baseline while generating four times as many pixels. The supplied web results are unrelated and do not corroborate the release, licensing, benchmark conditions, or comparable image quality; the case names Sahil Chopra and Manu Chopra but does not establish their roles.

Why it matters to Scott

Scott’s BRIA 3.2 spike concerns local inference, not diffusion training; the hits establish neither a training-efficiency position nor a project whose costs this release would change, so the overlap is topical rather than actionable. The radar tracks related text-to-image training developments but not JiT-DDT itself, and the supplied evidence does not corroborate Linum’s claimed savings or comparable image quality.
radar:jasper-from-scratch-t2i-kitradar:concept.training-efficiencyradar:concept.text-to-image
queries asked of Scott's wikis
  • diffusion image generation training projects
  • training compute economics algorithmic efficiency
  • pixel-space versus latent-space model architectures
  • open code weights permissive licensing reproducibility
  • GPU-hour benchmarks quality resolution tradeoffs

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 626h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-15 14:00⭐ origin echo-reconstructedLinum reports 3.6× fewer GPU-hours at 4× the pixels against its prior image-only baseline and releases JiT-DDT code and weights under Apache
Sahil Chopra and Manu Chopra on blog (echo) · attributed from hn.story.49729816
—
09-16 16:58first on hacker news · published · +27.0hTraining Text-to-Image Models 3.6× Faster
schopra909
—
09-16 16:58amplified on hacker news 👑hn.story.49729816
schopra909
peak 56 · 10 comments · 65% of case engagement
10-06 20:24amplified on hacker newshn.story.49983582
schopra909
peak 26 · 9 comments · 35% of case engagement
09-16 21:20our radar first saw it · +31.4hdiscovery anchor: hn.story.49729816—
pace: p62 vs 1032 stories at the 336h mark (now 626h old) — ahead of anthropic-context-compaction-cost-reversal (1.1x), behind cache-tax-idle-session-warming (1.0x)

Evidence (3) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnTraining Text-to-Image Models 3.6× Faster
Retrieved article excerpt

Open article · Retrieved 2026-09-16T21:22:30.387641+00:00

Research · 40 min read

# Training Text-to-Image Models 3.6× Faster

Beating latent diffusion models with pixel-space encoder-decoders

[Sahil Chopra](https://thesahilchopra.com/) & [Manu Chopra](https://www.linkedin.com/in/manu-chopra-50360b170/)

[Co-CEOs @ Linum](https://linum.ai) · September 16, 2026

Road to Linum v3· Issue 02previously: [data filtering](https://www.linum.ai/field-notes/data-filtering-gen-video)

TL;DR

Linum v2 was bottlenecked by the enormous size of its attention context window. A 720p, 5 second clip cost a whopping 110K tokens. To put that in perspective, LLMs see samples with fewer than 8K tokens for [97% of their pretraining](https://arxiv.org/pdf/2512.13961). Attention is quadratic in cost, so the biggest lever we have to accelerate model training is pruning the context window down.

Most generative image and video systems are Latent Diffusion Models (LDMs). They split compression and generation into independently trained modules: the Variational Autoencoder (VAE) and the DiT (Diffusion Transformer). Recently, pixel-space models like the JiT have shown to be a promising alternative. It reduces two models into one and allows the diffusion model to construct a latent space specifically for generation, rather than rely on one built for reconstruction.

When trained on our (image, caption) dataset, the JiT seems to struggle to produce finegrained details. We propose a novel encoder-decoder architecture (JiT-DDT) that recovers this detail and trains much more efficiently than its LDM counterpart. Against our Linum v2 baseline, the JiT-DDT trains a text-to-image model with 3.6× fewer GPU-hours, even though it generates images with 4× the pixels.

3.6× faster to train, at 4× the pixels

Linum v2 (ours, previous)\* sample, 256×256

**Linum v2 (ours, previous)\*** · 256×256

2.0B latent-space DiT + VAE

256 latent tokens

\* image-only checkpoint

JiT-DDT (ours, new) sample, 512×512

**JiT-DDT (ours, new)** · 512×512

2.5B active pixel-space DiT

320 pixel tokens = 64 encoder + 256 decoder

GPU-hours

0

3.6× fewer

0

samples seen

0M

4.2× fewer

0M

For more comparisons, see [Appendix](https://www.linum.ai/field-notes/jit-ddt#appendix)

Research release

[GitHubModel code](https://github.com/Linum-AI/jit-ddt)[Hugging FaceModel weights](https://huggingface.co/Linum-AI/jit-ddt)

JiT-DDT code and model weights are available under the Apache 2.0 license. We hope that by sharing our findings with the broader community, we can encourage others to also explore more efficient training methods. This should be treated as a research artifact, not a full model release. Stay tuned for more research checkpoints like this, en route to Linum v3.

## Hitting the VAE compression wall

Almost all generative image and video models are Latent Diffusion Models (LDMs). These have two key components, a Variational Auto Encoder (VAE) for compression and a Diffusion Transformer (DiT) for generation.

Operating in raw pixels is too expensive (especially for video), so we first need to find a way to reduce RGB pixels into a smaller amount of tokens for the DiT. This is where the VAE comes in. It's trained for compression and reconstruction. Specifically, it pushes our pixel-space samples through a probabilistic encoder, spits out -dimensional tokens, and then pushes these latent tokens through a probabilistic decoder to land back in pixel-space.

The VAE is trained to compress and reconstruct

replay

input x

Encoder

encoder

μ = [?, ?]

σ = [?, ?]

μ, σ

z ∈ ℝᵏ

sample z

Decoder

decoder

output x̂

‖x − x̂‖²

+ β · KL(q‖N)

loss

‹ready›

gradient (purple) reaches every weight\* simplified: in practice the KL term is ≈ 0, previously we trained a σ-VAE with an L1 reconstruction loss plus LPIPS and GAN losses; see [our VAE post](https://www.linum.ai/field-notes/vae-reconstruction-vs-generation).

When building a LDM, you train the VAE separately and then freeze it (i.e. no gradient flow from the DiT into the VAE). This way the latent space stays static throughout the course of DiT training. You run the VAE's encoder to embed your data, train the DiT to traverse the VAE's latent space, and then transform the DiT-generated latent tokens into pixel space using the VAE's decoder.

The VAE is trained once and frozen;  
the DiT learns to move through its latent space

TrainingInferencereplay

input x

❄Encoder

encoder

= E(x)

z

(1−t)·z

+ t·ε

INTERPOLATE

zₜ

DiT

trainable

dit

v̂

pred

‖v̂ − v‖²

v = ε − z

loss

ε ~ N(0, I)

gaussian

ε

SAMPLE ε

t ~ LogitNormal

sample t

‹ready›

VAE frozen (dashed) · DiT trainable (purple) · gradient stops at the DiT · t = 0 clean image, t = 1 pure Gaussian noise

We want to eke out as much token compression as possible from the VAE, so that we can curb the cost of attention in our DiT. But if you take a survey of the popular open source text-to-image models like FLUX, Ideogram, and Z-Image, you'll notice that they all cap out at 16×16 token reduction. This aligns with [our experiments on Image-Video VAEs from a few years ago](https://www.linum.ai/field-notes/vae-reconstruction-vs-generation). Unfortunately, it seems like there is an empirical ceiling on the amount of compression we can get out of a standard CNN VAE without degrading the reconstructions.

## Unlocking aggressive compression with a unified model

Last fall, Tianhong Li and Kaiming He published a paper ([JiT](https://arxiv.org/pdf/2511.13720)) that achieves 32×32 token reduction by throwing away the VAE altogether and pushing the compression task into the DiT itself.

Patchify: 4×4-pixel patches → 48-dim tokens → linear bottleneck to 12

replay

16×16 pixels, 3 channels (RGB) each

cut into 4×4-pixel patches (16 patches)

each patch is tokenized independently:
16 pixels × 3 channels become one 48-dim token

a linear layer W ∈ ℝ12×48 projects 48 dims down to 12

‹ready›

Illustrative. In JiT at 512px we use 32×32 patches, so a 512×512 image becomes 256 tokens, each starting at 32·32·3 = 3,072 dims; the bottleneck maps that to 256.

This approach to reducing token counts isn't particularly new. It was invented for vision transformers ([ViT](https://arxiv.org/pdf/2010.11929)) half a decade ago, and it's pretty commonly paired with a VAE to further condense token sequences before they enter the DiT.In Linum v2, our VAE gave us 8×8 (h×w) compression and 16-dimensional latents. At the base of the DiT, we applied 2×2 patchification to get 16×16 token compression and 64-dimensional latents. We used it in Linum v2 and so do models like FLUX.

*So, why hasn't anyone tried this before?* This feels like a free lunch. You get a (potentially) lossless way to cut down attention cost, and it's bone-dead simple.

In early 2025, papers like [VA-VAE](https://arxiv.org/pdf/2501.01423) demonstrated that DiTs struggle to learn from high dimensional inputs.There are small hacks like using an external model as a regularizer during VAE training (e.g. DINOv3) that (likely) enabled models like [FLUX-2](https://bfl.ai/research/representation-comparison) to make the leap from 64 latent dimensions to 128 latent dimensions for their DiT. But, these strategies just kick the can down the road on a clear learnability problem within the DiT. Aggressive patchification explicitly pushes information into the channel dimension, so it triggers this instability. But as it turns out, this is not intrinsic to the architecture. Rather, it's downstream of the v-prediction, v-loss flow matching objective that everyone's been using to train diffusion models these past few years.

### A quick refresher on flow matching

In old school 2022-era denoising diffusion ([DDPM](https://arxiv.org/pdf/2006.11239)), we [iteratively noise a sample](https://lilianweng.github.io/posts/2021-07-11-diffusion-models/#what-are-diffusion-models) and train a neural network to remove the noise. This way at inference time we can use our neural network to transform Gaussian noise into a sample from our data distribution over a sequence of steps. This formulation has a host of issues (e.g. [oversaturation in generation](https://arxiv.org/pdf/2305.08891), [unstable learning](https://arxiv.org/pdf/2312.02696), [distillation collapse](https://arxiv.org/pdf/2202.00512)), so in the intervening years the field has shifted away from it towards flow matching.

In [flow matching](https://arxiv.org/pdf/2210.02747), we construct a straight line path between every sample in our data distribution and a sample of Gaussian noise:The path between noise and samples does not have to be straight. But in practice, we all do it.

At , we recover . At , we get , where .We follow the DDPM convention throughout this post:  is data,  is noise. Some flow matching papers run the other way, with  as noise and  as data. The two formulations are equivalent. Then we train a network to approximate the velocity along that path:

We call this v-prediction, v-loss because the neural network is explicitly predicting velocity and it's trained on the MSE between its velocity prediction and the ground-truth, conditional velocity field.

### V-prediction and the curse of dimensionality

If you're training a flow matching model you don't necessarily need to train your neural network to predict and regress velocity. The three terms are linearly re-arrangeable; so you can mix and match , , and  across prediction and regression targets:

Nine ways to write one objective

Three targets, each linear in the other two

Rearrange one identity to fill each off-diagonal cell

Pick what the network predicts (columns) and what the loss measures (rows). Each off-diagonal cell is one of the three identities above, rearranged to turn the prediction into the loss target. x-prediction with v-loss is the cell we use. Table after Li & He (2025).

In JiT, Li and He revisited the v-prediction, v-loss decision that the field's been making since the inception of flow matching. They took a toy distribution (points on a spiral) and then projected these points from 2D to different high dimensional spaces of increasing size. For each of these spaces, they trained flow matching models with x-prediction, -prediction, and velocity-prediction; and found that the x-prediction was the only model type to accurately generate samples from the spiral distribution at large  dimensions. DiTs have been struggling to learn from high-dimensional inputs because of the curse of dimensionality.

Grid of spiral samples at D=2, 8, 16 and 512 for x-, epsilon- and v-prediction; only x-prediction holds up at D=512

A 2D spiral buried in a D-dimensional space by a random projection. As D grows, epsilon- and v-prediction collapse while x-prediction keeps recovering the spiral. Figure 2 from Li & He (2025).

Velocity is . When we do v-prediction, our neural network has to implicitly learn the signal () and noise (). Noise is a random Gaussian that will cover the entire -dimensional space. So, as we scale  the problem of fitting noise (within the velocity term) becomes exponentially harder. This is why aggressive patchification failed in the past and why LDMs have been struggling to learn from high-dimensional VAE latents. As we grow the channel dimension, we end up in the degenerate case where our DiT is struggling to learn high dimensional Gaussian noise.

By switching to x-prediction, we can try to side-step the curse of dimensionality. If we believe that images and videos naturally lie on a low dimensional manifold, we should be able to have our models predict  effectively even with high .

Noise fills the whole D-dimensional ball; images sit on a thin sliver of it

replay

ε ~ N(0, ID)

n = **0** · intrinsic dim = **512**

x₀ ~ image manifold

n = **0** · intrinsic dim = **3**

1 dot = 1 sample · D = 512 · each dot is plotted at coordinates 1, 2 and 3 of its 512same sphere, same n on both sides · the manifold is illustrative, not a real image set

In high-dimensional space (D = 512), noise is truly random. It spreads across th
schopra9095610
🟧 echo.blog ⭐Linum reports 3.6× fewer GPU-hours at 4× the pixels against its prior image-only baseline and releases JiT-DDT code and weights under ApacheSahil Chopra and Manu Chopra——
🟧 hnTraining Text-to-Image Models Without a VAEschopra909269

Interpretation history

Decision trace