Skip to content
Clothsy AI Talk to us

Efficient visual AI7 min readFabricVTON Research

Fast enough for a storefront

Diffusion models are slow by design. How step distillation works, what it costs in detail, and why speed is a research problem for try-on.

A try-on happens in the middle of shopping. The person has a product page open and a decision to make. If the image takes too long they move on, and if every image costs too much a store cannot offer try-on to every visitor. So speed and cost are not details to fix after the research is done. They decide whether the research reaches anyone at all, and the fastest ways to run a model change what it produces.

Why diffusion is slow

A diffusion model generates an image by starting from noise and removing it step by step. Each step is a full pass through a large neural network, and standard samplers take many steps. Latent diffusion made each step cheaper by working in a compressed latent space rather than on pixels [1], but the iterative structure remains.

Try-on adds more work per step. Many strong systems process the garment in its own network branch alongside the person, like the parallel UNets of TryOnDiffusion [2] or IDM-VTON’s garment network [3], and high output resolutions multiply the cost again.

Original sampler: dozens of stepsDistilled student: a handful of stepsNoiseImage
Figure 1Why speed is a research problem. A diffusion model turns noise into an image over many denoising steps. Distillation trains a student that gets to a comparable image in a handful of steps, and the cost of a try-on falls with the step count.

Fewer steps: distillation

The most effective lever is to take fewer steps. Distillation trains a fast “student” model to match what a slow “teacher” produces with many steps. A few landmarks:

  • Progressive distillation repeatedly trains a student to do in one step what the teacher did in two, halving the step count each round [4].
  • Consistency models learn to map any point along the denoising path directly to its end, which allows generation in a single step or a few [5].
  • Latent consistency models brought that idea to latent diffusion, producing high-resolution images in about two to four steps [6].
  • Adversarial diffusion distillation combines distillation with an adversarial loss, sampling in one to four steps [7].

What speed costs

Fewer steps are not free. Few-step students can lose fine detail and variety compared with their teachers, and for try-on, fine detail is exactly where fidelity lives: the print, the logo, the weave. A distilled model that looks just as good on average and quietly softens every logo is a worse product.

So distillation has to be judged with the same fidelity measures as the original model, on the same held-out test set, with particular attention to garment regions and printed detail. We described those measures in Evaluating texture fidelity. A speed-up only counts if fidelity holds.

Beyond step count

Step count is the biggest lever, not the only one:

  • Smaller students. The student does not have to be the same size as the teacher.
  • Lower precision. Running weights and activations at reduced numerical precision cuts memory and time, if quality is checked after the change.
  • Batching. Serving many requests together keeps hardware busy under real traffic.
  • Reuse what does not change. In try-on, many shoppers try the same product. Work that depends only on the garment is an opportunity to compute once and reuse.

Fast, affordable inference is one of our four open problems, and the one that most directly decides who gets to use everything else we build.

References

  1. Rombach et al. High-Resolution Image Synthesis with Latent Diffusion Models. CVPR, 2022.
  2. Zhu et al. TryOnDiffusion: A Tale of Two UNets. CVPR, 2023.
  3. Choi et al. Improving Diffusion Models for Authentic Virtual Try-on in the Wild. ECCV, 2024.
  4. Salimans & Ho Progressive Distillation for Fast Sampling of Diffusion Models. ICLR, 2022.
  5. Song et al. Consistency Models. ICML, 2023.
  6. Luo et al. Latent Consistency Models: Synthesizing High-Resolution Images with Few-Step Inference. arXiv, 2023.
  7. Sauer et al. Adversarial Diffusion Distillation. ECCV, 2024.

JOIN THE RESEARCH TEAM

Work on the open problems with us.

Students, researchers and engineers. The application takes about two minutes.