An image model sees a shirt as pixels. A person sees a material: cotton or silk, heavy or light, a woven check or a printed logo, cut close to the body or falling loose. Virtual try-on sits between the two. It has to produce pixels, but it is only useful if those pixels behave like the material would. That is why we think of try-on as a physical problem that happens to be solved in image space, and why fabric is harder than it looks.
Four properties do most of the damage when they go wrong.
Print and logos
Printed detail is high-frequency, it is small, and it is rigid in meaning. A stripe can bend with the fabric; a letter cannot turn into a different letter. A shopper will forgive a slightly wrong shadow long before they forgive a misspelled brand name.
Modern generators make this harder in a specific way. Latent diffusion models do not work on pixels directly. An autoencoder first compresses the image into a smaller latent grid, eight times smaller in each dimension in the widely used Stable Diffusion setup, and generation happens there [1]. That compression is what makes high-resolution generation affordable. It is also a bottleneck for exactly the detail try-on needs to carry: a small logo may span only a few latent cells before the generator does anything at all.
Weave and sheen
How fabric looks depends on its structure and on light. A twill weave throws diagonal highlights, satin reflects like a soft mirror, and a knit scatters light in its loops. Graphics describes this with reflectance models, and learned methods can now estimate those reflectance maps from a single photograph of a surface [2].
A product photo gives one view under one lighting setup. The try-on image usually has different light, a different angle and a curved surface. A model that has only learned “this garment is dark blue” will get the colour right and the fabric wrong: a flat, plastic-looking result where there should be the soft sheen of wool or the crisp surface of cotton.
Drape and folds
Drape is how a fabric hangs, and it is governed by physics: how much the material stretches, how stiffly it resists bending, and how much it weighs. The same dress falls differently on different bodies, and a heavy coat and a light blouse fall differently on the same body.
Cloth simulation has modelled this for decades. Implicit integration made simulation stable enough to take large time steps [3], and it remains a foundation of the field. Learning-based methods now model garment dynamics too: SNUG trains neural garment models with physics-based losses instead of simulated data [4], and HOOD uses a hierarchical graph network that generalizes to garments it has not seen [5].
Image-based try-on has none of the inputs these methods expect. There is no 3D sewing pattern and no material parameters, only a photograph. The model has to infer, from appearance alone, roughly how the fabric would behave on this particular body. When it cannot, it tends to copy the folds from the product photo, which is why so many try-ons look like the garment was pasted on.
Layers and occlusion
Real outfits overlap. A shirt is tucked or left out, a jacket hangs open, hair falls over a collar, a hand rests in a pocket, a bag strap crosses the chest. The model has to decide what is in front of what, and how contact changes shape. Warping-based systems addressed occlusion explicitly when fitting the garment [6], and newer work lets a user control how several garments are layered and styled [7]. Getting it right on arbitrary photos of real people is still open.
What this means
For us, taking fabric seriously has three consequences.
- Evaluate the garment, not just the image. Scores averaged over a whole picture barely notice a wrong logo. We look at the garment region, and at printed detail specifically. More in Evaluating texture fidelity.
- Carry detail explicitly. Fine detail should have its own path through the model instead of relying on a compressed latent to preserve it.
- Bring physics back in. Priors from simulation and from material understanding can inform drape where a single photo cannot.
These map directly onto two of our open problems: fidelity through warping, and drape.
References
- Rombach et al. High-Resolution Image Synthesis with Latent Diffusion Models. CVPR, 2022.
- Deschaintre et al. Single-Image SVBRDF Capture with a Rendering-Aware Deep Network. ACM Transactions on Graphics (SIGGRAPH), 2018.
- Baraff & Witkin Large Steps in Cloth Simulation. SIGGRAPH, 1998.
- Santesteban, Otaduy & Casas SNUG: Self-Supervised Neural Dynamic Garments. CVPR, 2022.
- Grigorev et al. HOOD: Hierarchical Graphs for Generalized Modelling of Clothing Dynamics. CVPR, 2023.
- Lee et al. High-Resolution Virtual Try-On with Misalignment and Occlusion-Handled Conditions. ECCV, 2022.
- Zhu et al. M&M VTO: Multi-Garment Virtual Try-On and Editing. CVPR, 2024.
