STREAM: Stochastic Riemannian Flow Matching with Anisotropic Decoder for Digital Histopathology Image Generation

DEEPNOID Inc.

NeurIPS 2026 Spotlight

    STREAM uses histopathology vision foundation models (VFMs) like UNI1 as the generative latent space itself, not as conditioning.

    TL;DR

    Existing state-of-the-art histopathology generative models2,3 use VFMs only as conditioning, and most of their output diversity comes from that signal. STREAM instead makes the ℓ2-normalized4,5 patch tokens of a frozen VFM the latent space itself and learns their distribution there, with Riemannian flow matching (RFM)6 on the hypersphere. STREAM makes two novel contributions: a stochastic bridge added to Diffusion Transformer (DiT)7 training and anisotropic noise added to decoder training. With these two contributions, and with every baseline comparison method using the same UNI encoder, STREAM achieves state-of-the-art generation Fréchet Inception Distance (gFID)8 on The Cancer Genome Atlas (TCGA) breast (TCGA-BRCA)9 and colorectal (TCGA-COADREAD)10 cohorts and on SPIDER-skin11.

    Why generate in VFM token space?

    Histopathology VFMs12,13 like UNI are widely used for downstream tasks such as classification, segmentation and retrieval. However, they are rarely used in histopathology image generation.

    Synthetic histopathology images can potentially address the growing data demands of foundation models14,15 and patient-privacy concerns16.

    Previous state-of-the-art generative models such as ZoomLDM and PixCell use VFMs only as conditioning. Therefore, they display conditioning-dominated diversity: on TCGA-BRCA, 62–75% of their output diversity comes from that conditioning signal rather than from the learned latent space.

    Conditioned generative modelse.g. ZoomLDM, PixCell

    1. VFM embedding  c
    2. Diffusion in a VAE latent
    3. VAE decoder
    4. Image

    Diversity comes mostly from c, and sampling needs a VFM embedding.

    STREAMunconditional

    1. Uniform noise on the sphere
    2. Riemannian flow in VFM token space
    3. Decoder
    4. Image

    No conditioning signal: sampling needs no image or embedding as input.

    The VFM token space is the right latent space

    VFM features carry much more semantic information compared to variational autoencoder (VAE) features. A linear probe on SPIDER-breast (benign vs malignant) reaches an area under the receiver operating characteristic curve (AUROC) ≥ 0.995 on VFM features and ≤ 0.640 on VAE features.

    Linear-probe AUROC on SPIDER-breast, benign vs malignant (paper Table 2)
    EncoderAUROC barAUROC
    VFM features
    UNI0.995
    UNI2-h0.997
    VAE features
    PixCell VAE0.640
    ZoomLDM VAE0.624
    Linear-probe AUROC on SPIDER-breast (benign vs malignant), higher is better.

    This shows that the VAE features used by previous latent diffusion models17,18 carry little semantic information, a trend that is similarly observed in previous works in the natural domain.19–21

    Why Riemannian Flow Matching for Histopathology Image Generation?

    We ℓ2-normalize the VFM tokens so that every token lies on the unit hypersphere. Measured on TCGA-BRCA, their tangent drift22 is high (above 0.72 at hop 1), so straight-line interpolants leave the manifold, and Riemannian flow matching on the hypersphere is the better-matched formulation.

    Eigenvalue spectrum of UNI and UNI2-h features on TCGA-BRCA: normalized eigenvalue on a log scale against eigenvalue index, decaying slowly over hundreds of directions; covariance effective rank 252.1 for UNI and 265.0 for UNI2-h. Tangent space analysis of UNI and UNI2-h: tangent drift (1 minus alignment) against graph hop distance 1 to 7, above 0.72 at hop 1 and rising to about 0.98 at hop 7, with the two encoders nearly identical.
    Eigenvalue spectrum (left) and tangent drift against graph-hop distance (right) of UNI and UNI2-h features on TCGA-BRCA.

    The eigenvalue spectrum shows why we use UNI: UNI2-h gains only 5% more effective rank (265 vs 252) for 50% more embedding dimension (1536 vs 1024).

    So a generative model should use the VFM token space itself as the latent space and learn the distribution there.

    And unconditional generation is the right fit: conditioning needs a real image's VFM embedding or labels at inference, especially since annotation is scarce in the histopathology domain.23,24

    Method at a glance

    Overview of STREAM in two stages. H&E patches pass through a frozen UNI encoder, marked with an ice-blue snowflake, into a grid of patch tokens. Stage 1: the tokens lie on a unit hypersphere, where a trainable DiT, marked with a fire-coloured flame, learns to move a point from x0 to x1 along the SLERP geodesic rather than the dashed Euclidean chord; an inset shows Gaussian noise added in the tangent plane at the geodesic point mu_t and mapped back onto the sphere by the exponential map (Exp) to give x_t. Stage 2: the SVD of the DiT's velocity-field Jacobian splits directions into high-response directions U_H (red) and low-response directions U_L (blue), which shape anisotropic noise on a tangent sheet, elongated along U_L, used to train a decoder, also marked with a flame. Generation: a noise ball of points on the sphere is transported by the DiT and decoded by the Decoder into the generated H&E tiles.
    Overview of STREAM: frozen VFM tokens on the unit hypersphere, a DiT trained along bridge-perturbed geodesics, and an anisotropic decoder.

    Stochastic perturbation bridge

    Standard Riemannian flow matching gives no general rectifiability guarantee25, so we perturb its geodesic interpolant with a stochastic bridge. Concretely, during training, each patch token travels from noise to data along the spherical linear interpolation (SLERP) geodesic on the hypersphere, perturbed by tangent Gaussian noise that the exponential map (Exp) carries back onto the sphere; its scale σ(t) = σmax sin(πt) vanishes at both ends. The perturbation gives every intermediate marginal full support on the sphere.

    μt = SLERP(x0, x1, t) xt = Expμt(σ(t) ε)

    Stochastic perturbation bridge Video coming soon
    Schematic on a 2-sphere (the model works per token on Sd−1, d = 1024), with the noise scale exaggerated.

    Anisotropic decoder

    Representation Autoencoder (RAE)20 is essentially the Euclidean baseline of our work as it utilizes VFM features as the diffusion latent. Its isotropic decoder that generates images treats every token direction equally, yet not all directions matter equally for generation: UNI's effective rank is 252 while its token dimension is d = 1024. We therefore train the decoder with anisotropic noise, choosing directions from the trained DiT's velocity-field Jacobian. The noise is shaped by the singular value decomposition (SVD) of that Jacobian: small noise along the high-response directions UH to preserve reconstruction fidelity, and large noise along the low-response directions UL, where the decoder spends its robustness budget.

    Σnoise = σH2 UHUH⊤ + σL2 ULUL⊤, σH ≪ σL

    Anisotropic decoder training Video coming soon
    Two of the 1,023 tangent directions of one token, drawn at the deployed ratio σH : σL = 1 : 10.

    State-of-the-art generation quality on three separate datasets

    Every method is retrained with the same frozen UNI encoder (baselines: ZoomLDM, PixCell, RAE, SVG26, REPA-E27). STREAM has the lowest generation FID on all three datasets.

    TCGA-BRCA

    • STREAM6.16
    • ZoomLDM7.43
    • RAE15.84
    • SVG17.37
    • REPA-E28.57
    • PixCell (off scale)104.18

    TCGA-COADREAD

    • STREAM7.68
    • ZoomLDM8.09
    • RAE19.64
    • SVG23.64
    • REPA-E13.41
    • PixCell (off scale)127.24

    SPIDER-skin

    • STREAM9.19
    • ZoomLDM11.14
    • RAE16.52
    • SVG22.24
    • REPA-E18.00
    • PixCell (off scale)453.28
    gFID, lower is better; PixCell, sampled without an input image, falls off the scale.
    Show as a table
    Generation FID (lower is better)
    MethodTCGA-BRCATCGA-COADREADSPIDER-skin
    STREAM6.167.689.19
    ZoomLDM7.438.0911.14
    RAE15.8419.6416.52
    SVG17.3723.6422.24
    REPA-E28.5713.4118.00
    PixCell104.18127.24453.28

    Efficient and fast inference

    On TCGA-BRCA, STREAM at a number of function evaluations (NFE) of 26 already beats RAE and SVG at NFE 50, and at NFE 50 it is also the cheapest of the three per image.

    gFID against NFE on TCGA-BRCA Mean gFID over 3 sampling seeds. STREAM: 49.70, 27.05, 9.31, 6.16 at NFE 6, 10, 26, 50. RAE: 55.03, 29.87, 18.27, 15.88. SVG: 55.80, 30.53, 18.82, 17.52. 0 20 40 60 6 10 26 50 NFE (function evaluations) gFID SVG, NFE 6: 55.80 ± 0.07 SVG, NFE 10: 30.53 ± 0.18 SVG, NFE 26: 18.82 ± 0.15 SVG, NFE 50: 17.52 ± 0.14 RAE, NFE 6: 55.03 ± 0.05 RAE, NFE 10: 29.87 ± 0.06 RAE, NFE 26: 18.27 ± 0.04 RAE, NFE 50: 15.88 ± 0.04 STREAM, NFE 6: 49.70 ± 0.09 STREAM, NFE 10: 27.05 ± 0.06 STREAM, NFE 26: 9.31 ± 0.02 STREAM, NFE 50: 6.16 ± 0.04 SVG 17.52 RAE 15.88 STREAM 6.16 9.31
    gFID against NFE on TCGA-BRCA, lower is better.
    Show as a table
    gFID against NFE on TCGA-BRCA, mean ± standard deviation (SD) over 3 sampling seeds (lower is better)
    NFESTREAMRAESVG
    649.70 ± 0.0955.03 ± 0.0555.80 ± 0.07
    1027.05 ± 0.0629.87 ± 0.0630.53 ± 0.18
    269.31 ± 0.0218.27 ± 0.0418.82 ± 0.15
    506.16 ± 0.0415.88 ± 0.0417.52 ± 0.14

    Latency, ms per image

    • STREAM94.8
    • RAE156.7
    • SVG175.0

    Compute, TFLOPs per image

    • STREAM12.14
    • RAE16.09
    • SVG24.16
    Latency and compute per generated image at NFE 50 (latency on one NVIDIA H200), lower is better.
    Show as a table
    Inference cost per generated image at NFE 50 (paper Table 17, lower is better)
    Modelms / imageTFLOPs / image
    STREAM94.812.14
    RAE156.716.09
    SVG175.024.16

    Ablations on STREAM's components

    Stochastic bridge and anisotropic decoder

    The two components are superadditive: together they improve reconstruction Fréchet Inception Distance (rFID) and gFID far more than the sum of their individual contributions. Both components are crucial: standard RFM with an isotropic decoder (top left) does not perform better than previous existing works like ZoomLDM.

    Ablation on TCGA-BRCA (UNI encoder) at 45K decoder steps: rFID and gFID, lower is better (paper Table 4)
    Isotropic decoderAnisotropic decoder
    No bridge rFID6.51gFID9.07 rFID5.02gFID8.27
    Bridge rFID6.51gFID9.37 rFID3.52gFID6.86STREAM
    rFID and gFID on TCGA-BRCA (UNI encoder, 45K decoder steps), lower is better.

    The decoder's directional preference tracks superadditivity

    We next examine how and why this superadditivity arises. It has an empirical explanation in the decoder's directional preference. Rdec = LUH / LUL compares the decoder's Learned Perceptual Image Patch Similarity (LPIPS)28 sensitivity along the high-response directions UH and the low-response directions UL. Anisotropic training tilts the decoder's preference toward UH, and STREAM (bridge + anisotropic decoder) ends up with the strongest UH preference, compared with no bridge or the isotropic decoder.

    STREAM, k = 64R_dec: isotropic 0.42, anisotropic 3.41
    STREAM, k∗ = 236R_dec: isotropic 1.12, anisotropic 2.31
    No bridge, k = 64R_dec: isotropic 0.62, anisotropic 2.50
    No bridge, k∗ = 232R_dec: isotropic 1.28, anisotropic 2.01
    Rdec on TCGA-BRCA, log scale; k = 64 is a fixed split and k∗ each model's own energy cut.
    Show as a table
    Directional-LPIPS ratio R_dec on TCGA-BRCA, 2,000 tiles (paper Table 5)
    RdecIsotropicAnisotropic
    STREAM (k = 64)0.423.41
    STREAM (k* = 236)1.122.31
    No bridge (k = 64)0.622.50
    No bridge (k* = 232)1.282.01

    With the bridge, anisotropic training shifts Rdec further toward UH than without it at both splits (0.42 → 3.41 vs 0.62 → 2.50 at k = 64, 1.12 → 2.31 vs 1.28 → 2.01 at k∗), so the largest shift belongs to STREAM, consistent with its superadditive rFID and gFID gains.

    Decoder response visualized in pixel space

    Under a perturbation inside the top-64 subspace (UH at k = 64), the anisotropic decoder separates from the isotropic one and responds more strongly. On UL (figure not shown here, refer to main text), no visual difference between the two decoders is observed.

    Grid of H&E decodes and amplified difference maps for one generated latent, perturbed inside the top-64 subspace at the clean latent and amplitudes 1x, 4x, 8x, 32x, 64x and 128x. The anisotropic decoder (upper block) shows brighter difference maps and visibly changed tissue at the largest amplitudes; the isotropic decoder (lower block) shows fainter differences.
    Anisotropic (upper) and isotropic (lower) bridge decoders decode the same latent perturbed in the top-64 block, not the deployed split, and |Δ| rows show the amplified difference from the clean decode.
    References (28)
    • 1Chen et al., Towards a general-purpose foundation model for computational pathology, Nature Medicine 2024.
    • 2Yellapragada et al., ZoomLDM: Latent Diffusion Model for multi-scale image generation, CVPR 2025. arXiv:2411.16969
    • 3Yellapragada et al., PixCell: A Generative Foundation Model for Digital Histopathology Images, arXiv preprint 2025. arXiv:2506.05127
    • 4Kumar and Patel, Learning on the Manifold: Unlocking Standard Diffusion Transformers with Representation Encoders, ECCV 2026. arXiv:2602.10099
    • 5Chang et al., HAE: Hyperspherical Autoencoder for High-Fidelity Image Reconstruction and Generation, arXiv preprint 2026. arXiv:2601.22904
    • 6Chen and Lipman, Flow matching on general geometries, ICLR 2024. arXiv:2302.03660
    • 7Peebles and Xie, Scalable diffusion models with transformers, ICCV 2023.
    • 8Heusel et al., GANs trained by a two time-scale update rule converge to a local Nash equilibrium, NeurIPS 2017.
    • 9The Cancer Genome Atlas Network, Comprehensive molecular portraits of human breast tumours, Nature 2012.
    • 10The Cancer Genome Atlas Network, Comprehensive molecular characterization of human colon and rectal cancer, Nature 2012. doi:10.1038/nature11252
    • 11Nechaev et al., SPIDER: A Comprehensive Multi-Organ Supervised Pathology Dataset and Baseline Models, arXiv preprint 2025. arXiv:2503.02876
    • 12Zimmermann et al., Virchow2: Scaling Self-Supervised Mixed Magnification Models in Pathology, arXiv preprint 2024. arXiv:2408.00738
    • 13Xu et al., A whole-slide foundation model for digital pathology from real-world data, Nature 2024.
    • 14Lu et al., A visual-language foundation model for computational pathology, Nature Medicine 2024.
    • 15Xiang et al., A vision-language foundation model for precision oncology, Nature 2025.
    • 16Wang et al., Self-improving generative foundation model for synthetic medical image generation and clinical applications, Nature Medicine 2025.
    • 17Rombach et al., High-resolution image synthesis with latent diffusion models, CVPR 2022.
    • 18Ho et al., Denoising diffusion probabilistic models, NeurIPS 2020.
    • 19Chen et al., Aligning Visual Foundation Encoders to Tokenizers for Diffusion Models, ICLR 2026.
    • 20Zheng et al., Diffusion Transformers with Representation Autoencoders, ICLR 2026. arXiv:2510.11690
    • 21Chen et al., Masked Autoencoders Are Effective Tokenizers for Diffusion Models, ICML 2025.
    • 22Xiong et al., Exploiting Low-Dimensional Manifold of Features for Few-Shot Whole Slide Image Classification, ICLR 2026.
    • 23Campanella et al., Clinical-grade computational pathology using weakly supervised deep learning on whole slide images, Nature Medicine 2019.
    • 24van der Laak et al., Deep Learning in Histopathology: The Path to the Clinic, Nature Medicine 2021.
    • 25Hertrich et al., On the relation between rectified flows and optimal transport, arXiv preprint 2025. arXiv:2505.19712
    • 26Shi et al., Latent Diffusion Model without Variational Autoencoder, arXiv preprint 2025. arXiv:2510.15301
    • 27Leng et al., REPA-E: Unlocking VAE for End-to-End Tuning with Latent Diffusion Transformers, ICCV 2025.
    • 28Zhang et al., The unreasonable effectiveness of deep features as a perceptual metric, CVPR 2018.

    BibTeX

    If you find our work useful, please cite it using the BibTeX below.

    @inproceedings{cho2026stream,
      title     = {{STREAM}: Stochastic Riemannian Flow Matching with Anisotropic Decoder for Digital Histopathology Image Generation},
      author    = {Cho, Won June and Jeong, Daeky and Lim, Hyeongyeol and Yoon, Hongjun},
      booktitle = {The Fortieth Annual Conference on Neural Information Processing Systems},
      year      = {2026},
      url       = {https://arxiv.org/abs/2606.07036}
    }