Abstract
Dynamic contrast-enhanced magnetic resonance imaging (DCE-MRI) is essential for breast cancer management, but reliance on gadolinium-based contrast agents (GBCAs) restricts use in contraindicated populations, prolongs scan protocols, and presents environmental toxicity concerns. Contrast synthesis offers a non-invasive alternative; however, existing approaches struggle to balance spatial realism with temporal continuity, suffer from slow iterative sampling, underutilize structural priors, and lack clinical validation. We propose a novel conditioned latent transport framework that predicts contrast enhancement in a single forward pass. By anchoring the latent trajectory to the pre-contrast anatomy and applying continuous time conditioning, the model synthesizes patient-specific contrast evolution at any given acquisition time. The proposed approach outperforms baseline and the state-of-the-art models across spatial, perceptual, temporal, and distributional metrics. Evaluated on an independent external cohort, the method demonstrates robustness to domain shifts induced by scanner noise as well as differing acquisition protocol. Furthermore, our synthetic contrast enhancement significantly improved downstream tumor segmentation performance, yielding a 22.4% relative increase in Dice coefficient (0.60 vs. 0.49 baseline pre-contrast, p < 0.01), reducing boundary segmentation error by over 39%, while outperforming all other generative model baselines. Finally, a reader study involving four breast radiologists evaluated the image quality, kinetic fidelity, and diagnostic viability of our synthesized sequences across 40 randomly selected cases. The results demonstrated that in 70% of cases, synthesized images provided sufficient clinical information to support the same management decisions as real DCE-MRI, suggesting a path toward safer and faster contrast-free or contrast-reduced imaging workflows.
Method
Contrast enhancement is predicted in one forward pass by transporting a latent that stays anchored to the patient's own pre-contrast anatomy.
-
1
Anchor the anatomy
A frozen encoder compresses the pre- and post-contrast images to latents. zpre becomes a fixed structural anchor, so the network never re-synthesizes baseline anatomy.
-
2
Predict only the change
Conditioned on zpre, a sinusoidal time embedding τ, and the interpolation step t, a U-Net regresses the latent subtraction map αΊenh ≈ zpost − zpre.
-
3
Decode with two-fold supervision
A frozen decoder returns to pixel space, where LPIPS enforces perceptual fidelity and a log-amplitude Fourier loss preserves fine micro-vascular detail.
Result 1 of 3
Synthesis
Anchoring the latent trajectory to pre-contrast anatomy and conditioning on continuous acquisition time produces contrast evolution that is both spatially faithful and temporally smooth, in a single forward pass, with no iterative sampling. The method leads baseline and state-of-the-art comparisons across pixel-level, perceptual, distributional and temporal metrics, and holds up on an independent external cohort whose pharmacokinetic wash-in is substantially faster than the training distribution.
A Dense trajectories from sparse acquisitions
A clinical protocol acquires only 4β5 post-contrast phases. Because τ is a continuous condition rather than a phase index, the model renders the enhancement curve at any requested time, turning a handful of discrete acquisitions into a dense trajectory.
B Temporal continuity under stochasticity
Injecting latent noise restores realistic texture, but sampling it independently per phase makes the sequence flicker non-physiologically. Holding a single patient-level noise map constant across phases leaves the temporal condition as the only driver of change, and yields smooth trajectories for both CCNet and our method. The difference is what happens across seeds: anchored to the pre-contrast latent, our anatomy stays fixed and the noise models epistemic uncertainty, whereas CCNet, initialized from pure noise, varies in tumor extent and hallucinates anatomy. That distinction is what makes the per-pixel uncertainty maps used in the reader study meaningful.
Result 2 of 3
Segmentation
Does the synthesized uptake carry real biological signal, or only the appearance of it? We feed synthetic images to a downstream tumor segmentation network and compare against real pre-contrast (Baseline) and real post-contrast (Upper Bound) inputs on the MAMA-MIA internal validation set.
Quantitative segmentation performance
| Method | Dice ↑ | HD95 ↓ |
|---|---|---|
| Segmentation model trained on post-contrast | ||
| Baseline | 0.17 (0.30) | 160.69 (104.48) |
| U-Net | 0.44 (0.36) | 82.20 (104.68)* |
| pix2pix | 0.22 (0.33) | 147.79 (108.24) |
| CCNet | 0.45 (0.35) | 76.81 (102.32) |
| TeNCA | 0.30 (0.35) | 122.31 (110.69) |
| Ours | 0.51 (0.37) | 68.78 (99.62) |
| Upper Bound | 0.68 (0.33) | 38.59 (77.73) |
| Segmentation model trained on pre-contrast | ||
| Baseline | 0.49 (0.37) | 71.48 (99.90) |
| U-Net | 0.56 (0.35) | 53.05 (89.01) |
| pix2pix | 0.44 (0.35) | 70.10 (96.55) |
| CCNet | 0.44 (0.35) | 78.51 (102.96) |
| TeNCA | 0.51 (0.36) | 62.59 (93.86) |
| Ours | 0.60 (0.33) | 43.38 (80.12) |
| Upper Bound | 0.63 (0.34) | 47.46 (84.74) |
Mean (standard deviation). Significance by paired Wilcoxon signed-rank test (p < 0.01); results that did not reach significance are marked *. Success on the post-contrast-trained network shows the synthesized uptake mirrors the ground truth; success on the pre-contrast-trained network shows the synthesis preserves the underlying morphology and texture rather than overwriting it.
Result 3 of 3
Clinical Reader Study
Four board-certified breast radiologists from three academic centers (5, 8, 13 and 15 years of sub-specialty experience) evaluated 40 cases stratified by tumor size, shape, and enhancement pattern, yielding 2,927 analyzable responses across four study sections. The full study platform, with the case presentation and questionnaires exactly as the readers saw them, is public.
Which synthesis did readers prefer?
Share of 160 blinded evaluations (40 cases Γ 4 readers) selecting each method as the superior image.
Preference for our method was consistent across all four readers and highly significant (p < 0.001); U-Net and TeNCA were statistically indistinguishable from one another (p = 1.0).
Data table
| Method | Votes | Share |
|---|---|---|
| Ours | 134 | 83.8% |
| U-Net | 13 | 8.1% |
| TeNCA | 13 | 8.1% |
| Total | 160 | 100% |
What happens to clinical management?
Impact of relying on the synthetic image instead of the real acquisition, across all side-by-side assessments.
Perceived synthetic appearance (OR 3.60) and diagnostic complexity (OR 2.14) independently predicted greater clinical impact; reader experience was protective (OR 0.54). Mean characterization agreement between synthetic and real images was 69.8% (median 66.7%). For every categorical task, GT-to-synthetic agreement (Cohen's κ) exceeded the inter-reader agreement (Fleiss' κ) already present among radiologists reading the real images.
Data table
| Impact on clinical management | Share of assessments |
|---|---|
| No change | 35.6% |
| Minor change, no impact on management | 34.4% |
| Major change, affects management or diagnosis | 30.0% |
What predicts a change in clinical management?
Proportional-odds ordinal logistic regression. Odds ratios above 1 indicate a higher likelihood of altered patient management; below 1 is protective. Log scale.
Image quality was not independently predictive once perceived appearance and diagnostic complexity were accounted for. A 95% confidence interval is drawn only for reader work experience, the one predictor for which the manuscript reports it (0.38β0.77).
Data table
| Predictor | Odds ratio | 95% CI | p |
|---|---|---|---|
| Perceived synthetic appearance | 3.60 | n/a | < 0.001 |
| Diagnostic complexity | 2.14 | n/a | < 0.001 |
| Image quality | 0.86 | n/a | 0.51 |
| Reader work experience | 0.54 | 0.38β0.77 | < 0.001 |
How do uncertainty maps shift reader confidence?
Share of evaluations in which the per-pixel uncertainty map raised or lowered diagnostic confidence, by case complexity. The balance in each row is "no change".
The map was informative in 49% of evaluations overall, peaking at 55% on the most complex cases and 67% on non-mass enhancement lesions. Its direction inverts with complexity (Kendall's τ = 0.218, p = 0.003). Downstream, this behaved as a safety mechanism: when it lowered confidence readers requested true contrast in 88% of instances; when it raised confidence, 82% proceeded without real contrast. Overall, readers would avoid or defer physical contrast injection in 64% of cases.
Caveat. Inter-reader agreement on the map's effect was low (Fleiss' κ = 0.02β0.10), and readers flagged cases where the map showed high confidence despite a clinically inaccurate synthetic image. It is a useful adjunct in aggregate, not a fail-safe.
Data table
| Case complexity | Confidence increased | Confidence decreased | Informative rate |
|---|---|---|---|
| All cases | 32% | 17% | 49% |
| Low complexity | 39% | 11% | 50% |
| High complexity | 14% | 41% | 55% |
Data Availability
- MAMA-MIA: 1,506 pre-treatment T1-weighted breast DCE-MRI studies with expert segmentations and harmonized acquisition times. Synapse ↗ (where geo-restricted: Health-RI XNAT ↗)
- Duke-Breast-Cancer-MRI: 922 patients, GE and Siemens scanners at 1.5T and 3T. TCIA ↗
- Karolinska Institutet cohort: 192 patients used for external validation. Private; the authors do not have permission to release it.
BibTeX
@misc{joshi2026densetemporalcontrastsynthesis,
title={Dense Temporal Contrast Synthesis via Conditioned Latent Transport},
author={Smriti Joshi and Apostolia Tsirikoglou and Daniel M. Lang and Richard Osuala and Noah MΓ‘rquez Varaa and Alejandro Guzman and Grzegorz Skorupko and Sebastian Ibarra Arregui and Lidia Garrucho and Akane Ohashi and Dimitra Ntoula and Eugen Divjak and OΔuz LafcΔ± and Jan C. Peeken and Julia A. Schnabel and Fredrik Strand and Oliver Diaz and Karim Lekadir},
year={2026},
eprint={2607.29394},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2607.29394},
}