Florence and Barcelona Researchers Run a Quantum Diffusion Model's Noise Step on IBM's 133-Qubit Torino for Blood-Cell, Brain MRI and Rib CT Images: What arXiv 2609.31070 Shows
Quentir Medicine Monitor
Evidence-based insights for quantum medicine. Published by Quentir Systems LLC · September 28, 2026.

Generative AI can now draw convincing medical images: blood-cell smears, brain scans, slices of a CT volume. Hospitals and research groups want such synthetic images to enlarge small training sets, to share data without exposing patients, and to test diagnostic software on rare cases. The most successful generators today are diffusion models, which learn to turn random noise back into a plausible image. A preprint posted on 25 September 2026 asks whether a real quantum computer can take over one half of that process.
The paper, "Quantum Diffusion Models for Medical Image Analysis" (arXiv 2609.31070), comes from Francesco Aldo Venturelli and Miguel A. González Ballester of Universitat Pompeu Fabra and the Barcelona Supercomputing Center, Stefano Martina, Marco Parigi and Filippo Caruso of the University of Florence, and Alba Cervera-Lierta of the Barcelona Supercomputing Center. They built a hybrid quantum diffusion model in which the step that adds noise to an image runs on IBM's 133-qubit ibm_torino processor, while a classical neural network learns to remove that noise again. They then used it to generate synthetic medical images from three public collections: blood-cell photographs, brain MRI slices and 3D rib CT volumes.
The authors report that their hybrid model is competitive with a classical diffusion model of the same design, and better on some measures, especially for the 3D rib volumes. The comparison is limited: 100 generated samples per dataset, small images, and a classical baseline the team reproduced itself. No radiologist judged the images, and no diagnostic task was run on them.
That leaves a narrow but real result. It is one of the few medical-imaging studies in quantum machine learning that runs part of its pipeline on physical quantum hardware with public medical-image datasets, and it deserves a precise reading of what the hardware did and what the numbers show.
How the Florence and Barcelona team split a diffusion model between IBM hardware and a classical network
A diffusion model has two halves. The forward process takes a real image and corrupts it step by step until only noise is left. The backward process is a neural network trained to reverse each step, so that after training it can start from pure noise and produce a new image. In classical practice both halves run on graphics processors.
Here the forward half is a discrete-time quantum walk, a quantum version of a random walk. Each possible brightness value of a pixel is a position on a ring, and the walk spreads the pixel's value around that ring over time. On ibm_torino the team used a small group of connected qubits, one of which acts as the "coin" that decides the direction of each step. For the blood-cell images, six qubits sufficed: the coin plus five that encode 32 brightness levels per color channel. For the brain MRI images, with 64 brightness levels, the team used seven: the coin plus six intensity qubits. The paper states that the hardware's own noise is part of the design. The authors write that quantum noise "is a constituent factor that is necessary for the convergence" of the forward process, so they ran without error correction and treated the device's imperfections as part of the diffusion.
The quantum computer does not process the images directly. The team ran the walk on the hardware to measure transition probabilities between brightness levels, then applied those measured probabilities independently to every pixel or voxel of the real images on a classical machine. The backward network was written and trained in PyTorch. This is how the method reaches image sizes that current quantum processors could never hold in their qubits: a few qubits describe the noise process, and classical computing does the rest. The design builds on the group's earlier hardware-based noisy quantum diffusion paper from 2025, and this version raises the number of intensity levels from 8 to 64.
Quantum pillar: computing. Technology readiness: TRL 3 of 9. The method has been shown as a proof of concept on public research image collections, with part of the computation on a real IBM quantum processor, and has not been tested on any clinical task, by any radiologist, or in any hospital workflow.
What the model generated from BloodMNIST, BraTS2020 brain MRI and FractureMNIST3D rib CT
The three test sets come from established public benchmarks. From MedMNIST, a lightweight benchmark collection for biomedical image analysis, the team took one class of BloodMNIST, color photographs of blood cells at 64 by 64 pixels, and FractureMNIST3D, 1,370 rib-fracture CT volumes of 28 by 28 by 28 voxels. From the BraTS2020 brain tumor segmentation challenge they extracted 484 single-slice grayscale brain MRI images cropped to 190 by 190 pixels with 64 brightness levels.
The classical comparison is a diffusion model on discrete states in the style of Austin and colleagues' 2021 structured denoising diffusion models, where a classical random walk replaces the quantum one. The team reports three standard measures on 100 generated samples per dataset: Kullback-Leibler divergence and Fréchet inception distance (FID), where lower is better, and structural similarity (SSIM), where higher is better. The quantum model appears in two noise-schedule settings.
On the 3D rib volumes the quantum model came out ahead on two measures: FID fell from 169.6 to 100.1 or 89.7, and Kullback-Leibler divergence fell from 0.047 to 0.013 or 0.011. SSIM matched the classical 0.564 in one setting and was slightly lower, 0.550, in the other. For these volumes the FID is an average over five 2D slices, so it describes slice quality more than the fidelity of the whole 3D volume. On blood-cell images both quantum settings beat the classical model on all three measures. On brain MRI the picture is mixed: one quantum setting had a lower FID (266.8 against 280.3) and a lower SSIM, and the other had a higher SSIM (0.676 against 0.646) and a worse FID of 297.7. The abstract's word for this is "competitive," and the table supports that word more than any stronger one.
Why 100 samples and FID scores from a general-purpose network limit the evidence
Three limits matter for anyone evaluating the claim. The first is sample size. Metrics computed on 100 generated images carry wide uncertainty, and the paper reports no confidence intervals or repeated runs in its main table. The second is the metric. FID measures distance using a network trained on everyday photographs, and the authors themselves note that it "can be a poor choice" for specialized medical data. That caveat applies in both directions: it may affect how the weaker FID values on brain MRI should be read, and it equally limits the weight of the stronger FID values on rib CT.
The third limit is image size. Sixty-four by 64 blood-cell images with 32 color levels and 28-voxel cubes are far below the resolution of routine clinical imaging. The authors compare their results with a 2025 hybrid quantum-classical latent diffusion study that worked at 1024 by 1024 pixels, and they acknowledge a higher FID than that work. The paper lists a GitHub repository for its code; when this Monitor checked the address on 28 September 2026, it returned a "not found" page, so outside groups cannot yet rerun the experiment from that link.
There is also no quantum advantage claim in the paper, and none should be read into it. The classical baseline is a simple discrete random walk. A stronger classical noise schedule might close the gap, and the paper does not test one. The result shows that a noisy quantum device can supply a usable forward process for real medical image data. It does not show that the quantum device is needed.
Readers who followed this Monitor's report on simulated quantum kernels that lost all 24 matched tests to a default classical SVM on brain MRI and breast ultrasound will recognize the pattern. Quantum machine learning papers on medical images now pass a first bar by running on real data. The next bar is a strong classical baseline, and after that a clinical task.
What synthetic-image generation on quantum hardware would need before a hospital could use it
For a hospital imaging department or a medical AI vendor, synthetic images are useful only when they improve something measurable: a classifier trained on augmented data that detects more fractures, a privacy-preserving data release that passes a re-identification test, or a rare-tumor dataset that radiologists accept as realistic. This study measures none of those outcomes. It measures how closely generated images resemble the originals on statistical scores.
Access to hardware is the other practical question. The forward probabilities came from an IBM cloud processor, and the approach needs only a few qubits, so it could run on many current machines. The European and US programs now funding quantum infrastructure, including the Department of Energy roadmap of 25 September 2026 that targets 50 to 100 or more logical qubits for 2028, aim at much larger error-corrected systems. This paper points the other way: it uses today's noisy devices and turns their noise into a feature. That is a sensible research direction for a group with access to both IBM hardware and the Barcelona Supercomputing Center. For clinical buyers it remains a laboratory method to watch, with no product, regulatory filing or clinical study attached.
Sources
Primary source: Francesco Aldo Venturelli, Stefano Martina, Marco Parigi, Filippo Caruso, Alba Cervera-Lierta and Miguel A. González Ballester (Universitat Pompeu Fabra, Barcelona Supercomputing Center, University of Florence), arXiv preprint 2609.31070, 25 September 2026, not yet peer reviewed. Also drawn on: the group's 2025 hardware-based quantum diffusion preprint; the MedMNIST and BraTS2020 dataset pages; Austin and colleagues' 2021 discrete diffusion paper; and the 2025 hybrid latent diffusion study the authors compare against; the readiness assessment is this Monitor's own.