OpenAI Connected ChatGPT to Epic Charts on 1 September 2026, With Read-Only Access
Quentir Medicine Monitor
Evidence-based insights for quantum medicine.
A clinician preparing for an afternoon appointment spends the first minutes of it reading: the last visit note, the newest labs, the medication list, whatever the specialist sent over. On 1 September 2026 OpenAI said that reading can be handed to ChatGPT, working from the patient's own Epic record.
Two capabilities ship together. Health organizations can connect their Epic electronic health record environments to ChatGPT for Healthcare, and a separate Healthcare Public Data plugin wires the same workspace to nine official public datasets, among them ClinicalTrials.gov, CMS Coverage, RxNorm, DailyMed and PubMed. Access to the chart is read-only. The model writes nothing back into the record.
Two things carry this deployment to the bedside, and a regulatory clearance is neither of them. The first is a business associate agreement, the contract HIPAA requires before an outside vendor may handle protected health information. The second is OpenAI's own physician evaluation, which reports 4,363 ratings across 27 clinical use cases, with 99.1 percent of responses graded safe.
Those figures come from evaluation results OpenAI published itself, and the company says so. The denominator is ratings, and the announcement does not say whether one response could receive several of them, so 99.1 percent of 4,363 leaves approximately 39 non-safe ratings rather than a countable set of unsafe answers. No definition of the safety rating is published: the announcement does not say what the scale was, what those 39 ratings concerned, how severity was graded, or whether reviewers were blinded to the source of the answer.
Where the product lands against FDA's four criteria for clinical decision support software, the fourth of which asks that a clinician can independently review the basis for a recommendation, is an open question that read-only access does not settle on its own, and the announcement reports no determination. The decision support interventions criterion at 45 CFR 170.315(b)(11) reaches interventions a developer supplies with its own health IT module, which leaves the scope question for a hospital to settle with its vendors.
Florida Atlantic's 90.26% Quantum Heart-Disease Result Came From a Simulator: the AI Paper of 21 May 2026 and the FAU Release of 27 August
Quentir Medicine Monitor
Evidence-based insights for quantum medicine.
Florida Atlantic University announced on 27 August 2026 that its engineers had built a quantum machine learning framework for heart disease prediction reaching more than 90 percent accuracy. The figure is precise and it is checkable. It is 90.26 percent, produced by a quantum support vector machine with angle encoding, averaged across five folds of a clinical file holding 918 patients.
The work behind that announcement appeared three months earlier. Muhammad Minoar Hossain, Md. Hasibul Hassan Himal and Arslan Munir published "A Comparative Study of Quantum Feature Maps and Quantum Classifiers for Heart Disease Prediction" in the MDPI journal AI on 21 May 2026, as article 180 of volume 7, under a Creative Commons license that lets anyone read the whole methods section. Section 2.6.2 of that paper states that every quantum experiment ran in a simulation-based environment rather than on a physical quantum processing unit. The university announcement does not carry that sentence, and neither does the coverage that followed it.
That difference decides what the study is evidence for. Simulated qubits establish whether an algorithm has promise in principle; a run on a physical processor establishes whether the machines that exist can deliver it, once noise, limited connectivity and readout error have had their say. Everything else in the paper holds up well under checking. Several of its numbers are more informative than the ones the announcement chose to lead with, and one of them disagrees with the paper's own abstract. The study is a careful piece of comparative work whose careful parts were the first thing lost in transmission.
Atrial Fibrillation Is Where Seoul Will Look for Quantum Advantage
A grant to find out whether quantum computing helps
On August 25, 2026 a Korean consortium was given two and a half years and 2.5 billion won to find out whether a quantum computer can compute cardiovascular blood flow on a schedule a clinic could live with. Seoul St. Mary's Hospital of the Catholic University of Korea, the University of Seoul and the medical software company Flownics were selected for a new 2026 challenge program run by Korea's Ministry of Science and ICT with the National Research Foundation. What the grant buys is a measurement rather than a product: the team has been funded to establish the conditions for quantum gain on one clinical calculation and to report those conditions as numbers, including how many circuits and how many measurements the answer cost.
What a hospital can already see, and what it cannot
A cardiologist reading a CT scan of a narrowed artery can measure the narrowing and cannot measure the flow. Velocity, pressure and wall shear stress are quantities that inform the risk of a clot forming or a muscle being starved, and none of them is in the picture. Computational fluid dynamics recovers them from the anatomy, and cardiology has begun buying it: pressure ratios computed from coronary CT are an established adjunct in stable chest pain assessment and have cleared a national payer's technology assessment in England. The trouble is what happens when a clinician asks for more. A finer mesh, honest nonlinear terms and patient-specific boundary conditions each multiply the arithmetic, and past a certain point the answer arrives too late to be part of a decision.
Atrial fibrillation first, adjudicated by imaging
The first target is atrial fibrillation, the most common sustained arrhythmia, present in roughly two to three percent of the population. In a fibrillating heart the left atrium stops emptying cleanly, blood stagnates in its appendage, and stagnation is where thrombus begins. Whatever the solver computes will be checked against real patients imaged with 4D Flow MRI, which records the speed and direction of the blood itself alongside the geometry of the vessels carrying it. The plan starts on classical hardware with a three-dimensional deep learning model that segments the heart, the left atrium and the aorta from CT, and a quantum hemodynamic model is built on top of that, stacking a variational algorithm, a Krylov subspace method, classical shadow measurement and non-Markovian error mitigation.
Ninety-five percent is a parity target
The headline figure deserves a careful reading. The stated aim is for quantum-based computational fluid dynamics to reach at least ninety-five percent of the precision of the classical method, which is a target to match the incumbent rather than to beat it. Any advantage the project finds will have to appear somewhere other than accuracy: in time to answer, in cost, or in problem sizes the classical solver cannot reach inside a clinical window. Stating it that way is unusually disciplined. The project fixes the accuracy at parity, names the baseline, names the disease, names the imaging modality that will adjudicate, and puts the burden of proof on resources. A hospital procurement officer has something to hold the team to in 2028, and a negative result in that year will be as informative as a positive one.
The Same Kind of AI Explanation Split the Room
One interface, two kinds of reader
A Nature Medicine study tested four forms of AI assistance in dermatology with 623 lay participants and 153 primary care physicians. The two groups completed different diagnostic tasks, yet the pattern across the experiments was clear: help from a strong model could improve average performance while exposing the most deferential users to larger errors when the model was wrong. For non-experts, a fluent LLM rationale carried particular force. They trusted the language whether the diagnosis was right or wrong, and vague or generic accounts could feel especially convincing. The result turns automation bias into an interface problem, not only a user-training problem.
The order of the screens changed the behavior
Participants were also assigned to different sequences. Some formed an initial diagnosis before seeing the model's suggestion; others received AI assistance before making their decision. The human-first workflow preserved more room for independent reasoning, while AI-first presentation produced stronger anchoring. Clinicians were more resilient to incorrect advice and gained the least diagnostic accuracy from LLM explanations, although prose could help their confidence track accuracy. The study therefore separates explanation quality from explanation placement. A sound model and readable rationale can still have a different effect when the model speaks first.
Why this belongs in hospital technology review
FDA, Health Canada, and MHRA principles already say that transparency for machine-learning medical devices depends on the audience, context, media, timing, and communication strategy. This study gives explanation timing a concrete clinical meaning. A hospital buys more than an algorithm: it adopts a sequence in which a nurse, physician, specialist, or patient encounters the output. The experiments do not prove that any named commercial product is unsafe, and they do not cover every care setting. They do show that an average accuracy gain can hide a failure mode concentrated among people with the least independent knowledge. No quantum technology was tested. The lesson will still matter if future quantum-assisted systems make medical models faster or more capable, because computation alone cannot decide when a human should see the answer.
The Sleep Study Still Has More to Say
A whole night becomes a few summary measures
An overnight sleep study records brain waves, eye movements, muscle tone, breathing, oxygen levels, heart rhythm and body position. Clinical interpretation often compresses that dense sleep physiology into familiar measures such as the apnea–hypopnea index. A Cleveland Clinic registry began with 10,000 studies. Preprocessing retained 9,608 for clustering, and 9,203 studies trained the foundation model. It identified five risk groups with different trajectories for mortality, cardiovascular disease and neurological disease. The highest-risk group had more than twice the mortality risk of the lowest group, while conventional apnea severity categories showed limited prognostic value.
The result survived a second cohort
The researchers tested the framework in the independent Sleep Heart Health Study, a population-based cohort collected with lower-resolution data. The model still distinguished higher- and lower-risk patients. That strengthens the result, but the analysis remains retrospective. It links physiological patterns to later outcomes and does not show that giving clinicians a risk group improves referral, monitoring, treatment or survival. The paper calls for replication in additional cohorts and prospective clinical trials before implementation.
The partnership includes quantum; this study does not
The team came together through the Cleveland Clinic–IBM Discovery Accelerator, a ten-year life-sciences partnership covering AI and quantum computing. IBM Research scientists are among the authors, and the Discovery Accelerator and the National Heart, Lung, and Blood Institute supported the research. This paper used classical artificial intelligence. No quantum computer, sensor, simulation or network appears in the method.
The hospital opportunity lies in data already collected. American sleep laboratories perform an estimated one million to four million studies each year. A reliable secondary analysis could extract more value from the same difficult night without adding another test. Adoption would still require cross-hospital validation, interpretable group definitions, clear consent and data-reuse rules, and a defined clinical response. The model has shown that the sleep study contains more prognostic structure than one familiar score preserves. Medicine has not yet shown how that extra message should change care.
What Happens When Medical AI Receives Conflicting Sources?
A direct answer can still rest on unstable ground
When a language model receives conflicting context on a health question, its answer can become less reliable than an answer drawn from its internal knowledge alone. That is the central result of the HealthContradict benchmark: 920 expert-verified instances pairing one health question and factual answer with two long documents that take opposing positions. Across the tested open models, mixed context reduced accuracy. The strongest biomedical model resisted the effect better than its general-purpose counterpart, but it did not escape it.
The source can change the answer
The more revealing comparison is not model against model. It is the same model under different source conditions. Give the strongest open biomedical system in the main comparison the correct document and accuracy rose to 91.1 percent. Give it only the incorrect document and performance fell by 21.6 percentage points from its no-context control. Give it both sides and accuracy still declined. In one smaller biomedical model, changing which document appeared later shifted accuracy by 5.9 points.
These are controlled benchmark results, not a clinical deployment trial. The core model suite ranged from 1 billion to 8 billion parameters, with additional GPT-4.1-mini and GPT-4o evaluations reported by the authors. The study did not test live retrieval pipelines, clinician use, or patient outcomes. Yet the experiment isolates a practical hazard: retrieving relevant material is not enough when the retrieved material disagrees.
A safer evaluation asks whether disagreement survives synthesis
Medical knowledge is not a static answer key. Studies differ by population, design, endpoint, and date; later work can challenge highly cited findings. A useful clinical AI therefore needs more than citation accuracy. It needs tests that preserve source order, provenance, study design, and the fact of disagreement itself.
For buyers and governance teams, the next benchmark should include source-order permutations, deliberately incorrect context, mixed-quality material, and an escalation path for unresolved conflict. The important question is not whether a model can produce one fluent answer. It is whether the model and the workflow around it can show why the source base does not yet support only one.