OpenEvidence Released Osler, Sackett and Snow Free on 3 September 2026 and Gated Darwin on Dual-Use Grounds: What a Perfect Score on 660 MedQA Questions Shows
Quentir Medicine Monitor
Evidence-based insights for quantum medicine.
OpenEvidence assigns its quickest model about five seconds per answer and its deeper modes thirty seconds and five minutes, and on 3 September 2026 the Miami company, whose search tool is used by a majority of US physicians on its own account, split its product along that line. Three models that differ in how long they think were released free to every verified clinician, and a fourth, described as its most capable, is one a researcher must apply to use. The company's stated reason for the fourth is that a model able to reason at the frontier of virology, immunology and human genetics could also, in the wrong hands, accelerate work the world tightly governs.
The release names Darwin as the withheld model and reports that it answered a physician-cleaned set of 660 MedQA questions without a single error, along with 72.8 percent on MedXpertQA, 82.7 percent on HealthBench Professional and 87.2 percent on the NOHARM harm benchmark. The dual-use rationale is stated in one sentence and the access rule in another, and the company has published its annotations and its model's full outputs for the four benchmarks. This Monitor reads what the 660 questions and the other three scores can support, how the company chose the comparators and the judge, and what the access decision looks like beside the framework the founder of this site published for exactly this kind of choice.
A Surgical Simulator Enters the Real-Time Loop
The loop is now fast enough for interaction
NVIDIA's Cosmos-H-Dreams takes an initial surgical video frame and a live stream of robotic actions, then generates the next scene in 12-frame blocks. The developers report roughly 160 frames per second on one RTX PRO 6000 GPU, up from about 10 frames per second for standard Cosmos-H-Surgical-Simulator inference. That throughput moves the system into real-time surgical simulation: a keyboard, Meta Quest controller, or learned robotic policy can act against the scene while the model keeps generating. The release specializes in tabletop suturing with the da Vinci Research Kit. It uses a 44-dimensional action format, causal attention, a streaming key-value cache, and a student model distilled from a bidirectional surgical-video teacher. This changes the experiment from a completed clip inspected later into a responsive environment that can be interrupted while a rollout is still unfolding.
Fidelity now becomes the demanding test
Interactivity supports closed-loop evaluation, but speed alone cannot show that the simulated robot behaves like the physical robot. The training material includes successful demonstrations along with needle drops, missed throws, unsuccessful knots, and out-of-distribution episodes. That is valuable because a simulator used for policy development must reproduce the consequences of poor actions as well as clean demonstrations. The developers themselves call for benchmarks covering tool-tip reach, pose accuracy, gripper cycles, idle stability, counterfactual actions, long-horizon drift, and agreement between simulated and physical policy outcomes.
The current release is a research and development platform for rehearsal, interactive demonstration, synthetic data, and robotic-policy testing. It does not report clinical performance, diagnostic value, prospective procedure studies, or validated transfer to patient care. Quentir reads it as an infrastructure advance whose medical value will depend on a precise simulation contract: the starting scene, action stream, generated response, time horizon, and physical comparator. The loop is fast enough to challenge in real time. Its fidelity remains open to independent testing.