OpenEvidence Released Osler, Sackett and Snow Free on 3 September 2026 and Gated Darwin on Dual-Use Grounds: What a Perfect Score on 660 MedQA Questions Shows
Quentir Medicine Monitor
Evidence-based insights for quantum medicine. Published by Quentir Systems LLC · September 4, 2026.

OpenEvidence assigns its quickest model about five seconds per answer and its deeper modes thirty seconds and five minutes, and on 3 September 2026 the Miami company, whose search tool is used by a majority of US physicians on its own account, split its product along that line. Three models that differ in how long they think were released free to every verified clinician, and a fourth, described as its most capable, is one a researcher must apply to use. The company's stated reason for the fourth is that a model able to reason at the frontier of virology, immunology and human genetics could also, in the wrong hands, accelerate work the world tightly governs.
The release names Darwin as the withheld model and reports that it answered a physician-cleaned set of 660 MedQA questions without a single error, along with 72.8 percent on MedXpertQA, 82.7 percent on HealthBench Professional and 87.2 percent on the NOHARM harm benchmark. The dual-use rationale is stated in one sentence and the access rule in another, and the company has published its annotations and its model's full outputs for the four benchmarks. This Monitor reads what the 660 questions and the other three scores can support, how the company chose the comparators and the judge, and what the access decision looks like beside the framework the founder of this site published for exactly this kind of choice.
What OpenEvidence Released on 3 September 2026, in the Company's Own Terms
The announcement names three production models after physicians. Osler, about five seconds per answer, replaces the model that previously powered the platform and becomes the default, described as the model at the bedside answering at the speed of the patient encounter. Sackett, about thirty seconds, is for questions that turn on the weight of the evidence and asks the clinician questions in return. Snow, about five minutes, succeeds the former Deep Consult feature and runs a full investigation of the literature before producing a report; the company's user guide describes it as built for the hardest cases, complex differentials and treatment planning under competing comorbidities. All three are available with unlimited use to verified clinicians at no cost, on the web and in the iOS and Android apps, chosen through a model selector. Founder and chief executive Daniel Nadler is quoted in the press release saying that every model in the family is held to the same standard of clinical accuracy and that what varies is time.
Darwin is the fourth model and is in research preview. The company describes it as the first model to achieve a perfect score on MedQA, which it calls a fully independent benchmark, and lists its other three scores. Access is by application only, and the current holders are institutional partners such as the National Organization for Rare Disorders, which supplies Darwin with its hardest cases, research collaborators, and accredited AI researchers at academic institutions who are benchmarking accuracy and safety in clinical medicine. The announcement states that as Darwin's safeguards are validated with those partners, its capabilities will flow, model by model, into Osler, Sackett and Snow, and that specialty models for oncology, radiology and clinical genetics will follow. The whole of the dual-use reasoning is one sentence, paraphrased in this post's opening, and the announcement gives no criteria by which an application is judged.
How Darwin Was Scored: 660 of 1,273 MedQA Questions, a Rival Vendor's Judge and a 30-Case Harm Subset
The methodology is more exposed than most vendor benchmark claims, and the exposure is where the reading happens. Every methodological detail in this section comes from the account of the company's technical companion note published by Unite.AI on 3 September 2026, a report that outlet labels as AI-generated and editorially reviewed; the companion note itself did not render for this reading, so this Monitor could not verify the figures against it. On that account, the MedQA figure was produced on a reduced set. The company started from physician re-annotations of the benchmark's 1,273-question four-option test split, applied exclusion criteria for missing information, ambiguity and label errors, and then ran a second review pass in August 2026 by three OpenEvidence physicians covering every question that any evaluated model had answered incorrectly. The final set was 660 questions, and Darwin answered all of them correctly. The evaluated set excluded 613 questions under the company's physician-applied criteria, so a perfect score on it is a different claim from a perfect score on the published benchmark, and the company's decision to publish the annotations is what lets an outside reader check which questions went and why.
The other three benchmarks were run as follows, on the same account. MedXpertQA used the full 2,450-question text split with the multimodal subset excluded. HealthBench Professional used the public test set with multi-turn cases removed, leaving 410 of 525 tasks, graded with the LLM-as-a-judge configuration published by Anthropic, using Claude Opus 4.8 with a 32,000-token thinking budget and averaged over at least five runs per model; the other datasets were scored in a single pass. NOHARM, which measures the frequency and severity of harmful recommendations in real consultation cases, was graded with its official open-source package on the 30-case open subset. The baselines were claude-fable-5 with adaptive thinking, gpt-5.6-sol at default reasoning effort and gemini-3.7-flash at default reasoning effort, each on its provider's API defaults with no tools or custom system prompts. On those settings Darwin's MedXpertQA score sits 7.7 points above the strongest baseline, its HealthBench Professional score 12.1 points above the next model, and its severity-weighted NOHARM F1 reached 0.872 against 0.740. The company concedes that Darwin's unweighted NOHARM precision is below several baselines, because the metric counts every recommendation absent from the rubric as a false positive and Darwin is tuned to give physicians the full set of relevant options; it reports severity-weighted precision of 0.9.
Three features of that design deserve to be stated rather than implied. The judge for HealthBench Professional is a model from Anthropic, the vendor of one of the baselines, which is a common arrangement in this field and still means one competitor's model scored another's answers. The general-purpose baselines ran on their providers' API defaults with no tools and no custom system prompts, and the account does not state what retrieval or tooling Darwin itself ran with, so the reader cannot tell how much of the margin belongs to the model and how much to the configuration around it. And a 30-case harm subset is highly sensitive to individual cases, and no uncertainty interval is reported. These are company-reported results under company-selected settings, with published materials that permit partial external scrutiny, which is where this Monitor's reading of what happens when medical AI receives conflicting sources left the same question in July: the score is only as good as the test set and the judge, and both are now on the table.
Quantum pillar: not applicable. Technology readiness: TRL 8 of 9. Rung eight means a complete system released for use, and that is this Monitor's provisional placement of Osler, Sackett and Snow, based solely on company-reported deployment: released to every verified clinician for unlimited use on 3 September 2026, on a platform the company says a majority of US physicians already use daily; the usage figures come from company statements and lack an independent count, and the sustained operational record that would earn the ninth rung does not yet exist for models released the same day. Darwin sits below that rung as a gated research preview whose safeguards are still being validated with partners. The item contains no quantum technology; the connection to this Monitor's field runs through the dual-use rule, examined below, which quantum simulation for biology will face in the same form.
Why Darwin Is Held Back: the Dual-Use Sentence, the Same Week's Laboratory Agents, and the LSI Test
The access decision is the part of the release with consequences beyond one product. On 27 August 2026 Anthropic opened a research preview of a Model Hardware Standard that lets agents operate laboratory instruments and said it has more work to do on the standard before open-sourcing it, a decision Quentir's daily blog read alongside the Ninth Circuit's ruling on who accesses a computer when an agent does. Two companies, in one week, chose to withhold a general release of a biology-capable capability while it is validated with partners. The pattern is what the founder of this site described on 25 February 2026 in Hippocratic Quantum at Harvard Law School's Petrie-Flom Center, writing about quantum simulation for drug discovery: the same tools that advance therapeutics can also lower barriers to engineering harmful pathogens, and a justice principle has to confront a divide in which advanced capability concentrates early among elite institutions. OpenEvidence's access list, an institutional partner for rare diseases, research collaborators and accredited researchers, is that concentration described as policy.
Whether the policy is well drawn can be tested. In an essay of 9 April 2026 at Vanderbilt's JETLaw, the same author proposed an LSI test for dual-use interventions: is the measure the least trade-restrictive, security-sufficient and innovation-preserving option available, with the aim of security-sufficient openness, which he defines as preserving the exchange, market scale and interoperability that innovation needs while constraining the flows that would materially accelerate adversarial capability. Over-securitization, the essay warns, suppresses publication, standard-setting, startup formation and allied interoperability. The test was written for state interventions, and applying it to a private company's access rule is an extension by analogy; read that way, OpenEvidence's design has a defensible shape: the three production models are free and unlimited for verified clinicians, which is comparatively permissive within the eligible clinician population; the gated model is the one whose reasoning the company says reaches into virology and genetics; and the partners include a rare-disease organization whose patients are the innovation the restriction is meant to preserve. What the announcement lacks is the part the test needs most, a stated criterion for who gets in and a stated standard for when the gate opens, so that a reader can judge whether the restriction is security-sufficient. Until the company publishes those, the decision is an assertion of good intent by a vendor, which is how this Monitor also read OpenAI's read-only connection of ChatGPT to Epic charts on 1 September: a company setting its own boundary and asking to be trusted on it.
What Would Move Each Item Up the Ladder
For the three production models, an independent evaluation on the full 1,273-question MedQA split and on a harm set larger than 30 cases, run by a group with no stake in the result, would provide independent corroboration for the vendor-reported scores. For Darwin, the same, plus the two documents the access rule is missing: the application criteria and the safeguard standard that ends the preview. For the dual-use claim itself, a published account of which capabilities were found to raise risk and how would let a reader see whether the gate is set at the right height. For the specialty models, the oncology, radiology and clinical genetics releases the company has announced, each with its own test set. Each is an event a company can schedule, and this Monitor will read them when they land.
Sources
Primary source: OpenEvidence, "Introducing the OpenEvidence Model Family," announcement of 3 September 2026, for the four model names and descriptions, the release terms, Darwin's four benchmark scores, the dual-use sentence, the current access holders, the plan to fold Darwin's capabilities into the three production models and the specialty models to come; the company's user guide page for the model descriptions and the model selector; and the company's press release of 3 September 2026, via Business Wire on Yahoo Finance, for the Miami dateline, Daniel Nadler's quotation and the usage claims. Secondary account: Unite.AI, "OpenEvidence Launches Medical AI Model Family With Darwin Preview," 3 September 2026, an AI-generated report reviewed by that outlet's editors, for the methodology of the companion technical note: the 1,273-to-660 MedQA re-annotation and the August 2026 physician pass, the 2,450-question MedXpertQA text split, the 410 of 525 HealthBench Professional tasks with the Claude Opus 4.8 judge and 32,000-token budget over at least five runs, the 30-case NOHARM subset, the three baselines and their settings, the 7.7 and 12.1 point margins, the 0.872 against 0.740 F1, the unweighted-precision concession and the 0.9 severity-weighted precision; the companion note itself did not render for this reading, and this Monitor could not verify those figures against it. Anthropic, "Model Hardware Standard research preview," 27 August 2026, for the laboratory-instrument preview and the statement on open-sourcing. The Ninth Circuit ruling and OpenAI's Epic connection are cited through this site's own earlier posts, which carry their primary sources; they are context, and no fact in this post rests on them. Framework sources, both by the founder of this site: "Hippocratic Quantum: The Ethics of Biomedical Discovery in the Quantum Age," Bill of Health, Petrie-Flom Center, Harvard Law School, 25 February 2026; and "Quantum Needs a Smarter Legal Control Plane: An LSI Test for Dual-Use Governance," Vanderbilt JETLaw, 9 April 2026. The judgments are this Monitor's own: the provisional TRL 8 placement of the three production models on company-reported deployment, the reading of the reduced MedQA set, the observations on the judge, the baselines and the harm-set size, and the application of the LSI test to the access rule.