AQCat Became a Claude Science Tool on August 19, 2026: The Evidence Behind SandboxAQ's Catalyst Model and the Contract Questions It Raises
In 1909 Fritz Haber demonstrated the synthesis of ammonia from its elements in a Karlsruhe laboratory, the work the Nobel Foundation's biography records as the basis of his 1918 prize. Between that demonstration and an industrial process stood the search for a catalyst that could be manufactured and run at scale, and the search for catalysts has stayed a bench job since. SandboxAQ puts the throughput of conventional laboratory methods at fewer than a hundred candidate materials a week.
On August 19, 2026 the company announced from Palo Alto that AQCat, its machine-learning model for predicting how candidate catalysts behave, is generally available on Claude Science through the Model Context Protocol. A researcher asks for a screening in plain language and receives the model's prediction as a tool result; the company's head of catalysis science says the point is that any researcher can run that calculation at scale without touching a line of code. On August 25 a second release announced AQCat in the AWS Marketplace, deployed through Amazon SageMaker inside the customer's own AWS account. This site reads announcements for what they document, and here the record is fuller than it first looks.
Practical takeaway. A scientific model reached through a tool interface is bought on three documents: the contract that says where the model runs, what crosses the boundary and what the vendor keeps; the logging requirement that preserves model version, query and answer for the life of the program; and the organization's own validation set that turns the vendor's accuracy figure into the organization's evidence. All three can be in place before an independent benchmark exists.
What the August releases document
Three things are stated by the company and can be followed to a source. First, availability on two channels: the Claude Science listing over the Model Context Protocol, announced August 19, and the AWS Marketplace listing as a SageMaker model package, announced August 25 and priced by inference host-hour and per request. Second, the training corpus: AQCat25, 13.5 million spin-polarized density-functional-theory calculations across roughly 47,000 intermediate-catalyst systems, released on Hugging Face in September 2025 under a Creative Commons non-commercial license, with model checkpoints and code, behind a terms-acceptance gate. Third, the science: a vendor-authored paper on AQCat25, its spin-aware models and their evaluations, published in the peer-reviewed journal npj Computational Materials. That is a stronger evidentiary base than most model announcements carry.
What the record still lacks
Two statements are performance claims. Accuracy approaching the field's most trusted methods and speed of up to 20,000 times faster than first-principles simulation rest on the company's own evaluations, run on the company's own data. The peer-reviewed paper is authored by the vendor; the releases cite no independent evaluation of the production service and no customer result a reader can reproduce; the two quoted endorsements, from an academic user at Texas Tech University and from a former founding director of DARPA's biotechnology office, describe capability in general terms. The precise distinction is the useful one: a vendor-authored peer-reviewed evaluation exists, and independent external validation of the production service and of the 20,000-times figure remains absent. Quentir's registers carry that distinction as a column, computed to date, by whom, against what. For AQCat the column reads: corpus public, vendor evaluation peer-reviewed, independent validation pending.
Two delivery routes, two data answers
The routes differ in the one place a research director cares about most. The AWS Marketplace listing states that the model deploys through SageMaker inside the customer's isolated AWS tenant, and that input data, inference payloads and chemical structures are never exposed to external networks, never shared with SandboxAQ and never used for retraining. The Claude Science route runs the model behind a tool interface that the assistant calls; where the model executes and who operates the endpoint on that route is a question for the vendor, and the answer belongs in the contract. Claude Science, for its part, states that every output carries an auditable history of how it was made, including the code, environment and message history behind a result; that is a product feature to cite, and a logging requirement to write down, since the Model Context Protocol itself specifies tool-call requests and results and leaves persistence to the client.
The diligence questions before the first screening run
Where does the model run, and who operates the endpoint, on each route the organization intends to use? Can the organization inspect the weights, or only the outputs? What does a catalyst query carry across the boundary, what does the vendor retain, for how long, and is any of it used for training? Is the model version recorded with each query and answer, and does the retention match the life of the research program? A catalyst query can carry the structures an organization is considering, which may be its most sensitive research intent, so the answers belong in the contract before the first run, on either route.
What the accuracy figure means for one organization's candidates
A predicted adsorption energy or an activity ranking is a model output with an error distribution. The vendor's accuracy statement describes the general case; the organization's use case is a specific one. An in-house validation on candidates whose behavior is already known converts the vendor's statement into the organization's own evidence, and it produces the one artifact a later reviewer will ask for: the comparison, on this organization's chemistry, on a stated date.
How Quentir Reads It
AQCat arrives with a public corpus, a vendor-authored peer-reviewed paper and two delivery routes whose data terms differ. The performance figures are the vendor's until an independent benchmark exists, and an honest register entry says so in one line. The checklist for a research director before the first query leaves the building: the route and its endpoint operator; weight access; data transit, retention and training use; per-query logging of version, query and answer; and the in-house validation set with its date. The Quentir Radar tracks model and vendor announcements of this kind as they land, with the evidence beside each entry.
Sources: SandboxAQ, "SandboxAQ Makes AQCat Generally Available on Claude for High-Throughput Catalyst Screening", PR Newswire, Palo Alto, August 19, 2026, a company release, for the Claude Science availability, the AQCat25 figures, the accuracy and speed characterizations, the fewer-than-a-hundred-per-week bench figure, the AQPotency announcement and the two quoted endorsements. SandboxAQ, "SandboxAQ Accelerates Fortune 500 Catalyst Discovery with AQCat, Now Generally Available in AWS Marketplace", PR Newswire, Palo Alto, August 25, 2026, a company release. AWS Marketplace, AQCat, AI Models for Catalyst Discovery, the listing, for the SageMaker deployment in the customer's isolated AWS tenant, the pricing model and the statement that input data, inference payloads and chemical structures are never shared with SandboxAQ or used for retraining. SandboxAQ and coauthors, "AQCat25: unlocking spin-aware, high-fidelity machine learning potentials for heterogeneous catalysis", npj Computational Materials, 2026, a vendor-authored peer-reviewed paper. SandboxAQ, AQCat25 dataset, Hugging Face, released September 10, 2025, CC BY-NC-SA 4.0, terms-gated, with model checkpoints at SandboxAQ/aqcat25-ev2. Anthropic, "Claude Science", June 30, 2026, for the auditable output history and the MCP connections. Model Context Protocol, logging specification, for what the protocol itself specifies. The Nobel Foundation, Fritz Haber, biographical, for the 1909 laboratory synthesis. The judgment that a vendor-authored peer-reviewed evaluation exists while independent validation of the production service is absent, and the diligence checklist, are Quentir's own reading.
Published intelligence, built to inform your own decisions. Published: September 3, 2026.