How the predictions are validated
Four measurements, each asking a different question, all run against data the engine has never been tuned on. The numbers below are the measured ones rather than the flattering ones, and where a figure is easy to misread it is stated with what it measures attached.
The measurements
56.5%
top-1, out of sample
On 394 reactions the engine has never been trained or tuned on, the product that was actually run and published is our first-ranked candidate 56.5% of the time; it is in the top five 59.3% of the time. This is a ranking measure on a test where several proposals are often chemically reasonable, so it understates how often the answer is useful and it is the honest headline anyway.
54 / 54
correctness audit
Reactions where a chemist can mark the answer right or wrong — SN2 stereochemistry, where a group goes on a ring, whether a reaction needs a reagent that is not present. This asks the question recall cannot: when the platform answers, is the answer wrong? All 54 currently pass.
32 / 32
blind set
Reactions written to catch families the engine must not confuse, each added after a failure rather than before. Every one of them failed when it was written; all 32 pass now.
43.5%
route planning, exact first
On 347 recorded reactions the engine has never seen, the recorded starting materials are the first proposal 43.5% of the time and appear at all 57.3% of the time. The right bond is identified first 44.7% of the time. Coverage is the limit here rather than ranking: where the recorded answer is proposed at all, it is ranked first in 96% of cases.
What 56.5% does not mean
It is not "right 56.5% of the time". It is how often the one product that happened to be run and written down is our first candidate, on reactions the engine has never seen. A proposal ranked second is frequently a reaction that would also work; the benchmark has no way to credit it, because the only thing it knows is what somebody chose to publish.
That is why the audits matter more for judging the tool. Recall asks whether the right answer is found; the correctness audit asks whether the answers given are wrong, which is the failure that actually costs somebody an afternoon. Both audits pass completely.
What these numbers do not cover
- Conditions are not taken into account. The platform predicts what reacts, not how temperature, time or equivalents change the outcome, so it cannot tell you a reaction will fail at the temperature you had in mind.
- Coverage is uneven. Predictions are on firmest ground for the most documented reaction types, such as cross-couplings and amide bond formation, and the evidence shown with each candidate is how you tell which kind of ground you are on.
- Directed and sp3 C–H chemistry is not covered. Neither is alkyl–alkyl coupling in route planning.
- No model output is a measurement. A candidate is a proposal with its precedent attached, and the platform will not report one as more than that.
When the numbers change
Numbers move as the engine changes. Every figure here is the current measurement, including when one has gone down.