Costing aspirin
One recorded run of a pharmaceutical costing engine, start to finish. A CAS number goes in; a quotable Excel sheet comes out. Every number below is the number the engine actually produced — including the one it labels as a placeholder, and the seven gaps it wrote into its own audit log.
The job
A contract manufacturer is asked to quote a molecule. Someone has to turn a synthetic route into a bill of materials, price every line, add the development and plant effort, and return a number per kilogram. Done by hand it takes days, and the arithmetic is spread across a spreadsheet nobody re-derives.
SciQ automates the business half of that, not the chemistry. It does not invent routes and it does not price by judgment. It takes a route, expands it into charges on a fixed template, and records where each figure came from. The design constraint that matters is the last one: a costing tool that cannot be audited is worth less than the chemist it replaced, because the chemist could at least be asked.
This page is a frozen run. Nothing on it computes; it replays a run recorded on 1 June 2026, and the two Excel files it produced are linked at the bottom so the formulas can be opened and checked.
Fifteen routes, and why a person still chooses
The first call is retrosynthesis: hand IBM RXN a target and get back candidate routes with confidence scores. For aspirin it returned fifteen. All fifteen scored above 0.995 and every one was tagged high-confidence; ten of them sit above 0.999.
That is not a ranking. The entire spread from best to worst is 0.0042, which is well inside the noise of a model scoring its own suggestions. A product that presented these as first, second and third choice would be inventing a distinction the numbers do not support.
Reading what the routes actually contain explains why. All fifteen are the same reaction — acetylation of salicylic acid — differing only in the additive:
| Acetyl donor | Additive that distinguishes the route | Routes |
|---|---|---|
| Acetic anhydride | pyridine, sulfuric acid, sodium hydroxide, sodium acetate, acetic acid, HCl, water, or nothing | 13 |
| Acetyl chloride | none | 1 |
[Ac]O[Ac] | none | 1 |
The last row is the interesting one. In SMILES, a bracketed Ac is
the element actinium, so [Ac]O[Ac] parses as an actinium oxide, not as
acetic anhydride. The model emitted a route whose starting material does not exist
in any catalogue, and scored it 0.9958 — a confidence
indistinguishable from the fourteen valid ones.
The run that was costed
The costed run took a different route from a separate retrosynthesis: a one-step oxidation of an aldehyde to the acid, from 2-acetoxybenzaldehyde, scored 0.964. From there everything is deterministic — no model touches a number again.
| Field | Value |
|---|---|
| Target | Aspirin · CAS 50-78-2 · MW 180.16 · PubChem CID 2244 |
| Route | Aldehyde to acid oxidation · 1 step · confidence 0.964 |
| Basis | 100 kg, contingency 0.2 → 120 kg charged through the stage |
| Procedure | add · add · add (dropwise) · quench · add · extract · wash · wash · dry |
| RXN synthesis | 6a1d4b487f6b9b3ec2b3da83 |
What actually gets charged
Nine materials. Three are costed on a mass basis, from molecular weight and equivalents. Six are costed on a volume basis, from density and a litres-per-kilogram multiple — because liquids are bought by volume, and a cost sheet that converts them to moles is answering a question nobody asked.
The chart makes the shape of a real process visible: 7,167 kg of material is charged to produce 120 kg of product — sixty kilograms in for every kilogram out, almost all of it solvent, water and brine. The substrate is the eighth-largest line on the sheet. Any costing tool that models only the reactants is modelling the smallest part of the bill.
| Material | CAS | MW | Density | Basis | Charge (kg) |
|---|---|---|---|---|---|
| 2-acetoxybenzaldehyde | 5663-67-2 | 164.16 | — | mass · 1.0 eq | 109.34 |
| Jones reagent | — | 256.15 | — | mass · 1.0 eq | 170.62 |
| Na2SO4 | 7757-82-6 | 142.04 | — | mass · 1.0 eq | 94.61 |
| Acetone | 67-64-1 | 58.08 | 0.7845 | volume · 10 L/kg | 941.40 |
| Isopropanol | 67-63-0 | 60.10 | 0.785 | volume · 10 L/kg | 942.00 |
| Water (reaction) | 7732-18-5 | 18.02 | 0.998 | volume · 10 L/kg | 1,197.60 |
| Ethyl acetate | 141-78-6 | 88.11 | 0.895 | volume · 10 L/kg | 1,074.00 |
| Water (wash) | 7732-18-5 | 18.02 | 0.998 | volume · 10 L/kg | 1,197.60 |
| Brine | — | 76.46 | 1.2 | volume · 10 L/kg | 1,440.00 |
Two formulas, written into the sheet
The engine does not compute quantities and paste values. It writes Excel formulas, so the sheet stays live in the customer's hands and every figure can be traced by clicking the cell. There are two, one per costing basis:
mass basis J = ($J$28 / $D$28) * D<row> * H<row> volume basis J = $J$28 * E<row> * H<row> $J$28 stage charge, kg 120 = 100 kg target + 20% contingency $D$28 product MW 180.16 D material MW 164.16 (2-acetoxybenzaldehyde) E density (volume rows only) H equivalents, or L/kg 1.0
Worked, for the substrate: (120 ÷ 180.16) × 164.16 × 1.0 = 109.34 kg. Six more formulas cover recovery, tax, line total, contribution and the roll-ups. Tax is 10% — the rate every historical sheet in the client corpus uses; earlier generic code had defaulted to 18%, which would have been wrong for these customers by a wide margin.
One thing the single-step case hides: the assumed molar yield (0.9 here) is recorded on the stage row but does not enter any quantity in this run. Yield propagates between stages, through a back-calculation from the stage above. With one stage there is nothing above it, so the target quantity enters directly and the yield is inert. It would bite on a four-stage route.
Where every number came from
This is the part that matters, and the reason the sheet is worth more than the number at the bottom of it. Each line carries the source it was resolved from, and each price will be routed to a source tier when the pricing layer is connected: a recent quote from this client's own history first, a catalogue price second, a bulk commodity price last.
| Material | Identity (CAS, MW) | Density | Price routes to |
|---|---|---|---|
| 2-acetoxybenzaldehyde | PubChem, by SMILES | n/a — mass basis | Client quote history |
| Jones reagent | PubChem, by SMILES (no CAS) | n/a — mass basis | Catalogue |
| Na2SO4 | PubChem, by name | n/a — mass basis | Catalogue |
| Brine | PubChem, by name (no CAS) | solvent table | Catalogue |
| Acetone | PubChem, by SMILES | solvent table | Bulk commodity |
| Isopropanol | PubChem, by name | solvent table | Bulk commodity |
| Ethyl acetate | PubChem, by name | solvent table | Bulk commodity |
| Water ×2 | PubChem, by name | solvent table | Bulk commodity |
Densities deserve a note. The retrosynthesis returned
none — it gives amounts in millilitres and leaves the conversion to you. The six
volume-costed lines were filled from a fixed table of solvent densities lifted
from the customer's own template, and written into the sheet as numbers rather than
as the template's original VLOOKUP. That is a deliberate trade: a
lookup that misses a name returns #N/A and the sheet arrives at the
customer broken. A resolved number cannot fail on their machine, and the audit log
carries the record of where it came from.
The gap ledger
Seven gaps were recorded during the run. Six were closed; one was not. Exactly one cell in the entire spreadsheet is shaded red — H15, the equivalents figure the engine had to assume.
| Gap | Materials | Resolution | Status |
|---|---|---|---|
| No density supplied | acetone, isopropanol, water ×2, ethyl acetate, brine | Filled from the canonical solvent-density table | Closed |
| No usable equivalent supplied | Na2SO4 | Written as 1.0, cell H15 shaded red, remark
EQ_ASSUMED_1.0 |
Open |
The drying agent is a real gap, not a rounding concern: at 1.0 equivalents it is charged at 94.61 kg, and the true figure depends on how wet the organic layer is — a number the model has no basis to produce. So the engine does not produce one. It writes the assumption, marks the cell, and puts the line in the audit log where a chemist will see it before the quote leaves the building.
Never let an assumed 1.0 pass as a real value.
That rule is the whole thesis in one line. The alternative — silently filling the number — produces a sheet that looks complete and is quietly wrong, which is the failure mode that makes chemists distrust automation in the first place.
From cost to price
The second sheet converts raw material cost into a quotable price: development and plant effort priced as full-time-equivalent weeks, then overhead and margin, then the fixed adders a plant actually bills for.
| Effort | Weeks | FTE | $/FTE-wk | USD |
|---|---|---|---|---|
| Process R&D | 5 | 2.0 | 750 | 7,500 |
| Analytical R&D | 8 | 0.4 | 750 | 2,400 |
| Process engineering | 8 | 0.2 | 750 | 1,200 |
| Project management | 4 | 0.5 | 250 | 500 |
| Manufacturing support | 25 | 1.0 | 750 | 18,750 |
| Manufacturing cost | 30,350 |
Price is cost ÷ (1 − overhead − margin), at 15% and 15%, which grosses 115,350 up to 164,785.71. Adders — capex, effluent treatment, analytical columns and shipment — bring it to 171,785.71, and the sheet rounds that up to the nearest thousand: $172,000, or $1,720 per kilogram.
How it was checked
The engine writes formulas, so the obvious failure is a sheet full of
#REF! that nobody opens until the customer does. The verification step
is to make something other than the engine do the arithmetic: each generated file
was recalculated headlessly in LibreOffice and read back.
Every cell resolved, no formula errors, and the recomputed values matched the engine's own — 109.342806394316 kg of substrate against a predicted 109.34, and $1,720 per kilogram. One discrepancy did surface and is worth naming: the engine's exported JSON records $1,717.86, because it omits the sheet's round-up. The Excel file is the source of truth by design, so $1,720 is the number — but the two artifacts disagreeing at all is exactly the kind of drift that only appears when you recompute rather than trust.
What this run does not show
One stage, one molecule, a public one. Aspirin was chosen because it is in the public domain and no client route is exposed by publishing it; the same engine reads four- and five-stage sheets, which are where yield propagation and stage roll-ups actually get tested.
Pricing is not connected. Sales-offer generation is not shown. And the architecture has since moved on: relying on a model to predict routes was demoted in June 2026 in favour of retrieving real, patent-cited reactions from a corpus — a decision this run helped make, since fifteen indistinguishable confidence scores and one actinium-based starting material are a poor foundation for a costing tool. The deterministic half below the route — the part on this page — is unchanged.
The artifacts
Both spreadsheets are the files the run produced, formulas intact. The JSON is the run record the two figures above are drawn from.