How much compute
A plant hands over a folder of exports. The compiler reads it, works out what each tag
is, what it is attached to and what is simply wrong, and writes a model. Below is how long
that takes as the plant gets bigger, measured on plants from 300 to 9,511 tags
with every size run in its own process so the memory figure is that size's own.
Log axes, so a straight line is a power law. The measured slope is n^2.232, which is worse than linear and is honest about the duplicate pass: it compares series against each other. At 9,511 tags a whole compile is 449 seconds. The two hollow points are published datasets rather than our own simulator: C-Town water network at 43 tags, HAI testbed at 86 tags, both of which land in well under a tenth of a second. The sweep stops there because building a synthetic plant twice that size needs more memory than the machine has, and a timing taken while a machine is swapping measures the disk.
What the compiler itself asks for, measured with an allocation tracer in a pass of its own because that tracer more than doubles the runtime and would have spoiled the timing above. It grows as n^1.027, so slightly slower than the plant does, and it tops out at 0.60 GB on a 9,511 tag site. That fits in the spare half of an ordinary industrial PC.
It does not want your cores
Pinned to a single core, with every threading library held to one thread, the same
compile takes 2.80 s against 2.81 s unpinned. The difference is inside the noise of the measurement, so the honest reading is that it does not need more than one core at all. That is the
useful number, because nobody buys a machine for us. Whatever we need has to fit beside
whatever that box is already doing.
Where the seconds actually go, from a profile rather than a guess. Time
spent inside numpy is pushed up to whichever of our own functions asked for it, because
knowing that a fifth of a compile is ndarray.partition tells you nothing and
knowing it went on describing the values tells you where to look. Finding duplicates is
the largest share and the reason the curve above bends: it compares series against each
other, so it grows faster than the plant does.
The part that is a language model
One step reads the paperwork: an instrument index, a datasheet, a 50 page manual. That is
the only place a language model runs, it runs on the same box, and it is small. Nothing is
sent anywhere, and no article body is ever fed to it, only headings and structured fields.
Every number it produces is checked against the fact the data layer computed before it is
allowed out.
Local models, measured reading the same paperwork. We ship ministral-3:3b at 2.95 GB, made by Mistral AI in the France under Apache 2.0. It reads the whole set in 80 seconds and phrases a finding in 8.6.
Where the model we ship comes from
We ship ministral-3:3b, made by Mistral AI in France under Apache 2.0. It is also the best reader we have measured on this job, which has not always been true and is worth saying plainly: an earlier version of this page argued that shipping a model from a country our buyers are comfortable with cost us real accuracy, and published the size of that cost. It was not true. The comparison behind it pitted the best Chinese model against one of the weakest models we had, and the trade-off was an artefact of that choice rather than a fact about where models are made. Every model measured stays in the table above, including the ones we did not pick.
What it would have cost to do this in a cloud
Not a competitor's invoice, which we cannot see. A volume, which is multiplication.
A plant of 1,469 tags sampled every second is 1,112 GB a year on the wire, a steady 0.28 Mbit a second that never stops. Sampled once a minute instead it is 18 GB, which is the number most people actually have in mind when they say this is cheap. A 20,000 tag site at a second is 15.1 TB a year. Running inside the fence, all of those are zero, and the security assessment has one fewer conduit to argue about.