Each one was written down before the data was opened, and is published whatever
it came out as.
Real plants, real data
On the two real Danish plants, against the variable tables their owners published: Avedøre went from 2 of 24 tags placed to 24 of 24, the Aalborg building from 0 of 188 to 151 of 188, every quantity placed is right, and the 37 the papers list in units we have no word for (lux, seconds) are asked, not guessed; the water and building connectors were written for these plants, so a third plant is the real test; the building's heat meter obeys the first law, power = 3953 x flow x temperature difference, 5% from what water carries in l/s and W, which is the units its owners published, worked out from the readings alone; and its supply air flow is not measured at all: it is 308 x the square root of the fan's inlet pressure, to 0.06%, a calculation found from the readings, so the flow is only as good as that one pressure tap; its ventilation meter reads 1000 times the two fans' power added up, to 100.0% of its movement, which gives the units and the wiring from the readings alone
Control loops
Which valve controls which reading, read from the data alone and preregistered: FALSIFIED on held-out periods, 4 of 6 answers right against an 80% bar (refusing the unsure ones still doubled the precision of the plain method, 67% against 33%); weather-compensated heating made two radiator valves follow the outdoor air rather than their rooms; a second attempt that first drops readings an actuator cannot move was also FALSIFIED on new held-out months, 4 of 7: in winter the valves follow the shared district heating flow, which they do move, more than their rooms, so neither is in the console; what does work, and is in the console, is checking the loop list a plant's names imply: SUPPORTED on never-opened months, 7 of 7 confirmed loops right and 197 of 198 deliberately wrong pairings rejected
Which placements are wrong, and proving how many
Proving how many tags are wrong from a few checks, preregistered: PARTIAL. An engineer checks 50 random placements and gets an exact 95% ceiling on the plant's error rate; across 10 held-out plants it held every time and averaged 7.5%, tighter than the textbook bound's 8.4%, so that bound, not a model, is what is now in the console. A model that learns which placements are wrong found 72% of them in the riskiest tenth on simulated plants, against 22% for the compiler's own confidence, but on the one real plant it ranked worse than that confidence (0.62 against 1.00), so it stays out until it has learned from real plants; a bound that leans on its predictions held only 69% of the time, not 95%
Attacks, and the relations a plant must keep
Catching attacks on a real testbed (HAI) with the relations a plant must keep, a valve's position against its command, a meter against its twin, preregistered: PARTIAL. On 38 attacks never opened before, the relations alone caught 22, the strongest plain method 19, at a similar F1 (0.44 against 0.47) and with the broken relation named; the bar of 80% of attacks was missed, so we claim no detection rate; the relations now run in the console's morning report, each break named with the tags that parted
Asking in your own words
Asking the plant in your own words, Norwegian or English, preregistered: PARTIAL. A small model on the same machine reads what is asked and picks only a question type and a machine from the plant's own list; the answer is computed from the data and never written by the model. On 80 questions never seen before it read 81% right, against 20% for the keywords that shipped, but only 80% of the Norwegian ones against a bar of 85%, so we do not claim it understands Norwegian yet; it is in the console, and it shows how it read each question
Laws found from data, and virtual sensors
The laws a plant follows, found from its own data and written so a person can read them, and an estimate when a reading dies, preregistered: SUPPORTED. On a real testbed, held out, a law of at most 5 terms was only 13% less accurate than a regression on everything, and an estimate for a dead reading was 38% as wrong as the frozen last value a control system shows; afterwards we found some laws leaned on a second output of the same meter, and with those excluded one claim falls just short, which we publish; the laws and the estimates are in the console, every estimate marked
Asking back when unsure
Asking back when it is unsure what a question means, preregistered: PARTIAL. When two readers on the same machine disagree, the console offers both readings and the operator picks one, which is kept with the plant; on 60 fresh questions it asked 16 times and offered the right reading 14 times, but 6 of 11 wrong readings went by without a question, mostly where both readers made the same mistake, and we say so
A guarantee the plant chooses
A guarantee the plant chooses instead of a confidence score, preregistered: SUPPORTED. From a spot check its own people answer, the console proves that of the tags it keeps, at most a chosen share are wrong; on 20 simulated plants never used before, 200 checks guaranteed at most 5% wrong while keeping 99% of tags, 100 checks at most 10% while keeping 99%, and the guarantee held in every audit; it is bought with checks, not with a better score
A virtual sensor across a change of season
The virtual sensor tested where it is hardest, a real office building across a change of season, preregistered: FALSIFIED. Fitted on winter and spring, tested on summer and autumn, neither our readable law nor a dense regression beat simply holding the last value: the law did on 8% of readings, the regression on 0%, and the regression, better on our testbed, was 9 times worse here; so the console now offers an estimate only where its law clearly beat the last value on recent history, and says which period the law was found in