Skip to content

Checking a run against what you meant

A scenario says what plant to build and what happens to it. It can also say what the run should show: the controlled temperature stays quiet while the valve absorbs the disturbance, a grade change settles within six hours, a fault shows on the valve within two hours. These statements go in an expect section. They do not change the plant (and not the config_hash), but homeostat diagnose simulates the scenario and checks each one, with the numbers behind the verdict.

This is how a person or an AI assistant finds out that a run is not what was intended: a setpoint step that never reached the valve, a controller that saturates, a fault too small to see.

expect:
  - id: loop_holds_temperature
    kind: stat
    signal: TIC-101.cv
    stat: std
    at_most: 0.3
    note: in AUTO the controller suppresses the feed disturbance on the controlled temperature
  - {id: valve_absorbs_it, kind: stat, signal: TIC-101.out, stat: std, at_least: 0.5}
  - {id: valve_not_saturated, kind: saturation, loop: TIC-101, at_most: 0.01}
homeostat diagnose plant.yaml
3 of 3 expectations hold.
  PASS  loop_holds_temperature  std of TIC-101.cv over the whole run is 0.09247 (expected at most 0.3)
  PASS  valve_absorbs_it        std of TIC-101.out over the whole run is 1.526 (expected at least 0.5)
  PASS  valve_not_saturated     TIC-101.out is at a limit for 0.0% of the samples over the whole run (expected at most 1.0%)

Every entry has a kind, an optional id (the name in the report; default expect[<position>]) and an optional note: the reason it should hold, in words. A note is shown when the expectation fails, so it carries the intent to whoever reads the failure. The shipped examples (homeostat examples) each end with an expect section.

The kinds

homeostat describe stat (or any kind) prints the fields of one; the scenario schema lists them all.

Kind It checks that Uses
stat the mean, standard deviation, minimum or maximum of a signal, over a window, is within at_least and at_most any signal
shift a signal's mean moved between two windows by at least or at most some amount any signal
saturation a controller output is at one of its limits for at most a fraction of the samples a loop
settles loops settle on the new setpoint within a time after a production transition the settle_times labeler
visible a fault becomes detectable on a signal within a time, or stays hidden (visible: false) the visibility labeler
feasible every loop starts on its setpoint with its output inside its limits the feasibility labeler
expect:
  # the controlled reading holds its setpoint, but the true value does not (a drifted transmitter)
  - {kind: stat, signal: TT-101, stat: mean, window: [24d, 29d], at_least: 74.8, at_most: 75.2}
  - {kind: stat, signal: TIC-101.cv, stat: mean, window: [24d, 29d], at_most: 74.0}
  # a change reaches a signal: the valve moves 5 % or more after the setpoint step
  - {kind: shift, signal: TIC-101.out, before: [0s, 2h], after: [3h, 4h], at_least: 5}
  # and a change does not reach one: the signal moved by less than 0.3
  - {kind: shift, signal: TIC-101.cv, before: [0s, 2h], after: [3h, 4h], at_most: 0.3, absolute: true}
  # every grade change settles within 6 hours, or only one of them
  - {kind: settles, within: 6h}
  - {kind: settles, transition: "grade_C->grade_A", loop: TIC-101, within: 12h}
  # a fault shows on the valve within 2 hours; another is too small to stand out from the reading's own variation
  - {kind: visible, fault: valve_stiction, signal: FIC-101.out, within: 2h}
  - {kind: visible, fault: sensor_noise, signal: TT-101, visible: false}

Details that decide whether an expectation says what you mean:

  • A window is [from, to], in time since the start of the run ([1d, 6d]). Without one, the whole run counts. Samples that are NaN (an analyzer between samples) are ignored.
  • std is the population standard deviation of the recorded samples. shift is the mean after the first window minus the mean before; give at_least for a rise and at_most for a fall (a negative number), or use absolute: true to judge only its size.
  • lane (default 0) picks the lane of a multi-lane scenario, for example the manual-mode twin of a twin run.
  • settles is judged on the true regulated variable and counts from the start of the transition, as the settle_times labeler does; band and hold set how close and for how long.
  • visible runs the scenario again with one twin lane per fault (the scenario must have one lane), and judges visibility on observed signals only (what a historian records, so measured tags and each controller's setpoint, mode and output). Every fault that matches fault (and target, when several units have it) must satisfy the expectation. It is an oracle test: a fault it calls hidden is hidden from any detector. A fault is often visible on a transmitter reading only for a short transient, because the loop then removes it from the reading; within and visible: false are about the first visibility, so judge a reading-versus-truth effect with stat on both signals as above.

Reading a diagnosis

Each expectation ends as one of three:

  • PASS: the run shows what was expected.
  • FAIL: the run does not. The line says what was measured and what was expected, and the evidence in the JSON has the numbers (the value, the count of samples, the settle times, the detection delays).
  • ERROR: the check could not be made, for example a visibility check on a signal that is not observed.

homeostat diagnose exits 0 when every expectation holds, 3 when one does not (1 when the scenario is invalid), so it can gate a script. With --json it prints {ok, checked, passed, failed, errors, checks: [...], warnings, reproducibility}.

A failure has two possible causes, and the fix differs: the plant is not what you meant (change the scenario), or the expectation was wrong (change it, knowing why). Do not loosen a bound just to make a check pass. When a failure is surprising, homeostat explain says what plant was described, and homeostat diff shows what an edit changed.

homeostat validate checks that every name an expectation uses exists: a misspelt signal, loop, transition, fault or lane is an error with a suggestion, found before any simulation. Expectations are checked for one plant: a scenario with a family of several plants must be drawn plant by plant (homeostat.prepare(source, plant=n)).

In Python:

from homeostat.ai.diagnose import diagnose

diagnosis = diagnose("plant.yaml")
diagnosis.ok                         # True when every expectation holds
diagnosis.checks[0].evidence         # the numbers behind the first verdict
print(diagnosis.render())