Skip to content

Measuring how well an assistant writes scenarios

Every improvement to Homeostat for AI assistants (clearer errors, explain, diagnose) is a guess until something measures it. The eval set does: it gives an assistant a request in plain language ("simulate a flow loop with an unmeasured pressure disturbance, record every 10 seconds") and scores the scenario it writes by what that scenario does.

homeostat eval list                 # the tasks, easiest first
homeostat eval show setpoint_step   # the request to give an assistant

The tasks cover a single loop (level 1), a fault and what it does to the data (2), a plan, a maintenance history or a twin experiment (3), and several units wired together (4).

How a scenario is scored

Each task has acceptance checks, written as ordinary expectations. They belong to the task, not to the candidate: the harness ignores any expect section the candidate wrote (a scenario cannot grade itself), simulates the candidate's plant, and checks the task's expectations on it. The measures:

  • valid: the scenario validates, with no errors.
  • passed / total: how many of the task's checks hold. A check that names something the scenario does not have (a signal, a loop, a fault) fails; it does not make the scenario invalid, because the scenario may be a perfectly good plant that does not do what was asked.
  • success: valid, and every check holds.
  • as written: the file validates exactly as the candidate wrote it, with its own expect section too. It does not decide success (the candidate's expectations are ignored), but it shows friction that success hides: a guessed key or a misspelt name that the tool had to correct.

A scenario that needs more than five million steps is not run and does not succeed: the run and any burn_in, divided by dt, times the lanes, plus the twin lanes that a visible check adds (one per fault, and one). A request for a month of data does not need a base step of one second. A scenario that validates but cannot be simulated (the plant never settles) stays valid; it is reported as not run.

homeostat eval score setpoint_step my_solution.yaml
setpoint_step: ok, 6 of 6 checks hold
  PASS  starts_on_setpoint: all 1 loops start on their setpoint with the output inside its limits
  PASS  the_setpoint_steps: shift of TIC-101.sp (mean 80 from 0s to 50min, mean 90 from 3h to 4h) is 10 (expected between 9.9 and 10.1)
  ...

Measuring an assistant

  1. Make an empty directory for the assistant to work in, with homeostat installed.
  2. Give it a task's request (homeostat eval show <task>) and the command line. Do not give it the acceptance checks (eval show --checks prints them, as the answer key) or point it at this package's files.
  3. Ask it to save its first draft as <task>.first.yaml before it validates anything, and its final scenario as <task>.yaml, after it has used validate, explain and diagnose as it sees fit.
  4. Score the directory:
homeostat eval run solutions/
task                  L  first draft        final
flow_disturbance      1  -                  ok
level_disturbance     1  -                  ok
setpoint_step         1  -                  ok
bias_hidden_by_loop   2  -                  ok
stiction_visible      2  -                  ok
fouling_and_cleaning  3  -                  ok
three_grades          3  -                  ok
twin_manual           3  -                  ok
product_quality       4  -                  ok

final:       9 of 9 tasks succeed (9 valid), 45 of 45 checks hold
by level:    L1 3/3, L2 2/2, L3 3/3, L4 1/1

This is the report for the nine reference solutions, which the repository's tests keep solving their tasks; it takes a few seconds. A failing task is listed below the table with the errors or the failing checks and their evidence, and a task with first drafts also shows first draft: 8 of 9 tasks succeed (8 valid), 7 of 9 validate exactly as written, 1 repaired afterwards (the numbers of your own run), and a first draft that needed a correction shows it as ok, 1 err as written or invalid (1 error).

The first draft columns say how good the documentation, the schema and the error messages are before any feedback: the share valid on the first try, the share that succeeds. The final columns say what the validate-explain-diagnose loop achieves. The difference, the drafts that were repaired, is what the feedback is worth. Run the same directory format against two versions of Homeostat (or two assistants) and compare. eval run exits 3 unless every final scenario succeeds, and --json prints the scores per task and per check.

The harness does not defend against an assistant that reads the answer key: the acceptance checks ship in the package, and the reference solutions (which prove each task can be solved, and which a set of near misses must fail) live only in the repository's tests. Keep the assistant out of both.

The Claude Code skill

homeostat skill prints a skill that teaches the whole workflow: what to run in what order, and the mistakes that cost the most time (parameters are keys of the unit, a regulated variable is changed through its setpoint, a loop hides a fault from the reading, YAML reads on as a boolean). To install it for a project:

homeostat skill --install       # writes .claude/skills/homeostat/SKILL.md

Claude Code then loads it when a request is about simulating a plant, generating process data or fixing a scenario.

Adding a task

A task is a YAML file in src/homeostat/ai/evals/tasks/ with an id, a title, a level, a prompt and a list of expect entries, and a reference solution with near misses in tests/evals/. The request must give every identifier the checks use (unit ids, durations, numbers). The tests check that the reference passes every check and that each near miss (a scenario that is close but wrong: no disturbance, a bias in the wrong direction, a missing grade) fails, so a check is shown to tell a solution from an almost-solution. Calibrate the thresholds on the reference, with margin for the other seeds and step sizes a correct solution may use.