Skip to content

integration-tests-model: the real-LLM harness suite fails nondeterministically and blocks merges #1052

Description

@data-bot-coasys

perspective.runInterpretationWithHarness (WS + real LLM) in tests/js/tests/model/run-interpretation-harness.test.ts passes and fails on identical trees. It has now blocked a merge, and it will keep doing so.

The controlled observation

tree job result
d9eed5ae4 (#1013) 29483 216 passing / 0 failing
d9eed5ae4 (#1013) 29491 216 passing / 0 failing
5852cfd65 (#988) 29558 215 passing / 1 failing

git diff d9eed5ae4 5852cfd65 is empty — #988's head is the same content as the tree that passed twice, reached by a different merge path. Same runner pool, same suite, same assertions.

Failure: derives an intention from existing beliefs and links it back to them → "harness pass must produce at least one instance" (run-interpretation-harness.test.ts:162). The model returned nothing usable at all — that is a property of the model's output, not of a code path.

Ruled out, with evidence

  • Model residency. The known failure mode (a foreign model resident on the 24GB card — see the gemma3:12b 104s-cold vs 2.2s-warm behaviour) does not apply: at the time of the failure ollama ps on the runner showed gemma3:12b (10GB) and nomic-embed-text resident, nvidia-smi 15.6/24.5 GB used, ~9GB headroom. The right model was loaded and warm.
  • A regression in the merged tree. Excluded by the empty diff above.

Prior history of the same suite

Two earlier failures on #1013's pre-fix tree, on two different test names: job 29461 (basedOn: [] — the relation was left empty) and job 29462 (same "must produce at least one instance" assertion, on the basedOn/contradicts test). Those had a real cause (#1005 declared a many cardinality with no example demonstrating one) and were fixed by d9eed5ae4. This one is on the fixed tree. So the suite has at least two distinct failure signatures, only one of which was a genuine defect.

Why it matters more than a normal flake

The assertion is all-or-nothing: did the model emit any instance. Every test in the file is one sample from a nondeterministic generator, and a single unlucky sample fails the job and blocks a merge on a required check. There is no retry, and nothing distinguishes "the model had an off sample" from "the prompt regressed" — which is exactly the discrimination #1013 needed and got only by running the suite four times by hand.

Suggested direction (not yet decided)

  1. Separate the axes. The prompt-contract assertions that can be deterministic (cardinality shape, key names, array-vs-bare) belong in Rust unit tests over interpretation_examples() — some_example_demonstrates_a_many_relation_as_an_array is the model to copy. What is left in the real-LLM suite is genuinely statistical.
  2. Give the statistical part a retry budget (n attempts, assert at least one produces instances) or move it off the required-check path onto the nightly LLM job, where a failure is information rather than a merge block.
  3. Do not simply widen a timeout — the model answered, it just answered emptily.

Filed while merging #988; the failed job is being re-run. Related: #1005 (the real defect this suite did catch), and the nightly llm-e2e job as the natural home for sampling-based assertions.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions