perspective.runInterpretationWithHarness (WS + real LLM) in tests/js/tests/model/run-interpretation-harness.test.ts passes and fails on identical trees. It has now blocked a merge, and it will keep doing so.
The controlled observation
| tree |
job |
result |
d9eed5ae4 (#1013) |
29483 |
216 passing / 0 failing |
d9eed5ae4 (#1013) |
29491 |
216 passing / 0 failing |
5852cfd65 (#988) |
29558 |
215 passing / 1 failing |
git diff d9eed5ae4 5852cfd65 is empty — #988's head is the same content as the tree that passed twice, reached by a different merge path. Same runner pool, same suite, same assertions.
Failure: derives an intention from existing beliefs and links it back to them → "harness pass must produce at least one instance" (run-interpretation-harness.test.ts:162). The model returned nothing usable at all — that is a property of the model's output, not of a code path.
Ruled out, with evidence
- Model residency. The known failure mode (a foreign model resident on the 24GB card — see the
gemma3:12b 104s-cold vs 2.2s-warm behaviour) does not apply: at the time of the failure ollama ps on the runner showed gemma3:12b (10GB) and nomic-embed-text resident, nvidia-smi 15.6/24.5 GB used, ~9GB headroom. The right model was loaded and warm.
- A regression in the merged tree. Excluded by the empty diff above.
Prior history of the same suite
Two earlier failures on #1013's pre-fix tree, on two different test names: job 29461 (basedOn: [] — the relation was left empty) and job 29462 (same "must produce at least one instance" assertion, on the basedOn/contradicts test). Those had a real cause (#1005 declared a many cardinality with no example demonstrating one) and were fixed by d9eed5ae4. This one is on the fixed tree. So the suite has at least two distinct failure signatures, only one of which was a genuine defect.
Why it matters more than a normal flake
The assertion is all-or-nothing: did the model emit any instance. Every test in the file is one sample from a nondeterministic generator, and a single unlucky sample fails the job and blocks a merge on a required check. There is no retry, and nothing distinguishes "the model had an off sample" from "the prompt regressed" — which is exactly the discrimination #1013 needed and got only by running the suite four times by hand.
Suggested direction (not yet decided)
- Separate the axes. The prompt-contract assertions that can be deterministic (cardinality shape, key names, array-vs-bare) belong in Rust unit tests over
interpretation_examples() — some_example_demonstrates_a_many_relation_as_an_array is the model to copy. What is left in the real-LLM suite is genuinely statistical.
- Give the statistical part a retry budget (n attempts, assert at least one produces instances) or move it off the required-check path onto the nightly LLM job, where a failure is information rather than a merge block.
- Do not simply widen a timeout — the model answered, it just answered emptily.
Filed while merging #988; the failed job is being re-run. Related: #1005 (the real defect this suite did catch), and the nightly llm-e2e job as the natural home for sampling-based assertions.
perspective.runInterpretationWithHarness (WS + real LLM)intests/js/tests/model/run-interpretation-harness.test.tspasses and fails on identical trees. It has now blocked a merge, and it will keep doing so.The controlled observation
d9eed5ae4(#1013)d9eed5ae4(#1013)5852cfd65(#988)git diff d9eed5ae4 5852cfd65is empty — #988's head is the same content as the tree that passed twice, reached by a different merge path. Same runner pool, same suite, same assertions.Failure:
derives an intention from existing beliefs and links it back to them→ "harness pass must produce at least one instance" (run-interpretation-harness.test.ts:162). The model returned nothing usable at all — that is a property of the model's output, not of a code path.Ruled out, with evidence
gemma3:12b104s-cold vs 2.2s-warm behaviour) does not apply: at the time of the failureollama pson the runner showedgemma3:12b(10GB) andnomic-embed-textresident,nvidia-smi15.6/24.5 GB used, ~9GB headroom. The right model was loaded and warm.Prior history of the same suite
Two earlier failures on #1013's pre-fix tree, on two different test names: job 29461 (
basedOn: []— the relation was left empty) and job 29462 (same "must produce at least one instance" assertion, on thebasedOn/contradictstest). Those had a real cause (#1005 declared amanycardinality with no example demonstrating one) and were fixed byd9eed5ae4. This one is on the fixed tree. So the suite has at least two distinct failure signatures, only one of which was a genuine defect.Why it matters more than a normal flake
The assertion is all-or-nothing: did the model emit any instance. Every test in the file is one sample from a nondeterministic generator, and a single unlucky sample fails the job and blocks a merge on a required check. There is no retry, and nothing distinguishes "the model had an off sample" from "the prompt regressed" — which is exactly the discrimination #1013 needed and got only by running the suite four times by hand.
Suggested direction (not yet decided)
interpretation_examples()—some_example_demonstrates_a_many_relation_as_an_arrayis the model to copy. What is left in the real-LLM suite is genuinely statistical.Filed while merging #988; the failed job is being re-run. Related: #1005 (the real defect this suite did catch), and the nightly
llm-e2ejob as the natural home for sampling-based assertions.