Problem
run_harness treats an empty tool_calls[] as "the model gave its final answer" and returns:
// rust-executor/src/ai_service/harness/mod.rs:240
if completion.tool_calls.is_empty() {
// Model returned a plain answer — done.
return Ok(completion.content);
}
Small local models regularly emit a well-formed tool call in the content channel instead of the native tool_calls[] channel. When they do, the harness discards a perfectly good call, buffers zero ops, and produces zero instances — indistinguishable, from the outside, from "the model did not understand the task".
Evidence
CircleCI job 29558 (workflow 3e6f2467, PR #988, commit 5852cfd65). The test
run-interpretation-harness.test.ts → "derives an intention from existing beliefs and links it back to them" failed all 8 internal attempts, each one like this:
harness: round=1 calls_used=0/15 tools_offered=11 tool_calls=[] content_preview="```json
{ "tool_call": { "name": "extintention_create",
"arguments": { "type": "ExtIntention",
"title": "Let's make it the intention going into the sprint." } } }
```"
harness: pass complete, ops_buffered=0 classes_offered=2
harness: apply_with_overlay produced 0 bases
The model chose the right tool, the right class and the right title on all 8 attempts. Only the channel was wrong. Variation across attempts was cosmetic ("type" as ExtIntention / Intention / intention; object vs. single-element array), so it was sampling normally.
Not an environment problem: [relation-hint-e2e] and [tool-events-e2e] both produced native tool calls and passed in the same job, and every completion returned in ~2s.
Proposed fix
When completion.tool_calls is empty, before returning the content as a final answer, try to parse a tool call out of it:
- strip an optional
```json / ``` fence
- accept either a single object or a single-element array
- accept the
{"tool_call": {"name", "arguments"}} shape seen here, and the bare {"name", "arguments"} shape
- require
name to match an offered tool; otherwise fall through to the existing "final answer" path unchanged
- log at WARN when the fallback fires, so the wrong-channel behaviour stays visible rather than becoming silent
Unit-testable without an LLM: feed HarnessCompletion { tool_calls: vec![], content: <the fenced string> } and assert one call is extracted with the right name and arguments. Include a negative case (prose content that merely contains a code fence) that must still be treated as a final answer.
Why this matters beyond the test
Local-first small models are the target of this design, not an edge case. Any deployment running a small model through the interpretation harness hits this and sees "the LLM produced nothing" with no indication that the call was made and thrown away.
Notes
Problem
run_harnesstreats an emptytool_calls[]as "the model gave its final answer" and returns:Small local models regularly emit a well-formed tool call in the
contentchannel instead of the nativetool_calls[]channel. When they do, the harness discards a perfectly good call, buffers zero ops, and produces zero instances — indistinguishable, from the outside, from "the model did not understand the task".Evidence
CircleCI job 29558 (workflow
3e6f2467, PR #988, commit5852cfd65). The testrun-interpretation-harness.test.ts→ "derives an intention from existing beliefs and links it back to them" failed all 8 internal attempts, each one like this:The model chose the right tool, the right class and the right title on all 8 attempts. Only the channel was wrong. Variation across attempts was cosmetic (
"type"asExtIntention/Intention/intention; object vs. single-element array), so it was sampling normally.Not an environment problem:
[relation-hint-e2e]and[tool-events-e2e]both produced native tool calls and passed in the same job, and every completion returned in ~2s.Proposed fix
When
completion.tool_callsis empty, before returning the content as a final answer, try to parse a tool call out of it:```json/```fence{"tool_call": {"name", "arguments"}}shape seen here, and the bare{"name", "arguments"}shapenameto match an offered tool; otherwise fall through to the existing "final answer" path unchangedUnit-testable without an LLM: feed
HarnessCompletion { tool_calls: vec![], content: <the fenced string> }and assert one call is extracted with the right name and arguments. Include a negative case (prose content that merely contains a code fence) that must still be treated as a final answer.Why this matters beyond the test
Local-first small models are the target of this design, not an edge case. Any deployment running a small model through the interpretation harness hits this and sees "the LLM produced nothing" with no indication that the call was made and thrown away.
Notes
relation-hint-e2eneeded 7 attempts in the same job), so more attempts is the wrong lever.