Skip to content

harness: tool call emitted in content channel is silently discarded (small local models) #1053

Description

@data-bot-coasys

Problem

run_harness treats an empty tool_calls[] as "the model gave its final answer" and returns:

// rust-executor/src/ai_service/harness/mod.rs:240
if completion.tool_calls.is_empty() {
    // Model returned a plain answer — done.
    return Ok(completion.content);
}

Small local models regularly emit a well-formed tool call in the content channel instead of the native tool_calls[] channel. When they do, the harness discards a perfectly good call, buffers zero ops, and produces zero instances — indistinguishable, from the outside, from "the model did not understand the task".

Evidence

CircleCI job 29558 (workflow 3e6f2467, PR #988, commit 5852cfd65). The test
run-interpretation-harness.test.ts → "derives an intention from existing beliefs and links it back to them" failed all 8 internal attempts, each one like this:

harness: round=1 calls_used=0/15 tools_offered=11 tool_calls=[] content_preview="```json
{ "tool_call": { "name": "extintention_create",
                 "arguments": { "type": "ExtIntention",
                                "title": "Let's make it the intention going into the sprint." } } }
```"
harness: pass complete, ops_buffered=0 classes_offered=2
harness: apply_with_overlay produced 0 bases

The model chose the right tool, the right class and the right title on all 8 attempts. Only the channel was wrong. Variation across attempts was cosmetic ("type" as ExtIntention / Intention / intention; object vs. single-element array), so it was sampling normally.

Not an environment problem: [relation-hint-e2e] and [tool-events-e2e] both produced native tool calls and passed in the same job, and every completion returned in ~2s.

Proposed fix

When completion.tool_calls is empty, before returning the content as a final answer, try to parse a tool call out of it:

  • strip an optional ```json / ``` fence
  • accept either a single object or a single-element array
  • accept the {"tool_call": {"name", "arguments"}} shape seen here, and the bare {"name", "arguments"} shape
  • require name to match an offered tool; otherwise fall through to the existing "final answer" path unchanged
  • log at WARN when the fallback fires, so the wrong-channel behaviour stays visible rather than becoming silent

Unit-testable without an LLM: feed HarnessCompletion { tool_calls: vec![], content: <the fenced string> } and assert one call is extracted with the right name and arguments. Include a negative case (prose content that merely contains a code fence) that must still be treated as a final answer.

Why this matters beyond the test

Local-first small models are the target of this design, not an edge case. Any deployment running a small model through the interpretation harness hits this and sees "the LLM produced nothing" with no indication that the call was made and thrown away.

Notes

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions