Skip to content

DEP: Complete allowed-tools semantics for Responses function tools #14999

Description

@xianlubird

Status

Draft

Area

Frontend

Summary

Implement first-class tool_choice.allowed_tools semantics for function tools on the Responses API without changing the advertised tool catalog between turns.

PR #14990 is the first implementation PR. It provides an immediately useful, function-only compatibility path by validating the requested subset, filtering the backend-visible tools, preserving auto or required, and preventing disallowed calls from reaching clients. Follow-up PRs will separate the complete function catalog from the per-turn allowed subset, enforce that subset throughout preprocessing and output conversion, and preserve the stable tool-definition prefix needed for prompt-cache reuse.

Motivation

Agent clients may send a stable function-tool catalog while allowing only a smaller subset on a particular turn. Dynamo's Responses-to-Chat adapter previously rejected this choice. PR #14990 makes those requests usable, but it implements the subset by removing non-allowed definitions before dispatch.

Filtering definitions is correct for tool availability, but it changes the rendered tool catalog whenever the allowed subset changes. In multi-turn agent workloads this can reduce prompt-cache reuse and makes the request representation conflate two different concepts:

  • the functions known to the model for the session; and
  • the functions permitted on the current turn.

The complete implementation should represent those concepts independently and apply one policy consistently across request validation, prompt rendering, guided decoding, parsing, unary conversion, and streaming conversion. It must continue to fail closed: an unsupported selector or a backend-generated disallowed call must never silently widen the allowed set.

Proposal

1. Define the function-only contract

For this DEP:

  • tools is the complete function catalog advertised by the request.
  • tool_choice.allowed_tools.tools is a per-turn subset of that catalog, resolved by function name.
  • mode: "auto" permits either assistant text or one or more calls from the subset.
  • mode: "required" requires one or more calls from the subset.
  • parallel_tool_calls: false retains its existing single-call constraint after the allowlist is applied.

An empty subset, an unknown function, an ambiguous flattened name, or a non-function selector returns an actionable 4xx before backend dispatch. Duplicate selectors are treated as one function. Hosted, MCP, custom, and other non-function tool execution remain outside this DEP and continue under issue #12859.

2. Preserve catalog and policy separately

Add a canonical internal representation that carries both:

available_functions: complete validated definitions
allowed_functions: optional per-turn set of canonical function identities
allowed_mode: auto | required

The representation must survive Responses conversion, unified request handling, preprocessing, remote request serialization where applicable, and response conversion. Validation and name resolution should live in one shared policy object rather than being reimplemented by unary and streaming paths.

For namespaced tools, resolve the selector to the same canonical identity used by the flattened backend representation. Reject collisions that cannot be represented unambiguously by the function-only selector shape.

3. Enforce the subset while retaining a stable tool prefix

Render the complete function catalog in a stable order so two requests with the same catalog but different allowed subsets retain the same tool-definition prefix. Encode the per-turn subset and mode separately, after the stable catalog portion or through an equivalent backend-neutral control field.

Preprocessing must derive tool constraints from the allowed definitions, not from the full catalog:

  • required-mode guided decoding or structural tags must admit only allowed functions;
  • auto mode may produce text or an allowed call, while disallowed calls are rejected by the parser/output policy;
  • named, none, ordinary auto, and ordinary required choices retain their existing behavior;
  • tool parsers that cannot enforce the representation must report an explicit capability error rather than silently ignoring the subset.

The exact prompt/control encoding may vary by supported template or parser, but it must satisfy the same observable contract. Any compatibility fallback that temporarily filters definitions must be explicit and measurable; it is not considered completion of the prompt-cache objective.

4. Centralize output enforcement

Apply the canonical allowlist to every function-call source:

  • structured unary tool calls;
  • text-parsed unary tool calls;
  • streaming tool-call deltas, including identity split across chunks;
  • parallel calls and resumed reasoning around tool calls.

A backend or parser violation must not leak a disallowed call. Unary requests return a server-side contract error. Streaming requests close open output items as incomplete, emit one response.failed, mark the terminal event before yielding it, and immediately drop the upstream stream. The same policy decides function identity in every path.

5. Deliver incrementally

PR Scope Completion criterion
PR 1 — compatibility behavior (#14990) Accept function-only allowed_tools, validate and filter the backend-visible definitions, preserve mode, and enforce the subset in unary and streaming responses. Agent clients can use selected function tools safely. The known limitation is that subset changes also change the rendered tool catalog.
PR 2 — canonical protocol and preprocessing policy Introduce the separate catalog/allowlist representation, shared validation and name resolution, request propagation, and allowed-subset constraints for preprocessing and parsing. The complete catalog reaches the prompt path unchanged while required/auto semantics use only the selected subset. Existing tool-choice modes do not regress.
PR 3 — stable-prefix integration and capability rollout Place per-turn selection outside the stable tool-definition prefix, add capability handling for supported templates/parsers, and measure prompt-cache reuse. Requests with the same catalog and different subsets share an identical tool-definition prefix; unsupported paths fail explicitly.
PR 4 — parity, observability, and documentation Consolidate unary/stream enforcement, add violation metrics, complete backend-neutral integration coverage, and document support and fallbacks. All supported serving paths expose the same behavior and the compatibility filtering fallback is no longer the default complete path.

PR 1 can merge independently because it is fail-closed and useful today. PRs 2–4 complete the architectural and prompt-cache goals without expanding this DEP to hosted-tool execution.

Requirements and acceptance criteria

  • Function-only allowed_tools supports both auto and required modes.
  • Request validation rejects empty, unknown, ambiguous, and non-function selections before dispatch.
  • The full function catalog and per-turn subset remain distinct after Responses conversion.
  • Changing only the allowed subset does not change the serialized/rendered tool-definition prefix or its cache identity.
  • Required-mode constraints admit only allowed function schemas.
  • Auto mode never exposes a disallowed function call; assistant text and allowed calls remain valid.
  • parallel_tool_calls: false, namespaced functions, reasoning items, finish reasons, and usage accounting retain their existing contracts.
  • Unary and streaming paths produce equivalent output and error semantics.
  • Unsupported template, parser, or backend combinations fail explicitly and identify the missing capability.
  • Deterministic tests cover conversion, validation, preprocessing constraints, parsed and structured outputs, chunk-split streaming identities, terminal failure ordering, and prompt-prefix stability.
  • Integration coverage exercises supported tool parsers with both modes and verifies that changing the subset preserves prompt-cache eligibility.
  • Metrics distinguish invalid client selections from backend/parser allowlist violations.

Alternate solutions

Keep filtering the tool catalog

This is PR 1 and remains a valid compatibility fallback. It is simple and fail-closed, but subset changes alter the prompt and reduce cache reuse, so it does not satisfy the complete design.

Add a prompt-only instruction

A textual allowlist can preserve the catalog, but a model may ignore it. Without parser, decoding, and output enforcement it is not an API contract.

Filter only the response

This prevents leakage but wastes generation, cannot faithfully implement required mode, and delays failure until after backend work. It remains defense in depth, not the primary mechanism.

Implement separate backend-specific adapters

Backend-specific controls may be needed below the shared policy, but defining semantics independently in each backend would create drift across Responses, preprocessing, and output conversion. The canonical policy belongs in the shared frontend/protocol path.

Decisions requested

  1. Should the canonical allowlist representation live in the shared protocol crate or in the LLM frontend request layer until the wire contract stabilizes?
  2. Which prompt/template and parser combinations can place per-turn policy after a stable tool-definition prefix without changing model behavior?
  3. Should an unsupported path use the explicit PR 1 filtering fallback or return a capability error once the complete path is available?
  4. Which prompt-cache measurement should gate removal of the default filtering fallback?

References

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    dep:draftDEP in draft statusdynamo-llmRelates to dynamo-llm componentfrontend`python -m dynamo.frontend` and `dynamo-run in=http|text|grpc`language::rustIssues/PRs that reference Rust codetool-calling

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions