Skip to content

Bound total in-flight agentic requests (roots AND sub-agent forks) #1358

Description

@thomas-primalabs

Problem

AIPerf has no knob that bounds the total number of requests in flight against the engine. The existing controls each bound a different thing:

  • --prefill-concurrency caps concurrent prefills but releases at first token, so it does not bound requests that are past TTFT and still streaming.
  • --concurrency caps concurrent session trees; within a tree, sub-agent forks can oversubscribe, so total in-flight requests can exceed the number. This is reflected in captured metrics but not in the live dashboard view and not in the concurrency specification (via YAML or CLI), which can lead to some confusion.

For agentic / DAG workloads with fan-out, this means the number of requests an engine actually sees concurrently is not directly controllable — which makes it hard to reproduce a specific in-flight load or to protect an engine from fork-driven oversubscription.

Proposed change

Add a request_concurrency dimension that caps every wire request — roots and sub-agent forks alike — from dispatch until the entire response terminates (not at TTFT). Surface:

  • --request-concurrency N (int ≥ 1): max total in-flight requests (until completed, cancelled, or errored); unset = unbounded (today's behavior).
  • --warmup-request-concurrency N (int ≥ 1): the same cap for the warmup phase; falls back to --request-concurrency when unset.
  • Phase-config fields request_concurrency and (AGENTIC_REPLAY only) agentic_warmup_request_concurrency, so the auto-synthesized cache-warmup burst can be throttled independently of profiling intensity.
  • Validator: prefill_concurrency <= request_concurrency (mirrors the existing prefill_concurrency <= concurrency check).
  • Sweep / dotted-path routing so request_concurrency is sweepable like the other concurrency dimensions.

Out of Scope

  • Per-worker or per-endpoint request caps
  • Changing the semantics of existing --concurrency or --prefill-concurrency
  • Rate-based control mechanisms (already exist via --request-rate; this is only about total in-flight requests)

Implementation sketch

Because I need the behavior for my own use cases, I already have a feature preview branch on https://github.com/PrimaLabs-AI/aiperf/tree/thomas-primalabs/request-concurrency . I am creating an upstream issue now instead of a PR because I believe this change is too major to directly PR without proper consideration of the AIPerf community first.

High level: A per-phase request-slot semaphore on ConcurrencyManager (acquire_request_slot / try_acquire_request_slot / release_request_slot), acquired in the credit issuer (both the blocking and non-blocking issue paths, with rollback of the session slot on failure) and released in the credit callback on every return except no_request virtual credits. Backward compatible (unset = unbounded).

Implementation Minutiae

  • A no_request virtual credit neither acquires nor releases a request slot so that the semaphore never underflows.
  • ConcurrencyManager: acquire_request_slot(phase_key, can_proceed_fn) (async, blocking), try_acquire_request_slot(...) (sync, non-blocking), release_request_slot(phase_key); a per-phase semaphore sized to request_concurrency to cover both synchronous and asynchronous paths.
  • CreditIssuer._issue_credit_ready / try_issue_credit: acquire request slot with session-slot rollback on failure; skip when turn.no_request. Slot acquisition is factored into an ordered acquire-with-LIFO-rollback helper pair, _acquire_slots (async) / _try_acquire_slots (sync), so the session -> request -> prefill chain (and its rollback) lives in one place.
  • CreditCallbackHandler.on_credit_return: if not credit.no_request: concurrency.release_request_slot(phase).

Notes for maintainers (two convention points)

While building my feature branch, I bumped into at least two problems that would matter before upstream adoption.

  1. BasePhaseConfig field count: the two new phase-config fields (request_concurrency, agentic_warmup_request_concurrency) push BasePhaseConfig to 31 fields, one over the check-ergonomics soft limit of 30 ("split into sub-models"). I manually bumped the ergonomics baseline entry to keep this PR single-concern, but it merits proper consideration on how the interface should be presented. Are 31 fields acceptable here, or would you prefer a specific field migrate into a sub-model (e.g. a concurrency or agentic-warmup group)? A model split would change the config access path and the YAML/CLI/schema shape.
  2. C901 complexity (resolved): adding the request-slot acquire/rollback tipped CreditIssuer._issue_credit_ready and try_issue_credit over the repo's C901 gate (check-ruff-baselined, threshold 10). Rather than ratchet the baseline, the slot-acquisition-with-LIFO-rollback ladder is extracted into a helper pair (_acquire_slots / _try_acquire_slots — an async and a sync variant, since the two issue paths differ only in awaiting vs. non-blocking acquire); both functions are now under the limit.

If it is better for me to directly open a PR for discussion of implementation specifics, I can do that.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions