Problem
AIPerf has no knob that bounds the total number of requests in flight against the engine. The existing controls each bound a different thing:
--prefill-concurrency caps concurrent prefills but releases at first token, so it does not bound requests that are past TTFT and still streaming.
--concurrency caps concurrent session trees; within a tree, sub-agent forks can oversubscribe, so total in-flight requests can exceed the number. This is reflected in captured metrics but not in the live dashboard view and not in the concurrency specification (via YAML or CLI), which can lead to some confusion.
For agentic / DAG workloads with fan-out, this means the number of requests an engine actually sees concurrently is not directly controllable — which makes it hard to reproduce a specific in-flight load or to protect an engine from fork-driven oversubscription.
Proposed change
Add a request_concurrency dimension that caps every wire request — roots and sub-agent forks alike — from dispatch until the entire response terminates (not at TTFT). Surface:
--request-concurrency N (int ≥ 1): max total in-flight requests (until completed, cancelled, or errored); unset = unbounded (today's behavior).
--warmup-request-concurrency N (int ≥ 1): the same cap for the warmup phase; falls back to --request-concurrency when unset.
- Phase-config fields
request_concurrency and (AGENTIC_REPLAY only) agentic_warmup_request_concurrency, so the auto-synthesized cache-warmup burst can be throttled independently of profiling intensity.
- Validator:
prefill_concurrency <= request_concurrency (mirrors the existing prefill_concurrency <= concurrency check).
- Sweep / dotted-path routing so
request_concurrency is sweepable like the other concurrency dimensions.
Out of Scope
- Per-worker or per-endpoint request caps
- Changing the semantics of existing
--concurrency or --prefill-concurrency
- Rate-based control mechanisms (already exist via
--request-rate; this is only about total in-flight requests)
Implementation sketch
Because I need the behavior for my own use cases, I already have a feature preview branch on https://github.com/PrimaLabs-AI/aiperf/tree/thomas-primalabs/request-concurrency . I am creating an upstream issue now instead of a PR because I believe this change is too major to directly PR without proper consideration of the AIPerf community first.
High level: A per-phase request-slot semaphore on ConcurrencyManager (acquire_request_slot / try_acquire_request_slot / release_request_slot), acquired in the credit issuer (both the blocking and non-blocking issue paths, with rollback of the session slot on failure) and released in the credit callback on every return except no_request virtual credits. Backward compatible (unset = unbounded).
Implementation Minutiae
- A
no_request virtual credit neither acquires nor releases a request slot so that the semaphore never underflows.
ConcurrencyManager: acquire_request_slot(phase_key, can_proceed_fn) (async, blocking), try_acquire_request_slot(...) (sync, non-blocking), release_request_slot(phase_key); a per-phase semaphore sized to request_concurrency to cover both synchronous and asynchronous paths.
CreditIssuer._issue_credit_ready / try_issue_credit: acquire request slot with session-slot rollback on failure; skip when turn.no_request. Slot acquisition is factored into an ordered acquire-with-LIFO-rollback helper pair, _acquire_slots (async) / _try_acquire_slots (sync), so the session -> request -> prefill chain (and its rollback) lives in one place.
CreditCallbackHandler.on_credit_return: if not credit.no_request: concurrency.release_request_slot(phase).
Notes for maintainers (two convention points)
While building my feature branch, I bumped into at least two problems that would matter before upstream adoption.
BasePhaseConfig field count: the two new phase-config fields (request_concurrency, agentic_warmup_request_concurrency) push BasePhaseConfig to 31 fields, one over the check-ergonomics soft limit of 30 ("split into sub-models"). I manually bumped the ergonomics baseline entry to keep this PR single-concern, but it merits proper consideration on how the interface should be presented. Are 31 fields acceptable here, or would you prefer a specific field migrate into a sub-model (e.g. a concurrency or agentic-warmup group)? A model split would change the config access path and the YAML/CLI/schema shape.
- C901 complexity (resolved): adding the request-slot acquire/rollback tipped
CreditIssuer._issue_credit_ready and try_issue_credit over the repo's C901 gate (check-ruff-baselined, threshold 10). Rather than ratchet the baseline, the slot-acquisition-with-LIFO-rollback ladder is extracted into a helper pair (_acquire_slots / _try_acquire_slots — an async and a sync variant, since the two issue paths differ only in awaiting vs. non-blocking acquire); both functions are now under the limit.
If it is better for me to directly open a PR for discussion of implementation specifics, I can do that.
Problem
AIPerf has no knob that bounds the total number of requests in flight against the engine. The existing controls each bound a different thing:
--prefill-concurrencycaps concurrent prefills but releases at first token, so it does not bound requests that are past TTFT and still streaming.--concurrencycaps concurrent session trees; within a tree, sub-agent forks can oversubscribe, so total in-flight requests can exceed the number. This is reflected in captured metrics but not in the live dashboard view and not in theconcurrencyspecification (via YAML or CLI), which can lead to some confusion.For agentic / DAG workloads with fan-out, this means the number of requests an engine actually sees concurrently is not directly controllable — which makes it hard to reproduce a specific in-flight load or to protect an engine from fork-driven oversubscription.
Proposed change
Add a
request_concurrencydimension that caps every wire request — roots and sub-agent forks alike — from dispatch until the entire response terminates (not at TTFT). Surface:--request-concurrency N(int ≥ 1): max total in-flight requests (until completed, cancelled, or errored); unset = unbounded (today's behavior).--warmup-request-concurrency N(int ≥ 1): the same cap for the warmup phase; falls back to--request-concurrencywhen unset.request_concurrencyand (AGENTIC_REPLAY only)agentic_warmup_request_concurrency, so the auto-synthesized cache-warmup burst can be throttled independently of profiling intensity.prefill_concurrency <= request_concurrency(mirrors the existingprefill_concurrency <= concurrencycheck).request_concurrencyis sweepable like the other concurrency dimensions.Out of Scope
--concurrencyor--prefill-concurrency--request-rate; this is only about total in-flight requests)Implementation sketch
Because I need the behavior for my own use cases, I already have a feature preview branch on https://github.com/PrimaLabs-AI/aiperf/tree/thomas-primalabs/request-concurrency . I am creating an upstream issue now instead of a PR because I believe this change is too major to directly PR without proper consideration of the AIPerf community first.
High level: A per-phase request-slot semaphore on
ConcurrencyManager(acquire_request_slot/try_acquire_request_slot/release_request_slot), acquired in the credit issuer (both the blocking and non-blocking issue paths, with rollback of the session slot on failure) and released in the credit callback on every return exceptno_requestvirtual credits. Backward compatible (unset = unbounded).Implementation Minutiae
no_requestvirtual credit neither acquires nor releases a request slot so that the semaphore never underflows.ConcurrencyManager:acquire_request_slot(phase_key, can_proceed_fn)(async, blocking),try_acquire_request_slot(...)(sync, non-blocking),release_request_slot(phase_key); a per-phase semaphore sized torequest_concurrencyto cover both synchronous and asynchronous paths.CreditIssuer._issue_credit_ready/try_issue_credit: acquire request slot with session-slot rollback on failure; skip whenturn.no_request. Slot acquisition is factored into an ordered acquire-with-LIFO-rollback helper pair,_acquire_slots(async) /_try_acquire_slots(sync), so the session -> request -> prefill chain (and its rollback) lives in one place.CreditCallbackHandler.on_credit_return:if not credit.no_request: concurrency.release_request_slot(phase).Notes for maintainers (two convention points)
While building my feature branch, I bumped into at least two problems that would matter before upstream adoption.
BasePhaseConfigfield count: the two new phase-config fields (request_concurrency,agentic_warmup_request_concurrency) pushBasePhaseConfigto 31 fields, one over thecheck-ergonomicssoft limit of 30 ("split into sub-models"). I manually bumped the ergonomics baseline entry to keep this PR single-concern, but it merits proper consideration on how the interface should be presented. Are 31 fields acceptable here, or would you prefer a specific field migrate into a sub-model (e.g. a concurrency or agentic-warmup group)? A model split would change the config access path and the YAML/CLI/schema shape.CreditIssuer._issue_credit_readyandtry_issue_creditover the repo's C901 gate (check-ruff-baselined, threshold 10). Rather than ratchet the baseline, the slot-acquisition-with-LIFO-rollback ladder is extracted into a helper pair (_acquire_slots/_try_acquire_slots— an async and a sync variant, since the two issue paths differ only in awaiting vs. non-blocking acquire); both functions are now under the limit.If it is better for me to directly open a PR for discussion of implementation specifics, I can do that.