Replies: 4 comments
|
For your tool-using task, a task-start delay and a provider-request limiter solve different problems. One task can issue several model calls, and retries or evaluators may consume the same provider quota. Even I would put the provider limiter immediately before each outbound attempt, shared by callers using the same quota scope. For a nominal 15 requests/minute, four-second spacing is only a simple starting model: the provider's actual window/burst rules and token limits still matter. A process-local limiter also cannot coordinate multiple experiment workers or other applications using the same quota. For an SDK feature, a useful split would be a task-admission hook for dataset pacing and a separately supplied request limiter in the model client. Cancellation should remove a queued admission, and retries should pass through the limiter again rather than bypassing it. That avoids presenting “tasks/minute” as an API-quota guarantee. On measurement, record queue wait and active execution separately. A parent experiment/task duration can honestly include waiting; a child model-call span should measure the actual attempt. Sleeping inside a fake evaluator makes the evaluator metric misleading and does not control intermediate tool calls. A deterministic test with a fake clock and one task making three calls would expose this distinction quickly. I haven't executed a provider-backed experiment for this check. |
|
Appreciate the feedback. Will be considered going forward. For anyone facing similar issues, please upvote this issue to help get it prioritized. |
|
Looking at @vynazevedo comment, a request per minute limit really seems to be the cleanest solution for rate limiting. But IMO such request-level control should be implemented in the LLM calling lib, not in the experiments runner. Langchain already supports that - "requests_per_second". Maybe it's possible to design some solution around that, e.g. create a The problem with solution above is that Langchain is not the only framework Langfuse can work with. And if I understand correctly, every framework has different approaches on per-request control, and that's a lot of work to do In my situation the task-level rate limiting would be enough because I explicitly control the amount of tool calls my agent can make by mentioning max amount in a prompt and having a hard cap via |
|
For reference, here's how it's implemented in promptfoo evals - https://www.promptfoo.dev/docs/configuration/rate-limits/#fixed-delay |
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Describe the feature or potential improvement
Didn't find the thread that matches my problem so decided to give it a try:
In summary: it would be cool to add an ability to control the rate of experiment task execution (e.g. "max tasks per minute" or a delay setting in
run_experiment)Such functionality would be useful for running experiments on free/budget hosted model providers with strict rate limiting like Google Gemini. Right now you either rely on retries from LLM calling libraries like Langchain or come up with other workarounds (which is not ideal, see "Current workarounds" section)
There is already a
max_concurrencyproperty that controls parallel execution in run_experiment. Setting it to "1" decreases the load, but for tight limits and complex tasks with tools, sending multiple LLM requests, it's not enough. Adding something liketask_seconds_delaywould offer more flexibilityMy example:
gemini-3.1-flash-litefree tier - 15 requests per minutemax_concurrency=1With this setup, I am getting rate limited after 20 requests on average
Current workarounds
python-genai- depends on circumstances. By the time you get rate-limited, some providers already enable additional restrictions. Some tools also may just not retry at reasonable intervals like in my caseAdditional information
Feature is missing both in JS and Python SDK
All reactions