Rate Limits
How did we calculate rate and budget limits for each organization?
Two separate questions, with different answers. An organisation’s limits are derived from its contract, so that what it bought and what it may spend per day agree.
The unit that connects euros to requests
A budget is measured in euros and a rate limit in requests, so translating one into the other needs a conversion factor. Derived from real usage statistics at GWDG: An average request consist of roughly 10,000 tokens, which at around 1 € per million tokens comes to about 0.01 € per request.
This is a statistical average used for rate-limiting, not a billing rule. You are always billed the token-exact cost of the model you actually used. Its only job is to make a budget expressible as a number of requests.
From a daily budget to a request rate limit
The request limit protects the service from overload. It is derived from the daily budget (see Budgets) through the average request above:
Worked through for a daily budget of 400 €:
| Step | Calculation | Result |
|---|---|---|
| Requests the daily budget is worth | 400 € ÷ 0.01 € | 40,000 requests |
| Spread over one hour | 40,000 ÷ 3,600 s | 11.1 requests per second |
| Per 10-second window | × 10 | 111 requests / 10 s |
The step worth explaining is dividing by 3,600 seconds (one hour) rather than 86,400 (one day). Spreading a day’s budget evenly over a day would give about 4.6 requests per 10 seconds. Spreading it over an hour instead gives 111, which is 24 times more generous, and that is the intention: the rate limit is not there to enforce the budget. The budget enforces itself, token-exact, as you spend it. The rate limit exists only to keep any one organisation from overwhelming shared infrastructure, so it is sized to let bursts through and throttle only sustained abuse.
Sizing by headcount instead
Where an organisation has a flat contingent for a number of users rather than a euro budget, the limit is sized from headcount:
An organisation of 60,000 users therefore gets 60 requests per 10 seconds. The divisor assumes that only a small fraction of an institution is ever making a request in the same ten seconds, which is what real usage looks like.
Limits are reviewed
For free-tier use, the exact limits are adjusted periodically according to the capacity actually available. Free access runs on fair-share nodes, which is spare capacity by definition. Limits are also kept sensible relative to pay-per-token pricing, so that a rate limit never makes a contract more expensive to use than paying per token would have been.