Request
How a request is checked
Every request passes three gates before it reaches a model: the API key it was made with, the user account that key belongs to, and the organisation that account belongs to. Each gate holds its own set of limits, split into the two kinds above: a request rate, counted in requests, and a budget, counted in euros. The request is measured against every limit that is set at that level, and it is rejected the moment any one of them is out of allowance, without the later gates being consulted. Only when all three gates pass does the request reach a model, and its cost is then written back to the budgets of all three levels at once.
Where each gate’s limits come from differs. An API key and an organisation are subject only to limits set on them directly, so if none is set, that gate simply lets the request through. Your user account is the one gate that always has a limit: your own if you have one, otherwise your organisation’s default for its members, otherwise the platform default.
- A level with no limit set simply adds no restriction; it does not exempt you from the levels after it. What you can actually do is whichever of the three limits is smallest, so a limit on your API key can only restrict you further. It can never grant you more than your account or your organisation allows.
- From the ChatAI the first box does not apply: no API key is involved, so the request starts at your user account.
- Request rate limits are evaluated across all three levels before budgets are, so a caller who is out of budget is still paced rather than free to retry.
- On a rejection, the
x-ratelimit-scopeheader names which of the three stopped you.