Skip to main content

AI coding · Billing

Rate Limit

Also called: Rate limiting, 429 error, Too Many Requests, Quota

The service limits how many requests you can send or how many tokens you can use in a period of time. Go over and it temporarily turns you away, usually with a 429 error code.

In detail

A rate limit is like a busy restaurant limiting orders. The kitchen can only make so much, so each table can only order so much per hour, and if you order more, they ask you to wait. Large model services usually limit requests per minute, tokens per minute, daily usage, and so on. When you go over, they return a 429 status code or a message like "rate limit exceeded."

You might run into it in two situations. One is when your AI coding tool runs out of quota and you have to wait a bit or switch models. The other is when the product you built calls a large model API and hits the limit once you have more users.

In the second case, your code needs to handle it. On a 429, wait a bit and retry, making the wait longer each time (exponential backoff). Show users a friendly message in the UI. And add Debounce on the frontend so users clicking over and over don't fire off a pile of requests.

Developer info
Term ID
ai-rate-limit
DOM selectors
No DOM cues. This concept isn't detected directly on a page.
Priority
1 · when several match at the same level, the higher priority wins
Version
v1 · updated Sep 29, 2026