# Rate limits & quotas

Every metered endpoint is governed by two independent limits from your organization's plan:

- A quota: the total number of requests allowed in the current billing period.
- A requests-per-minute (RPM) limit: how fast you can send them.

Both are set per endpoint, not for the API as a whole.

## Per-resource buckets

Each endpoint is metered against its own resource, named after its path with dots. Examples:

| Endpoint | Resource |
| --- | --- |
| `POST /v1/search` | `api.v1.search.query` |
| `POST /v1/chat` | `api.v1.chat.message` |
| `GET /v1/datasets` | `api.v1.datasets.list` |
| `POST /v1/datasets/{datasetId}/items/upload` | `api.v1.datasets.items.upload` |
| `GET /v1/datasets/{datasetId}/sources/{sourceId}/sync/jobs` | `api.v1.datasets.sources.syncJobs.list` |

Each endpoint's reference page names its resource, and every response tells you which resource was charged in the `X-RateLimit-Resource` header.

Buckets don't share usage. Running out of search quota doesn't affect chat, and a burst of dataset reads doesn't slow uploads. The streaming and non-streaming forms of an endpoint share one resource.

A few endpoints are not metered by your plan at all: image delivery, events and session-authenticated image upload. They have their own limits, described on their reference pages.

## Quota

Your quota for a resource is a request count for the current billing period. Each request that passes authentication and entitlement checks uses one unit, whether it then succeeds or fails. When the quota reaches zero, the resource returns `429` until the period resets.

A plan can also mark a resource as unlimited. In that case the quota headers read `unlimited` and no quota is used.

## Requests per minute

The RPM limit caps how many requests a resource accepts in a one-minute window. For your own keys the window is counted per organization, so every key and server in the organization shares it.

The window starts with the first request after a quiet period and lasts 60 seconds. Once it expires, the count starts again from zero.

## Order of checks

A request is checked in this order:

1. Authentication.
2. Entitlement: whether your plan includes this resource and allows access.
3. Quota: rejected with `429` if exhausted, otherwise one unit is used.
4. RPM: rejected with `429` if over the limit.
5. The endpoint runs.

Because quota is used before the RPM check, a request rejected for exceeding RPM still counts against your quota. Pacing requests below your RPM limit protects your quota as well as your latency.

## Quota refunds

If a request fails on our side — any `5xx` response, such as `500 Internal server error.` or a `502` when an upstream model fails — the unit of quota it used is refunded.

`4xx` responses are not refunded: validation errors, not-found errors and the like describe the request itself. A streaming request that fails partway through with an `error` event also keeps its charge, because the stream had already started with a `200`.

## Response headers

Metered endpoints return these headers on successful and failed responses alike, once the resource is known:

| Header | Value |
| --- | --- |
| `X-RateLimit-Resource` | The resource charged, for example `api.v1.search.query`. Always present. |
| `X-RateLimit-Quota-Limit` | Requests allowed this billing period, or `unlimited`. |
| `X-RateLimit-Quota-Remaining` | Requests left this billing period, or `unlimited`. |
| `X-RateLimit-Quota-Reset` | Unix time in seconds when the billing period resets. Omitted for unlimited quotas. |
| `X-RateLimit-RPM-Limit` | Requests allowed per minute for this resource. |
| `X-RateLimit-RPM-Remaining` | Requests left in the current window. |
| `X-RateLimit-RPM-Reset` | Unix time in seconds, set to 60 seconds after this response. |
| `Retry-After` | Seconds to wait. Only on `429` responses. |

Headers that don't apply are left out. A resource with no RPM limit sends no RPM headers. Responses rejected during authentication carry only `X-RateLimit-Resource`.

`X-RateLimit-RPM-Reset` is always 60 seconds from the current response, not the true end of the running window. Treat it as an upper bound: the window will have reset by then, and may reset sooner.

Some errors returned before an endpoint's limits are evaluated, such as an unparseable request body, carry no rate-limit headers.

## 429 responses

A `429` body names the limit that was hit and the resource:

```json
{ "error": "Monthly quota exceeded for \"api.v1.search.query\" (10000/10000 used)." }
```

```json
{ "error": "RPM limit exceeded for \"api.v1.search.query\" (61/60 used)." }
```

| Limit | Message format | `Retry-After` |
| --- | --- | --- |
| Quota | `Monthly quota exceeded for "<resource>" (<quota>/<quota> used).` | Seconds until the billing period ends. This can be days. |
| RPM | `RPM limit exceeded for "<resource>" (<used>/<limit> used).` | `60` |

The quota message always says "Monthly", but the quota resets with your billing period. Use `X-RateLimit-Quota-Reset` or `Retry-After` for the actual reset time.

Don't retry an exhausted quota in a loop. Check whether `Retry-After` is small enough to wait for, and if not, fail gracefully or upgrade your plan.

## Related 403 responses

These `403` errors come from the same checks. They mean your plan is not set up for the request, not that you are sending too fast, so waiting won't help:

| Message | Meaning |
| --- | --- |
| `No active subscription for this organization.` | The organization has no active subscription. |
| `No quota found for organization "<organizationId>".` | No quota exists for the current billing period. |
| `No entitlement for "<resource>".` | The plan doesn't include this endpoint. |
| `Access denied to "<resource>".` | The plan includes this endpoint but has it switched off. |

If two requests race for the last unit of quota, the loser receives the same `429` quota-exceeded response as any other request over the limit.

## Demo key limits

Requests made with the demo key use the demo plan's limits instead of yours. Their RPM window is counted per client IP address rather than per organization, and their quota is shared by everyone using the demo key. See [Demo mode](demo-mode).

## Backoff

Pace requests below your RPM limit, and back off when you get a `429`, `503` or transient `500`.

JavaScript:

```javascript
const sleep = (ms) => new Promise((resolve) => setTimeout(resolve, ms));

async function withBackoff(request, { maxAttempts = 5, maxWaitSeconds = 120 } = {}) {
  for (let attempt = 1; ; attempt++) {
    const res = await request();
    const retryable = res.status === 429 || res.status === 503 || res.status >= 500;
    if (!retryable || attempt === maxAttempts) return res;

    const retryAfter = Number(res.headers.get("Retry-After"));
    const backoff = Math.min(2 ** attempt, 30) + Math.random();
    const waitSeconds = Number.isFinite(retryAfter) && retryAfter > 0 ? retryAfter : backoff;

    // An exhausted quota can have a Retry-After of days. Give up instead of waiting.
    if (waitSeconds > maxWaitSeconds) return res;

    await sleep(waitSeconds * 1000);
  }
}

const res = await withBackoff(() =>
  fetch("https://stylor.ai/api/v1/search", {
    method: "POST",
    headers: {
      Authorization: `Bearer ${process.env.STYLOR_API_KEY}`,
      "Content-Type": "application/json",
    },
    body: JSON.stringify({ query: "navy wool blazer", limit: 10 }),
  })
);

console.log(res.status, res.headers.get("X-RateLimit-RPM-Remaining"));
```

Python:

```python
import os
import random
import time
import requests


def with_backoff(send, max_attempts=5, max_wait_seconds=120):
    for attempt in range(1, max_attempts + 1):
        res = send()
        retryable = res.status_code in (429, 503) or res.status_code >= 500
        if not retryable or attempt == max_attempts:
            return res

        retry_after = res.headers.get("Retry-After")
        backoff = min(2 ** attempt, 30) + random.random()
        wait_seconds = int(retry_after) if retry_after and retry_after.isdigit() else backoff

        # An exhausted quota can have a Retry-After of days. Give up instead of waiting.
        if wait_seconds > max_wait_seconds:
            return res

        time.sleep(wait_seconds)


res = with_backoff(lambda: requests.post(
    "https://stylor.ai/api/v1/search",
    headers={"Authorization": f"Bearer {os.environ['STYLOR_API_KEY']}"},
    json={"query": "navy wool blazer", "limit": 10},
    timeout=60,
))
print(res.status_code, res.headers.get("X-RateLimit-RPM-Remaining"))
```

Retrying a write can repeat its effect. See the retry advice in [Errors](errors) before wrapping uploads or deletes.
