Rate limits
Este conteúdo não está disponível em sua língua ainda.
The BenchKey API enforces rate limits to protect platform stability and ensure fair usage across tenants. Limits apply per API key.
Limits
Section titled “Limits”The API uses a token-bucket model per API key: a burst capacity of 100 requests that refills at 10 requests per second. This allows short bursts up to 100 in-flight requests while sustaining no more than 10 req/s over time. Both live (bk_live_) and test (bk_test_) keys have independent limits.
Rate-limit headers
Section titled “Rate-limit headers”Every response to a request made with a valid API key includes headers that describe your current usage. Requests rejected before the key is verified (a missing or invalid key, for example) and the unauthenticated discovery endpoints do not carry them.
| Header | Description |
|---|---|
RateLimit-Limit | Bucket capacity (maximum burst size) |
RateLimit-Remaining | Tokens remaining in the bucket right now |
RateLimit-Reset | Seconds until the bucket is full again (on a 429, the same value as Retry-After) |
Retry-After | Seconds to wait before retrying (only present on 429 responses) |
Example response headers
HTTP/1.1 200 OKRateLimit-Limit: 100RateLimit-Remaining: 87RateLimit-Reset: 2When you are rate-limited
Section titled “When you are rate-limited”If you exceed the limit the API returns 429 Too Many Requests:
{ "error": { "type": "rate_limit_error", "code": "rate_limited", "message": "API rate limit exceeded for this key.", "request_id": "req_01J8X4zzzzzz" }}The Retry-After header tells you exactly how many seconds to wait.
Exponential backoff
Section titled “Exponential backoff”For any 429 or transient 5xx response, use exponential backoff with jitter. Do not hammer the API in a tight retry loop.
curl (shell script)
Section titled “curl (shell script)”BENCHKEY_API_KEY="bk_live_acme_xxxxxxxxxxxx"MAX_ATTEMPTS=5ATTEMPT=0BACKOFF=1HEADERS=$(mktemp)
while [ $ATTEMPT -lt $MAX_ATTEMPTS ]; do RESPONSE=$(curl -s -D "$HEADERS" -w "\n%{http_code}" \ https://app.benchkey.com/api/v1/customers \ -H "Authorization: Bearer $BENCHKEY_API_KEY")
HTTP_CODE=$(echo "$RESPONSE" | tail -n 1) BODY=$(echo "$RESPONSE" | sed '$d')
if [ "$HTTP_CODE" -eq 200 ]; then echo "$BODY" break elif [ "$HTTP_CODE" -eq 429 ] || [ "$HTTP_CODE" -ge 500 ]; then RETRY_AFTER=$(grep -i '^retry-after:' "$HEADERS" | tr -dc '0-9') WAIT=${RETRY_AFTER:-$BACKOFF} echo "Rate limited or server error. Waiting ${WAIT}s..." >&2 sleep "$WAIT" BACKOFF=$((BACKOFF * 2)) ATTEMPT=$((ATTEMPT + 1)) else echo "Fatal error $HTTP_CODE: $BODY" >&2 break fidone
rm -f "$HEADERS"Node.js
Section titled “Node.js”async function fetchWithRetry(url, options, maxAttempts = 5) { let backoff = 1000; // ms
for (let attempt = 0; attempt < maxAttempts; attempt++) { const res = await fetch(url, options);
if (res.ok) return res.json();
if (res.status === 429 || res.status >= 500) { const retryAfter = res.headers.get("Retry-After"); const wait = retryAfter ? parseInt(retryAfter, 10) * 1000 : backoff; const jitter = Math.random() * 200; await new Promise((r) => setTimeout(r, wait + jitter)); backoff = Math.min(backoff * 2, 30_000); continue; }
const { error } = await res.json(); throw new Error(`[${error.code}] ${error.message}`); }
throw new Error("Max retry attempts reached");}
// Usageconst data = await fetchWithRetry( "https://app.benchkey.com/api/v1/customers", { headers: { Authorization: `Bearer ${process.env.BENCHKEY_API_KEY}`, }, });console.log(data);Best practices
Section titled “Best practices”- Batch reads: use
limit=100(the maximum) to reduce the number of list requests. - Cache aggressively: avoid re-fetching resources that haven’t changed. Store the
idand re-read only when needed. - Parallelize carefully: fan-out requests in parallel, but cap concurrent in-flight requests to avoid burst spikes (aim for ≤ 10 concurrent).
- Check headers proactively: if
RateLimit-Remainingdrops near zero, slow down before hitting429. - Use idempotency keys on writes: if a write request times out and you retry, an Idempotency-Key prevents duplicate side effects.
Bulk imports and backfills
Section titled “Bulk imports and backfills”There are no bulk endpoints. A batch write is a loop of single requests with persisted operation keys and results. This lets you track each record’s outcome and resume within the idempotency replay limits:
- Persist an
Idempotency-Keyand the request for each operation before sending (e.g.import-cust-8841-create). Reuse them after a crash. Completed responses are normally retained for 24 hours, so identical retries within that window replay the saved result. Outside it, reconcile your saved results and the destination records before retrying; restarting from the top is not guaranteed to avoid duplicates. Use a new key for a distinct update to the same source record. - Pace by the headers. Watch
RateLimit-Remainingand ease off as it approaches zero. Use the returned limit and reset values to pace the backfill; completion time depends on the requests and available capacity. - On
429, waitRetry-Afterseconds and resend the same request with the same key.
For a full migration from another system (RepairDesk, RepairShopr, CSV), use the in-app importer instead, it handles dedupe, staging review, and per-entity selection in ways a REST loop can’t.
Self-service usage introspection
Section titled “Self-service usage introspection”GET /api/v1/usageReturns request volume, error rate, and latency for the API key making the call over a configurable time window. It never includes your shop’s other keys, even with by_key=true. Useful for an integration’s own dashboards and alerting.
Usage across all of your shop’s keys is not available through the public API. Owners and admins can see it, broken down by key, on Settings → API keys in BenchKey.
Auth required: any valid API key. No specific scope is needed, and this endpoint stays available even when your plan or account status blocks other resources, so you can still diagnose your own traffic.
What is counted
Section titled “What is counted”Counts come from BenchKey’s per-request API log. A request is logged only after it passes API-key authentication and the rate limiter. For the calling key, the log records:
- every write (
POST,PATCH,PUT,DELETE), whatever its outcome; - every
4xxand5xxresponse returned after those checks, such as validation, scope, plan, and server errors; - successful reads. All of them are logged by default; if read sampling is turned on for the server, successful reads are counted from a sample.
These are not logged, so they never appear in total_requests, error_count, or error_rate:
- requests rejected with
429by the rate limiters: the per-key token bucket described above, and the per-IP limit on repeated authentication failures; - requests rejected before or during API-key authentication: a missing, malformed, revoked, or expired key (
401), or a write body sent withoutContent-Type: application/json(415); - the unauthenticated discovery endpoints (
/api/v1,/api/v1/health,/api/v1/openapi.json,/api/v1/openapi.yaml).
An error_rate of 0 does not mean the key was never throttled. Watch the RateLimit-* and Retry-After headers to detect throttling.
Query parameters
Section titled “Query parameters”| Parameter | Type | Default | Description |
|---|---|---|---|
window_days | integer | 30 | Lookback in days, ending at until. Ignored when since is set |
since | integer (epoch ms) | none | Window start (inclusive). Overrides window_days |
until | integer (epoch ms) | now | Window end (exclusive) |
by_key | boolean | false | true or 1 adds a by_key array for the calling key |
limit | integer | 100 | Maximum by_key entries (at most 1000). by_key only covers the calling key, so it never has more than one entry |
Malformed values are ignored and the default applies; they do not return an error.
# Last 7 days for the key making the callcurl "https://app.benchkey.com/api/v1/usage?window_days=7" \ -H "Authorization: Bearer bk_live_<tenantId>_<secret>"
# Explicit window, plus the calling key's by_key entrycurl "https://app.benchkey.com/api/v1/usage?since=1748822400000&until=1749081600000&by_key=true" \ -H "Authorization: Bearer bk_live_<tenantId>_<secret>"Node.js
Section titled “Node.js”const res = await fetch("https://app.benchkey.com/api/v1/usage?window_days=7", { headers: { Authorization: `Bearer ${process.env.BENCHKEY_API_KEY}` },});const usage = await res.json();console.log(`${usage.total_requests} requests, error rate ${usage.error_rate}`);Response
Section titled “Response”{ "object": "api_usage", "tenant_id": "acme", "window": { "since": "2026-06-07T00:00:00.000Z", "until": "2026-06-14T00:00:00.000Z" }, "total_requests": 4821, "by_status_class": { "2xx": 4710, "4xx": 108, "5xx": 3, "other": 0 }, "error_count": 111, "error_rate": 0.023, "duration_ms": { "total": 721500, "avg": 150, "max": 2340 }, "distinct_keys": 1}With by_key=true the response also includes a by_key array. It holds at most one entry, for the calling key, and is empty when that key has no logged requests in the window. The entry has fewer fields than the top level:
{ "by_key": [ { "api_key_id": "ak_a1b2c3d4e5f6a1b2c3d4e5f6", "total_requests": 4821, "error_count": 111, "error_rate": 0.023, "avg_duration_ms": 150 } ]}| Field | Type | Description |
|---|---|---|
window.since / window.until | string (ISO 8601) | Resolved window boundaries (start inclusive, end exclusive) |
total_requests | integer | Logged requests made with this key in the window |
by_status_class | object | Those requests by HTTP status class. other counts any status outside 2xx, 4xx, and 5xx |
error_count | integer | 4xx + 5xx responses among them |
error_rate | number | error_count / total_requests (0–1), rounded to 4 decimal places |
duration_ms.total | integer | Summed server handling time in milliseconds |
duration_ms.avg | integer | Mean server handling time in milliseconds, rounded |
duration_ms.max | integer | Slowest logged request in the window, in milliseconds |
distinct_keys | integer | 1 when this key has logged requests in the window, otherwise 0 |
by_key[].api_key_id | string | The calling key’s ID, the same id that GET /api/v1/me returns |
by_key[].total_requests, error_count, error_rate | integer, integer, number | Same meaning as the top-level fields |
by_key[].avg_duration_ms | integer | Mean server handling time in milliseconds, rounded |
Fresh tenant / no traffic: A window with no logged requests returns zero counts rather than an error. If the usage log cannot be read at all (for example, on a brand-new shop), the endpoint still answers
200with zero counts andnullwindow boundaries, so dashboards can poll it immediately after provisioning. If the rollup runs past its time budget, the endpoint answers504with codeapi_usage_timeout; retry with a shorter window.