Pular para o conteúdo

Rate limits

Este conteúdo não está disponível em sua língua ainda.

The BenchKey API enforces rate limits to protect platform stability and ensure fair usage across tenants. Limits apply per API key.

The API uses a token-bucket model per API key: a burst capacity of 100 requests that refills at 10 requests per second. This allows short bursts up to 100 in-flight requests while sustaining no more than 10 req/s over time. Both live (bk_live_) and test (bk_test_) keys have independent limits.

Every response to a request made with a valid API key includes headers that describe your current usage. Requests rejected before the key is verified (a missing or invalid key, for example) and the unauthenticated discovery endpoints do not carry them.

HeaderDescription
RateLimit-LimitBucket capacity (maximum burst size)
RateLimit-RemainingTokens remaining in the bucket right now
RateLimit-ResetSeconds until the bucket is full again (on a 429, the same value as Retry-After)
Retry-AfterSeconds to wait before retrying (only present on 429 responses)

Example response headers

HTTP/1.1 200 OK
RateLimit-Limit: 100
RateLimit-Remaining: 87
RateLimit-Reset: 2

If you exceed the limit the API returns 429 Too Many Requests:

{
"error": {
"type": "rate_limit_error",
"code": "rate_limited",
"message": "API rate limit exceeded for this key.",
"request_id": "req_01J8X4zzzzzz"
}
}

The Retry-After header tells you exactly how many seconds to wait.

For any 429 or transient 5xx response, use exponential backoff with jitter. Do not hammer the API in a tight retry loop.

Terminal window
BENCHKEY_API_KEY="bk_live_acme_xxxxxxxxxxxx"
MAX_ATTEMPTS=5
ATTEMPT=0
BACKOFF=1
HEADERS=$(mktemp)
while [ $ATTEMPT -lt $MAX_ATTEMPTS ]; do
RESPONSE=$(curl -s -D "$HEADERS" -w "\n%{http_code}" \
https://app.benchkey.com/api/v1/customers \
-H "Authorization: Bearer $BENCHKEY_API_KEY")
HTTP_CODE=$(echo "$RESPONSE" | tail -n 1)
BODY=$(echo "$RESPONSE" | sed '$d')
if [ "$HTTP_CODE" -eq 200 ]; then
echo "$BODY"
break
elif [ "$HTTP_CODE" -eq 429 ] || [ "$HTTP_CODE" -ge 500 ]; then
RETRY_AFTER=$(grep -i '^retry-after:' "$HEADERS" | tr -dc '0-9')
WAIT=${RETRY_AFTER:-$BACKOFF}
echo "Rate limited or server error. Waiting ${WAIT}s..." >&2
sleep "$WAIT"
BACKOFF=$((BACKOFF * 2))
ATTEMPT=$((ATTEMPT + 1))
else
echo "Fatal error $HTTP_CODE: $BODY" >&2
break
fi
done
rm -f "$HEADERS"
async function fetchWithRetry(url, options, maxAttempts = 5) {
let backoff = 1000; // ms
for (let attempt = 0; attempt < maxAttempts; attempt++) {
const res = await fetch(url, options);
if (res.ok) return res.json();
if (res.status === 429 || res.status >= 500) {
const retryAfter = res.headers.get("Retry-After");
const wait = retryAfter ? parseInt(retryAfter, 10) * 1000 : backoff;
const jitter = Math.random() * 200;
await new Promise((r) => setTimeout(r, wait + jitter));
backoff = Math.min(backoff * 2, 30_000);
continue;
}
const { error } = await res.json();
throw new Error(`[${error.code}] ${error.message}`);
}
throw new Error("Max retry attempts reached");
}
// Usage
const data = await fetchWithRetry(
"https://app.benchkey.com/api/v1/customers",
{
headers: {
Authorization: `Bearer ${process.env.BENCHKEY_API_KEY}`,
},
}
);
console.log(data);
  • Batch reads: use limit=100 (the maximum) to reduce the number of list requests.
  • Cache aggressively: avoid re-fetching resources that haven’t changed. Store the id and re-read only when needed.
  • Parallelize carefully: fan-out requests in parallel, but cap concurrent in-flight requests to avoid burst spikes (aim for ≤ 10 concurrent).
  • Check headers proactively: if RateLimit-Remaining drops near zero, slow down before hitting 429.
  • Use idempotency keys on writes: if a write request times out and you retry, an Idempotency-Key prevents duplicate side effects.

There are no bulk endpoints. A batch write is a loop of single requests with persisted operation keys and results. This lets you track each record’s outcome and resume within the idempotency replay limits:

  1. Persist an Idempotency-Key and the request for each operation before sending (e.g. import-cust-8841-create). Reuse them after a crash. Completed responses are normally retained for 24 hours, so identical retries within that window replay the saved result. Outside it, reconcile your saved results and the destination records before retrying; restarting from the top is not guaranteed to avoid duplicates. Use a new key for a distinct update to the same source record.
  2. Pace by the headers. Watch RateLimit-Remaining and ease off as it approaches zero. Use the returned limit and reset values to pace the backfill; completion time depends on the requests and available capacity.
  3. On 429, wait Retry-After seconds and resend the same request with the same key.

For a full migration from another system (RepairDesk, RepairShopr, CSV), use the in-app importer instead, it handles dedupe, staging review, and per-entity selection in ways a REST loop can’t.


GET /api/v1/usage

Returns request volume, error rate, and latency for the API key making the call over a configurable time window. It never includes your shop’s other keys, even with by_key=true. Useful for an integration’s own dashboards and alerting.

Usage across all of your shop’s keys is not available through the public API. Owners and admins can see it, broken down by key, on Settings → API keys in BenchKey.

Auth required: any valid API key. No specific scope is needed, and this endpoint stays available even when your plan or account status blocks other resources, so you can still diagnose your own traffic.

Counts come from BenchKey’s per-request API log. A request is logged only after it passes API-key authentication and the rate limiter. For the calling key, the log records:

  • every write (POST, PATCH, PUT, DELETE), whatever its outcome;
  • every 4xx and 5xx response returned after those checks, such as validation, scope, plan, and server errors;
  • successful reads. All of them are logged by default; if read sampling is turned on for the server, successful reads are counted from a sample.

These are not logged, so they never appear in total_requests, error_count, or error_rate:

  • requests rejected with 429 by the rate limiters: the per-key token bucket described above, and the per-IP limit on repeated authentication failures;
  • requests rejected before or during API-key authentication: a missing, malformed, revoked, or expired key (401), or a write body sent without Content-Type: application/json (415);
  • the unauthenticated discovery endpoints (/api/v1, /api/v1/health, /api/v1/openapi.json, /api/v1/openapi.yaml).

An error_rate of 0 does not mean the key was never throttled. Watch the RateLimit-* and Retry-After headers to detect throttling.

ParameterTypeDefaultDescription
window_daysinteger30Lookback in days, ending at until. Ignored when since is set
sinceinteger (epoch ms)noneWindow start (inclusive). Overrides window_days
untilinteger (epoch ms)nowWindow end (exclusive)
by_keybooleanfalsetrue or 1 adds a by_key array for the calling key
limitinteger100Maximum by_key entries (at most 1000). by_key only covers the calling key, so it never has more than one entry

Malformed values are ignored and the default applies; they do not return an error.

Terminal window
# Last 7 days for the key making the call
curl "https://app.benchkey.com/api/v1/usage?window_days=7" \
-H "Authorization: Bearer bk_live_<tenantId>_<secret>"
# Explicit window, plus the calling key's by_key entry
curl "https://app.benchkey.com/api/v1/usage?since=1748822400000&until=1749081600000&by_key=true" \
-H "Authorization: Bearer bk_live_<tenantId>_<secret>"
const res = await fetch("https://app.benchkey.com/api/v1/usage?window_days=7", {
headers: { Authorization: `Bearer ${process.env.BENCHKEY_API_KEY}` },
});
const usage = await res.json();
console.log(`${usage.total_requests} requests, error rate ${usage.error_rate}`);
{
"object": "api_usage",
"tenant_id": "acme",
"window": {
"since": "2026-06-07T00:00:00.000Z",
"until": "2026-06-14T00:00:00.000Z"
},
"total_requests": 4821,
"by_status_class": {
"2xx": 4710,
"4xx": 108,
"5xx": 3,
"other": 0
},
"error_count": 111,
"error_rate": 0.023,
"duration_ms": {
"total": 721500,
"avg": 150,
"max": 2340
},
"distinct_keys": 1
}

With by_key=true the response also includes a by_key array. It holds at most one entry, for the calling key, and is empty when that key has no logged requests in the window. The entry has fewer fields than the top level:

{
"by_key": [
{
"api_key_id": "ak_a1b2c3d4e5f6a1b2c3d4e5f6",
"total_requests": 4821,
"error_count": 111,
"error_rate": 0.023,
"avg_duration_ms": 150
}
]
}
FieldTypeDescription
window.since / window.untilstring (ISO 8601)Resolved window boundaries (start inclusive, end exclusive)
total_requestsintegerLogged requests made with this key in the window
by_status_classobjectThose requests by HTTP status class. other counts any status outside 2xx, 4xx, and 5xx
error_countinteger4xx + 5xx responses among them
error_ratenumbererror_count / total_requests (0–1), rounded to 4 decimal places
duration_ms.totalintegerSummed server handling time in milliseconds
duration_ms.avgintegerMean server handling time in milliseconds, rounded
duration_ms.maxintegerSlowest logged request in the window, in milliseconds
distinct_keysinteger1 when this key has logged requests in the window, otherwise 0
by_key[].api_key_idstringThe calling key’s ID, the same id that GET /api/v1/me returns
by_key[].total_requests, error_count, error_rateinteger, integer, numberSame meaning as the top-level fields
by_key[].avg_duration_msintegerMean server handling time in milliseconds, rounded

Fresh tenant / no traffic: A window with no logged requests returns zero counts rather than an error. If the usage log cannot be read at all (for example, on a brand-new shop), the endpoint still answers 200 with zero counts and null window boundaries, so dashboards can poll it immediately after provisioning. If the rollup runs past its time budget, the endpoint answers 504 with code api_usage_timeout; retry with a shorter window.

Status do sistema