Toucora API
🌐 Language

API status & changelog

What uptime and response time to expect, the limits that apply today, and how we roll out changes. The machine-readable version is at /api/status.json.

Current status: loading live status…

Latency & error SLOs

These are the service-level objectives we hold ourselves to. They are targets, not a contractual SLA: the API is in early access and provided as available. We alert on the objectives that have a live in-process signal (egress probe latency and consecutive extraction failures) and monitor the rest from request telemetry.

ObjectiveTargetWindowHow it is evaluated
API availability
Share of API requests that return a response other than a 5xx, over a 30-day rolling window.
99.5% of requests succeed 30-day rolling non_5xx_ratio >= 99.5%
Response time — cached and metadata endpoints
Cached transcripts, /api/health, /api/plans and other metadata reads. Liveness must never block on a third-party call.
p95 ≤ 500 ms 30-day rolling p95 <= 500 ms
Response time — uncached extraction
A first-time transcript extraction (YouTube caption fetch) for a single public video. Cached repeats are served from the response-time objective above.
p95 ≤ 8 s 30-day rolling p95 <= 8 s
YouTube egress probe latency
Latency of the background egress probe reported at /api/health and /api/health/egress. A rising probe is an early warning that extraction will degrade.
p95 ≤ 2,000 ms rolling probe probe latency <= 2,000 ms
Consecutive extraction failures
Consecutive failed extractions before an operator alert fires. Any run of three is treated as a dependency incident, not a user error.
0 (alert at 3) live consecutive_failures < 3
Server error rate
Share of API requests that return a 5xx. Expected extraction failures (blocked/unavailable videos) return a mapped 4xx and are not counted here.
< 1% of requests are 5xx 30-day rolling 5xx_ratio < 1%
Liveness response time
GET /api/health and GET /api/health/ready must answer immediately; dependency state is reported, not awaited, on the liveness path.
≤ 100 ms live liveness responds <= 100 ms

Health endpoints: GET /api/health is a liveness check that always answers immediately and never blocks on a third-party call. GET /api/health/ready is a readiness check that returns 503 when the YouTube egress is unavailable, so orchestrators can take a broken instance out of rotation. GET /api/health/egress runs one bounded egress probe on demand.

Current limits

Rate-limit policy

Rate limits are per API key (or per visitor/IP for anonymous traffic) and are returned on every response:

HeaderMeaning
X-RateLimit-LimitThe request ceiling for the current minute.
X-RateLimit-RemainingRequests left in the current window.
X-RateLimit-ResetUnix time when the window resets.
Retry-AfterSeconds to wait, sent with a 429.

When you exceed the limit the API returns 429 RATE_LIMITED with a Retry-After header. We recommend exponential backoff with jitter and honoring Retry-After. A burst above the limit is rejected, not queued.

Deprecation policy

Changelog

Incidents

This page and /api/status.json are the source of truth for availability. If something looks wrong, include the X-Request-ID from the failing response and send us a message — that is usually enough to reproduce it.