Toucora API
🌐 Language
Reference · v1

API documentation

Base URL: https://api.example.com/v1. All endpoints accept and return JSON unless a format is requested.

Introduction

Toucora turns YouTube captions into a normalized transcript object. The same engine powers the consumer website, so extraction, caching, and normalization behave identically for both.

{
  "video_id": "VIDEO_ID",
  "title": "Video title",
  "channel": "Channel name",
  "url": "https://youtube.com/watch?v=VIDEO_ID",
  "duration": 1234,
  "language": "en",
  "source": "youtube_auto",
  "segments": [{ "start": 12.42, "duration": 3.2, "text": "Example transcript text." }]
}

Quickstart

  1. Create a free account on the consumer site.
  2. Open the dashboard and create an API key.
  3. Call the API:
curl "https://api.example.com/v1/transcripts/dQw4w9WgXcQ" \
  -H "Authorization: Bearer YOUR_API_KEY"

Authentication

Send your key in the Authorization header. X-API-Key is also accepted.

Authorization: Bearer YOUR_API_KEY

Keys are hashed at rest and shown only once at creation. You can keep multiple keys per account and revoke any of them. Never expose keys in client-side code.

Single video

GET/v1/transcripts/{videoId}

ParameterInDescription
videoIdpathYouTube video id or a full YouTube URL.
langqueryLanguage code, e.g. en, es.
formatqueryjson (default), txt, srt, vtt, markdown.
timestampsquerytrue prefixes each txt line with its timestamp. txt drops timestamps by default.
cleanquerytrue strips musical note symbols and non-speech cues such as [Music] and (laughs).
Tip — response headers include X-Transcript-Language, X-Transcript-Source, and X-Credits-Charged.

Batch

POST/v1/transcripts with up to 50 ids. Each item returns its own status; one failure never invalidates the batch.

curl -X POST "https://api.example.com/v1/transcripts" \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -d '{"ids":["video1","video2"],"lang":"en"}'

Languages

GET /v1/transcripts/{videoId}/languages lists the caption tracks available. If the requested language is unavailable the API returns LANGUAGE_NOT_AVAILABLE and never silently switches languages.

Formats

Append ?format=srt to any single-transcript request to get subtitles. SRT and VTT preserve the original timing.

For plain text, format=txt returns one line per caption segment with no timestamps. Add timestamps=true to prefix every line with [mm:ss], and clean=true to remove musical notes (♪ ♫ 🎶) and non-speech cues like [Music], [Applause] or (laughs). Both options also apply to markdown, srt and vtt, and to the JSON response body and the batch endpoint.

curl "https://api.example.com/v1/transcripts/VIDEO_ID?format=txt&timestamps=true&clean=true" \
  -H "Authorization: Bearer YOUR_API_KEY"

Playlists

POST /v1/playlists/{playlistId}/transcripts resolves the playlist and returns an asynchronous job. Deleted, private, or captionless videos are recorded as failed items while the job continues.

Channels

POST /v1/channels/{channelId}/transcripts accepts a channel id, handle, or URL with a configurable limit.

Jobs

Create a job with POST /v1/jobs and poll GET /v1/jobs/{jobId}. Results are at GET /v1/jobs/{jobId}/results. Cancel an in-flight job with POST /v1/jobs/{jobId}/cancel and re-queue failed items with POST /v1/jobs/{jobId}/retry.

{ "job_id": "job_123", "status": "processing" }

States: queued, processing, completed, completed_with_errors, failed, cancelled.

Webhooks

Pass webhook_url when creating a job, or register a webhook in the dashboard. Events are job.completed and job.failed. Payloads are signed with X-Toucora-Signature when a secret is configured.

Ask

POST/v1/ask — grounded question answering with timestamp citations.

curl -X POST "https://api.example.com/v1/ask" \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -d '{"video_id":"VIDEO_ID","question":"What does the speaker say about AI agents?"}'

Summarize

POST /v1/summarize with an optional length of short, medium, or detailed.

Chapters

POST /v1/chapters returns timestamped chapters. Every timestamp is a real position in the transcript.

POST /v1/search with a query returns timestamped matches. Search combines exact, fuzzy, and semantic matching.

Quotes

POST /v1/quotes extracts notable standalone quotes from a transcript, each with its start time and a timestamp. Quotes are copied verbatim from the captions.

Repurpose

POST /v1/repurpose turns a transcript into another format. Set format to one of blog, newsletter, linkedin, x_thread, youtube_description, or show_notes.

curl -X POST "https://api.example.com/v1/repurpose" \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -d '{"video_id":"VIDEO_ID","format":"blog"}'

Clips

POST /v1/clips finds candidate short-form clips with an opening line, a plain-language label, why it works, and its timestamp range.

Translate

POST /v1/translate translates a transcript into one of the supported languages. Set mode to translated (default), bilingual, or original. When no AI provider is configured the original transcript is returned with "fallback": true.

AI status

GET /v1/ai/status reports whether AI-backed responses are enabled on this deployment. When AI is disabled, the AI endpoints still return useful extractive fallbacks with "fallback": true.

ask, summarize, chapters and repurpose accept an optional output_lang (e.g. en, es, de) to write the response in a language different from the transcript, exactly like the consumer site.

Usage & limits

GET /v1/usage returns credits remaining, requests, success/failure counts, and usage by kind. Rate-limited responses use HTTP 429 with Retry-After and X-RateLimit-* headers.

Errors

All errors share one shape and include a request_id.

{
  "error": {
    "code": "VIDEO_NOT_FOUND",
    "message": "The requested YouTube video could not be found.",
    "request_id": "req_123"
  }
}

Common codes: INVALID_API_KEY, RATE_LIMITED, INSUFFICIENT_CREDITS, INVALID_VIDEO_ID, VIDEO_UNAVAILABLE, TRANSCRIPT_NOT_AVAILABLE, LANGUAGE_NOT_AVAILABLE, JOB_NOT_FOUND, INVALID_REQUEST.

Examples & SDKs


    

Official SDKs for JavaScript, Python, and Go are on the roadmap after the API stabilizes.

MCP server

The API also speaks the Model Context Protocol over Streamable HTTP, so MCP clients can drive the same transcript and AI tools with your API key.

PropertyValue
Endpointhttps://api.example.com/mcp
TransportStreamable HTTP (JSON-RPC 2.0)
AuthAuthorization: Bearer YOUR_API_KEY

Tools: get_transcript, summarize, ask, chapters, search, quotes, clips, repurpose, translate. Each call is metered against your plan like the REST endpoints.

Claude Code

claude mcp add --transport http toucora "https://api.example.com/mcp" \
  --header "Authorization: Bearer YOUR_API_KEY"

Codex

Add a server to ~/.codex/config.toml:

[mcp_servers.toucora]
url = "https://api.example.com/mcp"
http_headers = { "Authorization" = "Bearer YOUR_API_KEY" }

Claude Desktop

Add a connector in Settings → Connectors → Add custom connector with the URL https://api.example.com/mcp and a request header Authorization: Bearer YOUR_API_KEY. For clients without native remote-header support, bridge it with mcp-remote:

{
  "mcpServers": {
    "toucora": {
      "command": "npx",
      "args": ["mcp-remote", "https://api.example.com/mcp", "--header", "Authorization: Bearer YOUR_API_KEY"]
    }
  }
}

Once connected, ask your client to “summarize this video: <YouTube link>” or “quote the exact lines about pricing in <link>”. The client calls get_transcript and the AI tools automatically.