by samson-art
Provides transcripts, chapters, metadata, subtitles and single video frames from YouTube and ten other video platforms to AI assistants via a Model Context Protocol (MCP) server.
Transcriptor MCP is an MCP server that extracts textual and visual data from videos—such as clean transcripts, raw subtitles, chapter lists, metadata, and single-frame images—and makes them available to AI assistants like Claude, ChatGPT, Cursor, and any other MCP‑compatible client. It supports eleven video platforms, including YouTube, TikTok, Twitch, Vimeo, and more.
https://transcriptor.gateway.mcpal.io/mcp (or add the URL in the client’s connector settings).claude mcp add --transport http transcriptor https://transcriptor.gateway.mcpal.io/mcp then invoke /mcp in Claude.@Transcriptor in a conversation.artsamsonov/transcriptor-mcp:latest (port 4200) and point any MCP client to http://localhost:4200/mcp.get_transcript, get_video_info, get_video_frame, etc.) in your prompts; responses include plain text and structured JSON.get_transcript, get_raw_subtitles, get_available_subtitles, get_video_info, get_video_chapters, get_video_frame, get_playlist_transcripts, search_videos.yt-dlp, ffmpeg, and optional Whisper transcription./health and /metrics (Prometheus format).Q: Which video platforms are supported? A: Eleven platforms – YouTube, Twitter/X, Instagram, TikTok, Twitch, Vimeo, Facebook, Bilibili, VK, Dailymotion, and Reddit.
Q: Do I need an API key to use the hosted server? A: No. The public endpoint is open; authentication can be added via a reverse proxy if desired.
Q: Can the server provide audio‑only transcription when subtitles are missing?
A: Yes. Set WHISPER_MODE to local or api and provide the necessary Whisper configuration variables.
Q: How is caching handled?
A: By default caching is off. Set CACHE_MODE=redis and provide CACHE_REDIS_URL to enable Redis‑based caching of subtitles and metadata.
Q: What limits exist for concurrent requests?
A: YT_DLP_MAX_CONCURRENCY (default 4) controls parallel yt‑dlp/ffmpeg processes; excess calls are rejected with a “server busy” error.
Q: How do I run the server locally without Docker?
A: Clone the repo, install dependencies (npm ci), then run npm run start:mcp for stdio or npm run start:mcp:http for HTTP on port 4200.
Q: Where can I find the OpenAPI/Swagger docs for the REST version?
A: Run the transcriptor-mcp-api Docker image (port 3000) and open http://localhost:3000/docs.
Connect one server. Then ask Claude, ChatGPT or etc about a video: the transcript, the chapters, the metadata, or a single frame. It works with 11 platforms, not only YouTube.
Connect · What to ask · Widgets · Platforms · Self-host
The hosted endpoint is:
https://transcriptor.gateway.mcpal.io/mcp
claude mcp add --transport http transcriptor https://transcriptor.gateway.mcpal.io/mcp
Then run /mcp and approve the sign-in in the browser. After this, claude mcp list shows ✔ Connected.
| Client | What to do |
|---|---|
| Claude (web and desktop) | Open Settings → Customize → Connectors. Select Add → Add custom connector, paste https://transcriptor.gateway.mcpal.io/mcp, then select Add. |
| ChatGPT | Open Transcriptor in the ChatGPT plugin directory and select Install plugin; sign in when asked. Or in ChatGPT open Plugins, search Transcriptor, select Install plugin. Then mention @Transcriptor in a chat. |
| Codex | Same directory, one install: ChatGPT and Codex share it. In a Codex task open Sources → Use plugins → Transcriptor; in the CLI, /plugins. |
Note: a new directory listing can take up to 6 hours to appear in Codex (Plugins in ChatGPT and Codex).
If your client is not in the list above, add the server with this configuration:
{
"mcpServers": {
"transcriptor": {
"url": "https://transcriptor.gateway.mcpal.io/mcp"
}
}
}
If you want to run the server yourself, read Self-host. The tools are the same and you need no account.
| Ask for this | Tool |
|---|---|
| "Summarize this video for me" | get_transcript |
| "Give me the subtitles as an SRT file" | get_raw_subtitles |
| "Is there a German track for this video?" | get_available_subtitles |
| "Who published this and how many views?" | get_video_info |
| "Go to the part about pricing" | get_video_chapters |
| "Show me the screen at 4:12" | get_video_frame |
| "Get English transcripts for the first 5 videos in this playlist" | get_playlist_transcripts |
| "Find recent videos about X" | search_videos (YouTube) |
Long transcripts come in parts. Each response gives a cursor for the next part, so no text is lost.
Each tool that takes a video accepts url. This is a link from a supported platform or a plain YouTube ID. Each tool returns content (text for the chat) and structuredContent (typed JSON for your code).
get_transcriptClean plain text, without timestamps, HTML, or speaker names. Without lang, the tool returns the track in the video's original language. Most platforms other than YouTube do not say which language a video is in; when the tool cannot tell which track that is, it answers with the list of tracks, and you call it again with type and lang. The inputs are the same as for get_raw_subtitles.
Response: videoId, url (the video page, as the server resolved it), type, lang, text, is_truncated, total_length, start_offset, end_offset. When more text is available, the response also has next_cursor.
get_raw_subtitlesRaw SRT or VTT content, in parts.
Input:
type — official or auto. Without lang, the tool picks a track of this typelang — a language code or track name, as get_available_subtitles lists it. Without it, the video's original language, as for get_transcriptresponse_limit — default 50000, minimum 1000, maximum 200000next_cursor — the cursor of the previous responseResponse: the fields of get_transcript, plus format (srt or vtt) and content.
get_available_subtitlesResponse: official and auto. Each field is a sorted list of language codes. Use this tool first, then give type and lang to the tools above.
get_video_infoExtended metadata from yt-dlp:
videoId, title, description, webpageUrluploader, uploaderId, channel, channelId, channelUrlduration, uploadDate, viewCount, likeCount, commentCounttags, categories, liveStatus, isLive, wasLive, availabilitythumbnail and thumbnailsget_video_chaptersResponse: chapters. Each item has startTime, endTime, and title. When the video has no chapters, the list is empty.
get_video_frameInput:
timecode — "MM:SS" or "HH:MM:SS.mmm"seconds — an alternative to timecode. Give one of the two, not bothformat — jpeg (default) or pngwidth — default 1280, maximum 1920, never larger than the sourcequality — 2 to 31, for jpeg onlyResponse: an image block, plus url, timestampSeconds, timestamp, mimeType, sizeBytes, and width. This tool needs ffmpeg. The Docker image includes it.
get_playlist_transcriptsInput:
url — a playlist URL, or a watch URL with list=type, lang, format — the same as get_raw_subtitles, except that lang is required: the original language is picked only for one video at a timeplaylistItems — a yt-dlp -I value such as 1:5, 1,3,7, or -1maxItems — the maximum number of videosResponse: results. Each item has videoId and text.
search_videosInput:
query — the search textlimit — default 10, maximum 50offset — the number of results to skipuploadDateFilter — hour, today, week, month, or yearresponse_format — json (default) or markdownResponse: results. Each item has videoId, title, url, duration, uploader, viewCount, and thumbnail.
Four tools have an interactive interface: get_transcript, get_video_info, get_video_frame, and search_videos. Clients that support MCP Apps and the ChatGPT Apps SDK show this interface in the chat. Other clients get the same data as text and JSON.
YouTube · Twitter/X · Instagram · TikTok · Twitch · Vimeo · Facebook · Bilibili · VK · Dailymotion · Reddit
Each tool that takes a video accepts a link from these 11 platforms. The tool search_videos works with YouTube only, through yt-dlp ytsearch.
The server does not download video or audio files for you. It returns text, metadata, and single frames.
The tools are the same as on the hosted endpoint. You need no account.
Run the server with Docker. The image serves Streamable HTTP on port 4200:
docker run --rm -p 4200:4200 artsamsonov/transcriptor-mcp:latest
Then point your client at http://localhost:4200/mcp.
For stdio, give the image an explicit command:
docker run --rm -i artsamsonov/transcriptor-mcp:latest npm run start:mcp
{
"mcpServers": {
"transcriptor": {
"command": "docker",
"args": ["run", "--rm", "-i", "artsamsonov/transcriptor-mcp:latest", "npm", "run", "start:mcp"]
}
}
}
The server starts with no environment variables. Each variable below is optional.
| Variable | Default | Function |
|---|---|---|
MCP_PORT and MCP_HOST |
4200 and 0.0.0.0 |
The HTTP listener |
COOKIES_FILE_PATH |
— | A Netscape cookies file for videos that need an account. See cookies.example.txt |
WHISPER_MODE |
off |
Set local or api to transcribe the audio when a video has no subtitles. Then set WHISPER_BASE_URL or WHISPER_API_KEY. WHISPER_MAX_DURATION_SECONDS skips longer videos and live streams; a video whose length the platform does not report is measured with ffprobe after the audio download |
CACHE_MODE |
off |
Set redis and CACHE_REDIS_URL to cache subtitles and metadata |
YT_DLP_MAX_CONCURRENCY |
4 |
How many yt-dlp/ffmpeg processes may run at once. YT_DLP_MAX_QUEUE (8) is how many calls may wait; beyond that a call is refused at once with "server busy". A call peaks at ~40 MiB, so the cap bounds platform throttling and latency, not memory |
SUBTITLES_RATE_LIMIT_HOLD_MS |
600000 |
After a platform answers 429 to a subtitle download, the server stops asking that platform for subtitles for this long and answers rate_limited right away. Each repeat doubles the wait, up to an hour; a successful download clears it. Metadata is not held back |
CANARY_INTERVAL_MS |
900000 |
How often the HTTP server fetches one transcript to prove the path still works. 0 turns it off; CANARY_URL picks the video |
YT_DLP_* |
— | Timeouts, proxy, and JS runtimes. See .env.example |
The same port serves GET /health and GET /metrics. The metrics are in Prometheus format and include the mcp_* counters.
Transport. The server accepts POST /mcp only. GET and DELETE return 405. The server is stateless and sends no Mcp-Session-Id.
The Node process does not check bearer tokens. Put a reverse proxy or a gateway in front of it for authentication and TLS. The hosted endpoint works this way.
REST API. A second image gives the same extraction over plain HTTP:
docker run --rm -p 3000:3000 artsamsonov/transcriptor-mcp-api:latest
The Swagger interface is at http://localhost:3000/docs. For a full stack with the API and the MCP server, read docker-compose.example.yml.
Development.
npm ci
npm run build
npm run start:mcp # stdio
npm run start:mcp:http # Streamable HTTP on port 4200
npm test
You need Node.js 22 or later (20 still works, but it reached end of life in April 2026), and yt-dlp in your PATH. Frame capture needs ffmpeg, and WHISPER_MAX_DURATION_SECONDS needs ffprobe (both ship in the same package, and in the Docker image). Other scripts: lint, type-check, format, test:coverage, test:e2e:api, and test:e2e:mcp.
Releases. The maintainer cuts them. The steps are in .claude/skills/release/SKILL.md. The version comes from package.json at runtime, through src/version.ts. Pushing a v* tag makes CI build both images and publish the MCP Registry entry from server.json.
Layout. src/mcp.ts (stdio entry), src/mcp-http.ts (Streamable HTTP), src/mcp-core.ts (tools, prompts, widgets), src/youtube.ts (yt-dlp), src/whisper.ts, src/cache.ts, src/index.ts (REST API), load/ (k6), and src/e2e/ (Docker smoke tests).
Pull requests are welcome. Read CONTRIBUTING.md first: it describes the cycle from issue to review, for people and for coding agents.
The hosted endpoint at transcriptor.gateway.mcpal.io is governed by the Terms of Service and the Privacy Policy.
A server you host yourself is not covered by those documents. It is governed by the MIT License only.
MIT © 2026 samson-art. Read LICENSE.
Please log in to share your review and rating for this MCP.
Explore related MCPs that share similar capabilities and solve comparable challenges
by headroomlabs-ai
Compress tool outputs, logs, files, RAG chunks, and conversation history before they reach the LLM, keeping answers identical while saving up to 95% of tokens for JSON payloads.
by modelcontextprotocol
A Model Context Protocol server for Git repository interaction and automation.
by zed-industries
A high‑performance, multiplayer code editor designed for speed and collaboration.
by modelcontextprotocol
Model Context Protocol Servers
by modelcontextprotocol
A Model Context Protocol server that provides time and timezone conversion capabilities.
by cline
An autonomous coding assistant that can create and edit files, execute terminal commands, and interact with a browser directly from your IDE, operating step‑by‑step with explicit user permission.
by upstash
Provides up-to-date, version‑specific library documentation and code examples directly inside LLM prompts, eliminating outdated information and hallucinated APIs.
by daytonaio
Provides a secure, elastic infrastructure that creates isolated sandboxes for running AI‑generated code with sub‑90 ms startup, unlimited persistence, and OCI/Docker compatibility.
by continuedev
Enables faster shipping of code by integrating continuous AI agents across IDEs, terminals, and CI pipelines, offering chat, edit, autocomplete, and customizable agent workflows.