Video to text for LLMs — 100% local AI video analysis
Turn any video into a package an LLM can read — transcript, keyframes, manifest — without the video ever leaving your machine.
黃志弘 Leo
Published 2026-07-12 · 1 readers
Why "video to text" needs more than a transcript
Most video-to-text tools give you a transcript and stop. That carries the words — but a product demo, a tutorial, or a fast-cut reel communicates visually. An LLM reading only the transcript is analyzing a radio show.
And most tools that go further do it in someone's cloud: you upload your footage — client work, unreleased products, internal recordings — to a third-party server just to get frames back.
The alternative: a local pipeline that produces both the text and the visuals, as plain files, on your own hardware.
The local pipeline
claude-real-video (crv) is a free, MIT-licensed command-line tool (2k+ GitHub stars). One command produces everything an LLM needs:
# install (Python 3.10+; macOS/Windows/Linux)
brew install ffmpeg
pip install "claude-real-video[whisper]"
crv lecture.mp4 -o out --lang en
→ out/frames/*.jpg # scene-change keyframes, deduplicated
→ out/transcript.txt # plain-text transcript (any LLM reads it)
→ out/transcript.json # timestamped segments for your own tools
→ out/MANIFEST.txt # tells the model how to read the package- Frames: extracted at every scene change plus a density floor — not a fixed 1-per-second quota — then deduplicated against a sliding window, so a repeated shot is sent once and a 10-minute static slide collapses to one image.
- Text: if the video ships subtitles, those are used (faster, more accurate); otherwise Whisper transcribes locally.
--whisper-model turbois a pruned large-v3 — much faster than large, with a minor quality trade-off. - Audio (optional):
--keep-audioalso saves the full soundtrack for audio-capable models. - No terminal needed: run
crv-webfor a local web page — paste a link, click Analyze, open the result viewer.
What "local" actually means here
| step | where it happens |
|---|---|
| Download / read the video | your machine (yt-dlp for URLs, or a local file) |
| Frame extraction + dedup | your machine (ffmpeg) |
| Transcription | your machine (Whisper) |
| Analysis by the LLM | your choice — paste into a cloud LLM, or feed a self-hosted model and nothing leaves your machine |
The source video never gets uploaded. If you later paste frames or transcript into ChatGPT, Claude, or Gemini, only that pasted content goes to that provider — with a self-hosted open-source vision model, the whole loop stays on your hardware.
Feeding the output to any model
The output is deliberately boring: JPEG images and plain text. That's what makes it universal.
- ChatGPT — drag frames + transcript into a chat. Full walkthrough: how to make ChatGPT watch a video
- Claude — same, or install it as a Claude Code skill so Claude runs the pipeline itself: Claude video analysis guide
- Self-hosted models — point your vision model at the frames and manifest; nothing ever leaves your machine
Tip: --grid tiles consecutive keyframes into 3x3 contact sheets — the model reads an ordered sequence, and you attach 9x fewer images.
When text + frames still isn't enough
Keyframes and transcript tell a model what is on screen and what was said. They can't carry how the camera moves, how fast the edit cuts, or how the speaker's voice shifts. For that, crv Pro ($29 one-time) adds a motion pass — camera-move classification, editing rhythm, action-burst frame sequences, and a timestamped perception timeline — computed locally like everything else, written into the same manifest as plain text.
FAQ
How do I convert a video to text for an LLM?
One command: crv "url-or-file". You get a transcript, timestamped JSON, deduplicated keyframes and a manifest — all local.
Can I do AI video analysis without uploading the video?
Yes — that's the point. Processing is entirely on your machine; you choose what (if anything) to share with a cloud model.
What formats and sources work?
Local files, plus URLs from YouTube, Instagram, TikTok and other yt-dlp-supported sites. macOS, Windows and Linux.
Is it really free?
The core tool is MIT-licensed and free. The optional Pro motion/emotion analysis is a one-time add-on — $29.