让 Claude (或任何 LLM) 实际观看视频 — 场景感知、去重帧 + 记录文本,从
▶ The 60-second pixel film — sound on (mp4 on GitHub) · an AI agent searches "how can an LLM truly understand video?", finds a key, and unlocks vision. 60-second real demo — real install, real run, real viewer.
Let Claude — or any LLM — actually watch a video.
pip install "claude-real-video[whisper]"
npx skills add HUANGCHIHHUNGLeo/claude-real-video # one command, installs the skill into Claude Code, Cursor, Codex, Copilot, Gemini CLI & 50+ agent hosts
Claude Code plugin marketplace (enable auto-update in /plugin → Marketplaces if you want it):
/plugin marketplace add HUANGCHIHHUNGLeo/claude-real-video
/plugin install claude-real-video@claude-real-video
Then paste a video link into your agent and ask about it. (CLI-only use? crv "<url>" works with just the pip install.)
Naming: crv is the short name for claude-real-video (the PyPI package). The paid add-on, crv Pro, is sold on Capafy under the listing name "llm-real-video Pro".
▶ New: the 40-second film — my AI agent learned to watch videos (and stopped working)
Same 58-second clip: fixed 1 fps sampling = 58 frames. crv keeps the 26 that actually differ — and
--gridpacks them into 3 contact sheets. Fewer tokens, nothing missed.
This free version lets your AI see the video. crv Pro lets it understand it — how it was shot (cut rhythm, camera moves) plus a timestamped timeline of what frames can't show: gestures, expressions, voice pitch shifts, emotion, sound events. One-time price $29 — get it on Capafy or buy with card via Lemon Squeezy.
Most AI tools don't really see a video. Paste a YouTube link into ChatGPT and it reads the transcript, not the picture. Claude won't take a video file at all. Even Gemini, which can read video natively, has to send it up to Google and samples frames at a fixed interval (1 fps by default), so fast cuts slip past.
claude-real-video does it differently, and the processing runs locally: point it at a URL or a
file, and it pulls the frames that actually matter (every scene change, not a
fixed quota), throws away the near-duplicates, transcribes the audio, and hands
you a clean folder any LLM can read. All the processing happens on your own machine — what gets sent anywhere is only the frames/text you choose to paste into an LLM afterwards.
crv "https://www.youtube.com/watch?v=..."
# → crv-out/frames/*.jpg + frames.json (per-frame timestamps) + transcript.txt/.json + MANIFEST.txt
Then drop the frames + MANIFEST.txt into Claude / ChatGPT / Gemini and ask away.
No terminal needed — run crv-web and a local page opens (Traditional Chinese / Simplified Chinese / English): paste a YouTube or Reels link or a file path, click Analyze, open the result viewer. Video analysis and output generation run on your machine — the source video never gets uploaded. (If you then paste the extracted frames or transcript into a cloud LLM, that data goes to that provider.)
Want to eyeball what the model will see first? Add --viewer — it writes a local viewer.html (video + keyframe grid + transcript) you can double-click open. No network, no extra installs.
Only part of a video matters (a 10-minute screen share inside a 90-minute call): --from 28:00 --to 43:00. ffmpeg seeks instead of decoding the whole file, Whisper only hears the window, and the frame budget is spent inside it — but every timestamp crv reports is still a source timecode you can quote to a colleague.
The meaning is small text (a terminal, a spreadsheet, an IDE): --frame-width 1600. Frame selection is the hard part and crv already does it; at 640px on a 1920-wide screen recording the right moment gets found and then the detail that made it worth finding is thrown away.
Slow-changing content (animation tutorials, gradual morphs, slow pans): add --adaptive — frames are picked against their rolling neighbourhood instead of a fixed threshold, so a 2-3s squash-and-stretch that never spikes any single frame still gets captured.
Text-heavy content (lecture slides, screen recordings, talking-head explainers): add --text-anchors — extra frames are forced at subtitle-cue timestamps, so each spoken segment gets a matching visual even when the scene barely changes. Needs a sidecar .srt/.vtt or an embedded subtitle track — captions burned into the pixels can't be detected. At most one forced frame per second; scene detection is untouched.
Multi-speaker content (interviews, podcasts, meetings): add --speakers — every transcript line gets a speaker label ([SPEAKER_00], [SPEAKER_01], …) so the model can follow who said what. Runs a local diarization model (45 MB, downloads once, no account or token needed). Install with pip install "claude-real-video[speakers]".
Not doing LLM work? It also works as a general-purpose video keyframe extractor — scene-change detection + dedup, no ML models to download.
Using Claude Code — or any coding agent? One command installs the skill (works with Claude Code, Cursor, Codex, Copilot, Gemini CLI and other agentskills.io-compatible hosts):
pip install "claude-real-video[whisper]"
npx skills add HUANGCHIHHUNGLeo/claude-real-video
Then just paste a video link into your agent and ask about it.
Manual install (clone + copy)git clone https://github.com/HUANGCHIHHUNGLeo/claude-real-video.git
mkdir -p ~/.claude/skills && cp -r claude-real-video/skills/claude-real-video ~/.claude/skills/
Tell it why you're watching, and keep what it finds:
crv "https://youtu.be/..." --why "find the pricing strategy" --kb ~/notes
--why makes the analysis focus on what you care about instead of a generic summary;
--kb saves the result as a dated note in your own notes folder, so it doesn't die in crv-out.
New in 0.10.x — analyse only the part that matters:
crv long-meeting.mp4 --from 28:00 --to 43:00
--from / --to cut a window out of a long video: ffmpeg seeks instead of decoding
the whole file, the transcript and frame budget follow the window, and every reported
timestamp is still a source timecode you can quote back to the original.
Real run on a 3-minute 640x360 video (benchmark/jfk-rice.mp4), Mac mini M4, local CPU, frames + dedup only (--no-transcribe). Image tokens estimated with Anthropic's (width x height) / 750 — 307 tokens/frame at 640x360.
--max-frames 80
80
23.4 s
~25k
--adaptive (catches slow morphs)
270
36.8 s
~83k
Dedup v0.7.16 — small-subject fast action no longer disappears. A percentage comparator is structurally blind to a subject that covers <1% of the frame (it can never change 8% of the pixels). Found in a user's 2,181-video batch run; fixed with a third "action channel". Synthetic repro — static 1280x720 shot, a 40x90 px subject (0.4% of frame) moves fast only in the last 10 of 65 frames:
Frames kept Action frames survived v0.7.15 2 1 / 10 v0.7.16 11 10 / 10 — full trajectoryMost "let an LLM watch a video" scripts (and Gemini's own pipeline) grab frames
at a fixed interval — e.g. one per second. That over-samples a static
screencast and under-samples a fast-cut reel. claude-real-video is smarter:
You feed the model fewer, more meaningful frames — cheaper context, better understanding.
pip install "claude-real-video[whisper]" # recommended: frames + dedup + audio transcription
pip install claude-real-video # core only (frames + dedup)
pip extras never install themselves — without [whisper] there is no speech-to-text
(videos that ship their own subtitles still get a transcript).
ffmpeg / ffprobe are used for frame extraction and audio, and aren't
pip-installable. Install them once:
brew install ffmpeg
Linux
sudo apt install ffmpeg (or your distro's package manager)
Windows
winget install Gyan.FFmpeg — or choco install ffmpeg — or download a build and add its bin\ folder to your PATH
Verify it's on your PATH:
ffmpeg -version
Transcription uses the whisper CLI (installed by the [whisper] extra, or
pip install openai-whisper). Whisper also relies on ffmpeg.
Faster + hallucination-proof transcripts (recommended): install the [fast]
extra and crv automatically switches to
faster-whisper — same models, same
output files, several times faster, and gated by Silero VAD (voice-activity
detection): music-only or silent audio yields an honest "no speech" note instead
of whisper's classic invented caption. No new flags to learn:
pip install 'claude-real-video[fast]'
If both are installed, faster-whisper wins; if it ever fails, crv falls back
to the whisper CLI on its own.
Apple Silicon (M1–M4): GPU transcription. Install the [mlx] extra and crv
runs Whisper on the Mac's GPU through
mlx-whisper —
a 21-minute talk that takes ~6 minutes on faster-whisper finishes in about a
minute on an M4, with the same transcript. The Silero VAD gate still runs
first, so music or silence never turns into invented captions; if the gate or
mlx cannot run, crv drops back to faster-whisper, then the CLI. Contributed by
@blazejp83 (#28).
pip install 'claude-real-video[mlx]'
Works on macOS, Windows, and Linux — Python 3.10+.
# A YouTube / Instagram / TikTok / ... link
crv "https://www.instagram.com/reel/XXXX/"
# A local file, English transcript, output to ./out
crv lecture.mp4 -o out --lang en
# Frames only, no transcription
crv clip.mp4 --no-transcribe
# A login-gated video (your own / authorised use): pass a Netscape cookie file
crv "https://..." --cookies cookies.txt
python -m claude_real_video ... works as an alias for crv too.
-o, --out
crv-out
output directory
--overwrite
off
replace a previous analysis living in the output directory (without this, a non-empty output dir is refused to avoid mixing videos)
--scene
0.30
scene-change sensitivity (lower = more frames)
--fps-floor
1.0
at least one fr
暂无开放 Issues,或尚未同步最近议题。