Skills

Give Claude Eyes: A Free Skill That Lets It Watch Any Video

7 minute readUpdated October 2026Explore more

TL;DR

Claude Code can look at images, but it can't play a video file. Hand it a link and the best it can do is read the page text or a transcript, which misses everything on screen. This free skill fixes that with three free tools on your own computer: yt-dlp grabs the video, ffmpeg pulls frames from the first 3 seconds and at every cut, and Whisper writes a timestamped transcript. Claude then reads the pictures and the words together and breaks down the hook, the cuts, the captions and the call to action. Copy the SKILL.md below into ~/.claude/skills/watch-video/ and say: break down this video.

Ask Claude why a video worked and it usually answers from the words alone. But most of what makes a short video work is on screen: the first frame, how fast it cuts, the captions, the b-roll. This skill turns a video into what Claude can actually read, a set of frames with timestamps plus a timestamped transcript, then has it write a breakdown you can learn from.

What the skill does

  1. 1Downloads the video at up to 720p with yt-dlp, or uses a file you already have.
  2. 2Pulls two frames a second from the first 3 seconds, because that is where the hook lives.
  3. 3Pulls one frame at every cut using ffmpeg scene detection, and saves the time of each cut.
  4. 4Builds a contact sheet of the first 16 cuts so Claude can see the whole edit in one image.
  5. 5Transcribes the audio with Whisper into a timestamped subtitle file.
  6. 6Reads the frames and the transcript together and writes breakdown.md: hook, cuts, on-screen text, structure, why it works, and ideas you can reuse.

Why frames? Anthropic's Claude Code docs say the Read tool returns PNG, JPG and other image files as visual content Claude can see. A video file isn't one of those, so it has to become images first. That is the whole trick.

What you need (all free)

yt-dlp

Free, open source video downloader (Unlicense). Over 195,000 GitHub stars on October 3, 2026.

FFmpeg

Free, open source tool for cutting video into frames and audio (LGPL, with some optional parts under GPL).

Whisper (OpenAI)

Free, open source speech-to-text that runs on your own computer (MIT). Over 109,000 GitHub stars on October 3, 2026.

bash# macOS (Homebrew)
brew install ffmpeg yt-dlp
pip install -U openai-whisper

# Windows (winget)
winget install Gyan.FFmpeg
winget install yt-dlp.yt-dlp
pip install -U openai-whisper

Whisper's README says it is expected to work with Python 3.8 to 3.11. The first time it runs, it downloads the model you pick. The base model used below was 145 MB on our machine.

Step 1: Install the skill

Make a folder for it, then save the file below as ~/.claude/skills/watch-video/SKILL.md. Skills in ~/.claude/skills/ load in every project on your computer.

bashmkdir -p ~/.claude/skills/watch-video
markdown---
name: watch-video
description: Lets Claude "watch" a video by pulling frames from the first 3 seconds and at every cut, plus a timestamped transcript, then breaking down the hook, cuts, on-screen text and call to action. Use when the user shares a video file or a public video link and asks what happens in it, why it worked, or how to reuse its structure.
---

# Watch a video

Claude Code reads images, not video files. Turn the video into frames plus a timestamped transcript, then read both together.

## Rules
- Only download public videos the user has the right to save, for private study. Never re-upload or redistribute them.
- If the user gives a local file, skip the download and copy it to video-notes/video.mp4.
- Describe only what is in the frames and the transcript. If something can't be seen in a frame, say so instead of guessing.
- Keep every file inside ./video-notes/ and run every command from inside that folder.

## Steps
1. Check the tools: `ffmpeg -version`, `yt-dlp --version`, `whisper --help`. If one is missing, give the user the install command (macOS: `brew install ffmpeg yt-dlp` and `pip install -U openai-whisper`) and stop.

2. Get the video and its length:
   mkdir -p video-notes/frames && cd video-notes
   yt-dlp -f "bv*[height<=720]+ba/b[height<=720]/b" --merge-output-format mp4 -o "video.%(ext)s" "VIDEO_URL"
   ffprobe -v error -show_entries format=duration -of default=nw=1:nk=1 video.mp4

3. Hook frames, two per second for the first 3 seconds:
   ffmpeg -hide_banner -loglevel error -y -t 3 -i video.mp4 -vf "fps=2,scale=720:-2" -q:v 3 frames/hook_%02d.jpg

4. One frame at every cut, with its timestamp:
   ffmpeg -hide_banner -y -i video.mp4 -vf "select='gt(scene,0.3)',showinfo,scale=720:-2" -fps_mode vfr -q:v 3 frames/cut_%04d.jpg 2> frames/cuts.log
   grep -o 'pts_time:[0-9.]*' frames/cuts.log | cut -d: -f2 > frames/cut_times.txt
   Line N of cut_times.txt is the time in seconds of cut_000N.jpg.
   If you get fewer than 3 cuts, also take one frame every 2 seconds:
   ffmpeg -hide_banner -loglevel error -y -i video.mp4 -vf "fps=1/2,scale=720:-2" -q:v 3 frames/every2s_%04d.jpg

5. Contact sheet of the first 16 cuts, for a fast overview:
   ffmpeg -hide_banner -loglevel error -y -pattern_type glob -i 'frames/cut_*.jpg' -vf "scale=320:-2,tile=4x4" -frames:v 1 frames/contact_sheet.jpg

6. Transcript with timestamps (skip if the video has no speech):
   ffmpeg -hide_banner -loglevel error -y -i video.mp4 -vn -ac 1 -ar 16000 audio.wav
   whisper audio.wav --model base --output_format srt --output_dir .
   This writes audio.srt.

7. Read every hook frame, the contact sheet and the cut frames with the Read tool. If there are more than 40 cut frames, read every second one. Then read audio.srt.

## Output
Write the breakdown to video-notes/breakdown.md and show it:
- **Hook (0 to 3s):** what is on screen in each hook frame, the first spoken line, and any on-screen text word for word.
- **Cuts:** how many, the average seconds between cuts, and a table of time, what changes on screen, and what is being said.
- **On-screen text and captions:** style, position, and how often they change.
- **Structure:** hook, setup, payoff and call to action, with timestamps.
- **Why it works:** 3 to 5 reasons, each tied to a specific frame or line. Label anything that is your interpretation.
- **Reuse this:** 3 ideas the user could apply to their own video without copying the creator's words or footage.

Start a new Claude Code session and type / to check that watch-video shows up. If it doesn't, ask Claude: What skills are available?

Step 2: Drop in a video

Ask in plain words, or type /watch-video. Some prompts that work well:

promptBreak down this video and tell me why the hook works: [video link]

Use watch-video on ~/Desktop/my-reel.mp4. How many cuts are there, and where does the pace slow down?

Watch these two videos and compare their first 3 seconds frame by frame: [link 1] [link 2]

Everything lands in a video-notes folder: the frames, the cut times, the contact sheet, the transcript, and breakdown.md. Delete the folder when you're done with it.

Tune it

  • Missing cuts? Lower the scene threshold in step 4 from 0.3 to 0.2. Getting a frame for every flicker? Raise it to 0.4.
  • Better transcript: change --model base to --model small. It's slower but more accurate.
  • Video already has captions? yt-dlp can save them instead of running Whisper: add --write-auto-subs --sub-langs en --skip-download to the yt-dlp command.
  • Long videos: for anything over a few minutes, use the every-2-seconds command in step 4 and point Claude at the part you care about.
calesthio/OpenMontage

Open source agentic video production system. Its video-understand skill is the original. Over 62,000 GitHub stars on October 3, 2026.

Now you just drop in a video and Claude breaks down exactly what's on screen. Inside the Claude Code Club we share the skills, prompts and workflows we actually run every day, and help each other set them up. It's nine dollars a month at https://www.skool.com/claudecodeclub/about. Everything on this page works without it.

Common questions

  • Can Claude watch videos?

    Not directly in Claude Code. Its Read tool shows Claude images, such as PNG and JPG files, not video files. This skill turns a video into frames and a timestamped transcript, which Claude can read together.

  • Is this skill free?

    Yes. The skill is a text file, and the three tools it uses (yt-dlp, FFmpeg and Whisper) are free and open source. Whisper runs on your own computer, so there's no API bill for the transcript.

  • Is it legal to download videos for this?

    It depends on the video, the platform's terms and where you live. Use it on your own videos, or on public videos you have the right to save, for private study only, and never re-upload them.

  • Does it work on Windows?

    The tools all run on Windows. The contact sheet step uses a file pattern that works on macOS and Linux, so on Windows ask Claude to build the contact sheet a different way or skip it.

  • How long does it take?

    For a short video, the frames take seconds. Whisper is the slow part and depends on your computer and the model you choose. The base model is a good balance of speed and accuracy.

Want the skills we use to study what's working?

Get the other 8 in the skills stack, plus 8,000+ members - $9/mo, cancel anytime.

Join the Club