The bradautomates/claude-video project enables video analysis and interaction through Claude, an AI agent, providing capabilities for content analysis, bug diagnosis, and video summarization.
Source: Official README and linked documentation View on GitHub →This project is being discussed because its innovative approach to video analysis using AI, addressing the need for efficient video content processing and understanding. Its unique technical choices, such as scene-aware frame extraction and integration with Claude, set it apart in the market.
Source: Official README and linked documentationThe project allows for the analysis of videos by extracting frames, transcribing audio, and providing a timestamped transcript, enabling detailed examination of video content.
Source: Official README and linked documentationIt extracts frames based on scene changes, which can be adjusted for efficiency or fidelity, providing a balance between speed and detail.
Source: Official README and linked documentationThe project integrates with Claude, an AI agent, to provide detailed analysis and answers based on the video content, enhancing the user experience.
Source: Official README and linked documentationThe architecture is modular, with separate components for video downloading, frame extraction, transcription, and interaction with Claude. It uses `yt-dlp` for video downloading, `ffmpeg` for frame extraction, and Whisper API for transcription. The project is structured into various directories for different components, such as `.agents`, `.skills`, and `hooks`.
Source: Official repository tree and dependency filesyt-dlpffmpegWhisper APIThe project is suitable for developers, content analysts, and anyone needing to analyze video content efficiently. It is useful for diagnosing bugs from video recordings, summarizing long videos, analyzing competitors' content, and extracting insights from video data.
Source: Official README and linked documentationThe latest official version is v0.2.0, changelog date is 2026-06-29, and GitHub Release was released on 2026-07-01. Mainly added four levels of detail, default frame deduplication, automatic blocking when Whisper exceeds 25 MB, fixed-point timestamps, no-whisper, and reorganized into self-contained skills to support more Agent Skills hosts.
Source: github.com/bradautomates/claude-video/releases/tag/v0.2.0;https://git…status: docs-only | checked_at: 2026-09-23 | environment: Check the official README, skills/watch/SKILL.md, CHANGELOG and v0.2.0 Release | limitations: skill, ffmpeg or yt-dlp not installed; video not downloaded, framing, calling Groq/OpenAI Whisper, reading image or reviewing token/time cost | observed_output: Skill frontmatter annotation 0.2.0; README lists ffmpeg, yt-dlp, fourth gear detail and first installation process; changelog date is 2026-06-29, GitHub Release released on 2026-07-01
Source: github.com/bradautomates/claude-video/blob/main/skills/watch/SKILL.md…The Google Gemini API officially supports direct video submission, static mode pressing 1 FPS and audio processing, and also provides agentic mode for long videos; uploading, parsing, and pushing hosted APIs are completed. Claude Video preprocesses locally with yt-dlp and ffmpeg, and then hands over the selected frames and subtitles to the hosting agent. It can directly control the time window and frame budget, but it must maintain dependence and process download permissions. You want to evaluate the Gemini API with hosted native video input; you want to add a checkable local preprocessing process to your existing coding proxy to evaluate Claude Video.
Source: ai.google.dev/gemini-api/docs/video-understanding;https://github.com/…The bradautomates/claude-video project is a tool for anyone needing to analyze video content efficiently and interact with it using AI. Its integration with Claude and configurable frame extraction options make it a valuable asset for developers and content analysts, though it may require significant computational resources and more detailed documentation.
Skill frontmatter annotation 0.2.0; README lists ffmpeg, yt-dlp, fourth gear detail and first installation process; changelog date is 2026-06-29, GitHub Release released on 2026-07-01 skill, ffmpeg or yt-dlp not installed; video not downloaded, framing, calling Groq/OpenAI Whisper, reading image or reviewing…
The review environment was: Check the official README, skills/watch/SKILL.md, CHANGELOG and v0.2.0 Release claude-video behavior outside that scope still needs a local check.