Describe Video with AI
Get a text description of every second of your video — an AI model describes each frame locally in your browser. No upload, no signup.
Drop a video here
or click to select a file
MP4, WebM, MOV — processed locally, never uploaded
Requirements & limitations
The AI model runs entirely in your browser — that keeps your video private, but it means your hardware does the work. Know what to expect before a long run:
Recommended setup
- Best: desktop Chrome or Edge with WebGPU — runs SmolVLM (~1–3 s per frame), which uses your optional context hint. The tool detects this automatically, and the results panel says which model was used.
- Fallback: any modern desktop browser — a lighter model via WebAssembly, shorter captions, ~5–15 s per frame.
- Memory: 8 GB RAM recommended; close heavy tabs for long videos.
- Download: one-time model download — ~400 MB (WebGPU) or ~50 MB (fallback), plus ~100 MB if you enable speech transcription — cached by your browser afterwards.
- Phones: not recommended — limited memory and usually no WebGPU. On laptops, plug in for long runs.
Limitations
- Describes sampled moments, not motion — it sees "a person near a door", not "a person opens the door".
- Small on-device models make mistakes a cloud model wouldn't, and won't recognize niche subjects on their own — use the context field to tell the model what it's looking at. Treat the output as a fast draft to edit, not a final product.
- 150 frames max per run — pick a longer interval to cover long videos end to end.
- Speech transcription is optional (checkbox, +~100 MB download). Non-speech audio — music, sound effects — is never described.
- The video must be a format your browser can decode — MP4 (H.264) and WebM work everywhere; MOV support varies.
What is Describe Video with AI?
The Describe Video with AI tool turns a video into timestamped text. It samples one frame per interval (every second by default), describes each frame with an on-device vision model, and can optionally transcribe the speech with Whisper — merging what is shown and what is said into one scene-by-scene transcript you can copy or download as TXT, VTT, or SRT. Because the models run in your browser, nothing is uploaded — useful for accessibility audio-description drafts, video search notes, and dataset labeling.
How to use
- Drop a video file (MP4, WebM, or MOV) or click to select one.
- Pick how often to sample a frame, optionally tick "Also transcribe speech" to include what's said, and click Describe.
- Watch the timestamped descriptions stream in, then copy the transcript or download it as TXT, VTT, or SRT.
Frequently asked questions
How does the AI video describer work?
It grabs a frame at a fixed interval and runs each one through a vision AI model in your browser. With WebGPU it uses SmolVLM (~400 MB one-time download, then cached), which accepts an optional context hint — tell it "a padel match" and frames are described as padel, not just "a game". Otherwise a lighter ~50 MB captioning model with generic captions. You get a timestamped description of what happens on screen.
Is my video uploaded to a server?
No. Frame extraction and the AI model run entirely on your device. The model is downloaded once and cached — your video never leaves your browser.
What is the difference between this and video-to-text transcription?
Video to Text transcribes the audio — what is said. This tool describes the visuals — what is shown, frame by frame. Tick "Also transcribe speech" to get both in one timeline, merged by timestamp. Only speech is transcribed — music and sound effects are not described.
How long does it take?
Roughly 1–3 seconds per frame with WebGPU (Chrome/Edge desktop), 5–15 seconds per frame on the WebAssembly fallback. A desktop with 8 GB RAM is recommended; phones are not. Use a longer interval for long videos, and Stop ends the run early while keeping everything described so far.
Last updated
Powered by maratool