Describe Video with AI

Get a text description of every second of your video — an AI model describes each frame locally in your browser. No upload, no signup.

Drop a video here

or click to select a file

MP4, WebM, MOV — processed locally, never uploaded

Requirements & limitations

The AI model runs entirely in your browser — that keeps your video private, but it means your hardware does the work. Know what to expect before a long run:

Recommended setup

  • Best: desktop Chrome or Edge with WebGPU — runs SmolVLM (~1–3 s per frame), which uses your optional context hint. The tool detects this automatically, and the results panel says which model was used.
  • Fallback: any modern desktop browser — a lighter model via WebAssembly, shorter captions, ~5–15 s per frame.
  • Memory: 8 GB RAM recommended; close heavy tabs for long videos.
  • Download: one-time model download — ~400 MB (WebGPU) or ~50 MB (fallback), plus ~100 MB if you enable speech transcription — cached by your browser afterwards.
  • Phones: not recommended — limited memory and usually no WebGPU. On laptops, plug in for long runs.

Limitations

  • Describes sampled moments, not motion — it sees "a person near a door", not "a person opens the door".
  • Small on-device models make mistakes a cloud model wouldn't, and won't recognize niche subjects on their own — use the context field to tell the model what it's looking at. Treat the output as a fast draft to edit, not a final product.
  • 150 frames max per run — pick a longer interval to cover long videos end to end.
  • Speech transcription is optional (checkbox, +~100 MB download). Non-speech audio — music, sound effects — is never described.
  • The video must be a format your browser can decode — MP4 (H.264) and WebM work everywhere; MOV support varies.

What is Describe Video with AI?

The Describe Video with AI tool turns a video into timestamped text. It samples one frame per interval (every second by default), describes each frame with an on-device vision model, and can optionally transcribe the speech with Whisper — merging what is shown and what is said into one scene-by-scene transcript you can copy or download as TXT, VTT, or SRT. Because the models run in your browser, nothing is uploaded — useful for accessibility audio-description drafts, video search notes, and dataset labeling.

How to use

  1. Drop a video file (MP4, WebM, or MOV) or click to select one.
  2. Pick how often to sample a frame, optionally tick "Also transcribe speech" to include what's said, and click Describe.
  3. Watch the timestamped descriptions stream in, then copy the transcript or download it as TXT, VTT, or SRT.

Frequently asked questions

How does the AI video describer work?

It grabs a frame at a fixed interval and runs each one through a vision AI model in your browser. With WebGPU it uses SmolVLM (~400 MB one-time download, then cached), which accepts an optional context hint — tell it "a padel match" and frames are described as padel, not just "a game". Otherwise a lighter ~50 MB captioning model with generic captions. You get a timestamped description of what happens on screen.

Is my video uploaded to a server?

No. Frame extraction and the AI model run entirely on your device. The model is downloaded once and cached — your video never leaves your browser.

What is the difference between this and video-to-text transcription?

Video to Text transcribes the audio — what is said. This tool describes the visuals — what is shown, frame by frame. Tick "Also transcribe speech" to get both in one timeline, merged by timestamp. Only speech is transcribed — music and sound effects are not described.

How long does it take?

Roughly 1–3 seconds per frame with WebGPU (Chrome/Edge desktop), 5–15 seconds per frame on the WebAssembly fallback. A desktop with 8 GB RAM is recommended; phones are not. Use a longer interval for long videos, and Stop ends the run early while keeping everything described so far.

Last updated

Powered by maratool