Skip to main content
Not yet available for batches created through the developer API. The keyframe index this searches is built only by the fastdrop.io subscriber pipeline. A batch you create with POST /v1/batch is never indexed, so POST /v1/batch/{id}/visual-search accepts your key, finds the batch, and returns zero results — not an error. The endpoint is documented here because it exists and will work once API batches are indexed; until then, treat this page as a preview and use transcript search, which does work for API-created batches.For the same reason, visual search is not exposed as an MCP tool. See MCP Setup for the tools that are.
Visual search is in beta. It works well for broad visual descriptions but not for fine-grained details like specific colors or small text.
Visual search lets you find videos by describing what’s in them. It uses CLIP to match your text description against keyframe thumbnails extracted from each video.

How it works

  1. When a batch is processed, keyframe thumbnails are extracted from each video
  2. Each thumbnail is embedded using CLIP (ViT-B/32) into a 512-dimensional vector
  3. When you search, your text query is embedded using the same CLIP model
  4. FastDrop finds thumbnails with the highest cosine similarity to your query
Step 2 is the one that doesn’t run for API-created batches today. Thumbnails are extracted, but nothing embeds them, so there is nothing for step 4 to match against. Visual search requires an API key. It costs 0 credits per query — the CLIP embedding cost is included with thumbnail extraction.

Code examples

Example response

Response fields

Example queries

Works well

These types of descriptions produce reliable results:
  • "outdoor establishing shot" — landscape and wide shots
  • "person at desk" — interview or office setups
  • "whiteboard with writing" — presentation or lecture frames
  • "close-up of hands" — detail shots
  • "drone aerial view" — aerial footage
  • "dark room with screen" — screen recording or demo setups

Less reliable

CLIP matches broad visual concepts, not fine details:
  • "blue shirt" — color-specific queries are inconsistent
  • "text says 'hello'" — can’t read specific text
  • "exactly 3 people" — can’t count reliably
  • "logo in top-right corner" — spatial positioning is weak
For the most comprehensive results, use both search types. The unified search service fuses transcript and visual scores (0.6 transcript + 0.4 visual) to find videos that match on both content and visuals.

Error responses