AI & automationTestingBuilt by LabsLast updated 29 July 2026

Scene FinderNatural-language search across a video scenes

Finding one specific moment in a video archive means scrubbing through hours of footage, and the person doing it is usually the one who can least afford the time. Scene Finder searches spoken content and visual scenes together and matches by meaning rather than by keyword, so a plain description of what you remember is enough to land on the frame.

Semantic Video SearchMulti-Modal RetrievalPrecise Clip NavigationRAG-powered Search
Scene Finder logo

Try the demo

Select a Collection

The problem this solves

Archives grow faster than the teams that use them. A producer who needs the moment a particular goal was scored, or the shot where a sponsor's product appeared, has three options: remember roughly where it was, ask someone who might, or scrub. All three cost hours per request, and the cost repeats every time the same footage is wanted for a different purpose.

Keyword search over transcripts only helps when somebody happened to say the words being searched for. It fails entirely on anything purely visual, which in video is most of what matters.

Who this is for

This is relevant if you are:

Broadcaster

A broadcaster whose production teams pull archive footage daily and wait on a librarian to locate it

Sports organisation

A sports organisation that needs a specific passage of play found on request, during or straight after a fixture

Rights holder

A rights holder licensing clips who has to find them from a customer's description rather than a timecode

How it works

  1. 1Every video in a collection is indexed twice: once from its transcript, and once from generated descriptions of what is visible on screen.
  2. 2Your query is rewritten automatically into the form each index needs, so a single plain-language question searches spoken and visual content at the same time.
  3. 3Results from both indexes are merged - union to widen the net, intersection to require a match on both - and ranked by meaning rather than by keyword overlap.
  4. 4Each result opens in a player positioned at the exact timestamp, so a match is confirmed or dismissed without leaving the results list.

Semantic Search

Meaning-Based Retrieval

Understands the intent behind your query, finding relevant content even when exact words don't match.

Multi-Modal Search

Transcript + Visual

Searches spoken content and visual scenes in the same query, so a moment can be found by what was said, what was shown, or both.

Query Rewriting

Intelligent Query Analysis

Automatically optimizes your search query for both spoken and visual content, extracting key concepts.

Collection Management

Organized Video Libraries

Browse and search within organized video collections or search across all content at once.

Clip Player

Precise Playback

Results play from the matched timestamp, with controls built for reviewing many short passages in sequence.

Result Merging

Flexible Combinations

Union (OR) or intersection (AND) modes let you control how transcript and visual results combine.

What makes it different

Most video search is transcript search with a better interface: it finds what was said and is blind to what was shown. Scene Finder treats the visual track as a first-class index and lets you require agreement between the two, which is what makes it usable on sport, events and any footage where the important thing is never spoken aloud. It is designed to sit alongside an existing MAM or archive rather than replace it, because that is the only way it can be adopted without a migration project first.

Current status & next steps

TestingReviewed 29 July 2026

Scene Finder is in controlled testing. The demo on this page runs against a prepared set of collections and is opened to selected partners on invitation, which is exactly what the Testing status commits us to.

Retrieval quality rests on transcript accuracy and on the quality of the scene descriptions produced at indexing time; visual search in particular works best where scenes were properly captioned when they were indexed. Large collections take slightly longer to search comprehensively, results are capped at the top ten matches per query to keep the interface responsive, and intersection mode returns fewer results wherever the two modes do not overlap.

The next step is a pilot against a customer archive, so retrieval quality can be measured on real material and real queries instead of on content we picked ourselves.

Want to try this on your own archive?

Tell us roughly how much footage you hold and what people search for most often, and we will tell you what a pilot against your material would involve.