AI & automationReady to pilotBuilt by LabsLast updated 29 July 2026

Sport HighlightsPublishable highlight clips minutes after the final whistle

Highlights that arrive hours after the final whistle miss the audience that was waiting for them. Sport Highlights detects goals, overtakes, penalties and other key events from the finished broadcast using speech recognition and large language models, then produces chapter markers and standalone, social-ready clips within minutes.

Automatic Highlight DetectionLanguage-Agnostic ProcessingLLM / VLMMulti-Sport SupportStandalone Clip Generation
Sport Highlights logo

Try the demo

Browse processed sports broadcasts and their automatically detected highlights

⚠️

Failed to load video library

The problem this solves

The window for sports highlight distribution is narrow. Fans who want to see the decisive moment will find it through a competitor's feed or a pirated clip if the rights holder has not published within the first hour. Broadcast teams that depend on manual clip selection and export workflows cannot move that fast: the full match finishes, someone watches it back, someone else trims and exports, and by the time the clip is approved for distribution the moment has passed.

The underlying content is rich with signal - commentators name the events, scoreboards display the score, graphics confirm the player - but extracting that signal has always required a human to watch the footage first. A pipeline that reads that signal directly from audio and on-screen graphics can produce a publishable package in the time it currently takes to find the right section of the timeline.

Who this is for

This is relevant if you are:

Sports organisation

A sports organisation that needs a shareable highlight package ready within minutes of the final whistle, before audience attention shifts

Broadcaster

A broadcaster running multi-event coverage across several simultaneous fixtures who cannot staff a human clipper on every feed

Rights holder

A rights holder distributing to platforms that require clips rather than full-match VOD and currently outsourcing that clipping work

How it works

  1. 1The finished broadcast is uploaded to the pipeline. Audio is extracted and passed to AWS Transcribe, which returns a time-stamped transcript of the commentary.
  2. 2A large language model analyses the transcript to identify high-impact moments - goals, penalties, red cards, crashes, podium finishes - and assigns a timestamp range to each.
  3. 3A vision-language model inspects frames around each candidate moment to verify the event against on-screen graphics: scoreboards, leaderboards and pop-up overlays, correcting commentary errors against visual facts.
  4. 4Each confirmed highlight is clipped by FFmpeg to a standalone file with a title and short summary generated from the combined audio and visual evidence.
  5. 5The output is a structured package of chapter markers for the full VOD and standalone clip files ready for platform upload or social distribution.

Large Language Models (LLMs)

Highlight Detection & Analysis

Intelligently identifies high-impact moments like goals, crashes, and game-changing plays from transcript context.

AWS Transcribe

Speech-to-Text Processing

Accurate transcription with precise timestamps for reliable chapter marker positioning.

Vision-Language Models (VLMs)

Object Detection & OCR

Analyzes broadcast screenshots to detect scoreboards, leaderboards, and event pop-ups - extracting accurate player names, scores, and rankings via OCR. This corrects transcription errors by grounding the LLM output in visual facts from on-screen graphics.

Multi-Lingual Support

Language-Agnostic Processing

Supports broadcasts in multiple languages including English, German, Spanish, and more - with language-aware highlight titles and summaries.

FFmpeg

Video Processing & Clip Generation

Creates standalone highlight clips with precise start/end times for social sharing and VOD enrichment.

Multi-Part S3 Upload

Large File Handling

Reliable upload of large video files with resume capability and progress tracking.

What makes it different

Most automated highlight tools work from the audio track alone, which means they fail the moment the commentary does not describe what is happening on screen. This POC grounds every detection in both tracks: the LLM reads what was said, and the VLM reads what the broadcast graphic showed. When a commentator's language is ambiguous - or when the relevant event is a score correction rather than a spoken announcement - the visual track provides the fact. The result is a clip package that reflects what actually happened rather than what the commentary happened to make explicit.

Current status & next steps

Ready to pilotReviewed 29 July 2026

Sport Highlights is ready to pilot. The demo on this page processes full broadcast files end-to-end against real match footage, and the pipeline is in the state we would hand to a client for a structured evaluation.

Detection quality depends on commentary quality and language clarity, and the pipeline is optimised for single-language broadcasts - switching languages within one feed reduces accuracy. Processing time scales with the duration of the source, clips are capped at ten to thirty seconds so each one stays focused on a single moment, and OCR accuracy varies with the quality of the broadcast graphics and how much visual information a frame carries.

The next step is running against a fixture from your own rights portfolio so that detection quality can be measured on the sport and production style you actually distribute.

Want highlights ready before the post-match discussion ends?

Tell us what sport you cover, how you currently clip highlights, and what your distribution window looks like, and we will tell you what a pilot against your own broadcasts would involve.