Streaming platformReady to pilotBuilt by LabsLast updated 29 July 2026

StreamtokShort-form video feed built from content you already own

Most of a long-form catalogue is never opened, because deciding what to watch is work and a browser full of thumbnails asks for that decision up front. Streamtok extracts the moments that carry a programme on their own and assembles them into a continuous personalised feed, so the same catalogue reaches the viewer who wants to browse and the viewer who wants to be shown, without any new production spend.

Automatic Highlight GenerationPersonalized Content DeliveryShort-form Video Clips
Streamtok logo

Watch the demo

Click to unmute

The problem this solves

A platform built around long-form content puts a decision between the viewer and the content. Someone who opens an app with ten minutes to spare, or with no particular intent at all, is shown a grid of thumbnails and asked to commit to forty-five minutes. That is not a generational problem: it is what a browse-first interface does to anyone who arrives without a title already in mind. The engagement gap is not about content quality, it is about the format the catalogue is offered in.

Commissioning short-form separately is the obvious response, but the economics rarely work: short-form production costs are not proportionally lower than long-form, the rights situation for purpose-made clips is often more complicated, and the resulting feed is thin until a significant library accumulates. The catalogue already exists; the problem is that no one has extracted its short-form potential.

Who this is for

This is relevant if you are:

Broadcaster

A broadcaster whose long-form catalogue is fully cleared for streaming but whose platform analytics show short sessions and high early drop-off

Sports organisation

A sports organisation with a deep archive of match footage that wants to offer a highlight feed without a dedicated editorial team cutting clips manually

How it works

  1. 1Each VOD asset is transcribed using OpenAI Whisper, producing a segment-level transcript with timestamps that the extraction pipeline can reason over.
  2. 2A large language model analyses the transcript in zero-shot mode, identifying the passages most likely to work as standalone short clips - moments of high information density, narrative tension or audience reaction.
  3. 3Voice Activity Detection aligns each identified passage to a natural speech pause, so clip boundaries do not cut mid-sentence or mid-word.
  4. 4The recommendation engine uses enriched metadata - sourced via the InsysGo API - alongside LLM reasoning to order clips for each viewer, so the feed is personalised rather than chronological.
  5. 5Clips are capped at 30 seconds and assembled into a continuous feed, which the platform surfaces in whichever short-form pattern its own product already uses.

Large Language Models (LLMs)

Zero-Shot Highlight Extraction

Pinpoints captivating 5-30s clips directly from the transcript without training data.

Retrieval-Augmented Generation (RAG)

Content Recommendation Engine

Uses enriched metadata and LLM reasoning for highly personalised, contextual recommendations.

OpenAI Whisper

High-Precision Transcription

Accurate transcription of VOD content with segment-level timestamps; performance may vary based on audio complexity.

Voice Activity Detection (VAD)

Clip Boundary Alignment

Aligns clip boundaries to natural speech pauses for high-quality, non-disruptive cuts.

InsysGo API

Metadata Extraction

Provides metadata for recommendation enrichment.

What makes it different

Most automated clipping tools identify highlights by energy signals - crowd noise, music changes, volume spikes - which works for sport but fails on commentary, interview, documentary and drama. Streamtok works from the semantic content of the transcript, so it can identify a compelling moment in a conversation as reliably as it identifies a goal. The recommendation layer then treats the viewer's engagement history as a signal, ordering clips to extend the session rather than simply replaying the most-watched content from the archive.

Current status & next steps

Ready to pilotReviewed 29 July 2026

Streamtok is ready to pilot. The demo on this page shows the full pipeline output - automatic clip extraction, feed assembly and personalised ordering - running against a prepared set of VOD assets.

Clips are capped at thirty seconds to hold the short-form format. Clip analysis currently works from transcription text alone, so visual and richer audio features play no part in the decision, and there is no speaker diarization: several speakers in one conversational clip are not told apart. Reliance on OpenAI Whisper introduces variability in transcription speed and accuracy, and highlight generation is only as fast as the underlying language model.

The next step is a pilot against a customer catalogue: real assets, real audience segments, real engagement targets.

Want to turn your catalogue into a short-form feed?

Tell us roughly how much VOD you hold and which audience you are trying to reach. We will show you what Streamtok would extract from a sample of your content.