AI & automationReady to pilotBuilt by LabsLast updated 29 July 2026

Voice CloningSpeaker voices recreated in languages they never recorded

Localising a presenter-led format means recasting the voice in every market and losing the one viewers recognise. Voice Cloning recreates a speaker from a short audio sample and has them speak naturally in languages they never recorded, so the same voice carries the programme across every territory without a studio session per language.

Voice CloningMultilingual SupportControllable GenerationText-to-Speech
Voice Cloning logo

Hear the demo

Listen to the reference audio and compare it with the cloned voices in English and German.

Speaker 0
Reference
English
German
Speaker 1
Reference
English
German
Speaker 2
Reference
English
German
Speaker 3
Reference
English
German
Raquel
Reference
English
German
Verena
Reference
English
German

The problem this solves

Multi-language distribution of presenter-led content is a casting problem as much as a translation one. A dubbed track with a different voice changes the perceived identity of the programme. For formats built on a specific presenter - documentary series, news analysis, instructional content - that substitution degrades the product in a way that audiences notice even if they cannot articulate why.

The traditional route is to have the original presenter record in each target language, which requires scheduling, a studio, a session fee and a director in every market. For high-value properties that cost is accepted. For anything else, the language version either does not get made or it gets dubbed by someone else, which is a different product.

Who this is for

This is relevant if you are:

Broadcaster

A broadcaster or rights holder distributing presenter-led formats across multiple language markets

Buyer

A content owner localising a documentary, instructional series or branded content without the budget for multi-territory recording sessions

How it works

  1. 1A short reference audio sample from the target speaker is used to extract a voice profile - the tonal and prosodic characteristics that make the voice recognisable.
  2. 2A translation of the source script is prepared in the target language, preserving sentence rhythm where the target language allows.
  3. 3The Chatterbox TTS model generates speech in the target language using the extracted voice profile, with controllable parameters for temperature, expressiveness and quality.
  4. 4The output is reviewed against the reference to confirm voice identity is preserved across the language boundary.

Chatterbox TTS

Core Voice Synthesis Engine

State-of-the-art text-to-speech model with native voice cloning capabilities.

Multilingual Models

Multi-Language Voice Generation

Support for English, German, Spanish, and more languages via specialized multilingual models.

Controllable Parameters

Fine-Tuned Speech Control

Adjust temperature for variation, exaggeration for expressiveness, and CFG weight for quality vs diversity balance.

GPU Acceleration

High-Performance Processing

CUDA-optimized for fast generation with CPU fallback support for accessibility.

What makes it different

Standard TTS dubbing replaces the original voice with a neutral synthetic one: the presenter's identity is gone. This system clones the voice first, so the target-language version is recognisably the same person. The controllable parameters - temperature, expressiveness, quality weighting - exist specifically to let the output be tuned to match the register and delivery style of the source material rather than defaulting to a generic synthetic vocal character.

Current status & next steps

Ready to pilotReviewed 29 July 2026

Voice Cloning is ready to pilot. The demo on this page shows cloned voices for several speakers across English and German, which represents the current state of the build.

Clone quality follows the quality and the length of the reference audio provided. A CUDA GPU is recommended for usable throughput; CPU-only operation works but is significantly slower. Cross-language cloning is still work in progress, so reference audio in the same language as the target text gives the best results today.

The next step is a pilot against a real localisation project - an existing piece of content that needs a language version - so clone quality can be evaluated against the source material under production conditions.

Interested in a pilot localisation?

Tell us which format you would like to localise and which language markets matter most, and we will tell you what a pilot would involve.