AI-powered video dubbing โ from YouTube URL to English-dubbed MP4 in one command.
Dubify is a fully automated multilingual video dubbing pipeline. Given a YouTube URL, it:
- Downloads the video using
yt-dlp - Extracts the audio track via FFmpeg
- Transcribes speech in any language using Faster-Whisper (auto-detects source language)
- Adapts the script into natural spoken English using Google Gemini 2.5 Flash
- Synthesises English audio with Microsoft Edge TTS neural voices
- Time-aligns each clip to the original speaker's timing via FFmpeg's
atempofilter - Exports the finished MP4 โ original video quality preserved, audio replaced
No re-encoding of the video stream. No per-segment API calls for translation. One command.
+--------------------+
| YouTube Video |
+---------+----------+
|
v
+--------------------+
| Download (yt-dlp) |
+---------+----------+
|
v
+--------------------+
| Audio Extraction |
| (FFmpeg โ WAV) |
+---------+----------+
|
v
+--------------------+
| Faster-Whisper |
| Transcription |
+---------+----------+
|
v
+--------------------+
| Gemini Script |
| Adaptation |
+---------+----------+
|
v
+--------------------+
| Edge-TTS |
| + atempo Alignment |
+---------+----------+
|
v
+--------------------+
| Audio Merge |
| (adelay + amix) |
+---------+----------+
|
v
+--------------------+
| Final Dubbed Video |
+--------------------+
| Module | Class | Responsibility |
|---|---|---|
src/models.py |
Segment, VideoProject, PipelineStage |
Shared data contracts โ no logic |
config.py |
Config |
All env vars in one place โ injected everywhere |
src/downloader.py |
VideoDownloader |
yt-dlp download with Rich progress bar |
src/extractor.py |
AudioExtractor |
FFmpeg โ 16 kHz mono WAV for Whisper |
src/transcriber.py |
Transcriber |
Faster-Whisper speech-to-text, auto language detection |
src/translator.py |
BaseTranslator, GeminiTranslator |
Batch dubbing script adaptation via Gemini 2.5 Flash |
src/synthesizer.py |
BaseSynthesizer, EdgeTTSSynthesizer |
Edge TTS + atempo time-stretching |
src/audio_composer.py |
AudioComposer |
adelay + amix timeline, copy-mux with video |
src/pipeline.py |
DubbingPipeline |
Orchestration, timing, caching, progress display |
src/reporter.py |
โ | Terminal summary and processing_report.json |
src/utils.py |
โ | Logging, file helpers, cleanup |
main.py |
โ | Thin CLI only โ zero business logic |
- ๐ Automatic language detection โ Whisper identifies the source language without configuration
- ๐ฆ Single batch translation โ all segments sent to Gemini in one API call (10โ50ร cheaper than per-segment calls)
- โก Concurrent TTS synthesis โ up to 8 Edge-TTS requests in-flight simultaneously (semaphore-limited)
- ๐๏ธ Per-segment time alignment โ each clip is speed-adjusted via
atempoto match original timing - ๐๏ธ Stage caching & resumability โ skip completed stages on reruns with
--resume(default) - ๐ Retry logic โ transient Edge-TTS failures retried 3ร with exponential back-off
- ๐ Processing reports โ human-readable terminal summary + machine-readable
processing_report.json - ๐ฌ Lossless video quality โ video stream copied with
-c:v copy, only audio replaced
- Python 3.11+
- FFmpeg (must be on PATH or configured via
.env)- Windows:
winget install ffmpegor download from ffmpeg.org - macOS:
brew install ffmpeg - Linux:
sudo apt install ffmpeg
- Windows:
- Google Gemini API key โ get one free at aistudio.google.com
# 1. Clone the repository
git clone https://github.com/your-username/Dubify.git
cd Dubify
# 2. Create and activate a virtual environment
python -m venv .venv
.venv\Scripts\activate # Windows
# source .venv/bin/activate # macOS / Linux
# 3. Install dependencies
pip install -r requirements.txt
# 4. Configure environment variables
cp .env.example .env
# Edit .env โ at minimum set GEMINI_API_KEYpython main.pyDubify will prompt for a YouTube URL interactively.
python main.py "https://www.youtube.com/watch?v=dQw4w9WgXcQ"| Flag | Description |
|---|---|
url |
YouTube video URL (positional, optional โ prompts if omitted) |
--resume |
Use cached stage results (default behaviour) |
--force |
Ignore all caches and rerun the full pipeline |
--from STAGE |
Restart from a specific stage: download / extract / transcribe / translate / synthesize / compose |
# Force full rerun, ignoring all caches
python main.py "https://youtube.com/watch?v=..." --force
# Resume from the TTS stage (re-synthesise but reuse transcription/translation)
python main.py "https://youtube.com/watch?v=..." --from synthesizedownloads/ Downloaded source videos
temp/ Intermediate files (auto-cleaned unless KEEP_TEMP=True)
outputs/ Final dubbed videos + processing_report.json + segments.json
logs/ Daily log files (YYYY-MM-DD.log)
All settings are read from a .env file (copy .env.example to get started).
| Variable | Default | Description |
|---|---|---|
GEMINI_API_KEY |
(required) | Google Gemini API key |
WHISPER_MODEL |
base |
Whisper model: tiny / base / small / medium / large / large-v3 |
GEMINI_MODEL |
gemini-2.5-flash |
Gemini model identifier |
EDGE_TTS_VOICE |
en-US-AndrewNeural |
Edge TTS voice โ run edge-tts --list-voices for all options |
OUTPUT_FORMAT |
mp4 |
Output container format |
KEEP_TEMP |
False |
Set True to keep intermediate files (useful for debugging) |
FFMPEG_BIN |
ffmpeg |
Path to ffmpeg binary |
FFPROBE_BIN |
ffprobe |
Path to ffprobe binary |
MAX_CONCURRENT_TTS |
8 |
Max simultaneous Edge-TTS requests |
MIN_ATEMPO |
0.75 |
Lower bound for TTS time-stretching ratio |
MAX_ATEMPO |
1.35 |
Upper bound for TTS time-stretching ratio |
SKIP_DURATION |
0.15 |
Drop segments shorter than this (seconds) |
TARGET_SEGMENT_DURATION |
1.5 |
Target merged-segment minimum duration (seconds) |
MAX_SEGMENT_DURATION |
8.0 |
Hard cap on a merged segment group (seconds) |
MAX_SEGMENT_GAP |
0.5 |
Gap threshold for starting a new merge group (seconds) |
Dubify/
โโโ main.py # CLI entry point (no business logic)
โโโ config.py # Environment-based configuration
โโโ requirements.txt
โโโ .env.example
โ
โโโ src/
โ โโโ models.py # Segment, VideoProject, PipelineStage
โ โโโ pipeline.py # Stage orchestration and timing
โ โโโ reporter.py # Terminal summary + JSON report generation
โ โโโ downloader.py # yt-dlp video download
โ โโโ extractor.py # FFmpeg audio extraction
โ โโโ transcriber.py # Faster-Whisper transcription
โ โโโ translator.py # Gemini batch translation
โ โโโ synthesizer.py # Edge TTS + atempo time-stretching
โ โโโ audio_composer.py # adelay + amix + mux
โ โโโ cache.py # Stage caching and resumability
โ โโโ utils.py # Logging, file helpers, text utilities
โ
โโโ downloads/ # Downloaded source videos
โโโ temp/ # Intermediate files (auto-cleaned)
โโโ outputs/ # Final dubbed videos + reports
โโโ logs/ # Daily log files
| Component | Technology |
|---|---|
| Video download | yt-dlp |
| Audio extraction | FFmpeg |
| Transcription | Faster-Whisper (CTranslate2-optimised) |
| Translation / Adaptation | Google Gemini 2.5 Flash |
| Text-to-Speech | Microsoft Edge TTS (neural voices, free) |
| Time alignment | FFmpeg atempo filter (chained for arbitrary ratios) |
| Audio composition | FFmpeg adelay + amix filtergraph |
| CLI / progress | Rich |
| Configuration | python-dotenv |
Single VideoProject state object โ every stage reads from and writes to the same object rather than passing tuples or dictionaries between functions. Adding a new field (e.g. speaker) requires zero changes to stage signatures.
Batch translation โ all segments are sent to Gemini in a single JSON round-trip, not one call per segment. This is 10โ50ร cheaper on longer videos and avoids rate limits.
Dubbing script adaptation, not literal translation โ Gemini is prompted as a dubbing expert, not a translator. It is instructed to produce English that sounds like it was originally spoken in English and naturally occupies the same duration as the original segment. This dramatically reduces the amount of atempo stretching needed downstream.
Per-segment atempo time-stretching โ after TTS generation, each clip is compared to the original segment duration and speed-adjusted via FFmpeg's atempo filter. This is the single biggest factor in dubbing quality and synchronisation.
No video re-encoding โ the composer uses -c:v copy to preserve original video quality. Only the audio stream is replaced.
Abstract provider interfaces โ BaseTranslator and BaseSynthesizer decouple the pipeline from specific vendors. Swap Gemini for DeepL, or Edge TTS for ElevenLabs, by implementing one class and changing one constructor argument.
Chunked audio composition โ segments are mixed in groups of 60 to stay within FFmpeg's input limits and Windows command-line length constraints on videos with many segments.
Stage caching โ each pipeline stage serialises its output to disk. On subsequent runs for the same video, completed stages are restored from cache instantly, enabling rapid iteration on any single stage.
- Translation accuracy on very long videos โ batching 200+ segments in a single Gemini call may occasionally produce misaligned translations. Consider splitting very long videos at chapter boundaries.
- Atempo range โ FFmpeg's
atempoaccepts[0.5, 2.0]per filter instance. Dubify chains multiple filters automatically, but extreme ratios (very fast or very slow speakers) may still sound unnatural. - CPU Whisper is slow โ the
basemodel takes ~2ร real-time on CPU. Uselarge-v3on a GPU for production-quality transcription at faster-than-real-time speed. - Edge TTS requires internet โ the TTS stage makes network calls to Microsoft's Edge speech service. Offline operation requires switching to a local TTS provider.
- Single target language โ the current pipeline always targets English. Multi-target dubbing requires additional configuration.
The following features are architecturally prepared (abstract base classes in place) but not yet implemented:
| Feature | Notes |
|---|---|
| Speaker diarisation | Segment.speaker field is already present; add pyannote.audio |
| Voice cloning (XTTS) | Implement XTTSSynthesizer(BaseSynthesizer) |
| ElevenLabs TTS | Implement ElevenLabsSynthesizer(BaseSynthesizer) |
| Subtitle generation | Export .srt / .vtt from project.segments |
| Batch processing | Accept a list of URLs; VideoProject is fully serialisable |
| GPU acceleration | Pass device="cuda" to Whisper via config |
| Multiple translation providers | DeepLTranslator, IndicTranslator, etc. |
| Web UI | DubbingPipeline.run() is already async-friendly |
Benchmarked on a 10-minute Hindi YouTube video (CPU only, base Whisper model):
| Stage | Time |
|---|---|
| Download | ~8s |
| Audio Extraction | ~3s |
| Transcription (CPU, base) | ~2m 30s |
| Translation (Gemini batch) | ~45s |
| Speech Synthesis (8 concurrent) | ~3m 00s |
| Compose & Merge | ~45s |
| Total | ~7m 10s |
Realtime factor: ~0.7ร (faster than real-time on a 10-minute video).
Using large-v3 on a CUDA GPU reduces transcription to under 1 minute.
โญโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฎ
โ ๐ฌ Dubify โ AI Video Dubbing System โ
โฐโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฏ
โโโโโโโโโโโโโโโโโโโโ Pipeline โโโโโโโโโโโโโโโโโโโโโ
โ Downloading video (8s) โ video_title.mp4
โ Extracting audio (3s) โ video_title_audio.wav
โธ Transcribing speechโฆ
โโโโโโโโโโโโโโโโโโโโ 212 / 212 0:02:42
โ Transcribing speech (2m 42s) โ HI ยท 212 segments
โ Translating & adapting script (1m 12s) โ 212 segments โ English
โธ Synthesizing speechโฆ
โโโโโโโโโโโโโโโโโโโโ 191 / 212 0:06:38
โ Synthesizing speech (6m 38s) โ 212 clips generated
โ Merging dubbed audio & exporting (51s) โ outputs/video_title_dubbed.mp4
โญโโโโโโโโโโโโโโโโโโโ โ Processing Complete โโโโโโโโโโโโโโโโโโโโฎ
โ โ
โ Input Video https://youtube.com/watch?v=... โ
โ Output outputs/video_title_dubbed.mp4 โ
โ โ
โ Video Duration 11m 02s โ
โ Processing Time 12m 31s โ
โ Realtime Factor 1.13x โ
โ โ
โ โโ Stage Times โโ โ
โ Download 8s โ
โ Audio Extraction 3s โ
โ Transcription 2m 42s โ
โ Translation 1m 12s โ
โ Speech Synthesis 6m 38s โ
โ Compose & Merge 51s โ
โ โ
โ โโ Segments โโ โ
โ Total Segments 212 โ
โ Translated 212 โ
โ Synthesized 212 โ
โ Tempo Adjusted 174 โ
โ Avg Tempo Ratio 0.91 โ
โ โ
โ Detected Language HI โ
โฐโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฏ
๐ Report saved: outputs/processing_report.json
outputs/processing_report.json:
{
"input_video": "https://www.youtube.com/watch?v=...",
"output_video": "outputs/video_title_dubbed.mp4",
"video_duration_seconds": 662.0,
"processing_time_seconds": 751.4,
"realtime_factor": 1.1349,
"stage_times": {
"download": 8.1,
"extract_audio": 3.2,
"transcription": 162.4,
"translation": 72.1,
"tts": 398.3,
"time_alignment": 0.0,
"merge": 51.4
},
"segments": {
"total": 212,
"translated": 212,
"synthesized": 212,
"tempo_adjusted": 174,
"average_ratio": 0.9147
}
}MIT