Skip to content
shriza1991Public

About

AI-powered multilingual video dubbing system

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

ย 

History

10 Commits

Folders and files

Repository files navigation

Dubify

AI-powered video dubbing โ€” from YouTube URL to English-dubbed MP4 in one command.

Python 3.11+ License: MIT FFmpeg


Overview

Dubify is a fully automated multilingual video dubbing pipeline. Given a YouTube URL, it:

  1. Downloads the video using yt-dlp
  2. Extracts the audio track via FFmpeg
  3. Transcribes speech in any language using Faster-Whisper (auto-detects source language)
  4. Adapts the script into natural spoken English using Google Gemini 2.5 Flash
  5. Synthesises English audio with Microsoft Edge TTS neural voices
  6. Time-aligns each clip to the original speaker's timing via FFmpeg's atempo filter
  7. Exports the finished MP4 โ€” original video quality preserved, audio replaced

No re-encoding of the video stream. No per-segment API calls for translation. One command.


Architecture

                 +--------------------+
                 |  YouTube Video     |
                 +---------+----------+
                           |
                           v
                 +--------------------+
                 | Download (yt-dlp)  |
                 +---------+----------+
                           |
                           v
                 +--------------------+
                 | Audio Extraction   |
                 | (FFmpeg โ†’ WAV)     |
                 +---------+----------+
                           |
                           v
                 +--------------------+
                 | Faster-Whisper     |
                 | Transcription      |
                 +---------+----------+
                           |
                           v
                 +--------------------+
                 | Gemini Script      |
                 | Adaptation         |
                 +---------+----------+
                           |
                           v
                 +--------------------+
                 | Edge-TTS           |
                 | + atempo Alignment |
                 +---------+----------+
                           |
                           v
                 +--------------------+
                 | Audio Merge        |
                 | (adelay + amix)    |
                 +---------+----------+
                           |
                           v
                 +--------------------+
                 | Final Dubbed Video |
                 +--------------------+

Module responsibilities

Module Class Responsibility
src/models.py Segment, VideoProject, PipelineStage Shared data contracts โ€” no logic
config.py Config All env vars in one place โ€” injected everywhere
src/downloader.py VideoDownloader yt-dlp download with Rich progress bar
src/extractor.py AudioExtractor FFmpeg โ†’ 16 kHz mono WAV for Whisper
src/transcriber.py Transcriber Faster-Whisper speech-to-text, auto language detection
src/translator.py BaseTranslator, GeminiTranslator Batch dubbing script adaptation via Gemini 2.5 Flash
src/synthesizer.py BaseSynthesizer, EdgeTTSSynthesizer Edge TTS + atempo time-stretching
src/audio_composer.py AudioComposer adelay + amix timeline, copy-mux with video
src/pipeline.py DubbingPipeline Orchestration, timing, caching, progress display
src/reporter.py โ€” Terminal summary and processing_report.json
src/utils.py โ€” Logging, file helpers, cleanup
main.py โ€” Thin CLI only โ€” zero business logic

Features

  • ๐ŸŒ Automatic language detection โ€” Whisper identifies the source language without configuration
  • ๐Ÿ“ฆ Single batch translation โ€” all segments sent to Gemini in one API call (10โ€“50ร— cheaper than per-segment calls)
  • โšก Concurrent TTS synthesis โ€” up to 8 Edge-TTS requests in-flight simultaneously (semaphore-limited)
  • ๐ŸŽ›๏ธ Per-segment time alignment โ€” each clip is speed-adjusted via atempo to match original timing
  • ๐Ÿ—ƒ๏ธ Stage caching & resumability โ€” skip completed stages on reruns with --resume (default)
  • ๐Ÿ” Retry logic โ€” transient Edge-TTS failures retried 3ร— with exponential back-off
  • ๐Ÿ“Š Processing reports โ€” human-readable terminal summary + machine-readable processing_report.json
  • ๐ŸŽฌ Lossless video quality โ€” video stream copied with -c:v copy, only audio replaced

Installation

Prerequisites

Setup

# 1. Clone the repository
git clone https://github.com/your-username/Dubify.git
cd Dubify

# 2. Create and activate a virtual environment
python -m venv .venv
.venv\Scripts\activate        # Windows
# source .venv/bin/activate   # macOS / Linux

# 3. Install dependencies
pip install -r requirements.txt

# 4. Configure environment variables
cp .env.example .env
# Edit .env โ€” at minimum set GEMINI_API_KEY

Usage

Basic

python main.py

Dubify will prompt for a YouTube URL interactively.

Pass URL directly

python main.py "https://www.youtube.com/watch?v=dQw4w9WgXcQ"

CLI flags

Flag Description
url YouTube video URL (positional, optional โ€” prompts if omitted)
--resume Use cached stage results (default behaviour)
--force Ignore all caches and rerun the full pipeline
--from STAGE Restart from a specific stage: download / extract / transcribe / translate / synthesize / compose

Examples

# Force full rerun, ignoring all caches
python main.py "https://youtube.com/watch?v=..." --force

# Resume from the TTS stage (re-synthesise but reuse transcription/translation)
python main.py "https://youtube.com/watch?v=..." --from synthesize

Output directories

downloads/    Downloaded source videos
temp/         Intermediate files (auto-cleaned unless KEEP_TEMP=True)
outputs/      Final dubbed videos + processing_report.json + segments.json
logs/         Daily log files  (YYYY-MM-DD.log)

Configuration

All settings are read from a .env file (copy .env.example to get started).

Variable Default Description
GEMINI_API_KEY (required) Google Gemini API key
WHISPER_MODEL base Whisper model: tiny / base / small / medium / large / large-v3
GEMINI_MODEL gemini-2.5-flash Gemini model identifier
EDGE_TTS_VOICE en-US-AndrewNeural Edge TTS voice โ€” run edge-tts --list-voices for all options
OUTPUT_FORMAT mp4 Output container format
KEEP_TEMP False Set True to keep intermediate files (useful for debugging)
FFMPEG_BIN ffmpeg Path to ffmpeg binary
FFPROBE_BIN ffprobe Path to ffprobe binary
MAX_CONCURRENT_TTS 8 Max simultaneous Edge-TTS requests
MIN_ATEMPO 0.75 Lower bound for TTS time-stretching ratio
MAX_ATEMPO 1.35 Upper bound for TTS time-stretching ratio
SKIP_DURATION 0.15 Drop segments shorter than this (seconds)
TARGET_SEGMENT_DURATION 1.5 Target merged-segment minimum duration (seconds)
MAX_SEGMENT_DURATION 8.0 Hard cap on a merged segment group (seconds)
MAX_SEGMENT_GAP 0.5 Gap threshold for starting a new merge group (seconds)

Project Structure

Dubify/
โ”œโ”€โ”€ main.py                  # CLI entry point (no business logic)
โ”œโ”€โ”€ config.py                # Environment-based configuration
โ”œโ”€โ”€ requirements.txt
โ”œโ”€โ”€ .env.example
โ”‚
โ”œโ”€โ”€ src/
โ”‚   โ”œโ”€โ”€ models.py            # Segment, VideoProject, PipelineStage
โ”‚   โ”œโ”€โ”€ pipeline.py          # Stage orchestration and timing
โ”‚   โ”œโ”€โ”€ reporter.py          # Terminal summary + JSON report generation
โ”‚   โ”œโ”€โ”€ downloader.py        # yt-dlp video download
โ”‚   โ”œโ”€โ”€ extractor.py         # FFmpeg audio extraction
โ”‚   โ”œโ”€โ”€ transcriber.py       # Faster-Whisper transcription
โ”‚   โ”œโ”€โ”€ translator.py        # Gemini batch translation
โ”‚   โ”œโ”€โ”€ synthesizer.py       # Edge TTS + atempo time-stretching
โ”‚   โ”œโ”€โ”€ audio_composer.py    # adelay + amix + mux
โ”‚   โ”œโ”€โ”€ cache.py             # Stage caching and resumability
โ”‚   โ””โ”€โ”€ utils.py             # Logging, file helpers, text utilities
โ”‚
โ”œโ”€โ”€ downloads/               # Downloaded source videos
โ”œโ”€โ”€ temp/                    # Intermediate files (auto-cleaned)
โ”œโ”€โ”€ outputs/                 # Final dubbed videos + reports
โ””โ”€โ”€ logs/                    # Daily log files

Technology Stack

Component Technology
Video download yt-dlp
Audio extraction FFmpeg
Transcription Faster-Whisper (CTranslate2-optimised)
Translation / Adaptation Google Gemini 2.5 Flash
Text-to-Speech Microsoft Edge TTS (neural voices, free)
Time alignment FFmpeg atempo filter (chained for arbitrary ratios)
Audio composition FFmpeg adelay + amix filtergraph
CLI / progress Rich
Configuration python-dotenv

Design Decisions

Single VideoProject state object โ€” every stage reads from and writes to the same object rather than passing tuples or dictionaries between functions. Adding a new field (e.g. speaker) requires zero changes to stage signatures.

Batch translation โ€” all segments are sent to Gemini in a single JSON round-trip, not one call per segment. This is 10โ€“50ร— cheaper on longer videos and avoids rate limits.

Dubbing script adaptation, not literal translation โ€” Gemini is prompted as a dubbing expert, not a translator. It is instructed to produce English that sounds like it was originally spoken in English and naturally occupies the same duration as the original segment. This dramatically reduces the amount of atempo stretching needed downstream.

Per-segment atempo time-stretching โ€” after TTS generation, each clip is compared to the original segment duration and speed-adjusted via FFmpeg's atempo filter. This is the single biggest factor in dubbing quality and synchronisation.

No video re-encoding โ€” the composer uses -c:v copy to preserve original video quality. Only the audio stream is replaced.

Abstract provider interfaces โ€” BaseTranslator and BaseSynthesizer decouple the pipeline from specific vendors. Swap Gemini for DeepL, or Edge TTS for ElevenLabs, by implementing one class and changing one constructor argument.

Chunked audio composition โ€” segments are mixed in groups of 60 to stay within FFmpeg's input limits and Windows command-line length constraints on videos with many segments.

Stage caching โ€” each pipeline stage serialises its output to disk. On subsequent runs for the same video, completed stages are restored from cache instantly, enabling rapid iteration on any single stage.


Known Limitations

  • Translation accuracy on very long videos โ€” batching 200+ segments in a single Gemini call may occasionally produce misaligned translations. Consider splitting very long videos at chapter boundaries.
  • Atempo range โ€” FFmpeg's atempo accepts [0.5, 2.0] per filter instance. Dubify chains multiple filters automatically, but extreme ratios (very fast or very slow speakers) may still sound unnatural.
  • CPU Whisper is slow โ€” the base model takes ~2ร— real-time on CPU. Use large-v3 on a GPU for production-quality transcription at faster-than-real-time speed.
  • Edge TTS requires internet โ€” the TTS stage makes network calls to Microsoft's Edge speech service. Offline operation requires switching to a local TTS provider.
  • Single target language โ€” the current pipeline always targets English. Multi-target dubbing requires additional configuration.

Future Improvements

The following features are architecturally prepared (abstract base classes in place) but not yet implemented:

Feature Notes
Speaker diarisation Segment.speaker field is already present; add pyannote.audio
Voice cloning (XTTS) Implement XTTSSynthesizer(BaseSynthesizer)
ElevenLabs TTS Implement ElevenLabsSynthesizer(BaseSynthesizer)
Subtitle generation Export .srt / .vtt from project.segments
Batch processing Accept a list of URLs; VideoProject is fully serialisable
GPU acceleration Pass device="cuda" to Whisper via config
Multiple translation providers DeepLTranslator, IndicTranslator, etc.
Web UI DubbingPipeline.run() is already async-friendly

Performance

Benchmarked on a 10-minute Hindi YouTube video (CPU only, base Whisper model):

Stage Time
Download ~8s
Audio Extraction ~3s
Transcription (CPU, base) ~2m 30s
Translation (Gemini batch) ~45s
Speech Synthesis (8 concurrent) ~3m 00s
Compose & Merge ~45s
Total ~7m 10s

Realtime factor: ~0.7ร— (faster than real-time on a 10-minute video). Using large-v3 on a CUDA GPU reduces transcription to under 1 minute.


Sample Terminal Output

โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚    ๐ŸŽฌ  Dubify  โ€”  AI Video Dubbing System                โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ

โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ Pipeline โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€

  โœ“ Downloading video                (8s)  โ€” video_title.mp4
  โœ“ Extracting audio                 (3s)  โ€” video_title_audio.wav
  โ–ธ Transcribing speechโ€ฆ
    โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆ 212 / 212  0:02:42
  โœ“ Transcribing speech             (2m 42s)  โ€” HI ยท 212 segments
  โœ“ Translating & adapting script   (1m 12s)  โ€” 212 segments โ†’ English
  โ–ธ Synthesizing speechโ€ฆ
    โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–‘โ–‘ 191 / 212  0:06:38
  โœ“ Synthesizing speech             (6m 38s)  โ€” 212 clips generated
  โœ“ Merging dubbed audio & exporting (51s)  โ€” outputs/video_title_dubbed.mp4

โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ โœ“  Processing Complete โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚                                                              โ”‚
โ”‚  Input Video        https://youtube.com/watch?v=...          โ”‚
โ”‚  Output             outputs/video_title_dubbed.mp4           โ”‚
โ”‚                                                              โ”‚
โ”‚  Video Duration     11m 02s                                  โ”‚
โ”‚  Processing Time    12m 31s                                  โ”‚
โ”‚  Realtime Factor    1.13x                                    โ”‚
โ”‚                                                              โ”‚
โ”‚  โ”€โ”€ Stage Times โ”€โ”€                                           โ”‚
โ”‚    Download         8s                                       โ”‚
โ”‚    Audio Extraction 3s                                       โ”‚
โ”‚    Transcription    2m 42s                                   โ”‚
โ”‚    Translation      1m 12s                                   โ”‚
โ”‚    Speech Synthesis 6m 38s                                   โ”‚
โ”‚    Compose & Merge  51s                                      โ”‚
โ”‚                                                              โ”‚
โ”‚  โ”€โ”€ Segments โ”€โ”€                                              โ”‚
โ”‚    Total Segments   212                                      โ”‚
โ”‚    Translated       212                                      โ”‚
โ”‚    Synthesized      212                                      โ”‚
โ”‚    Tempo Adjusted   174                                      โ”‚
โ”‚    Avg Tempo Ratio  0.91                                     โ”‚
โ”‚                                                              โ”‚
โ”‚  Detected Language  HI                                       โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ

๐Ÿ“„ Report saved: outputs/processing_report.json

Sample Processing Report

outputs/processing_report.json:

{
    "input_video": "https://www.youtube.com/watch?v=...",
    "output_video": "outputs/video_title_dubbed.mp4",

    "video_duration_seconds": 662.0,
    "processing_time_seconds": 751.4,
    "realtime_factor": 1.1349,

    "stage_times": {
        "download": 8.1,
        "extract_audio": 3.2,
        "transcription": 162.4,
        "translation": 72.1,
        "tts": 398.3,
        "time_alignment": 0.0,
        "merge": 51.4
    },

    "segments": {
        "total": 212,
        "translated": 212,
        "synthesized": 212,
        "tempo_adjusted": 174,
        "average_ratio": 0.9147
    }
}

License

MIT

About

AI-powered multilingual video dubbing system

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages