DeepDub

We live in a highly connected world. Humans around the world, of all cultures, are consuming entertainment at a pace faster than ever before in history. Media is quickly being imported and exported between countries via the Internet. In this media, the visual medium is of great power. It can

2025-06-28 16:26:05 - Adil Khan

Project Title

DeepDub

Project Area of Specialization Artificial IntelligenceProject Summary

We live in a highly connected world. Humans around the world, of all cultures, are consuming entertainment at a pace faster than ever before in history. Media is quickly being imported and exported between countries via the Internet.

In this media, the visual medium is of great power. It can transfer a state of human presence as-is would be seen from the human eyes and ears. However it brings with itself a classic issue since even before all this technology; the clash of languages and culture. However much English may be considered as a “default” medium, only 25% of the world has an even faint understanding of what “English” is to begin with. So keeping that in mind, English entertainment is not the silver bullet for global entertainment, as it may not suit viewers of other cultures.

Our Final Year Project, named “DeepDub”, aims to provide seamless translations of videos. It will achieve this by:

This will make dubbed videos much more entertaining and relatable to viewers of different languages. It will also reduce the cost of localization exponentially, as it leads to easier retakes and adjustments. In the best case, there wouldn’t even be a voice actor; hence also saving HR costs and time.

Project Objectives
  1. Create a working system/engine which can translate videos from one language to another

  2. Support Urdu, English and Turkish

  3. Create practical applications (web app and/or mobile app) which demonstrate the power of the engine

  4. Produce high quality results that are acceptable to a general audience

  5. Write at least 2 research papers on the idea and the implementation (Research Gap)

  6. Create an End-to-End system for quicker results

Project Implementation Method

Our proposed implementation requires a video, the source language and a target language as in input from the user. The proposed implementation comprises multiple independent modules which work in a sequential order. Following is a brief overview of each module:

  1. The first module is an Automatic Speech Recognition model which recognises and translates speech utterances made in the input video. The current implementation uses the latest open source ASR model wav2vec2. The output of this module is a raw transcription file containing speech utterances the ASR model has predicted.

  2. This raw transcription file acts as an input for the next module which is a Forced Aligner the purpose of which is to timestamp the speech utterances relative to the time they were spoken at in the input video. The output of this module is a word by word Time Stamped Transcript file. The current implementation uses Aeneas open source model to perform this transcription.

  3. The time stamped transcript file is then given as input to our custom built Clustering module which aims to cluster the transcripted words into transcripted sentences by calculating the difference between each timestamped word and clustering those which have specifically close differences between them.

  4. At this point, we obtain timestamped sentences which essentially is an SRT file. This SRT file is then given input to a Text-to-Text translation module which converts the original source text into the target language the user has provided as input. The current implementation currently uses Opus open source Text-to-Text translation module which has a wide range of supported languages. The output of this module is a Translated Time stamped Transcript (SRT) file.

  5. This Translated SRT file along with the input video act as an input to the next module which is the Text-to-Speech module which aims to convert the text translated into audio. The current implementation uses an open source module named Real Time Voice Cloning (an unofficial implementation of the research paper “Transfer Learning from Speaker Verification to Multispeaker Text-To-Speech Synthesis” ) which not only performs text to speech but also clones the speaker's voice from the. The output of this module is an synthesized Audio file based on the translated SRT file.

  6. The final module is a GAN based Speech-to-Lip ( Lip synchronization model ) which takes in the synthesized audio and the input video as its input. It utilizes these inputs to synchronize the lips of the actor present in the video according to the audio file provided. This module is integrated into the pipeline to make the output much more immersive and natural. The current implementation uses an open source model Wav2Lip based on the research “A Lip Sync Expert Is All You Need for Speech to Lip Generation In the Wild” paper. The output of this module is a completely translated and visually synchronized video according to the user's specified target language.

Benefits of the Project Statistics, Observations and Benefits

We observed that (Globally):

Target Industries: 

This project is aimed to facilitate faster, cost efficient and higher quality dubbed video content production. Therefore the applications of this project covers a broad range of video content production industries. These industries include:

Technical Details of Final Deliverable

Project architecture:

FASTAPI is used to connect the frontend with the backend

Final Deliverable of the Project Software SystemCore Industry ITOther Industries Media , Telecommunication Core Technology Artificial Intelligence(AI)Other TechnologiesSustainable Development Goals Industry, Innovation and Infrastructure, Responsible Consumption and ProductionRequired Resources
Item Name Type No. of Units Per Unit Cost (in Rs) Total (in Rs)
Total in (Rs) 68000
Inno3D Geforce GTX 1650 TWIN X2 OC 4GB GDDR6 Video Graphics Card Equipment16800068000

More Posts