← Back to VOLUME 15, ISSUE 8, AUGUST 2026
This work is licensed under a Creative Commons Attribution 4.0 International License.
AuraVoice: An End-to-End Multilingual Voice Cloning and Video Dubbing Pipeline Integrating OpenVoice V2, Whisper, and MeloTTS
Sagar M, Rakshith G C, Dr. Dharani N V
π 7 viewsπ₯ 5 downloads
Abstract: Creating multilingual voice content today typically requires re-recording narration in every target language or stitching together separate tools for transcription, translation, and speech synthesis. This paper presents AuraVoice, an end-to-end web-based platform that unifies voice cloning, cross-lingual speech translation, and automated video dubbing in a single pipeline. The system combines OpenVoice V2 for zero-shot tone-color conversion, faster-whisper for automatic speech recognition, MeloTTS for multilingual base speech synthesis, and FFmpeg for audio/video processing, orchestrated through a FastAPI backend with a lightweight browser-based frontend. Given a short reference audio sample, AuraVoice extracts speaker tone characteristics and reproduces them in synthesized speech across six languages while preserving the original speaker's identity. The video dubbing module extends this pipeline to full videos by extracting the source audio, translating and re-synthesizing it in the target language, and re-muxing it with the original visual track. We describe the system architecture, module-level implementation, and lazy-loading/caching strategy used to manage GPU memory across multiple deep learning models, and discuss the functional testing performed on each module. The platform demonstrates that recent advances in zero-shot voice cloning and open automatic speech recognition can be composed into a practical, integrated tool for multilingual content creation, reducing the manual effort required compared to disjoint single-purpose tools.
Keywords: Voice Cloning, Tone-Color Conversion, OpenVoice, Text-to-Speech, Speech Translation, Video Dubbing, Whisper, FastAPI, Speech Synthesis
Keywords: Voice Cloning, Tone-Color Conversion, OpenVoice, Text-to-Speech, Speech Translation, Video Dubbing, Whisper, FastAPI, Speech Synthesis
How to Cite:
[1] Sagar M, Rakshith G C, Dr. Dharani N V, βAuraVoice: An End-to-End Multilingual Voice Cloning and Video Dubbing Pipeline Integrating OpenVoice V2, Whisper, and MeloTTS,β International Journal of Advanced Research in Computer and Communication Engineering (IJARCCE), DOI: 10.17148/IJARCCE.2026.15803
