πŸ“ž +91-7667918914 | βœ‰οΈ ijarcce@gmail.com
International Journal of Advanced Research in Computer and Communication Engineering
International Journal of Advanced Research in Computer and Communication Engineering A monthly Peer-reviewed & Refereed journal
ISSN Online 2278-1021ISSN Print 2319-5940Since 2012
IJARCCE adheres to the suggestive parameters outlined by the University Grants Commission (UGC) for peer-reviewed journals, upholding high standards of research quality, ethical publishing, and academic excellence.
← Back to VOLUME 15, ISSUE 8, AUGUST 2026

AuraVoice: An End-to-End Multilingual Voice Cloning and Video Dubbing Pipeline Integrating OpenVoice V2, Whisper, and MeloTTS

Sagar M, Rakshith G C, Dr. Dharani N V

πŸ‘ 7 viewsπŸ“₯ 5 downloads
Share: 𝕏 f in ✈ βœ‰
Abstract: Creating multilingual voice content today typically requires re-recording narration in every target language or stitching together separate tools for transcription, translation, and speech synthesis. This paper presents AuraVoice, an end-to-end web-based platform that unifies voice cloning, cross-lingual speech translation, and automated video dubbing in a single pipeline. The system combines OpenVoice V2 for zero-shot tone-color conversion, faster-whisper for automatic speech recognition, MeloTTS for multilingual base speech synthesis, and FFmpeg for audio/video processing, orchestrated through a FastAPI backend with a lightweight browser-based frontend. Given a short reference audio sample, AuraVoice extracts speaker tone characteristics and reproduces them in synthesized speech across six languages while preserving the original speaker's identity. The video dubbing module extends this pipeline to full videos by extracting the source audio, translating and re-synthesizing it in the target language, and re-muxing it with the original visual track. We describe the system architecture, module-level implementation, and lazy-loading/caching strategy used to manage GPU memory across multiple deep learning models, and discuss the functional testing performed on each module. The platform demonstrates that recent advances in zero-shot voice cloning and open automatic speech recognition can be composed into a practical, integrated tool for multilingual content creation, reducing the manual effort required compared to disjoint single-purpose tools.

Keywords: Voice Cloning, Tone-Color Conversion, OpenVoice, Text-to-Speech, Speech Translation, Video Dubbing, Whisper, FastAPI, Speech Synthesis

How to Cite:

[1] Sagar M, Rakshith G C, Dr. Dharani N V, β€œAuraVoice: An End-to-End Multilingual Voice Cloning and Video Dubbing Pipeline Integrating OpenVoice V2, Whisper, and MeloTTS,” International Journal of Advanced Research in Computer and Communication Engineering (IJARCCE), DOI: 10.17148/IJARCCE.2026.15803

Creative Commons License This work is licensed under a Creative Commons Attribution 4.0 International License.