📞 +91-7667918914 | ✉️ ijarcce@gmail.com
International Journal of Advanced Research in Computer and Communication Engineering
International Journal of Advanced Research in Computer and Communication Engineering A monthly Peer-reviewed & Refereed journal
ISSN Online 2278-1021ISSN Print 2319-5940Since 2012
IJARCCE adheres to the suggestive parameters outlined by the University Grants Commission (UGC) for peer-reviewed journals, upholding high standards of research quality, ethical publishing, and academic excellence.
← Back to VOLUME 15, ISSUE 7, JULY 2026

Hybrid 3DCvT-BiGRU Model for Lip Reading Recognition

Manjunath Hosamani, Mohammed Zubair, Nithin C., Rakesh, and Shwetha Kamath

👁 21 views📥 5 downloads
Share: 𝕏 f in

Abstract. Lip reading — the process of understanding spoken language by visually interpreting lip movements — has gained significant importance in scenarios where audio is unclear, noisy, or unavailable. Despite considerable progress in deep-learning-based visual speech recognition, existing systems continue to suffer from inadequate fine-grained spatial detail, weak temporal modelling, limited generalisation across speakers, and prohibitive computational demands in real-time deployment settings. This paper presents a novel automatic lip-reading system that converts silent video into readable text through a hybrid deep-learning pipeline. The proposed architecture integrates a 3D Convolutional Vision Transformer (3DCvT) — which captures rich spatiotemporal features across both local motion patterns and global spatial context — with a Bidirectional Gated Recurrent Unit (BiGRU) network that models sequential lip-movement dependencies in both the forward and backward temporal directions. Connectionist Temporal Classification (CTC) is employed for label-sequence alignment during training and inference. Experimental evaluation on lip-movement video data demonstrates that the proposed model achieves 96% word-level recognition accuracy with low inference latency, outperforming a standalone 3D CNN baseline (87%) and a pure Transformer model (92%). The system is designed for deployment in real-world accessibility applications, including assistive communication tools for hearing-impaired individuals, speech recognition in acoustically challenging environments, and silent-video analysis for surveillance and security contexts. The paper further provides a detailed critical analysis of limitations in current models and outlines directions for future research including speaker adaptation, audio-visual fusion, and multilingual dataset construction. 

Keywords: Lip Reading · Visual Speech Recognition · 3D Convolutional Vision Transformer (3DCvT) · Bidirectional GRU · Deep Learning · Spatiotemporal Feature Extraction · Connectionist Temporal Classification · Silent Video Translation · Sequence Modelling · Accessibility Technology.

How to Cite:

[1] Manjunath Hosamani, Mohammed Zubair, Nithin C., Rakesh, and Shwetha Kamath, “Hybrid 3DCvT-BiGRU Model for Lip Reading Recognition,” International Journal of Advanced Research in Computer and Communication Engineering (IJARCCE), DOI: 10.17148/IJARCCE.2026.15739

Creative Commons License This work is licensed under a Creative Commons Attribution 4.0 International License.