Hybrid 3DCvT-BiGRU Model for Lip Reading Recognition
Manjunath Hosamani, Mohammed Zubair, Nithin C., Rakesh, and Shwetha Kamath
Abstract. Lip reading — the process of understanding spoken language by visually interpreting lip movements — has gained significant importance in scenarios where audio is unclear, noisy, or unavailable. Despite considerable progress in deep-learning-based visual speech recognition, existing systems continue to suffer from inadequate fine-grained spatial detail, weak temporal modelling, limited generalisation across speakers, and prohibitive computational demands in real-time deployment settings. This paper presents a novel automatic lip-reading system that converts silent video into readable text through a hybrid deep-learning pipeline. The proposed architecture integrates a 3D Convolutional Vision Transformer (3DCvT) — which captures rich spatiotemporal features across both local motion patterns and global spatial context — with a Bidirectional Gated Recurrent Unit (BiGRU) network that models sequential lip-movement dependencies in both the forward and backward temporal directions. Connectionist Temporal Classification (CTC) is employed for label-sequence alignment during training and inference. Experimental evaluation on lip-movement video data demonstrates that the proposed model achieves 96% word-level recognition accuracy with low inference latency, outperforming a standalone 3D CNN baseline (87%) and a pure Transformer model (92%). The system is designed for deployment in real-world accessibility applications, including assistive communication tools for hearing-impaired individuals, speech recognition in acoustically challenging environments, and silent-video analysis for surveillance and security contexts. The paper further provides a detailed critical analysis of limitations in current models and outlines directions for future research including speaker adaptation, audio-visual fusion, and multilingual dataset construction.
Keywords: Lip Reading · Visual Speech Recognition · 3D
Convolutional Vision Transformer (3DCvT) · Bidirectional GRU · Deep Learning · Spatiotemporal Feature Extraction · Connectionist
Temporal Classification · Silent Video Translation · Sequence Modelling ·
Accessibility Technology.
How to Cite:
[1] Manjunath Hosamani, Mohammed Zubair, Nithin C., Rakesh, and Shwetha Kamath, “Hybrid 3DCvT-BiGRU Model for Lip Reading Recognition,” International Journal of Advanced Research in Computer and Communication Engineering (IJARCCE), DOI: 10.17148/IJARCCE.2026.15739
