← Back to VOLUME 15, ISSUE 8, AUGUST 2026
This work is licensed under a Creative Commons Attribution 4.0 International License.
Mitro: Calibrated Multimodal Gesture, Voice, and LLM Control
Srijan Mani Tripathi, Vyom Pandey, Mohammad Rayyan Basha, Dr. Beenarani Manoj
π 5 viewsπ₯ 4 downloads
Abstract: This paper presents Mitro, a desktop human-computer interaction (HCI) system that fuses three input modalities into a single interaction pipeline: (1) computer-vision-based hand-gesture recognition for direct manipulation of the operating system (cursor control, clicks, scroll, volume, and brightness); (2) speech-based natural-language commanding through automatic speech recognition (ASR) and text-to-speech (TTS); and (3) a large language model (LLM) reasoning layer that replaces brittle keyword-matched command dispatch with open-domain natural-language understanding and tool/function calling. Mitro is built as an upgrade to an existing rule-based gesture-and-voice controller that suffered from three concrete limitations: order-sensitive substring-matched command routing, a single hardcoded audio energy threshold that did not generalize across acoustic environments, and a single hardcoded gesture-classification threshold with a fragile single-frame-reset debounce scheme. This work addresses all three limitations with a unified calibrate-once-then-lock design pattern applied consistently across the audio and vision pipelines, and generalizes command handling from rule-based matching to LLM-mediated tool dispatch while preserving deterministic, network- independent fast paths for safety-critical control flow. Gesture-state changes, voice transcripts, and LLM responses are surfaced into a single chronological interaction log, providing an observable, demonstrable trace of the system's multimodal behavior. Section IV outlines the evaluation methodology used to measure calibration benefit, gesture- smoothing benefit, command-routing robustness, and end-to-end latency, together with the qualitative findings obtained during development; the corresponding quantitative results are to be reported once the described experiments are executed on the deployed system.
Keywords: Multimodal Human-Computer Interaction, Gesture Recognition, Speech Recognition, Large Language Models, Tool Calling, Adaptive Calibration, Temporal Smoothing, Cross-Modal Fusion, MediaPipe, Groq
Keywords: Multimodal Human-Computer Interaction, Gesture Recognition, Speech Recognition, Large Language Models, Tool Calling, Adaptive Calibration, Temporal Smoothing, Cross-Modal Fusion, MediaPipe, Groq
How to Cite:
[1] Srijan Mani Tripathi, Vyom Pandey, Mohammad Rayyan Basha, Dr. Beenarani Manoj, βMitro: Calibrated Multimodal Gesture, Voice, and LLM Control,β International Journal of Advanced Research in Computer and Communication Engineering (IJARCCE), DOI: 10.17148/IJARCCE.2026.15805
