HyperAIHyperAI

Command Palette

Search for a command to run...

3 months ago

Visual Speech Recognition in a Driver Assistance System

{Alexey Karpov Alexandr Axyonov Alexey Kashevnik Dmitry Ryumin Denis Ivanko}

Visual Speech Recognition in a Driver Assistance System

Abstract

Visual speech recognition or automated lipreading is a field of growing attention. Video data proved its usefulness in multimodal speech recognition, especially when acoustic data is heavily noised or even inaccessible. In this paper, we present a novel method for visual speech recognition. We benchmark it on the famous LRW lip-reading dataset by outperforming the existing approaches. After a comprehensive evaluation, we adapt the developed method and test it on the collected RUSAVIC corpus we recorded in-the-wild for vehicle driver. The results obtained demonstrate not only the high performance of the proposed method, but also the fundamental possibility of recognizing speech only by using video modality, even in such difficult natural conditions as driving.

Benchmarks

BenchmarkMethodologyMetrics
lipreading-on-lip-reading-in-the-wildVosk + MediaPipe + LS + MixUp + SA + 3DResNet-18 + BiLSTM + Cosine WR
Top-1 Accuracy: 88.7

Build AI with AI

From idea to launch — accelerate your AI development with free AI co-coding, out-of-the-box environment and best price of GPUs.

AI Co-coding
Ready-to-use GPUs
Best Pricing
Get Started

Hyper Newsletters

Subscribe to our latest updates
We will deliver the latest updates of the week to your inbox at nine o'clock every Monday morning
Powered by MailChimp
Visual Speech Recognition in a Driver Assistance System | Papers | HyperAI