@dblp

Improving Audio-Visual Speech Recognition by Lip-Subword Correlation Based Visual Pre-training and Cross-Modal Fusion Encoder.

, , , , , , and . CoRR, (2023)

Links and resources

Tags