Inproceedings,

Improving Audio-Visual Speech Recognition by Lip-Subword Correlation Based Visual Pre-training and Cross-Modal Fusion Encoder.

, , , , , , and .
ICME, page 2627-2632. IEEE, (2023)

Meta data

Tags

Users

  • @dblp

Comments and Reviews