Using neural networks (CNNs and LSTMs) to detect talking faces in videos, using both visual and audio modalities.