GCPR-2021 AudioCLIP:Extending CLIP to Image, Text and Audio
AudioCLIP: tri-modal contrastive learning on video data
Abstract
Adds audio as a third modality alongside text and image, integrating the ESResNeXt audio model into the CLIP framework using audio datasets.
Model

Several video datasets provide aligned text, image, and audio within each clip; the architecture follows CLIP and incorporates all three modalities. Because the modalities are paired, contrastive learning is straightforward.
Result


<HR align=left color=#987cb9 SIZE=1>