GCPR-2021 AudioCLIP:Extending CLIP to Image, Text and Audio

· GCPR· · vision-language, transformer, contrastive-learning

Paper: AudioCLIP:Extending CLIP to Image, Text and Audio

Code: https://github.com/AndreyGuzhov/AudioCLIP

AudioCLIP: tri-modal contrastive learning on video data

Abstract

Adds audio as a third modality alongside text and image, integrating the ESResNeXt audio model into the CLIP framework using audio datasets.

Model

avatar

Several video datasets provide aligned text, image, and audio within each clip; the architecture follows CLIP and incorporates all three modalities. Because the modalities are paired, contrastive learning is straightforward.

Result

avatar
avatar

<HR align=left color=#987cb9 SIZE=1>