arXiv-2021 How Much Can CLIP Benefit Vision-and-Language Tasks?
Paper: How Much Can CLIP Benefit Vision-and-Language Tasks?
Code: https://github.com/clip-vil/CLIP-ViL
CLIP-ViL: An empirical study of CLIP on vision-language downstream tasks
Abstract
An empirical study showing that initializing multimodal models with CLIP can further improve accuracy on downstream vision-language tasks.
Introduction

Main contribution: the first large-scale empirical study that uses CLIP-pretrained weights to initialize the visual encoder and evaluates them on a broad set of downstream tasks.
Experiments






<HR align=left color=#987cb9 SIZE=1>