arXiv-2021 How Much Can CLIP Benefit Vision-and-Language Tasks?

· arXiv· · vision-language, contrastive-learning

Paper: How Much Can CLIP Benefit Vision-and-Language Tasks?

Code: https://github.com/clip-vil/CLIP-ViL

CLIP-ViL: An empirical study of CLIP on vision-language downstream tasks

Abstract

An empirical study showing that initializing multimodal models with CLIP can further improve accuracy on downstream vision-language tasks.

Introduction

avatar

Main contribution: the first large-scale empirical study that uses CLIP-pretrained weights to initialize the visual encoder and evaluates them on a broad set of downstream tasks.

Experiments

avatar
avatar
avataravatar
avatar
avatar

<HR align=left color=#987cb9 SIZE=1>