ICML-2021 Learning Transferable Visual Models From Natural Language Supervision
Paper: Learning Transferable Visual Models From Natural Language Supervision
CLIP: Replace unimodal positives with multimodal pairs; use text prompts for transfer learning

Abstract and Related Work
- A major reason CLIP is strong is that OpenAI collected 400 million image–text pairs, breaking away from fixed categorical labels.
-
It fully removes the constraint of categorical labels: neither training nor inference requires a predefined label list. Any image can be queried by feeding the model different text sentences to test whether objects of interest are present, which has spawned many interesting follow-ups.
2.1 StyleCLIP: CLIP + StyleGAN—text edits guide image generation, code
2.2 CLIPDraw: a CLIP pretrained model guides image generation; a few steps of gradient descent suffice to produce sketch-like images
2.3 Object detection and segmentation: detect novel categories—for example, not only recognizing a toy but also what color it is
- Pure zero-shot performance was initially too low (11.5%). Even though category-based pretraining was known to generalize poorly, few pursued the direction, so people looked for a middle ground between images and text.
- Examples include learning from Instagram data with hashtags, or from JFT-300 labels (10k+ categories)—all ways to learn visual features from natural language—but most required training at hundred-million to billion scale. That suggests the issue was not the CLIP-style approach itself but insufficient data. They tried eight models spanning two orders of magnitude in scale and found performance scales positively with scale.
- Even after freezing the pretrained backbone (linear probe) and training only a final classification head, results beat ImageNet-trained baselines while remaining compute-efficient.
Method

-
Why use natural language supervision:
1.1 No manual labeling of categories or heavy data cleaning—only image–text pairs. Because the supervision is free-form text rather than an n-way label, the model’s inputs and outputs have much higher degrees of freedom.
1.2 Training binds images and text together, so the learned representation is genuinely multimodal, not purely visual.
-
They built a large dataset (~400M pairs). With so much data, heavy augmentation is unnecessary—they use only random cropping and still struggle to overfit. Scale also makes hyperparameter tuning hard; temperature becomes a critical hyperparameter.
-
Human captions for the same image can differ wildly, so predictive pretraining (e.g., caption generation) is slow. Reframing the task as image–text matching is much easier: there is no need to predict text token-by-token, and training efficiency improves by about 4×.
-
Whether the projection head is nonlinear or linear barely matters (nonlinear heads helped SimCLR by ~10 points)—possibly because nonlinearity mainly helps unimodal image-only contrastive learning.
-
The visual encoder can be ResNet or Transformer.
Experiment
-
Prior unsupervised or self-supervised methods mainly studied representation quality—learning features that generalize (e.g., MoCo, SimCLR, DINO)—but downstream use still typically requires labeled fine-tuning. Can we avoid tuning altogether?
-
Zero-shot inference
Images and text pass through encoders; cosine similarity followed by softmax yields predictions. On ImageNet classification, this means 1000 text prompts—effectively asking the image, one by one, “Is this a dog?”, “Is this a cat?”, and so on.

-
Prompt templates
The authors use 80 prompt templates as textual anchors. Wording matters, which later motivated prompt engineering and prompt ensemble in CLIP.
3.1 Why prompt engineering or prompt ensemble?
① Polysemy: one word, many senses—using a single word as the text input is ambiguous. ② There is no dedicated classification head; training pairs one sentence with one image, so at inference a word (class name) must be expanded into a sentence to avoid unnecessary accuracy drops.
3.2 Prompt engineering
With more prior knowledge about a dataset, templates can be chosen more carefully to shrink the search space—for example, knowing a benchmark is all animals.
3.3 Prompt ensemble
Applying multiple templates and aggregating scores usually improves results.
-
They evaluate transfer on 27 datasets, showing broad applicability. CLIP does well on standard classification benchmarks, but on very hard tasks—counting (complex), tumor classification (domain knowledge)—zero-shot is weak; few-shot comparison is fairer. In practice, full fine-tuning still beats all alternatives by a large margin.


- The authors even compare humans vs. CLIP: volunteers classify 37 fine-grained dog/cat breeds without reference exemplars, alongside one-shot and two-shot human experiments.

- They also compare which categories are hard for humans vs. CLIP—classes humans find difficult tend to be hard for CLIP as well.
Limitation
-
Strong numbers are often vs. ResNet-50; compared with various SOTAs, CLIP still lags by a lot—perhaps needing ~1000× scale. New methods must improve compute and data efficiency.
-
Zero-shot is poor on some datasets (e.g., fine-grained classification) and on abstract or difficult tasks; in many domains CLIP can be near random guessing.
-
At inference, large train–test distribution shift hurts generalization—for example ~88% on MNIST, where a linear baseline can do better.
-
Zero-shot still chooses among a fixed set of categories. A more flexible setup is caption generation (generative models) so the model can produce novel outputs—raising whether a loss could unify contrastive and generative learning.
-
Data efficiency is low: training effectively sees on the order of 12.8 billion image exposures. Reducing data needs may require stronger augmentation, pseudo-labels, or self-supervision.
-
Using ImageNet as the recurring benchmark for downstream evaluation introduces bias—it is not truly zero-shot in the wild.
-
Web-scraped, minimally cleaned data may encode societal biases and enable misuse.
-
Many visual concepts are hard to describe in language; a few downstream examples (few-shot) can help—yet sometimes adding a few labeled images hurts accuracy compared with zero-shot.
<HR align=left color=#987cb9 SIZE=1>