CVPR-2020 Momentum Contrast for Unsupervised Visual Representation Learning
Paper: Momentum Contrast for Unsupervised Visual Representation Learning
MoCo v1: Contrastive learning as dictionary lookup, with a queue and momentum encoder for a large dictionary
- Strong writing craft: The first paragraph of the intro contrasts CV and NLP and explains why unsupervised learning has underperformed in CV; the second introduces contrastive learning without heavy detail, summarizes prior methods as dictionary lookup, and proposes MoCo under a unified CV/NLP framing—broadening the paper’s overall scope.
-
Two design constraints are validated: the dictionary should be large, and representations should be consistent.
A large dictionary forces the model to learn more discriminative features;
Consistency keeps learning in a shared, semantically aligned encoder space and helps avoid shortcut solutions.
- The encoder is updated by gradient backpropagation; the momentum encoder is updated with a large momentum coefficient. A queue structure links batch size to dictionary size so the dictionary can be made sufficiently large.
<HR align=left color=#987cb9 SIZE=1>