ECCV-2020 CMC: Contrastive Multiview Coding
Paper: Contrastive Multiview Coding
CMC: Maximize mutual information across views; the pretext task is multiview (multimodal)
- The abstract is compelling: whether you see a dog or hear it bark, you can recognize it as a dog—motivating learning invariance across views.
-
Positive pairs: although inputs come from different sensors (modalities), they still correspond to the same underlying image; their representations should be close in feature space.
Negative pairs: views drawn from other images.
- Different inputs across views may require multiple encoders—potentially one encoder per view—which increases computational cost; transformers may process multiple modalities jointly without modality-specific architectural changes.
Comment: Code: http://github.com/HobbitLong/CMC/
<HR align=left color=#987cb9 SIZE=1>