Vision — Paper Notes
Video understanding, multimodal vision-language models
2014
2015
2016
2018
2019
2020
ECCV-2020 CMC: Contrastive Multiview Coding
NeurIPS-2020 Big Self-Supervised Models are Strong Semi-Supervised Learners
NeurIPS-2020 Bootstrap your own latent:A new approach to self-supervised Learning
NeurIPS-2020 Unsupervised Learning of Visual Features by Contrasting Cluster Assignments
arXiv-2020 BYOL works even without batch statistics
ICML-2020 A Simple Framework for Contrastive Learning of Visual Representations
ECCV-2020 End-to-End Object Detection with Transformers
ICLR-2020 DivideMix:Learning with Noisy Labels as Semi-supervised Learning
CVPR-2020 Momentum Contrast for Unsupervised Visual Representation Learning
2021
NeurIPS-2021 VLMo:Unified Vision-Language Pre-Training with Mixture-of-Modality-Experts
NeurIPS-2021 Intriguing Properties of Vision Transformers
NeurIPS-2021 Align before Fuse:Vision and Language Representation Learning with Momentum Distillation
CVPR-2021 Masked Autoencoders Are Scalable Vision Learners
arXiv-2021 ActionCLIP:A New Paradigm for Video Action Recognition
ICCV-2021 Swin Transformer:Hierarchical Vision Transformer using Shifted Windows
arXiv-2021 How Much Can CLIP Benefit Vision-and-Language Tasks?
ICML-2021 Perceiver:General Perception with Iterative Attention
arXiv-2021 CLIP4Clip:An Empirical Study of CLIP for End to End Video Clip Retrieval
ICCV-2021 An Empirical Study of Training Self-Supervised Vision Transformers
ICML-2021 Learning Transferable Visual Models From Natural Language Supervision
ICML-2021 ViLT:Vision-and-Language Transformer Without Convolution or Region Supervision
GCPR-2021 AudioCLIP:Extending CLIP to Image, Text and Audio
ICLR-2021 An Image is Worth 16x16 Words:Transformers for Image Recognition at Scale
ICML-2021 Is Space-Time Attention All You Need for Video Understanding
2022
MM-2022 Can Language Understand Depth?
NeurIPS-2022 CoCa:Contrastive Captioners are Image-Text Foundation Models
SIGGRAPH-2022 CLIPasso:Semantically-Aware Object Sketching
CVPR-2022 GroupViT:Semantic Segmentation Emerges from Text Supervision
ECCV-2022 CDS:Contrastive Deep Supervision
CVPR-2022 PointCLIP:Point Cloud Understanding by CLIP
arXiv-2022 GLIPv2:Unifying Localization and Vision-Language Understanding
OpenAI-2022 Hierarchical Text-Conditional Image Generation with CLIP Latents
ICLR-2022 Open-vocabulary Object Detection via Vision and Language Knowledge Distillation
ICLR-2022 Perceiver IO:A General Architecture for Structured Inputs & Outputs
ICML-2022 BLIP:Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
CVPR-2022 Grounded Language-Image Pre-trainin