Vision — Paper Notes
Video understanding, multimodal vision-language models
2014
2015
2016
2018
DeepMind-2018 Representation Learning with Contrastive Predictive Coding
2018-07-10 · vision-language, contrastive-learning
CVPR-2018 Unsupervised Feature Learning via Non-Parametric Instance-level Discrimination
2018-05-01 · unsupervised-learning, contrastive-learning
CVPR-2018 Non-local Neural Networks
2018-04-13 · convolutional-neural-network
CVPR-2018 A Closer Look at Spatiotemporal Convolutions for Action Recognition
2018-04-12 · video, action-recognition
2019
2020
ECCV-2020 CMC: Contrastive Multiview Coding
2020-12-18 · contrastive-learning
NeurIPS-2020 Big Self-Supervised Models are Strong Semi-Supervised Learners
2020-12-06 · semi-supervised learning
NeurIPS-2020 Bootstrap your own latent:A new approach to self-supervised Learning
2020-12-06 · self-supervised learning
NeurIPS-2020 Unsupervised Learning of Visual Features by Contrasting Cluster Assignments
2020-12-01 · unsupervised-learning, contrastive-learning
arXiv-2020 BYOL works even without batch statistics
2020-10-20 · vision-language, transformer, contrastive-learning
ICML-2020 A Simple Framework for Contrastive Learning of Visual Representations
2020-07-01 · contrastive-learning
ECCV-2020 End-to-End Object Detection with Transformers
2020-05-28 · transformer, object-detection
ICLR-2020 DivideMix:Learning with Noisy Labels as Semi-supervised Learning
2020-04-26 · noisy label, semi-supervised, GMM, MixMatch
CVPR-2020 Momentum Contrast for Unsupervised Visual Representation Learning
2020-03-23 · vision-language, contrastive-learning
2021
NeurIPS-2021 VLMo:Unified Vision-Language Pre-Training with Mixture-of-Modality-Experts
2021-12-01 · vision-language, contrastive-learning, transformer
NeurIPS-2021 Intriguing Properties of Vision Transformers
2021-11-25 · vision-language, transformer
NeurIPS-2021 Align before Fuse:Vision and Language Representation Learning with Momentum Distillation
2021-11-10 · vision-language
CVPR-2021 Masked Autoencoders Are Scalable Vision Learners
2021-11-06 · vision-language, Transformer
arXiv-2021 ActionCLIP:A New Paradigm for Video Action Recognition
2021-09-17 · video, action-recognition, contrastive-learning
ICCV-2021 Swin Transformer:Hierarchical Vision Transformer using Shifted Windows
2021-08-17 · vision-language, transformer
arXiv-2021 How Much Can CLIP Benefit Vision-and-Language Tasks?
2021-07-06 · vision-language, contrastive-learning
ICML-2021 Perceiver:General Perception with Iterative Attention
2021-06-23 · vision-language
arXiv-2021 CLIP4Clip:An Empirical Study of CLIP for End to End Video Clip Retrieval
2021-05-08 · Video, Contrastive Learning
ICCV-2021 An Empirical Study of Training Self-Supervised Vision Transformers
2021-04-05 · vision-language
ICML-2021 Learning Transferable Visual Models From Natural Language Supervision
2021-02-26 · vision-language, contrastive-learning
ICML-2021 ViLT:Vision-and-Language Transformer Without Convolution or Region Supervision
2021-02-05 · vision-language, contrastive-learning, transformer
GCPR-2021 AudioCLIP:Extending CLIP to Image, Text and Audio
2021-01-24 · vision-language, transformer, contrastive-learning
ICLR-2021 An Image is Worth 16x16 Words:Transformers for Image Recognition at Scale
2021-01-13 · vision-language, transformer
ICML-2021 Is Space-Time Attention All You Need for Video Understanding
2021-01-09 · video, vision-language, transformer
2022
MM-2022 Can Language Understand Depth?
2022-10-10 · vision-language, contrastive-learning
NeurIPS-2022 CoCa:Contrastive Captioners are Image-Text Foundation Models
2022-08-27 · vision-language, contrastive-learning
SIGGRAPH-2022 CLIPasso:Semantically-Aware Object Sketching
2022-07-22 · contrastive-learning
CVPR-2022 GroupViT:Semantic Segmentation Emerges from Text Supervision
2022-07-18 · video, vision-language
ECCV-2022 CDS:Contrastive Deep Supervision
2022-07-12 · contrastive-learning
CVPR-2022 PointCLIP:Point Cloud Understanding by CLIP
2022-06-23 · vision-language, contrastive-learning
arXiv-2022 GLIPv2:Unifying Localization and Vision-Language Understanding
2022-06-05 · vision-language, contrastive-learning, object-detection
OpenAI-2022 Hierarchical Text-Conditional Image Generation with CLIP Latents
2022-04-13 · vision-language, contrastive-learning
ICLR-2022 Open-vocabulary Object Detection via Vision and Language Knowledge Distillation
2022-01-29 · vision-language, contrastive-learning, object-detection
ICLR-2022 Perceiver IO:A General Architecture for Structured Inputs & Outputs
2022-01-29 · vision-language
ICML-2022 BLIP:Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
2022-01-24 · vision-language, transformer, contrastive-learning
CVPR-2022 Grounded Language-Image Pre-trainin
2022-01-17 · LLM, NLP, vision-language
ICLR-2022 Language-driven Semantic Segmentation
2022-01-10 · LLM, NLP, vision-language
2023
CVPR-2020 Improved Baselines with Momentum Contrastive Learning
2023-03-09 · vision-language
ICLR-2023 AIM:Adapting Image Models for Efficient Video Action Recognition
2023-02-06 · video, action-recognition
CVPR-2023 Image as a Foreign Language:BEiT Pretraining for All Vision and Vision-Language Tasks
2023-01-15 · vision-language
WACV-2023 MixGen:A New Multi-Modal Data Augmentation
2023-01-09 · vision-language, data-augmentation