Vision — Paper Notes

Video understanding, multimodal vision-language models

2021

NeurIPS-2021 VLMo:Unified Vision-Language Pre-Training with Mixture-of-Modality-Experts

NeurIPS-2021 Intriguing Properties of Vision Transformers

NeurIPS-2021 Align before Fuse:Vision and Language Representation Learning with Momentum Distillation

CVPR-2021 Masked Autoencoders Are Scalable Vision Learners

arXiv-2021 ActionCLIP:A New Paradigm for Video Action Recognition

ICCV-2021 Swin Transformer:Hierarchical Vision Transformer using Shifted Windows

arXiv-2021 How Much Can CLIP Benefit Vision-and-Language Tasks?

ICML-2021 Perceiver:General Perception with Iterative Attention

arXiv-2021 CLIP4Clip:An Empirical Study of CLIP for End to End Video Clip Retrieval

ICCV-2021 An Empirical Study of Training Self-Supervised Vision Transformers

ICML-2021 Learning Transferable Visual Models From Natural Language Supervision

ICML-2021 ViLT:Vision-and-Language Transformer Without Convolution or Region Supervision

GCPR-2021 AudioCLIP:Extending CLIP to Image, Text and Audio

ICLR-2021 An Image is Worth 16x16 Words:Transformers for Image Recognition at Scale

ICML-2021 Is Space-Time Attention All You Need for Video Understanding

2022

MM-2022 Can Language Understand Depth?

NeurIPS-2022 CoCa:Contrastive Captioners are Image-Text Foundation Models

SIGGRAPH-2022 CLIPasso:Semantically-Aware Object Sketching

CVPR-2022 GroupViT:Semantic Segmentation Emerges from Text Supervision

ECCV-2022 CDS:Contrastive Deep Supervision

CVPR-2022 PointCLIP:Point Cloud Understanding by CLIP

arXiv-2022 GLIPv2:Unifying Localization and Vision-Language Understanding

OpenAI-2022 Hierarchical Text-Conditional Image Generation with CLIP Latents

ICLR-2022 Open-vocabulary Object Detection via Vision and Language Knowledge Distillation

ICLR-2022 Perceiver IO:A General Architecture for Structured Inputs & Outputs

ICML-2022 BLIP:Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

CVPR-2022 Grounded Language-Image Pre-trainin

ICLR-2022 Language-driven Semantic Segmentation

← Browse all notes