Vision — Paper Notes

Video understanding, multimodal vision-language models

2021

NeurIPS

NeurIPS-2021 VLMo:Unified Vision-Language Pre-Training with Mixture-of-Modality-Experts

2021-12-01 · vision-language, contrastive-learning, transformer

NeurIPS

NeurIPS-2021 Intriguing Properties of Vision Transformers

2021-11-25 · vision-language, transformer

NeurIPS

NeurIPS-2021 Align before Fuse:Vision and Language Representation Learning with Momentum Distillation

2021-11-10 · vision-language

CVPR

CVPR-2021 Masked Autoencoders Are Scalable Vision Learners

2021-11-06 · vision-language, Transformer

arXiv

arXiv-2021 ActionCLIP:A New Paradigm for Video Action Recognition

2021-09-17 · video, action-recognition, contrastive-learning

ICCV

ICCV-2021 Swin Transformer:Hierarchical Vision Transformer using Shifted Windows

2021-08-17 · vision-language, transformer

arXiv

arXiv-2021 How Much Can CLIP Benefit Vision-and-Language Tasks?

2021-07-06 · vision-language, contrastive-learning

ICML

ICML-2021 Perceiver:General Perception with Iterative Attention

2021-06-23 · vision-language

arXiv

arXiv-2021 CLIP4Clip:An Empirical Study of CLIP for End to End Video Clip Retrieval

2021-05-08 · Video, Contrastive Learning

ICCV

ICCV-2021 An Empirical Study of Training Self-Supervised Vision Transformers

2021-04-05 · vision-language

ICML

ICML-2021 Learning Transferable Visual Models From Natural Language Supervision

2021-02-26 · vision-language, contrastive-learning

ICML

ICML-2021 ViLT:Vision-and-Language Transformer Without Convolution or Region Supervision

2021-02-05 · vision-language, contrastive-learning, transformer

GCPR

GCPR-2021 AudioCLIP:Extending CLIP to Image, Text and Audio

2021-01-24 · vision-language, transformer, contrastive-learning

ICLR

ICLR-2021 An Image is Worth 16x16 Words:Transformers for Image Recognition at Scale

2021-01-13 · vision-language, transformer

ICML

ICML-2021 Is Space-Time Attention All You Need for Video Understanding

2021-01-09 · video, vision-language, transformer

2022

MM

MM-2022 Can Language Understand Depth?

2022-10-10 · vision-language, contrastive-learning

NeurIPS

NeurIPS-2022 CoCa:Contrastive Captioners are Image-Text Foundation Models

2022-08-27 · vision-language, contrastive-learning

SIGGRAPH

SIGGRAPH-2022 CLIPasso:Semantically-Aware Object Sketching

2022-07-22 · contrastive-learning

CVPR

CVPR-2022 GroupViT:Semantic Segmentation Emerges from Text Supervision

2022-07-18 · video, vision-language

ECCV

ECCV-2022 CDS:Contrastive Deep Supervision

2022-07-12 · contrastive-learning

CVPR

CVPR-2022 PointCLIP:Point Cloud Understanding by CLIP

2022-06-23 · vision-language, contrastive-learning

arXiv

arXiv-2022 GLIPv2:Unifying Localization and Vision-Language Understanding

2022-06-05 · vision-language, contrastive-learning, object-detection

OpenAI

OpenAI-2022 Hierarchical Text-Conditional Image Generation with CLIP Latents

2022-04-13 · vision-language, contrastive-learning

ICLR

ICLR-2022 Open-vocabulary Object Detection via Vision and Language Knowledge Distillation

2022-01-29 · vision-language, contrastive-learning, object-detection

ICLR

ICLR-2022 Perceiver IO:A General Architecture for Structured Inputs & Outputs

2022-01-29 · vision-language

ICML

ICML-2022 BLIP:Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

2022-01-24 · vision-language, transformer, contrastive-learning

CVPR

CVPR-2022 Grounded Language-Image Pre-trainin

2022-01-17 · LLM, NLP, vision-language

ICLR

ICLR-2022 Language-driven Semantic Segmentation

2022-01-10 · LLM, NLP, vision-language

← Browse all notes