arXiv-2022 GLIPv2:Unifying Localization and Vision-Language Understanding

· arXiv· · vision-language, contrastive-learning, object-detection

Paper: GLIPv2:Unifying Localization and Vision-Language Understanding

Code: https://github.com/microsoft/GLIP

GLIPv2: More Tasks and Datasets Built on GLIP

Abstract

The overall architecture remains GLIP; the extension is to unify more tasks and datasets within the same framework—for example, segmentation, detection, VQA, and image captioning.

Introduction

avatar

Images are still processed by a single encoder, while the text side supports a broader set of understanding tasks, followed by deep fusion between modalities.

GLIPv2: Unifying Localization and VL Understanding

avatar

Experiment

avatar
avatar

<HR align=left color=#987cb9 SIZE=1>