arXiv-2022 GLIPv2:Unifying Localization and Vision-Language Understanding
Paper: GLIPv2:Unifying Localization and Vision-Language Understanding
GLIPv2: More Tasks and Datasets Built on GLIP
Abstract
The overall architecture remains GLIP; the extension is to unify more tasks and datasets within the same framework—for example, segmentation, detection, VQA, and image captioning.
Introduction

Images are still processed by a single encoder, while the text side supports a broader set of understanding tasks, followed by deep fusion between modalities.
GLIPv2: Unifying Localization and VL Understanding

Experiment


<HR align=left color=#987cb9 SIZE=1>