Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models

ICCV 2015 Bryan A. PlummerLiwei WangChris M. CervantesJuan C. CaicedoJulia HockenmaierSvetlana Lazebnik

The Flickr30k dataset has become a standard benchmark for sentence-based image description. This paper presents Flickr30k Entities, which augments the 158k captions from Flickr30k with 244k coreference chains, linking mentions of the same entities across different captions for the same image, and associating them with 276k manually annotated bounding boxes... (read more)

PDF Abstract

Code


No code implementations yet. Submit your code now

Tasks


Results from the Paper


TASK DATASET MODEL METRIC NAME METRIC VALUE GLOBAL RANK RESULT LEADERBOARD
Image Retrieval Flickr30K 1K test HGLMM FV [email protected] 24.7 # 9
[email protected] 66.8 # 8
[email protected] 53.4 # 7