Evaluating Multimodal Representations on Visual Semantic Textual Similarity

4 Apr 2020Oier Lopez de LacalleAnder SalaberriaAitor SoroaGorka AzkuneEneko Agirre

The combination of visual and textual representations has produced excellent results in tasks such as image captioning and visual question answering, but the inference capabilities of multimodal representations are largely untested. In the case of textual representations, inference tasks such as Textual Entailment and Semantic Textual Similarity have been often used to benchmark the quality of textual representations... (read more)

PDF Abstract

Results from the Paper


  Submit results from this paper to get state-of-the-art GitHub badges and help the community compare results to other papers.

Methods used in the Paper