7 dataset results for Sentence Embeddings AND English

The Discovery datasets consists of adjacent sentence pairs (s1,s2) with a discourse marker (y) that occurred at the beginning of s2. They were extracted from the depcc web corpus.

9 PAPERS • 1 BENCHMARK

SemEval-2014 Task-10

SemEval 2014 is a collection of datasets used for the Semantic Evaluation (SemEval) workshop, an annual event that focuses on the evaluation and comparison of systems that can analyze diverse semantic phenomena in text. The datasets from SemEval 2014 are used for various tasks, including but not limited to:

6 PAPERS • NO BENCHMARKS YET

AuxAD

AuxAD is a a distantly supervised dataset for acronym disambiguation.

1 PAPER • NO BENCHMARKS YET

AuxAI

AuxAI is a distantly supervised dataset for acronym identification.

1 PAPER • NO BENCHMARKS YET

GeoCoV19

GeoCoV19 is a large-scale Twitter dataset containing more than 524 million multilingual tweets. The dataset contains around 378K geotagged tweets and 5.4 million tweets with Place information. The annotations include toponyms from the user location field and tweet content and resolve them to geolocations such as country, state, or city level. In this case, 297 million tweets are annotated with geolocation using the user location field and 452 million tweets using tweet content.

3 PAPERS • NO BENCHMARKS YET

PIT

PIT (Paraphrase and Semantic Similarity in Twitter)

Paraphrase and Semantic Similarity in Twitter (PIT) presents a constructed Twitter Paraphrase Corpus that contains 18,762 sentence pairs.

22 PAPERS • 1 BENCHMARK

WikiMatrix

WikiMatrix is a dataset of parallel sentences in the textual content of Wikipedia for all possible language pairs. The mined data consists of:

87 PAPERS • NO BENCHMARKS YET

Datasets

7 dataset results for Sentence Embeddings AND English