LCSTS

Introduced by Hu et al. in LCSTS: A Large Scale Chinese Short Text Summarization Dataset

LCSTS is a large corpus of Chinese short text summarization dataset constructed from the Chinese microblogging website Sina Weibo, which is released to the public. This corpus consists of over 2 million real Chinese short texts with short summaries given by the author of each text. The authors also manually tagged the relevance of 10,666 short summaries with their corresponding short texts 10,666 short summaries with their corresponding short texts.

Source: LCSTS: A Large Scale Chinese Short Text Summarization Dataset

Homepage