Priberam Compressive Summarization Corpus: A New Multi-Document Summarization Corpus for European Portuguese

LREC 2014 · Miguel B. Almeida, Mariana S. C. Almeida, Andr{\'e} F. T. Martins, Helena Figueira, Pedro Mendes, Cl{\'a}udia Pinto ·

In this paper, we introduce the Priberam Compressive Summarization Corpus, a new multi-document summarization corpus for European Portuguese. The corpus follows the format of the summarization corpora for English in recent DUC and TAC conferences. It contains 80 manually chosen topics referring to events occurred between 2010 and 2013. Each topic contains 10 news stories from major Portuguese newspapers, radio and TV stations, along with two human generated summaries up to 100 words. Apart from the language, one important difference from the DUC/TAC setup is that the human summaries in our corpus are {\textbackslash}emph{compressive}: the annotators performed only sentence and word deletion operations, as opposed to generating summaries from scratch. We use this corpus to train and evaluate learning-based extractive and compressive summarization systems, providing an empirical comparison between these two approaches. The corpus is made freely available in order to facilitate research on automatic summarization.

PDF Abstract