Grammarly
Source Sentence Detection in Abstractive Summarization
Pages
13
Time to read
38 mins
Publication
Language
English
Pages
13
Time to read
38 mins
Publication
Language
English
This technical report investigates the process of identifying source sentences in neural abstractive summarization models. It defines source sentences as those containing essential information for generating summaries and analyzes how these sentences contribute to the final output. The study employs datasets from CNN/DailyMail and XSum, annotating source sentences for both reference and system-generated summaries produced by the PEGASUS model. The report formulates a task for automatic source sentence detection and evaluates various methods, including perplexity-based and similarity-based approaches. Experimental results indicate that the perplexity-based method excels in highly abstractive contexts, while similarity-based methods are effective in more extractive scenarios. The findings contribute to the understanding of how summarization models utilize input information, enhancing the explainability and interpretability of generated summaries. The report also discusses the implications of these findings for future research in text summarization.