Multi-document Summarization by Information Distance

Date

2009

Journal Title

Journal ISSN

Volume Title

Publisher

Abstract

Fast changing knowledge on the Internet can be acquired more efficiently with the help of automatic document summarization and updating techniques. This paper described a novel approach for multi-document update summarization. The best summary is defined to be the one which has the minimum information distance to the entire document set. The best update summary has the minimum conditional information distance to a document cluster given that a prior document cluster has already been read. Experiments on the DUC 2007 dataset and the TAC 2008 dataset have proved that our method closely correlates with the human summaries and outperforms other programs such as LexRank in many categories under the ROUGE evaluation criterion.

Description

Keywords

DOCUMENT SUMMARIZATION, CLUSTER ANALYSIS, INFORMATION DISTANCE

Citation

Long, C., Huang, M. L., Zhu, X. Y., Li, M. (2009). Multi-document Summarization by Information Distance. Proceedings of the Ninth IEEE International Conference on Data Mining, 2009. ICDM '09, Miami, USA. (p. 866-871). doi: 10.1109/ICDM.2009.107

DOI