HyperAIHyperAI

Command Palette

Search for a command to run...

3 months ago

HowSumm: A Multi-Document Summarization Dataset Derived from WikiHow Articles

Odellia Boni Guy Feigenblat Guy Lev Michal Shmueli-Scheuer Benjamin Sznajder David Konopnicki

HowSumm: A Multi-Document Summarization Dataset Derived from WikiHow Articles

Abstract

We present HowSumm, a novel large-scale dataset for the task of query-focused multi-document summarization (qMDS), which targets the use-case of generating actionable instructions from a set of sources. This use-case is different from the use-cases covered in existing multi-document summarization (MDS) datasets and is applicable to educational and industrial scenarios. We employed automatic methods, and leveraged statistics from existing human-crafted qMDS datasets, to create HowSumm from wikiHow website articles and the sources they cite. We describe the creation of the dataset and discuss the unique features that distinguish it from other summarization corpora. Automatic and human evaluations of both extractive and abstractive summarization models on the dataset reveal that there is room for improvement.

Code Repositories

odelliab/HowSumm
Official
Mentioned in GitHub

Benchmarks

BenchmarkMethodologyMetrics
document-summarization-on-howsumm-methodCES (query: method + article titles)
ROUGE-1: 48.3
document-summarization-on-howsumm-methodLexRank (query: method + article titles)
ROUGE-1: 47.1
document-summarization-on-howsumm-methodCES (query: method + article + steps titles)
ROUGE-1: 52.2
document-summarization-on-howsumm-methodLexRank (query: method title)
ROUGE-1: 47.7
document-summarization-on-howsumm-methodGreedyRel (query: method title)
ROUGE-1: 43.4
document-summarization-on-howsumm-methodCES (query: method title)
ROUGE-1: 48.4
document-summarization-on-howsumm-methodGreedyRel (query: method + article titles)
ROUGE-1: 42.3
document-summarization-on-howsumm-methodLexRank (query: method + article + steps titles)
ROUGE-1: 53.5
document-summarization-on-howsumm-methodGreedyRel (query: method + article + steps titles)
ROUGE-1: 48.6
document-summarization-on-howsumm-stepBM25-HierSumm (query: step + method + article titles)
ROUGE-1: 21.9
document-summarization-on-howsumm-stepLexRank (query: step + method titles)
ROUGE-1: 38.2
document-summarization-on-howsumm-stepGreedyRel (query: step title)
ROUGE-1: 30.1
document-summarization-on-howsumm-stepLexRank (query: step title)
ROUGE-1: 39.6
document-summarization-on-howsumm-stepGreedyRel (query: step + method titles)
ROUGE-1: 30.3
document-summarization-on-howsumm-stepLexRank (query: step + method + article titles)
ROUGE-1: 36.3
document-summarization-on-howsumm-stepCES (query: step title)
ROUGE-1: 39.3
document-summarization-on-howsumm-stepCES (query: step + method + article titles)
ROUGE-1: 37.0
document-summarization-on-howsumm-stepCES (query: step + method titles)
ROUGE-1: 38.3
document-summarization-on-howsumm-stepBM25-HierSumm (query: step + method titles)
ROUGE-1: 23.0
document-summarization-on-howsumm-stepBM25-HierSumm (query: step title)
ROUGE-1: 22.3

Build AI with AI

From idea to launch — accelerate your AI development with free AI co-coding, out-of-the-box environment and best price of GPUs.

AI Co-coding
Ready-to-use GPUs
Best Pricing
Get Started

Hyper Newsletters

Subscribe to our latest updates
We will deliver the latest updates of the week to your inbox at nine o'clock every Monday morning
Powered by MailChimp
HowSumm: A Multi-Document Summarization Dataset Derived from WikiHow Articles | Papers | HyperAI