HyperAIHyperAI

Command Palette

Search for a command to run...

5 months ago

DebateSum: A large-scale argument mining and summarization dataset

Allen Roush; Arvind Balaji

DebateSum: A large-scale argument mining and summarization dataset

Abstract

Prior work in Argument Mining frequently alludes to its potential applications in automatic debating systems. Despite this focus, almost no datasets or models exist which apply natural language processing techniques to problems found within competitive formal debate. To remedy this, we present the DebateSum dataset. DebateSum consists of 187,386 unique pieces of evidence with corresponding argument and extractive summaries. DebateSum was made using data compiled by competitors within the National Speech and Debate Association over a 7-year period. We train several transformer summarization models to benchmark summarization performance on DebateSum. We also introduce a set of fasttext word-vectors trained on DebateSum called debate2vec. Finally, we present a search engine for this dataset which is utilized extensively by members of the National Speech and Debate Association today. The DebateSum search engine is available to the public here: http://www.debate.cards

Code Repositories

Benchmarks

BenchmarkMethodologyMetrics
extractive-document-summarization-onBERT-Large
ROUGE-L: 49.98
extractive-document-summarization-onLongformer-Base
ROUGE-L: 57.21
extractive-document-summarization-onGPT2-Medium
ROUGE-L: 53.23

Build AI with AI

From idea to launch — accelerate your AI development with free AI co-coding, out-of-the-box environment and best price of GPUs.

AI Co-coding
Ready-to-use GPUs
Best Pricing
Get Started

Hyper Newsletters

Subscribe to our latest updates
We will deliver the latest updates of the week to your inbox at nine o'clock every Monday morning
Powered by MailChimp
DebateSum: A large-scale argument mining and summarization dataset | Papers | HyperAI