HyperAIHyperAI

Command Palette

Search for a command to run...

3 months ago

SPGISpeech: 5,000 hours of transcribed financial audio for fully formatted end-to-end speech recognition

SPGISpeech: 5,000 hours of transcribed financial audio for fully formatted end-to-end speech recognition

Abstract

In the English speech-to-text (STT) machine learning task, acoustic models are conventionally trained on uncased Latin characters, and any necessary orthography (such as capitalization, punctuation, and denormalization of non-standard words) is imputed by separate post-processing models. This adds complexity and limits performance, as many formatting tasks benefit from semantic information present in the acoustic signal but absent in transcription. Here we propose a new STT task: end-to-end neural transcription with fully formatted text for target labels. We present baseline Conformer-based models trained on a corpus of 5,000 hours of professionally transcribed earnings calls, achieving a CER of 1.7. As a contribution to the STT research community, we release the corpus free for non-commercial use at https://datasets.kensho.com/datasets/scribe.

Benchmarks

BenchmarkMethodologyMetrics
speech-recognition-on-spgispeechConformer
Word Error Rate (WER): 5.7

Build AI with AI

From idea to launch — accelerate your AI development with free AI co-coding, out-of-the-box environment and best price of GPUs.

AI Co-coding
Ready-to-use GPUs
Best Pricing
Get Started

Hyper Newsletters

Subscribe to our latest updates
We will deliver the latest updates of the week to your inbox at nine o'clock every Monday morning
Powered by MailChimp
SPGISpeech: 5,000 hours of transcribed financial audio for fully formatted end-to-end speech recognition | Papers | HyperAI