HyperAIHyperAI

Command Palette

Search for a command to run...

3 months ago

Text-Free Prosody-Aware Generative Spoken Language Modeling

Text-Free Prosody-Aware Generative Spoken Language Modeling

Abstract

Speech pre-training has primarily demonstrated efficacy on classification tasks, while its capability of generating novel speech, similar to how GPT-2 can generate coherent paragraphs, has barely been explored. Generative Spoken Language Modeling (GSLM) \cite{Lakhotia2021} is the only prior work addressing the generative aspects of speech pre-training, which replaces text with discovered phone-like units for language modeling and shows the ability to generate meaningful novel sentences. Unfortunately, despite eliminating the need of text, the units used in GSLM discard most of the prosodic information. Hence, GSLM fails to leverage prosody for better comprehension, and does not generate expressive speech. In this work, we present a prosody-aware generative spoken language model (pGSLM). It is composed of a multi-stream transformer language model (MS-TLM) of speech, represented as discovered unit and prosodic feature streams, and an adapted HiFi-GAN model converting MS-TLM outputs to waveforms. We devise a series of metrics for prosody modeling and generation, and re-use metrics from GSLM for content modeling. Experimental results show that the pGSLM can utilize prosody to improve both prosody and content modeling, and also generate natural, meaningful, and coherent speech given a spoken prompt. Audio samples can be found at https://speechbot.github.io/pgslm. Codes and models are available at https://github.com/pytorch/fairseq/tree/main/examples/textless_nlp/pgslm.

Code Repositories

pytorch/fairseq
Official
pytorch

Benchmarks

BenchmarkMethodologyMetrics
language-modelling-on-salmonpGSLM
Background (Domain) Consistency: 57.0
Background (Random) Consistency: 66.0
Background Alignment: 53.5
Gender Consistency: 88.5
Room Consistency: 53.5
Sentiment Alignment: 55.5
Sentiment Consistency: 40.5
Speaker Consistency: 83.0

Build AI with AI

From idea to launch — accelerate your AI development with free AI co-coding, out-of-the-box environment and best price of GPUs.

AI Co-coding
Ready-to-use GPUs
Best Pricing
Get Started

Hyper Newsletters

Subscribe to our latest updates
We will deliver the latest updates of the week to your inbox at nine o'clock every Monday morning
Powered by MailChimp
Text-Free Prosody-Aware Generative Spoken Language Modeling | Papers | HyperAI