HyperAIHyperAI

Command Palette

Search for a command to run...

Back-to-School · Up to 20% top-up bonus + RTX 5090 GPU hours Learn More

CantTalkAboutThis Topic Control Dataset

Date

Organization

NVIDIA

Paper URL

2404.03820

License

CC BY NC 4.0

The "Can't Talk About This" topic control dataset, released by NVIDIA in 2024, focuses on dialogue topic management and safe alignment of large language models; related research findings can be found in the paper titled «CantTalkAboutThis: Aligning Language Models to Stay on Topic in Dialogues”, aimed at improving large language models’ ability to maintain topic focus in task-oriented dialogues and enhancing their robustness against distractions.

The dataset contains 1,080 synthetic dialogue samples covering nine domains: health, banking, travel, education, finance, insurance, law, real estate, and computer troubleshooting. Each dialogue includes distracting turns designed to test the model’s resistance to topic drift. Fine-tuning on this dataset significantly enhances performance in instruction-following and safety-related tasks, enabling effective identification of sensitive topics and handling of restricted content.

Dataset Composition

The dataset primarily consists of the following fields:

  • domain: The field or category the dialogue belongs to
  • scenario: The specific context or task being discussed
  • system_instruction: Dialogue strategy instructions provided to the model, typically containing complex sets of rules specifying allowed and prohibited discussion topics
  • conversation: Full dialogue content, including main-topic exchanges and distracting turns
  • distractors: List of distracting turns, comprising bot responses along with corresponding user replies intended as counter-responses to those bot turns
  • conversation_with_distractors: Complete dialogue incorporating all distracting elements

The dataset is split into training and testing subsets. The training set was synthetically generated using the GPT-4 Turbo model, while the test set comprises human-labeled evaluation data featuring more complex and realistic distractor examples for assessing model performance.

Citation```bibtex

@inproceedings{sreedhar2024canttalkaboutthis, title={CantTalkAboutThis: Aligning Language Models to Stay on Topic in Dialogues}, author={Sreedhar, Makesh and Rebedea, Traian and Ghosh, Shaona and Zeng, Jiaqi and Parisien, Christopher}, booktitle={Findings of the Association for Computational Linguistics: EMNLP 2024}, pages={12232--12252}, year={2024}, organization={Association for Computational Linguistics} }

Build AI with AI

From idea to launch — accelerate your AI development with free AI co-coding, out-of-the-box environment and best price of GPUs.

AI Co-coding
Ready-to-use GPUs
Best Pricing

HyperAI Newsletters

Subscribe to our latest updates
We will deliver the latest updates of the week to your inbox at nine o'clock every Monday morning
Powered by MailChimp