HyperAIHyperAI

Command Palette

Search for a command to run...

Reasoning Corpus Large Model Mind Chain Dataset

Date

6 hours ago

License

Apache 2.0

Reasoning Corpus is a large model thinking dataset released by SupraLabs, primarily used for supervised fine-tuning (SFT), model distillation, and instruction training. This dataset contains approximately 5 million samples, with each sequence length limited to 5,000 tokens. Each sample includes the original user prompt, inference trajectory, final assistant response, source information, estimated token length, and a pre-formatted ChatML representation. The data originates from over 60 public inference data warehouses, covering multiple fields such as science, mathematics, coding, logical reasoning, finance and economics, medicine, and multilingual STEM. It integrates thought chain data generated by mainstream large models such as DeepSeek-v4, DeepSeek-r1, Qwen3, and Gemma4.

Data fields:

  • repo_id: Identifier of the upstream data source repository
  • tok_len: Pre-computed estimate of the sample token length
  • user: Original user prompt or task command
  • thought_trace: The intermediate reasoning thought chain generated by the model
  • assistant: The final answer derived from the reasoning process.
  • ChatML: Preformatted standard ChatML formatted text

Build AI with AI

From idea to launch — accelerate your AI development with free AI co-coding, out-of-the-box environment and best price of GPUs.

AI Co-coding
Ready-to-use GPUs
Best Pricing

HyperAI Newsletters

Subscribe to our latest updates
We will deliver the latest updates of the week to your inbox at nine o'clock every Monday morning
Powered by MailChimp