HyperAIHyperAI

Command Palette

Search for a command to run...

Eureka Bench Logs Evaluation Log Dataset

Date

Organization

Microsoft Corporation

Paper URL

2409.10566

License

Apache 2.0

Eureka Bench Logs is a dataset released by Microsoft in 2024 that contains evaluation logs for large foundation models, with the related paper titled "Eureka: Evaluating and Understanding Large Foundation Models", aiming to provide in-depth understanding and assessment of the reasoning capabilities of large foundation models in complex tasks.

The dataset includes run logs from the Eureka ML Insights framework, recording detailed performance of different model families and specific models across various benchmarks. The data is organized by benchmark, model family, and model name, covering research results on inference-time scaling for models such as Phi-4-reasoning. This information helps researchers analyze performance differences and optimization directions for models on tasks such as image-to-text.

Dataset Composition

  • Benchmarks: Logs are categorized according to specific evaluation benchmarks, covering multimodal tasks such as image-to-text.
  • Model Families: Data is grouped by model family, facilitating comparison of progress across different models within the same architecture.
  • Model Names: Each log entry corresponds to a specific model instance, recording detailed inference and evaluation data.

These log data provide rich empirical material for evaluating and understanding large foundation models.

Build AI with AI

From idea to launch — accelerate your AI development with free AI co-coding, out-of-the-box environment and best price of GPUs.

AI Co-coding
Ready-to-use GPUs
Best Pricing

HyperAI Newsletters

Subscribe to our latest updates
We will deliver the latest updates of the week to your inbox at nine o'clock every Monday morning
Powered by MailChimp