HyperAIHyperAI

Command Palette

Search for a command to run...

Image Understanding Benchmark Dataset

Date

Organization

Microsoft Corporation

License

cdla-permissive-2.0

Image Understanding Benchmark is a synthetic dataset for image understanding evaluation released by Microsoft in 2024, designed to assess multimodal models' capabilities in image understanding, spatial reasoning, visual prompting, object recognition, and detection.

The dataset contains four subtasks: object detection, object recognition, spatial reasoning, and visual prompting, generating approximately 10,240 images in total. The data is procedurally generated, combining COCO objects and Places365 backgrounds, with random rotations, positions, and scaling variations, and can be used to evaluate the visual understanding capabilities of multimodal models.

Dataset Composition

The dataset includes 8 configurations, divided into single-object and pairs conditions, covering the following four subtasks:

  • Object Detection
  • Object Recognition
  • Spatial Reasoning
  • Visual Prompting

Each subtask generates 1,280 images under both conditions, totaling approximately 10,240 images. Each configuration mainly contains the following fields:

  • id: Unique identifier for each data sample.
  • image: Input image.
  • prompt: Prompt text used for model evaluation.
  • ground_truth: Standard answer for certain tasks.
数据集示例
数据集示例

Build AI with AI

From idea to launch — accelerate your AI development with free AI co-coding, out-of-the-box environment and best price of GPUs.

AI Co-coding
Ready-to-use GPUs
Best Pricing

HyperAI Newsletters

Subscribe to our latest updates
We will deliver the latest updates of the week to your inbox at nine o'clock every Monday morning
Powered by MailChimp