HyperAIHyperAI

Command Palette

Search for a command to run...

3 months ago

A Step Towards Worldwide Biodiversity Assessment: The BIOSCAN-1M Insect Dataset

A Step Towards Worldwide Biodiversity Assessment: The BIOSCAN-1M Insect Dataset

Abstract

In an effort to catalog insect biodiversity, we propose a new large dataset of hand-labelled insect images, the BIOSCAN-Insect Dataset. Each record is taxonomically classified by an expert, and also has associated genetic information including raw nucleotide barcode sequences and assigned barcode index numbers, which are genetically-based proxies for species classification. This paper presents a curated million-image dataset, primarily to train computer-vision models capable of providing image-based taxonomic assessment, however, the dataset also presents compelling characteristics, the study of which would be of interest to the broader machine learning community. Driven by the biological nature inherent to the dataset, a characteristic long-tailed class-imbalance distribution is exhibited. Furthermore, taxonomic labelling is a hierarchical classification scheme, presenting a highly fine-grained classification problem at lower levels. Beyond spurring interest in biodiversity research within the machine learning community, progress on creating an image-based taxonomic classifier will also further the ultimate goal of all BIOSCAN research: to lay the foundation for a comprehensive survey of global biodiversity. This paper introduces the dataset and explores the classification task through the implementation and analysis of a baseline classifier.

Code Repositories

zahrag/BIOSCAN-1M
Official
pytorch
Mentioned in GitHub

Benchmarks

BenchmarkMethodologyMetrics
classification-on-bioscan-1m-insect-datasetBIOSCAN_1M_order_classifier
Macro F1: 92.65
classification-on-bioscan-1m-insect-datasetBIOSCAN_1M_family_classifier
Macro F1: 91.45

Build AI with AI

From idea to launch — accelerate your AI development with free AI co-coding, out-of-the-box environment and best price of GPUs.

AI Co-coding
Ready-to-use GPUs
Best Pricing
Get Started

Hyper Newsletters

Subscribe to our latest updates
We will deliver the latest updates of the week to your inbox at nine o'clock every Monday morning
Powered by MailChimp
A Step Towards Worldwide Biodiversity Assessment: The BIOSCAN-1M Insect Dataset | Papers | HyperAI