HyperAIHyperAI

Command Palette

Search for a command to run...

4 months ago

Learning Deep Structure-Preserving Image-Text Embeddings

Liwei Wang; Yin Li; Svetlana Lazebnik

Learning Deep Structure-Preserving Image-Text Embeddings

Abstract

This paper proposes a method for learning joint embeddings of images and text using a two-branch neural network with multiple layers of linear projections followed by nonlinearities. The network is trained using a large margin objective that combines cross-view ranking constraints with within-view neighborhood structure preservation constraints inspired by metric learning literature. Extensive experiments show that our approach gains significant improvements in accuracy for image-to-text and text-to-image retrieval. Our method achieves new state-of-the-art results on the Flickr30K and MSCOCO image-sentence datasets and shows promise on the new task of phrase localization on the Flickr30K Entities dataset.

Benchmarks

BenchmarkMethodologyMetrics
image-retrieval-on-flickr30k-1k-testSPE
R@1: 29.7
R@10: 72.1
R@5: 60.1
phrase-grounding-on-flickr30k-entities-testDSPE
R@1: 43.89
R@10: 68.66
R@5: 64.46

Build AI with AI

From idea to launch — accelerate your AI development with free AI co-coding, out-of-the-box environment and best price of GPUs.

AI Co-coding
Ready-to-use GPUs
Best Pricing
Get Started

Hyper Newsletters

Subscribe to our latest updates
We will deliver the latest updates of the week to your inbox at nine o'clock every Monday morning
Powered by MailChimp
Learning Deep Structure-Preserving Image-Text Embeddings | Papers | HyperAI