6 months ago

Xinyu Geng Peng Xia Zhen Zhang Xinyu Wang Qiuchen Wang Ruixue Ding Chenxi Wang Jialong Wu Yida Zhao Kuan Li

Abstract

Web agents such as Deep Research have demonstrated superhuman cognitive abilities, capable of solving highly challenging information-seeking problems. However, most research remains primarily text-centric, overlooking visual information in the real world. This makes multimodal Deep Research highly challenging, as such agents require much stronger reasoning abilities in perception, logic, knowledge, and the use of more sophisticated tools compared to text-based agents. To address this limitation, we introduce WebWatcher, a multi-modal Agent for Deep Research equipped with enhanced visual-language reasoning capabilities. It leverages high-quality synthetic multimodal trajectories for efficient cold start training, utilizes various tools for deep reasoning, and further enhances generalization through reinforcement learning. To better evaluate the capabilities of multimodal agents, we propose BrowseComp-VL, a benchmark with BrowseComp-style that requires complex information retrieval involving both visual and textual information. Experimental results show that WebWatcher significantly outperforms proprietary baseline, RAG workflow and open-source agents in four challenging VQA benchmarks, which paves the way for solving complex multimodal information-seeking tasks.

Source PDF

Build AI with AI

From idea to launch — accelerate your AI development with free AI co-coding, out-of-the-box environment and best price of GPUs.

AI Co-coding

Ready-to-use GPUs

Best Pricing

Get Started View Pricing

HyperAI Newsletters

Subscribe to our latest updates

We will deliver the latest updates of the week to your inbox at nine o'clock every Monday morning

HyperAI

6 months ago

Visual Question Answering

Xinyu Geng Peng Xia Zhen Zhang Xinyu Wang Qiuchen Wang Ruixue Ding Chenxi Wang Jialong Wu Yida Zhao Kuan Li

Abstract

Source PDF

Build AI with AI

From idea to launch — accelerate your AI development with free AI co-coding, out-of-the-box environment and best price of GPUs.

AI Co-coding

Ready-to-use GPUs

Best Pricing

Get Started View Pricing

HyperAI Newsletters

Subscribe to our latest updates

We will deliver the latest updates of the week to your inbox at nine o'clock every Monday morning

HyperAI

6 months ago

Visual Question Answering

Xinyu Geng Peng Xia Zhen Zhang Xinyu Wang Qiuchen Wang Ruixue Ding Chenxi Wang Jialong Wu Yida Zhao Kuan Li

Abstract

Source PDF

Build AI with AI

From idea to launch — accelerate your AI development with free AI co-coding, out-of-the-box environment and best price of GPUs.

AI Co-coding

Ready-to-use GPUs

Best Pricing

Get Started View Pricing

HyperAI Newsletters

Subscribe to our latest updates

We will deliver the latest updates of the week to your inbox at nine o'clock every Monday morning

Command Palette

WebWatcher: Breaking New Frontiers of Vision-Language Deep Research Agent

Xinyu Geng Peng Xia Zhen Zhang Xinyu Wang Qiuchen Wang Ruixue Ding Chenxi Wang Jialong Wu Yida Zhao Kuan Li4 more

Abstract

Build AI with AI

HyperAI Newsletters

Command Palette

WebWatcher: Breaking New Frontiers of Vision-Language Deep Research Agent

Xinyu Geng Peng Xia Zhen Zhang Xinyu Wang Qiuchen Wang Ruixue Ding Chenxi Wang Jialong Wu Yida Zhao Kuan Li4 more

Abstract

Build AI with AI

HyperAI Newsletters

Command Palette

WebWatcher: Breaking New Frontiers of Vision-Language Deep Research Agent

Xinyu Geng Peng Xia Zhen Zhang Xinyu Wang Qiuchen Wang Ruixue Ding Chenxi Wang Jialong Wu Yida Zhao Kuan Li4 more

Abstract

Build AI with AI

HyperAI Newsletters

Xinyu Geng Peng Xia Zhen Zhang Xinyu Wang Qiuchen Wang Ruixue Ding Chenxi Wang Jialong Wu Yida Zhao Kuan Li

Xinyu Geng Peng Xia Zhen Zhang Xinyu Wang Qiuchen Wang Ruixue Ding Chenxi Wang Jialong Wu Yida Zhao Kuan Li

Xinyu Geng Peng Xia Zhen Zhang Xinyu Wang Qiuchen Wang Ruixue Ding Chenxi Wang Jialong Wu Yida Zhao Kuan Li