HyperAIHyperAI

Command Palette

Search for a command to run...

PANOVLN: TOWARDS EFFECTIVE PANORAMIC VISION-AND-LANGUAGE NAVIGATION

Zhen Wang Changpeng Wang Zhe Liu Zhangyang Qi Yuxiang Lu Zimo Zeng Donglian Qi Xi Chen

Abstract

Recent vision-language models (VLMs) have advanced vision-and-language navigation (VLN), enabling models to predict navigation actions from visual observations and language instructions. In this work, we explore VLN with panoramic observations and introduce PanoVLN. The motivation is straightforward: more complete visual context should enable better-informed navigation decisions. For example, a panorama can reveal a passage outside a perspective camera’s field of view, allowing the model to identify the intended route without additional exploration. However, we find that simply replacing perspective images with panoramas yields only limited gains. Our diagnosis suggests that fully exploiting wider visibility requires modifications to action prediction, training supervision, and visual representation. First, wider visibility supports longer-horizon action planning. We make the model predict longer action sequences, enabling larger turns and subsequent movement from a single panorama. Specifically, we introduce a confidence-guided execution (CGE) strategy that dynamically determines how many predicted actions to execute before replanning. Second, wider visibility also brings more complex route choices. We therefore construct training routes with frequent branching points and clear instructions to provide targeted supervision for route selection. Third, panoramic navigation requires understanding spatial relationships across viewing directions, beyond recognizing individual landmarks. We combine semantic and geometric features from RGB panoramas to capture both scene content and spatial layout without adding visual tokens. With a 4B backbone and RGB-only input, PanoVLN surpasses the previous SOTA by 11.9% and 8.7% in success rate on R2R-CE and RxR-CE Val-Unseen. Real-world experiments on a quadruped further demonstrate faster navigation with fewer pauses than prior VLN methods. Our project page is available at https://wangzhen-w.github.io/PanoVLN/.

One-sentence Summary

Researchers from Zhejiang University and The University of Hong Kong propose PanoVLN, a panoramic vision-and-language navigation model that uses confidence-guided execution for longer-horizon actions, branching-point training supervision, and combined semantic-geometric RGB panorama features, surpassing the previous state of the art by 11.9%11.9\%11.9% and 8.7%8.7\%8.7% in success rate on R2R-CE and RxR-CE Val-Unseen and enabling faster real-world quadruped navigation with fewer pauses.

Key Contributions

  • PanoVLN is a panoramic vision-and-language navigation method that predicts longer action sequences from a single panorama and uses confidence-guided execution to determine how many predicted actions to execute before replanning.
  • A decision-centric training dataset of 98K trajectories across 800 HM3D scenes provides frequent branching points and visually grounded instructions, with denser sampling around turns and stopping points for route selection and completion supervision.
  • The method combines semantic VLM features with geometric PanoVGGT features from the same RGB panorama; with a 4B RGB-only backbone, it achieves 77.3% success on R2R-CE Val-Unseen and 78.0% on RxR-CE Val-Unseen, improving over the previous state of the art by 11.9% and 8.7% and enabling faster real-world quadruped navigation with fewer pauses.

Introduction

Vision-and-language navigation (VLN) requires an agent to follow natural-language instructions through an environment, and recent vision-language models have improved this task by predicting navigation actions from visual observations. Most prior work uses perspective images, which limit the visual context available at each decision point. The authors investigate equirectangular panoramas, which provide a 360-degree view and can reveal passages, landmarks, and route alternatives. They find that simply replacing perspective images with panoramas does not improve performance under the same setup, so they propose PanoVLN, which adapts action prediction with longer horizons and confidence-guided execution, creates a 98K-trajectory decision-centric training dataset, and fuses semantic VLM features with panoramic geometric features. This approach achieves state-of-the-art success rates on R2R-CE and RxR-CE and enables faster real-world robot navigation with fewer pauses.

Dataset

Dataset composition and sources

  • The authors construct 98K navigation trajectories across 800 HM3D scenes.
  • Trajectories are designed to contain frequent branching points, where the agent must choose among multiple visible traversable paths.
  • Each trajectory is paired with an instruction that identifies the chosen path and stopping location.

Route construction and filtering

  • Walkable space is divided into connected areas using the navigation mesh.
  • A branching point is defined as having at least two visible, traversable paths to different areas, excluding the incoming path.
  • Endpoints are sampled in different areas, and routes passing through branching points are retained.
  • Rendering-quality checks remove candidates with mesh holes or incomplete geometry, followed by near-duplicate removal.
  • An expert converts remaining routes into primitive action sequences.
  • Replay verifies goal reachability and confirms that the chosen path and its alternatives are visible at each branching point.

Instruction construction and verification

  • Trajectories are divided into travel, branching, and arrival segments.
  • Travel segments use first-person video with the expert path marked on the ground.
  • Branching and arrival segments additionally use eight-view compass images.
  • Qwen3.8-27B describes movement, identifies the chosen path from visible cues, and specifies the stopping location.
  • Descriptions are combined in route order, with repetition removed and wording refined.
  • Verification uses clean videos and compass images without instruction or route overlays.
  • Three checks are applied: motion consistency, choice grounding, and stop grounding.
  • Mismatched segments are revised locally and reverified; only samples passing all three checks are retained.

Training sample construction and usage

  • The data is used to provide supervision for learning path selection from panoramic observations.
  • For H = 18, the authors use a stride-six grid to reduce overlap between adjacent grid targets from 17 to 12 actions.
  • Additional states are added at sustained-turn onsets and near termination to supervise turning and stopping.
  • Each state is paired with its H-step expert action sequence.
  • The grid preserves route coverage, while added states emphasize action transitions.
  • The provided section does not specify train/eval split or mixture ratios.

Method

The authors propose a method to fully exploit the complete visual context provided by panoramas in the Vision-and-Language Navigation task. Starting from a baseline model that takes panoramas as input, they introduce three key adaptations: longer action-sequence supervision and execution, decision-centric data construction, and geometry-aware visual representations.

To leverage the wider visibility of panoramic observations, the authors extend the action prediction horizon. Instead of predicting a single step, the policy is trained to predict a sequence of the next HHH expert actions. The training objective minimizes the negative log-likelihood of the expert action sequence using teacher forcing:

Lact=−1H∑i=1Hlog⁡pθ(at,i∗∣Ot,At,<i∗)\mathcal{L}_{\mathrm{act}} = - \frac{1}{H} \sum_{i=1}^{H} \log p_{\theta} \left(a_{t,i}^{*} \mid \mathcal{O}_{t}, \mathbf{A}_{t,<i}^{*}\right)Lact​=−H1​i=1∑H​logpθ​(at,i∗​∣Ot​,At,<i∗​)

where At,<i∗\mathbf{A}_{t,<i}^{*}At,<i∗​ contains the preceding expert actions and pθp_{\theta}pθ​ is the VLM next-token distribution.

During inference, the execution length is adapted based on prediction uncertainty through a mechanism called Confidence-Guided Execution. The uncertainty for a generated action is defined as the negative log probability of the predicted action. As shown in the figure below:

Mean uncertainty rises after an initial dip and exhibits substantial variation across policy calls. To handle this, the execution mechanism extends the executed prefix as long as the cumulative uncertainty Ut(k)U_t(k)Ut​(k) remains within a predefined budget BBB, ensuring at least Emin⁡E_{\min}Emin​ actions are executed:

Et=max⁡{k∈{1,…,H}:k≤Emin⁡ or Ut(k)≤B}E_{t} = \max \left\{k \in \{1, \dots, H\}: k \leq E_{\min} \text{ or } U_{t}(k) \leq B \right\}Et​=max{k∈{1,…,H}:k≤Emin​ or Ut​(k)≤B}

The agent executes this prefix and then reobserves the environment unless it predicts a stop action.

To provide better supervision for selecting the correct path among multiple visible options, the authors construct a decision-centric training dataset. They generate navigation trajectories across various scenes, specifically targeting branching points where at least two traversable paths are visible. After filtering out routes with rendering issues or near-duplicates, an expert converts the valid routes into primitive action sequences. Instructions are constructed by dividing trajectories into travel, branching, and arrival segments. A large language model describes the movement and identifies the chosen path using first-person video and compass images. The authors rigorously verify motion consistency and choice grounding, revising and retaining only the segments that pass all checks. To reduce overlap between adjacent training states, they employ a stride-six grid for sampling and add specific states at sustained-turn onsets and near termination to emphasize action transitions.

Finally, the authors develop a geometry-aware visual representation to better understand the spatial relationships within a panoramic observation. They allocate a larger number of tokens to the current equirectangular panorama and fewer tokens to each history frame. To fuse geometric information without adding extra visual tokens, a pretrained PanoVGGT encoder extracts geometric features from the current RGB panorama. These features are resampled in ERP coordinates and grouped to align with the VLM merged current tokens. A trainable MLP projects the aligned geometric groups into the visual-token embedding space for residual fusion:

Vˉt=Vt+αfψ(Gt)\bar{V}_{t} = V_{t} + \alpha f_{\psi}(G_{t})Vˉt​=Vt​+αfψ​(Gt​)

where α\alphaα is a fixed residual scale. This fusion combines semantics and geometry from corresponding ERP regions while preserving the token count and order, ultimately conditioning the action prediction alongside the instruction and history tokens. The VLM and projection layers are trained jointly, while the geometry encoder remains frozen.

Experiment

Experiments evaluate RGB-only PanoVLN on R2R-CE and RxR-CE Val-Unseen splits in Matterport3D using Habitat, with metrics including navigation error, success rate, SPL, and nDTW. Simulation results show PanoVLN achieves state-of-the-art success rates on both benchmarks and benefits from panoramic context, while real-world tests on a Unitree Go2 across hallway, office, and campus settings demonstrate reliable indoor route following, outdoor transfer, and efficient execution through longer predicted segments with confidence-guided execution. Ablations confirm that longer ERP prediction horizons, turn- and termination-aware sampling, confidence-guided execution, PanoVGGT panoramic features, and decision-centric training trajectories all improve navigation and stopping behavior.

PanoVLN achieves the highest success rate and SPL on both R2R-CE and RxR-CE, setting a new state of the art by large margins over prior panoramic navigation methods. A restricted-data version of PanoVLN also leads its training-data group in success rate and SPL. The results connect panoramic context and decision-centric trajectories to better generalization in unseen scenes. PanoVLN surpasses previous best success rates by 11.9 percentage points on R2R-CE and 8.7 percentage points on RxR-CE. PanoVLN trained without navigation data beyond R2R-CE and RxR-CE still leads the restricted-data group in SR and SPL on both benchmarks. Adding decision-centric trajectories further improves performance in unseen scenes.

Across the real-world routes, PanoVLN achieves the shortest navigation duration and the highest travel speed among all compared methods. It also spends the least time waiting for policy responses and records far fewer pauses and policy calls. Although its per-request inference latency is not the lowest, its overall execution flow is the most efficient. PanoVLN combines the shortest navigation duration with the highest travel speed, while JanusVLN is the slowest by a wide margin. PanoVLN has the lowest waiting fraction, fewest pauses, and fewest policy calls, despite not having the lowest per-request latency.

Compared with random-start sampling, the proposed turn- and termination-aware sampling lowers navigation error while keeping oracle success nearly unchanged. It also increases success rate and SPL by a clear margin, indicating more reliable termination and route completion when turn and stop states receive stronger supervision. The proposed sampling reduces navigation error relative to random-start sampling while maintaining comparable oracle success. It improves success rate and SPL, suggesting more reliable termination in the goal region and better supervision of turn and stop states.

Confidence-guided execution achieves the best navigation performance on both benchmarks, outperforming fixed execution lengths and random execution. Fixed execution at one action is the strongest fixed setting, while longer fixed horizons generally reduce success and path quality. The results indicate that adapting execution to model confidence is more reliable than committing to a preset or random number of actions. Confidence-guided execution records the lowest navigation error and the highest success and path-quality metrics on both benchmarks. Among fixed strategies, one-action execution is strongest, and longer fixed horizons degrade success and SPL, especially on RxR-CE. Random execution over one to eighteen actions underperforms confidence-guided execution and trails the one-action fixed strategy on RxR-CE.

Geometry encoder comparison under matched fusion and execution settings shows mixed effects for existing encoders. PanoVGGT achieves the highest success rate on both R2R-CE and RxR-CE, and also leads on R2R-CE oracle success and SPL and on RxR-CE nDTW. The results suggest its panoramic geometric features add spatial cues that support route selection. PanoVGGT leads all compared encoders in success rate on both benchmarks, with the best R2R-CE oracle success and SPL and the best RxR-CE nDTW. Alternative encoders have mixed effects: UniK3D improves R2R-CE navigation error and SPL but not RxR-CE success, while DA^2 and DAP tend to reduce success on both benchmarks.

The experiments benchmark PanoVLN on R2R-CE and RxR-CE, real-world navigation routes, and ablations of sampling, execution, and geometry encoders. PanoVLN sets new state-of-the-art success and SPL on both benchmarks by large margins, and its real-world runs achieve the shortest duration and highest speed with fewer pauses and policy calls despite not having the lowest per-request latency. Turn- and termination-aware sampling and confidence-guided execution improve success and path quality over random or fixed baselines, while among geometry encoders PanoVGGT gives the strongest overall results and other encoders yield mixed effects.


Build AI with AI

From idea to launch — accelerate your AI development with free AI co-coding, out-of-the-box environment and best price of GPUs.

AI Co-coding
Ready-to-use GPUs
Best Pricing

HyperAI Newsletters

Subscribe to our latest updates
We will deliver the latest updates of the week to your inbox at nine o'clock every Monday morning
Powered by MailChimp