Pith. sign in

Paper Citation Record · LEDGER

VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning

As of 18 August 2026, this Paper Citation Record lists 67 of 67 outbound references and 29 inbound Pith citation observations for arXiv:2506.17221.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.17221 v2

Coverage vector

measured 67 of 67 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T19:13:35.477535Z

measured 96 of 96 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 29 of 29 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T14:28:52.281460Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-05T11:41:02.489833Z

Reference resolution

67 of 67 outbound references displayed

  • verified exact0
  • verified fuzzy35
  • unresolved32
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 95820f54-6477-4737-adfc-a83fcceeb0bc · outbound

This paper cites On Evaluation of Embodied Navigation Agents.

VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning On Evaluation of Embodied Navigation Agents

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:34.028057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:34.028057Z digest=sha256:101f42571d1c91defdc67bf120dec6ff864896d5367c1f9eed59cb42e9e015b8

Observation 03b99020-06ec-4168-a455-16f04ced02dc · outbound

This paper cites Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments.

VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:38.375762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:13:34.119892Z digest=sha256:121d18b86e31aa481ecc952d653ed5e68908cad125ccc11863d8c5d4c1d9e830

Observation e5c9ecc8-645d-4305-b11c-2fbe24adca73 · outbound

This paper cites InternLM2 Technical Report.

VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning InternLM2 Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:34.123874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:34.123874Z digest=sha256:104529b9a77b6752e5d8ffc034b15bb55cdacfe7b1f2e7af7918f8c3e09e1961

Observation 39cce60e-1488-4877-8bc3-a49c2a7001e1 · outbound

This paper cites Matterport3D: Learning from RGB-D data in indoor environments.

VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning Matterport3D: Learning from RGB-D data in indoor environments

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:38.207351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:13:34.128700Z digest=sha256:7c8d81ee4fa21dcbb87b768d8dfb9150af6b697f3bd94724718516f4e0652aa1

Observation 0207b2e9-a086-4c2f-8608-a1fac5f3f3c1 · outbound

This paper cites Mapgpt: Map-guided prompting for unified vision-and-language navigation.

VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning Mapgpt: Map-guided prompting for unified vision-and-language navigation

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:38.174363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:13:34.133535Z digest=sha256:0579ae2c20825213c75cd53d2c90705992db3d18f58759198b0305dc70bf3bcc

Observation e78035b7-549d-4078-b256-330e318eab6b · outbound

This paper cites Topological planning with transformers for vision-and-language navigation.

VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning Topological planning with transformers for vision-and-language navigation

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:38.163714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:13:34.137422Z digest=sha256:d8352c0e3ef6caa6d514e3098a50deb5baaafb8109579eeb6249dfffa55f8f38

Observation bc4f4967-a47e-4ded-87b0-f95b28402055 · outbound

This paper cites Weakly-supervised multi-granularity map learning for vision-and-language navigation.

VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning Weakly-supervised multi-granularity map learning for vision-and-language navigation

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:38.152988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:13:34.142245Z digest=sha256:4abfa36ac7d028deb6a3bae4b2b380e8533b39b6cd6e7fd53bbaba4a89ef1bb4

Observation 1ea06167-e011-4855-9021-92a07169cb14 · outbound

This paper cites Action-aware zero-shot robot navigation by exploiting vision-and-language ability of foundation models.

VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning Action-aware zero-shot robot navigation by exploiting vision-and-language ability of foundation models

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:37.976398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:13:34.146675Z digest=sha256:e7caef4e4b337b05a7b4a662ae2384b08830009ea84443f6e7aee5dc755449c8

Observation d730a088-f690-4c8d-a931-221c2bd60bb4 · outbound

This paper cites History aware multimodal transformer for vision-and-language navigation.

VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning History aware multimodal transformer for vision-and-language navigation

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:37.964476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:13:34.149939Z digest=sha256:604efc9eb2aef88621620b67d147b087f9f1801b5ec364289be8631e0ec376c7

Observation b41ab028-19db-4a8e-a4b5-753ec7b0ef97 · outbound

This paper cites Uniter: Universal image-text representation learning.

VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning Uniter: Universal image-text representation learning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:37.952439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:13:34.153030Z digest=sha256:65b8aef1dafe27470a0561d2b5791b3a40bbbed38cde0881ecbd4e05b51d3f7a

Observation f4e8e106-83e4-4343-81d7-e9d8bec7007b · outbound

This paper cites CLIP-Nav: Using CLIP for Zero-Shot Vision-and-Language Navigation.

VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning CLIP-Nav: Using CLIP for Zero-Shot Vision-and-Language Navigation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:34.157333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:34.157333Z digest=sha256:6a4464215a432cafcc2b60fb2b732a352e56817495962f17c20a31aff74816f5

Observation c9741b62-7064-4284-a6f4-b7f06ec1a3c8 · outbound

This paper cites Cross-modal map learning for vision and language navigation.

VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning Cross-modal map learning for vision and language navigation

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:37.803921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:13:34.161576Z digest=sha256:d3123a0c40741fc6424af68c8c38fec4adccdd61501073af94d251a4d01b5a7b

Observation 697f58a8-9f65-4993-9ff7-9268d923dbbd · outbound

This paper cites Airbert: In-domain pretraining for vision-and-language navigation.

VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning Airbert: In-domain pretraining for vision-and-language navigation

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:37.723752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:13:34.167609Z digest=sha256:466c14d84c1a9f6664095e50f3c801c15965d97a153b5a5a1037524107f5217f

Observation 785dd256-8062-4a6d-a829-f1952d97f422 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:34.171189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:34.171189Z digest=sha256:e7cb5febdc46e45bb96baa5242c860d5c423297b67fa4fdb97e5ba59a9e73b29

Observation 3f99d3f0-6805-467f-94f1-0230da081561 · outbound

This paper cites Towards learning a generic agent for vision-and-language navigation via pre-training.

VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning Towards learning a generic agent for vision-and-language navigation via pre-training

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:37.711785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:13:34.175470Z digest=sha256:d9c81c4222bf188375b89fa5cec3329cff0e6c11594c8b1a8ccea4f5ef4a250d

Observation 17049884-3f30-4baf-af0b-8db308361993 · outbound

This paper cites Bridging the gap between learning in discrete and continuous environments for vision-and-language navigation.

VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning Bridging the gap between learning in discrete and continuous environments for vision-and-language navigation

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:37.596724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:13:34.179320Z digest=sha256:ae66f96110d619cbd1bbf516a18a1d0db34387a3c21d1daf81586f3baf57c647

Observation 806dff77-4bba-4208-a52a-4979530cd93b · outbound

This paper cites Visual language maps for robot navigation.

VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning Visual language maps for robot navigation

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:37.549894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:13:34.183373Z digest=sha256:63f3597c8f3db26d6782bb1f24faca5f07e7c2f5bcd91cebd5a9faf72c0e4af3

Observation c4658de4-9c2e-4aa0-b1f2-327923451d36 · outbound

This paper cites Qwen2.5-Coder Technical Report.

VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning Qwen2.5-Coder Technical Report

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:34.187501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:34.187501Z digest=sha256:97573c453c38d1535a6fa2228db280943141049634ae0fb44aea0afb1bd1da87

Observation 99160d84-0f57-43c3-937b-616320948f42 · outbound

This paper cites SASRA: Semantically-aware Spatio-temporal Reasoning Agent for Vision-and-Language Navigation in Continuous Environments.

VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning SASRA: Semantically-aware Spatio-temporal Reasoning Agent for Vision-and-Language Navigation in Continuous Environments

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:34.236495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:34.236495Z digest=sha256:9a29ed09ce71300b8f651e2ae638fdcc54451a669c809fb0d9e9867de05447c0

Observation e1840740-6108-41ca-a6b0-109baf4aaf77 · outbound

This paper cites Preference Optimization for Reasoning with Pseudo Feedback.

VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning Preference Optimization for Reasoning with Pseudo Feedback

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:34.302141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:34.302141Z digest=sha256:255e4ea963244137e20d7e03155c0310481158ab3a84180b329d90f6fd62f39c

Observation 802f7653-1b0d-4e2a-b34c-618cc36a8ec6 · outbound

This paper cites Waypoint models for instruction-guided navigation in continuous environments.

VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning Waypoint models for instruction-guided navigation in continuous environments

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:37.470184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:13:34.470076Z digest=sha256:12f6b8a33c6883e5200bf603de3d758e47b3b1d2c91bfbc3f68adcafeae30581

Observation 7a30ae0f-578e-42f2-9d6c-ca424061c678 · outbound

This paper cites Sim-2-sim transfer for vision-and-language navigation in continu- ous environments.

VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning Sim-2-sim transfer for vision-and-language navigation in continu- ous environments

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:37.400296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:13:34.568146Z digest=sha256:619cae58655b7aff806162c236447c6fbb743ce84a8428b7dec4f43f871c6a15

Observation 0a56a40f-7e5e-455b-b4f6-108cf6bc48f7 · outbound

This paper cites Beyond the nav-graph: Vision and language navigation in continuous environments.

VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning Beyond the nav-graph: Vision and language navigation in continuous environments

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:37.387544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:13:34.624553Z digest=sha256:db282ba21ebf0fe0d5589ba843b03ae3e448fdab05f0cfa2bbcc678643c0102e

Observation e520c7a6-6cfc-4087-9022-9f0c700f1ad7 · outbound

This paper cites Room-across- room: Multilingual vision-and-language navigation with dense spatiotemporal grounding.

VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning Room-across- room: Multilingual vision-and-language navigation with dense spatiotemporal grounding

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:37.262456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:13:34.628683Z digest=sha256:ba1e68e32e16bf0e72af2ee0897562bae6f6e94548dad9bb5b74c63f02c1cf10

Observation c940f246-21f2-45bc-9120-d0d5e31d92c4 · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:34.632894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:34.632894Z digest=sha256:50f26857caa28f376f6b3408818bb8109b8cc52854e35f96a28dff4ef3172cc9

Observation 294741cb-7211-46f6-9a0c-efdb784586c5 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning LLaVA-OneVision: Easy Visual Task Transfer

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:34.637627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:34.637627Z digest=sha256:1ff9719552f7a0ec111c5981609aa2ca9f06613d98183dfe868ce50309e91c66

Observation 88b44393-45b1-4300-b672-c7118fe9035f · outbound

This paper cites VisualBERT: A Simple and Performant Baseline for Vision and Language.

VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning VisualBERT: A Simple and Performant Baseline for Vision and Language

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:34.642564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:34.642564Z digest=sha256:431afe697a8a754cef0eca3023e472708dec165d30017675cd20e2bef5e10103

Observation 9d7b072b-2819-49f8-89ec-b0ca015c7125 · outbound

This paper cites Oscar: Object-semantics aligned pre-training for vision-language tasks.

VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning Oscar: Object-semantics aligned pre-training for vision-language tasks

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:37.179433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:13:34.647351Z digest=sha256:73e788014889e9a1a581b815923d83490d0c10e80187d28c5f5b68d657d40f2f

Observation 0da77dcf-3ab5-4fe2-9120-8f1fc6bbd410 · outbound

This paper cites DeepSeek-V3 Technical Report.

VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning DeepSeek-V3 Technical Report

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:34.651647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:34.651647Z digest=sha256:2c27a69ae6226691dab4fcb27ee625d7504ebcee1baf960f673b88ad8d08531f

Observation 7a3af02d-e39d-4fc3-9183-98dedf14e1f8 · outbound

This paper cites Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs.

VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:34.656773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:34.656773Z digest=sha256:2b756d3a5aeac1689e2c9cbc80bc7d0ac9b953d2f38fae575ac0eeeda69c7092

Observation 0a325bba-ec89-4346-a5f6-d4d744577944 · outbound

This paper cites Visual instruction tuning.

VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning Visual instruction tuning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:34.660965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:34.660965Z digest=sha256:0c4e1675abfbde1bd37ea84e2900cc2fcd0bf0bc11723603a30cb45a7e0a3574

Observation 67aede46-2b0c-45b4-87fc-9d58544cd12d · outbound

This paper cites Aligning Cyber Space with Physical World: A Comprehensive Survey on Embodied AI.

VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning Aligning Cyber Space with Physical World: A Comprehensive Survey on Embodied AI

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:34.665920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:34.665920Z digest=sha256:d0cde291d289c0e33ee75e4af75f0f5c21a1893225ead249239b44fc435344a3

Observation 3ffb0c44-ba32-4a86-9701-b58cf209230d · outbound

This paper cites Instructnav: Zero-shot system for generic instruction navigation in unexplored environment.

VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning Instructnav: Zero-shot system for generic instruction navigation in unexplored environment

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:37.161342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:13:34.670307Z digest=sha256:f7f3d785b138b0db6d99c26ce4467122ce343d4ada86a046456a9ae733893052

Observation 5f45575c-524a-4a69-a58f-e9dccd63ccaa · outbound

This paper cites Discuss Before Moving: Visual Language Navigation via Multi-expert Discussions.

VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning Discuss Before Moving: Visual Language Navigation via Multi-expert Discussions

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:34.674399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:34.674399Z digest=sha256:b06d3859a2502ddcfa45ec6cbb35ad5be396afc2667ad05ab59ff4d4aedbc9d2

Observation 1630eed5-2748-4040-ad58-ae41279ceef6 · outbound

This paper cites Openeqa: Embodied question answering in the era of foundation models.

VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning Openeqa: Embodied question answering in the era of foundation models

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:36.955064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:13:34.678838Z digest=sha256:a1f900805aac53d3a2cffb85360c775e773226a747e5176dc35b3704f71465f0

Observation 04aecad7-8264-4423-9a18-067a2b59e62e · outbound

This paper cites Improving vision-and-language navigation with image-text pairs from the web.

VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning Improving vision-and-language navigation with image-text pairs from the web

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:36.895040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:13:34.682526Z digest=sha256:ff010c657b439bcfb08bc2d5d89005d268735c22651315ad8ff82ef6ba0f3515

Observation 52a91510-ec62-4ba5-8dd6-26896011261a · outbound

This paper cites Core challenges of social robot navigation: A survey.ACM Transactions on Human-Robot Interaction, 2023.

VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning Core challenges of social robot navigation: A survey.ACM Transactions on Human-Robot Interaction, 2023

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:36.882612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:13:34.832303Z digest=sha256:e0d77e7b9f9e755600b61be7a2397bf7af6eaec804615fed2bd6098faaf4a812

Observation 2f3dc8c7-0ce2-4101-814d-4c1c5b642f5a · outbound

This paper cites WMNav: Integrating Vision-Language Models into World Models for Object Goal Navigation.

VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning WMNav: Integrating Vision-Language Models into World Models for Object Goal Navigation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:34.882988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:34.882988Z digest=sha256:af59ee53195124255a58f34f4497e399d2350b851e6176e9a9e3da57b9e8cdbe

Observation 5fee35b4-0900-41df-a076-6127f161c7c8 · outbound

This paper cites Hello gpt-4o, 2024.

VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning Hello gpt-4o, 2024

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:36.806819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:13:34.941894Z digest=sha256:cc8b57e85893a7c454ead59962218668e1e4a686175f1c8fb0f85e3598141ef7

Observation a353eb50-356f-489e-80fa-cd2efd9adf0a · outbound

This paper cites Habitat 3.0: A Co-Habitat for Humans, Avatars and Robots.

VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning Habitat 3.0: A Co-Habitat for Humans, Avatars and Robots

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:35.003844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:35.003844Z digest=sha256:3dbfc5a1d0ceee67b0d5d4f0a27967479c9c75045f28803ec313450adbbcbd75

Observation 92d978e1-ad7d-4697-b8b2-2741cb00972b · outbound

This paper cites Hop: History- and-order aware pre-training for vision-and-language navigation.

VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning Hop: History- and-order aware pre-training for vision-and-language navigation

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:36.736526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:13:35.024408Z digest=sha256:068ea9f97994623694e8a55d1548c450ad875e45e222b89f095b8e9723057168

Observation 9cc78092-9cd9-421b-86ac-0e845ac13106 · outbound

This paper cites March in chat: Interactive prompting for remote embodied referring expression.

VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning March in chat: Interactive prompting for remote embodied referring expression

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:36.723183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:13:35.028417Z digest=sha256:f757b61e246bca894302e1aa9147693c10e2a79fd195eecf1a1537d57752d5a7

Observation 7c358a7f-d137-4465-9834-efd9306310f0 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning Direct preference optimization: Your language model is secretly a reward model

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:36.710421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:13:35.032505Z digest=sha256:35830c4c02709f97e67b04573c5284105b78e3684564d8cf87e64a1a4d1c810a

Observation 5dc97792-0628-496d-88d0-8f3ffcae0902 · outbound

This paper cites ZeRO: Memory Optimizations Toward Training Trillion Parameter Models.

VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning ZeRO: Memory Optimizations Toward Training Trillion Parameter Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:35.036739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:35.036739Z digest=sha256:755ddc9a3e23163614a7f7a2fc993e4c7bbf6c00a21659feb8d51e033ad2b351

Observation 2a45ed90-b05c-4e60-aa7a-3d9eff51b244 · outbound

This paper cites Language-Aligned Waypoint (LAW) Supervision for Vision-and-Language Navigation in Continuous Environments.

VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning Language-Aligned Waypoint (LAW) Supervision for Vision-and-Language Navigation in Continuous Environments

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:35.040784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:35.040784Z digest=sha256:9f067b2252529bbf444e2370fc3061851111f62d499808ed74a1fc76f2d61501

Observation 03faa876-b8ed-4f59-9b04-28bda83bdc0e · outbound

This paper cites Habitat: A Platform for Embodied AI Research.

VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning Habitat: A Platform for Embodied AI Research

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:36.570681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:13:35.045064Z digest=sha256:b634f418043ec14cae0b9ba322c79ef997df8825b93e3a377283d7a34b71cad2

Observation af473092-c51e-41af-ae58-b89b1d093ac0 · outbound

This paper cites Proximal Policy Optimization Algorithms.

VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning Proximal Policy Optimization Algorithms

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:35.048732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:35.048732Z digest=sha256:1af962feff3c6230a01c9e746e21bbb3aaa8430a5a8b6d50e4ce7b8e62ec281c

Observation ea8bbc1f-f466-4dc8-ac43-b7f4575299b5 · outbound

This paper cites Lm-nav: Robotic navigation with large pre-trained models of language, vision, and action.

VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning Lm-nav: Robotic navigation with large pre-trained models of language, vision, and action

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:36.455953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:13:35.052892Z digest=sha256:e04950692031d94a01db861c4093b78da61b6d8c3519f5cb51b5d462c13e653a

Observation c4eda079-a033-49f0-bd35-e33ab1691b84 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:35.056673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:35.056673Z digest=sha256:3869027bcbc3ccf8d4b8b58f716312b73ebdd2766540ae0144769cc5247e0264

Observation cdf2fac2-3abc-4af1-a821-ff090a05bdb1 · outbound

This paper cites VL-BERT: Pre-training of Generic Visual-Linguistic Representations.

VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning VL-BERT: Pre-training of Generic Visual-Linguistic Representations

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:35.061441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:35.061441Z digest=sha256:e3ac93d43ccfd65c50eafa5a4d2076b36e02982585a8bc85f3a51655b2d9929a

Observation 20bac8e7-16be-4130-8bb0-a77188020fda · outbound

This paper cites Aligning Large Multimodal Models with Factually Augmented RLHF.

VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning Aligning Large Multimodal Models with Factually Augmented RLHF

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:35.069330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:35.069330Z digest=sha256:7d3908b6e155af208d62a6f96eb90191b421395288209e5ed2aeb3fef79d0ef1

Observation 0add378d-4fe4-4695-aec2-e461b438987f · outbound

This paper cites Lxmert: Learning cross-modality encoder representations from transformers.

VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning Lxmert: Learning cross-modality encoder representations from transformers

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:36.443423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:13:35.073962Z digest=sha256:da95fa959d7d94d7a58bd23b8b8f31f73b2322e27a279e6dac629999c6464aee

Observation 471ca004-79fa-4a4e-b5bd-95f5debe78f2 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:35.078104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:35.078104Z digest=sha256:865087cdb7001ed38b2dc3c85282d4ed71adf5e019affca170a846c5d9f049e0

Observation e0723114-6cbb-4338-a40e-6655eb42e865 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:35.082112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:35.082112Z digest=sha256:7143b1f6b252babbb70ce6360f815a32313ac28d8a448ceb0e8f3d4154c8b720

Observation eac053ab-16e0-427a-bbf0-39a34d8e9763 · outbound

This paper cites Cross-modal semantic alignment pre-training for vision-and-language navigation.

VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning Cross-modal semantic alignment pre-training for vision-and-language navigation

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:36.304409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:13:35.121601Z digest=sha256:46e849a90dda98b82758e92b446ca1500517e7641dc23630a9d463f925e881b1

Observation 2d2d3e94-8a51-46e0-9cf3-f9ef8b0d1122 · outbound

This paper cites RlHF-V: Towards trustworthy mllms via behavior alignment from fine-grained correctional human feedback.

VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning RlHF-V: Towards trustworthy mllms via behavior alignment from fine-grained correctional human feedback

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:36.190848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:13:35.255984Z digest=sha256:1aee08f95ba69988f93fe55de092a2f820630afb2a9a959e4b791eabe34ce74c

Observation e6d0083e-8dd8-4b57-aad8-aa2b0ffca1c3 · outbound

This paper cites RLAIF-V: Aligning mllms through open-source ai feedback for super gpt-4v trustworthiness.arXiv:2405.17220, 2024.

VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning RLAIF-V: Aligning mllms through open-source ai feedback for super gpt-4v trustworthiness.arXiv:2405.17220, 2024

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:35.317485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:35.317485Z digest=sha256:d450bbe6b76dfc10a7ef7555ac151006b03a2175f36bc067d0f25fdd7b2e4c4a

Observation 3c43d417-d420-4ebc-96e4-374abb7c4902 · outbound

This paper cites Uni-navid: A video-based vision-language-action model for unifying embodied navigation tasks.

VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning Uni-navid: A video-based vision-language-action model for unifying embodied navigation tasks

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:36.178360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:13:35.321801Z digest=sha256:f81c7b5802a47f6378c16c3892b92a8e28c42279211b2cabc94c860f4f919beb

Observation bd3a18ce-95ba-4e7a-add5-5bc2e405887d · outbound

This paper cites Navid: Video-based vlm plans the next step for vision-and-language navigation.

VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning Navid: Video-based vlm plans the next step for vision-and-language navigation

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:36.167393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:13:35.325678Z digest=sha256:88b975227cbe6715680f39ef53bcccf6477f15c252710d8b31ae8d9afe4b9d62

Observation b154de35-01e1-4a52-9abc-b5a2363c4374 · outbound

This paper cites CodeDPO: Aligning Code Models with Self Generated and Verified Source Code.

VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning CodeDPO: Aligning Code Models with Self Generated and Verified Source Code

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:35.330200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:35.330200Z digest=sha256:fd313b8cf92bb0a9aaca6ea46eefd7c4ec8c6f3d83b6499fe2741130865abeec

Observation 342689a5-54d0-457f-b725-6638bf4380a0 · outbound

This paper cites o1-Coder: an o1 Replication for Coding.

VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning o1-Coder: an o1 Replication for Coding

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:35.335052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:35.335052Z digest=sha256:23b09b0ca1d8d3d39710d7595c70ca6a8969af4f28fedfe22c2988c3f588a4dc

Observation 8fe4223b-2061-47d0-bd69-ad85a6f78bbc · outbound

This paper cites Beyond Hallucinations: Enhancing LVLMs through Hallucination-Aware Direct Preference Optimization.

VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning Beyond Hallucinations: Enhancing LVLMs through Hallucination-Aware Direct Preference Optimization

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:35.339035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:35.339035Z digest=sha256:f236de7a590c9ba01982e202e8f1d28ac2d0812a012c1d7227d496fd401b1d47

Observation c72e89a7-32b8-4e48-88b8-064a8684baec · outbound

This paper cites Towards learning a generalist model for embodied navigation.

VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning Towards learning a generalist model for embodied navigation

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:35.926368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:13:35.406823Z digest=sha256:e0dd2d81aa6a88be5a18914ebe82bd72bd111b7d1ff99fb4130965845738e133

Observation 214c657b-3da0-4473-be45-54f0f86dbf7f · outbound

This paper cites Navgpt-2: Unleashing navigational reasoning capability for large vision-language models.

VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning Navgpt-2: Unleashing navigational reasoning capability for large vision-language models

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:35.912524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:13:35.464627Z digest=sha256:ff6a376710957fca5ec7290a253c12c45887d405c08b76629b4c5ef793d5a142

Observation 8f2a9eee-70ce-4a39-81d1-78360dd66e04 · outbound

This paper cites NavGPT: Explicit Reasoning in Vision-and-Language Navigation with Large Language Models.

VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning NavGPT: Explicit Reasoning in Vision-and-Language Navigation with Large Language Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:35.468781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:35.468781Z digest=sha256:8201ed04eaab7c40b9aa55dbf3ebf2d533e530f9f50e3ee64cc2537e3401018e

Observation 69e6592b-45cb-4852-ae65-72397ef3a9b9 · outbound

This paper cites Aligning Modalities in Vision Large Language Models via Preference Fine-tuning.

VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning Aligning Modalities in Vision Large Language Models via Preference Fine-tuning

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:35.473688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:35.473688Z digest=sha256:8ed49bd0823d437a7db5d7e5d40e54b43c1a7926aa4dbb3d0acc717749f01e40

Observation 822b315f-3eaf-4ef9-a0ee-2a97b7514e12 · outbound

This paper cites Deep Learning for Embodied Vision Navigation: A Survey.

VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning Deep Learning for Embodied Vision Navigation: A Survey

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:35.477535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:35.477535Z digest=sha256:1af87c4c5031d81d599e1350ec9390fe6dd8eb5e75e3db989b3a2f7a5e005ef0

Pith citing papers

Observation 73c5e88f-7f8c-427d-a906-4abc614a6403 · inbound

Pseudo Depth Meets Gaussian: A Feed-forward RGB SLAM Baseline cites this paper.

Pseudo Depth Meets Gaussian: A Feed-forward RGB SLAM Baseline VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T23:56:04.195788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:56:04.195788Z digest=sha256:058eca714dee20d50a89030723208c6dd7044494107b099ce7c7f0d707eabaa6

Observation b85f0e97-f99f-4328-b6a6-3b51bc81b640 · inbound

Nav-R1: Reasoning and Navigation in Embodied Scenes cites this paper.

Nav-R1: Reasoning and Navigation in Embodied Scenes VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T17:31:42.089676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:31:42.089676Z digest=sha256:f7aacca1f7a83dcb2fc60ec023cd332716ee1a426efa9e19bffed2122a70ab97

Observation 5618983b-961c-4d25-8831-f4de5fb791b7 · inbound

DreamNav: A Trajectory-Based Imaginative Framework for Zero-Shot Vision-and-Language Navigation cites this paper.

DreamNav: A Trajectory-Based Imaginative Framework for Zero-Shot Vision-and-Language Navigation VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T17:00:45.020389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:00:45.020389Z digest=sha256:95f6560adcaeacb6096c2d0aad8d2a9bfc89485a7316d59aa625fa9e519ed665

Observation aa6ee4d9-00b3-47dd-9ceb-a3d06dbea990 · inbound

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle cites this paper.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning

Reference 140

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:37.828525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:37.828525Z digest=sha256:355d954cf48e41a7200136d81f2a014f1762f6de906dba47db33861ffff4b3fd

Observation 8f6d4493-0bc1-4a99-9994-6333f0c359cd · inbound

IRPO: Boosting Image Restoration via Post-training GRPO cites this paper.

IRPO: Boosting Image Restoration via Post-training GRPO VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-03T19:24:39.321499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:24:39.321499Z digest=sha256:f00f0b9e0de8966f73a0a00461b545a1f973fb8e58bcab032a7e6b99baaa1c4e

Observation 49bc9c5f-727f-4eb4-9bc8-98d0d5dd10a2 · inbound

Aerial Vision-Language Navigation with a Unified Framework for Spatial, Temporal and Embodied Reasoning cites this paper.

Aerial Vision-Language Navigation with a Unified Framework for Spatial, Temporal and Embodied Reasoning VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-16T23:58:42.589185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-16T23:58:40.992942Z digest=sha256:a7d5a6e5ee128f4f0eb42c42adb02243ac54787bdac3c71a87cb63d107e315e3

Observation 6ef87bd5-bd44-4240-8069-7b85f724ef06 · inbound

Token Warping Helps MLLMs Look from Nearby Viewpoints cites this paper.

Token Warping Helps MLLMs Look from Nearby Viewpoints VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-05-13T21:08:17.442855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-13T21:07:55.062113Z digest=sha256:66aa7cc9a57170b83ab36785554b64d00075a601b966c5ca01b653f424350158

Observation 818345e2-d07f-4a43-a906-1cca8202e0c9 · inbound

Hypothesis Graph Refinement: Hypothesis-Driven Exploration with Cascade Error Correction for Embodied Navigation cites this paper.

Hypothesis Graph Refinement: Hypothesis-Driven Exploration with Cascade Error Correction for Embodied Navigation VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-13T17:08:00.907219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-13T17:05:57.205365Z digest=sha256:4a2c6dbcf1d77145c70b57eb56173b55bf062ad74141be30b2e445514e131636

Observation fde93708-7efc-41f8-aaa4-f929867df094 · inbound

Enhancing MLLM Spatial Understanding via Active 3D Scene Exploration for Multi-Perspective Reasoning cites this paper.

Enhancing MLLM Spatial Understanding via Active 3D Scene Exploration for Multi-Perspective Reasoning VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:15:47.751455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T19:16:46.753641Z digest=sha256:f07bf4ecd95e94a6648f1ffdeaebd58451cca65003329aa78b6c69b60340af55

Observation 53443418-d67e-4f2a-9708-e9072ca1faf8 · inbound

HiRO-Nav: Hybrid ReasOning Enables Efficient Embodied Navigation cites this paper.

HiRO-Nav: Hybrid ReasOning Enables Efficient Embodied Navigation VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:41:01.531030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T17:01:13.350219Z digest=sha256:039c502218fe71cba42848ed6b92e993b3b6ecc9487a615a0cd227be47c8fabc

Observation 062ae992-5bef-4e9a-af06-a426058fbc1f · inbound

Think before Go: Hierarchical Reasoning for Image-goal Navigation cites this paper.

Think before Go: Hierarchical Reasoning for Image-goal Navigation VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:51:10.485653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-10T05:43:27.972164Z digest=sha256:56cd9e21ac6d504c75f4df578be55a094865b36de459735f130f53157fe04fcd

Observation d7e87bfa-b62d-4502-b272-c4e2f2ccce5b · inbound

Dual-Anchoring: Addressing State Drift in Vision-Language Navigation cites this paper.

Dual-Anchoring: Addressing State Drift in Vision-Language Navigation VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:51:45.828270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-10T06:50:34.310831Z digest=sha256:b67fee358656def06c0ee58c7af6b403bb4ddb6fb19383828bf5ad3fa654ed14

Observation 9f612723-deee-4adf-a40e-1a7f51c2744e · inbound

Steadily moving semi-infinite fracture in plane poroelasticity cites this paper.

Steadily moving semi-infinite fracture in plane poroelasticity VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning

Reference 77

Resolution
verified exact
local_arxiv, observed 2026-07-05T11:41:02.491147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-05T11:39:05.686584Z digest=sha256:b764cf21f5e8b68a4b37e655b6cb7e9ad3ef1c8b14f068a60d890f8a9520cc15

Observation c75ecfb5-fb62-42f1-a8d0-f4067bd19a30 · inbound

XEmbodied: A Foundation Model with Enhanced Geometric and Physical Cues for Large-Scale Embodied Environments cites this paper.

XEmbodied: A Foundation Model with Enhanced Geometric and Physical Cues for Large-Scale Embodied Environments VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:51:10.296777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T05:46:36.865150Z digest=sha256:b29fc5df1210e31e3012008c4903fb2ae14543c3102af884f9237451445e752b

Observation e37d385e-ef1d-407e-acb5-948dbc1a8cc6 · inbound

SpaAct: Spatially-Activated Transition Learning with Curriculum Adaptation for Vision-Language Navigation cites this paper.

SpaAct: Spatially-Activated Transition Learning with Curriculum Adaptation for Vision-Language Navigation VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:36:31.243113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-07T05:05:33.606975Z digest=sha256:4f74a7c2cf0570272acf8082b49f56053fffcca5de505337b65c76a4577db74c

Observation 79ce5e0b-fb83-47c5-bf81-7c5e0cf19292 · inbound

Beyond Thinking: Imagining in 360$^\circ$ for Humanoid Visual Search cites this paper.

Beyond Thinking: Imagining in 360$^\circ$ for Humanoid Visual Search VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:01:18.332757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-12T02:58:46.728868Z digest=sha256:a944b6fbd01887e49fcc243e13e8a84df2fda6e37a6c48560917adfe3f735c5e

Observation f66da04d-f31d-4ccc-9688-5ff8f158570e · inbound

WorldVLN: Autoregressive World Action Model for Aerial Vision-Language Navigation cites this paper.

WorldVLN: Autoregressive World Action Model for Aerial Vision-Language Navigation VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-20T17:48:48.633274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T17:47:32.953903Z digest=sha256:fc53d5a3299c6a50050073d77447be9fd8df6c265a2d7bfcb8a363b3b456aaac

Observation 0ce6c4d5-b991-4c1b-a393-9f1103944045 · inbound

AwareVLN: Reasoning with Self-awareness for Vision-Language Navigation cites this paper.

AwareVLN: Reasoning with Self-awareness for Vision-Language Navigation VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-22T04:46:04.568388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-22T04:46:03.020800Z digest=sha256:9877aec7b2689501efe60bd638c5a9582c57e1561a2cb2b82a05928402b19be7

Observation 8da6f638-a852-4621-a212-12dc15632643 · inbound

World Models as Group Actions cites this paper.

World Models as Group Actions VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:24:40.331396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T13:17:42.720825Z digest=sha256:1b3c3400e56f25c10a014616acac2d3aa49fab86722a2b5377e62fe3b7f713b1

Observation ed9be44b-6172-4d41-b974-7b2b4d67a4ed · inbound

Bridging the 2D-3D Gap: A Hierarchical Semantic-Geometric Map for Vision Language Navigation cites this paper.

Bridging the 2D-3D Gap: A Hierarchical Semantic-Geometric Map for Vision Language Navigation VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:34:01.336932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-29T22:33:32.671550Z digest=sha256:2524479d28396d0744f45217bf7c9ff5579a99cd17d8159a2be885d069dc0531

Observation b0a0df41-f434-4da9-90ef-c03458f72283 · inbound

Reasmory: 3D Reconstruction as Explicit Memory for VLMs Spatial Reasoning cites this paper.

Reasmory: 3D Reconstruction as Explicit Memory for VLMs Spatial Reasoning VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:46:13.683917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-28T17:46:21.821764Z digest=sha256:76ec2fa8d760611280738124a7787bd21aaa699da3b1fef4657fdbf5200fab4d

Observation 59072d98-3d15-47d9-b66f-d43b0980ef2a · inbound

Goal2Pixel: Grounding Goals to Pixels for Vision-Language Navigation cites this paper.

Goal2Pixel: Grounding Goals to Pixels for Vision-Language Navigation VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:16:15.779796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-28T15:39:13.803571Z digest=sha256:329663ed50f0c798768d4276d5d7fc0d2c033f7e4f15bc952a4cdcae7d8ff757

Observation 39828a1e-b0db-486a-b879-ad91aba9459b · inbound

Beyond Waypoints: A Trajectory-Centric Waypointing Paradigm for Vision-Language Navigation cites this paper.

Beyond Waypoints: A Trajectory-Centric Waypointing Paradigm for Vision-Language Navigation VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:07:17.765948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T21:40:58.329546Z digest=sha256:5b521a90e7443baa4f3fcf4edd5c51b2bd313acba80c4e2db8e6360375fc61a1

Observation 9ee761ed-2ee7-4e27-8aab-274aa4eee27c · inbound

Watch, Remember, Reason: Human-View Video Understanding with MLLMs cites this paper.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning

Reference 239

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:15.728048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:ac2e22575ce7b16fe2a7ebe02ccd2c40680cebc860063741f65e5a91a8f33a3c

Observation afdb57f2-0228-4d53-b770-cbc3316ed17e · inbound

Path-level Hindsight Instructions for Semantic Exploration in Vision-Language Navigation cites this paper.

Path-level Hindsight Instructions for Semantic Exploration in Vision-Language Navigation VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:08:21.390625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-03T14:04:46.400336Z digest=sha256:c5231de3d343b2738b6bdf88b83287aa79362a11cbd682667c5f003c4e189da6

Observation 43227267-f85c-427b-b2e8-587b41454996 · inbound

A Comprehensive Survey and Systematic Real-World Evaluation of Embodied Vision-and-Language Navigation cites this paper.

A Comprehensive Survey and Systematic Real-World Evaluation of Embodied Vision-and-Language Navigation VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-14T15:39:59.878169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:39:59.878169Z digest=sha256:a082945c851b3ee665a160145b3acb9ef2dc36c7875c5c3a13f22f1bda69587d

Observation 86424cd5-e548-4337-b63d-c8f01723c1ad · inbound

Joint On-and-Off Policy Learning for Vision-and-Language Navigation cites this paper.

Joint On-and-Off Policy Learning for Vision-and-Language Navigation VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T05:13:54.793452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:13:54.793452Z digest=sha256:92cc950fb935fa07819c8564fd832bb0472c5b8b6fa7341d9dbea3a45a8563e6

Observation 7b9be73c-fe2f-4bf4-b991-db29888e2848 · inbound

Towards Dual-Brain Minimal Sufficient Representation for Vision-Language Navigation cites this paper.

Towards Dual-Brain Minimal Sufficient Representation for Vision-Language Navigation VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T03:26:15.484911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:26:15.484911Z digest=sha256:04a3a344ed22586e079382e6bec01a358c8e5443f996fb5235c29e024c507cde

Observation 918d788f-3643-4bfe-b999-c2ffe3ed2f8f · inbound

RecoverFly: A Failure-Aware Reinforcement Learning Post-Training Framework for Aerial Vision-Language Navigation cites this paper.

RecoverFly: A Failure-Aware Reinforcement Learning Post-Training Framework for Aerial Vision-Language Navigation VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T14:28:52.281460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:28:52.281460Z digest=sha256:81bea9332edf2504ea6086c8cd2d731525e8e7904ebee3734cb1553bc58b1864