Pith. sign in

Paper Citation Record · LEDGER

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations

As of 23 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 2 inbound Pith citation observations for arXiv:2506.08566.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.08566 v1

Coverage vector

measured 65 of 65 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:12:12.349149Z

measured 67 of 67 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:34:35.053699Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T23:02:52.667737Z

Reference resolution

65 of 65 outbound references displayed

  • verified exact10
  • verified fuzzy21
  • unresolved25
  • parse uncertain0
  • malformed identifier7
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4b4c694c-d47f-4296-a2a7-f1778dedafb3 · outbound

This paper cites Spatio- temporal dynamics and semantic attribute enriched visual encoding forvideocaptioning,in:CVPR,pp.12487–12496.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Spatio- temporal dynamics and semantic attribute enriched visual encoding forvideocaptioning,in:CVPR,pp.12487–12496

Reference 1

Resolution
malformed identifier
no resolver link, observed 2026-08-07T05:12:12.021740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:12:12.021740Z digest=sha256:23df2f925a7c6f01d280d7d1ca2e117874ce5a008d43f5a5900fa05da9911714

Observation 96807f12-72d1-4b10-ae32-2c537c85d637 · outbound

This paper cites Bottom-up and top-down attention for image captioningandvisualquestionanswering,in:CVPR,pp.6077–6086.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Bottom-up and top-down attention for image captioningandvisualquestionanswering,in:CVPR,pp.6077–6086

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T05:12:12.027146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:12:12.027146Z digest=sha256:8f01a75c1bcd4cae369096a6edac8e79001d3ec1b781214e0d32dd090f1e9ba6

Observation eacef57d-aabb-412c-a4c9-80b3bdc53d23 · outbound

This paper cites Vision- and-language navigation: Interpreting visually-grounded navigation instructions in real environments, in: CVPR, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Vision- and-language navigation: Interpreting visually-grounded navigation instructions in real environments, in: CVPR, pp

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T05:12:12.032305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:12:12.032305Z digest=sha256:6f374baf56fcb34456773126ef95d21752f11b841bd8a3fa0de97fa08c8e3149

Observation a2ccec0e-972c-4408-81c0-c0f9a75152c0 · outbound

This paper cites NLTK: the natural language toolkit, in: Calzolari, N., Cardie, C., Isabelle, P.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations NLTK: the natural language toolkit, in: Calzolari, N., Cardie, C., Isabelle, P

Reference 4

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T05:12:14.075459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:12:12.036976Z digest=sha256:a05821099314ba71c30d05173f1d37d41d55654d7d4c9e666acff0da729c30eb

Observation 34cf6f45-2840-4dc4-8083-16501565f751 · outbound

This paper cites Matterport3d: LearningfromRGB-Ddatainindoorenvironments,in:3DV,pp.667–.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Matterport3d: LearningfromRGB-Ddatainindoorenvironments,in:3DV,pp.667–

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:12:14.654668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:12:12.041272Z digest=sha256:ef9559d35499e4c87caae2714af68d0fca8e8fac387665be079d4edcd35f0da6

Observation 9ac936c7-125b-458d-b108-78302e980c76 · outbound

This paper cites TOUCH- DOWN: natural language navigation and spatial reasoning in visual street environments, in: CVPR, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations TOUCH- DOWN: natural language navigation and spatial reasoning in visual street environments, in: CVPR, pp

Reference 6

Resolution
malformed identifier
no resolver link, observed 2026-08-07T05:12:12.050547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:12:12.050547Z digest=sha256:9988ebb1abdb7342fd2dcfcafc057d4e7ada387c6ddaf5c3c592630da364a74c

Observation 4c521f32-b680-4250-a58c-cdfd0af18af6 · outbound

This paper cites Historyawaremul- timodaltransformerforvision-and-languagenavigation,in:NeurIPS, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Historyawaremul- timodaltransformerforvision-and-languagenavigation,in:NeurIPS, pp

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:12:14.641035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:12:12.055232Z digest=sha256:1cf8f5467ae5970bc8a6491f82fffc405662141b2c6db2f1e8619349c4e4abc8

Observation 72778319-bebe-4dcc-9303-e3e108bda61f · outbound

This paper cites Learning from unlabeled 3d environments for vision-and-language navigation,in:ECCV,pp.638–655.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Learning from unlabeled 3d environments for vision-and-language navigation,in:ECCV,pp.638–655

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:12:12.059807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:12:12.059807Z digest=sha256:4cd5d8e0c708dda5b5995f3d5ccab4e06e8545ce1606d975cb23c686ea3b5e2f

Observation 4800851c-3b88-4622-9bab-b5fd9a8a98b2 · outbound

This paper cites Unifying vision-and- language tasks via text generation, in: ICML, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Unifying vision-and- language tasks via text generation, in: ICML, pp

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:12:14.627226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:12:12.068175Z digest=sha256:bbdb170a0f7922519a096f346340b0b6b1b920efe6da4298a7594c16c284b946

Observation c0100153-0501-4fb7-a125-7fce788e1b16 · outbound

This paper cites Auxiliary fine-grained alignment constraints for vision-and-language navigation, in: ICME, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Auxiliary fine-grained alignment constraints for vision-and-language navigation, in: ICME, pp

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T05:12:12.072754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:12:12.072754Z digest=sha256:d124911e863c01f0e734439e09a4707216bce8b3608f0d4c279389f57789d8c5

Observation 88915f0b-fdae-43db-8e09-6efb5b734a8c · outbound

This paper cites Meteor universal: Language specifictranslationevaluationforanytargetlanguage,in:Proceedings of the Ninth Workshop on Statistical Machine Translation, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Meteor universal: Language specifictranslationevaluationforanytargetlanguage,in:Proceedings of the Ninth Workshop on Statistical Machine Translation, pp

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:12:14.613994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:12:12.081842Z digest=sha256:eb8b9b867d4d3f497cbd3b37d24a13ea962c1c0fbc5bd970e3e9235fd3857315

Observation e6f20fcc-cd24-44e8-b6f8-e65156778454 · outbound

This paper cites Visionformobilerobotnavigation: A survey.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Visionformobilerobotnavigation: A survey

Reference 14

Resolution
verified exact
doi, observed 2026-08-07T05:12:12.517455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:12:12.090129Z digest=sha256:c1f824924abf3cb146eb3b79c8e4733f9d9c55e7cef4e1232621956233b8f066

Observation ea206fc4-b2f1-43d0-9691-7a816cbfdb65 · outbound

This paper cites Aerial vision-and-dialog navigation, in: Findings of the Association forComputationalLinguistics,pp.3043–3061.doi:10.18653/V1/2023.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Aerial vision-and-dialog navigation, in: Findings of the Association forComputationalLinguistics,pp.3043–3061.doi:10.18653/V1/2023

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:12:12.094146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:12:12.094146Z digest=sha256:80a8e247a3cd0f22e978e05efbda91e5ae46dcf569564ddaaa76b736786eb8e5

Observation 8da7b90f-f272-4703-9169-ab8ce3500280 · outbound

This paper cites an unresolved cited work.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Unresolved cited work

Reference 16

Resolution
verified exact
doi, observed 2026-08-07T05:12:12.495451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:12:12.098369Z digest=sha256:e48cf71ea7f49fed55b191074b5b5434c7e77dad7a93918f6b193d00a85b02e3

Observation 3689e419-a13b-4ab2-b2af-a6f0384b1ffc · outbound

This paper cites Speaker-follower models for vision-and-language navigation, in: NeurIPS, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Speaker-follower models for vision-and-language navigation, in: NeurIPS, pp

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:12:14.599570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:12:12.102791Z digest=sha256:498638809c99d4af180a5fd0e94d2be1007c5b06bf45489d799952ed78f35d80

Observation 025a00c6-43d4-4456-aa90-5ecf5772357c · outbound

This paper cites Vision- and-language navigation: A survey of tasks, methods, and future directions, in: ACL, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Vision- and-language navigation: A survey of tasks, methods, and future directions, in: ACL, pp

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:12:12.111576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:12:12.111576Z digest=sha256:1933286b5bf938778dfe5ffa5926d83152c0f23326826d0114d452dd2676b8d2

Observation 752621a0-ee27-464f-9529-1c43697b88c4 · outbound

This paper cites Landmark-rxr: Solving vision-and-language naviga- tion with fine-grained alignment supervision, in: NeurIPS, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Landmark-rxr: Solving vision-and-language naviga- tion with fine-grained alignment supervision, in: NeurIPS, pp

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:12:14.586568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:12:12.124012Z digest=sha256:1712d72a9d3eba36bc12dc3405526f3ce843593d6a0a871eb92108640191372b

Observation 8d989065-f2ba-40ba-97dd-c59da456644d · outbound

This paper cites Lan- guage and visual entity relationship graph for agent navigation, in: NeurIPS.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Lan- guage and visual entity relationship graph for agent navigation, in: NeurIPS

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:12:14.572587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:12:12.128097Z digest=sha256:b7af6328980f385c587ae63990a0f25cb0b10a79c64c9a92e4013aab72fec8d8

Observation 966cfa01-4b5d-4b6f-98cb-97e5af460351 · outbound

This paper cites Sub-instruction aware vision-and-language navigation, in: EMNLP, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Sub-instruction aware vision-and-language navigation, in: EMNLP, pp

Reference 24

Resolution
verified exact
doi, observed 2026-08-07T05:12:12.473819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:12:12.132307Z digest=sha256:40d3c6f17cbfa194b0b2e5646839a1d1bc80adf6d27a08561bbc3394d1460c24

Observation 5ea3fefe-2247-452c-847e-f97c09506236 · outbound

This paper cites VLNBERT: Arecurrentvision-and-languageBERTfornavigation,in:CVPR,pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations VLNBERT: Arecurrentvision-and-languageBERTfornavigation,in:CVPR,pp

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T05:12:12.136388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:12:12.136388Z digest=sha256:6b5ce6379c763e937601e8174c377f94e1c0b4db7e042359fba29157938faf6c

Observation 76cec562-f291-462e-a3fd-11258133fb05 · outbound

This paper cites Transferable representation learning in vision-and- language navigation, in: ICCV, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Transferable representation learning in vision-and- language navigation, in: ICCV, pp

Reference 26

Resolution
malformed identifier
no resolver link, observed 2026-08-07T05:12:12.140285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:12:12.140285Z digest=sha256:7179c17453de1f846df760ca5f565ea028a7008d14a2bbbeef1f31f9e248ecfb

Observation d90a169e-7aca-4fc8-a8c3-7a16064af85c · outbound

This paper cites an unresolved cited work.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Unresolved cited work

Reference 28

Resolution
malformed identifier
no resolver link, observed 2026-08-07T05:12:12.149291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:12:12.149291Z digest=sha256:fec911b5297a64e5945d7d67275f007ce4b7361ba064527b8f313072172835c4

Observation aaae132a-7902-4e99-9330-79976e2d8237 · outbound

This paper cites Simple and effective synthesisofindoor3dscenes,in:AAAI,pp.1169–1178.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Simple and effective synthesisofindoor3dscenes,in:AAAI,pp.1169–1178

Reference 29

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T05:12:14.559165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:12:12.153917Z digest=sha256:da34f563941963276c1e108b3f9373f59a8dc9ad7bc8dcc0dabd5784dc57a8bf

Observation 6057cf87-c913-48cb-8dda-9a4383e98b0a · outbound

This paper cites Waypoint models for instruction-guided navigation in continuous environments, in: ICCV, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Waypoint models for instruction-guided navigation in continuous environments, in: ICCV, pp

Reference 31

Resolution
verified exact
raw_fallback, observed 2026-08-07T05:12:13.262137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:12:12.161951Z digest=sha256:08b58417f9de6b2524ec1b1a95740d0545cccbc5200c90f17cf41d9b93cf3729

Observation f1893db0-14cf-41a9-a385-1fb30c8d9fcc · outbound

This paper cites Room- across-room:Multilingualvision-and-languagenavigationwithdense spatiotemporalgrounding,in:EMNLP,pp.4392–4412.doi:10.18653/ v1/2020.emnlp-main.356.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Room- across-room:Multilingualvision-and-languagenavigationwithdense spatiotemporalgrounding,in:EMNLP,pp.4392–4412.doi:10.18653/ v1/2020.emnlp-main.356

Reference 32

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T05:12:14.545422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:12:12.167048Z digest=sha256:f4b65d59667c342d07d059502ed33f34a716d5c3c9e82273c3258c883c21bb8e

Observation 7cbc4aa2-365c-4f2b-842e-70b1e23089ef · outbound

This paper cites Unicoder- vl: A universal encoder for vision and language by cross-modal pre- training, in: AAAI, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Unicoder- vl: A universal encoder for vision and language by cross-modal pre- training, in: AAAI, pp

Reference 33

Resolution
verified exact
doi, observed 2026-08-07T05:12:12.451747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:12:12.171799Z digest=sha256:65d892cee505639aca34fce374f95da8693ce64205ba71ac81b2029a159ec6eb

Observation 8be528bc-5df8-494e-98d0-835d505db6f9 · outbound

This paper cites Panogen: Text-conditioned panoramic envi- ronment generation for vision-and-language navigation, in: NeurIPS.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Panogen: Text-conditioned panoramic envi- ronment generation for vision-and-language navigation, in: NeurIPS

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:12:14.532183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:12:12.175858Z digest=sha256:c9acbdfedecbdf0f2dcc8004273ed36dc57b343910007f33fc02292493d5eb53

Observation dac5984a-7651-4248-a27b-e5b54a09d0e3 · outbound

This paper cites KERM: knowledge enhanced reasoning for vision-and-language navigation, in: CVPR, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations KERM: knowledge enhanced reasoning for vision-and-language navigation, in: CVPR, pp

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T05:12:12.184569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:12:12.184569Z digest=sha256:f3fe171d0a33465db2c4be7cf0f9733801b24b7bd41af2c41db52ac4c9a2508f

Observation 214a6609-ca0c-460b-8f6d-37c8d98f1a0e · outbound

This paper cites ROUGE: A package for automatic evaluation of summaries,in:TextSummarizationBranchesOut,Barcelona,Spain.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations ROUGE: A package for automatic evaluation of summaries,in:TextSummarizationBranchesOut,Barcelona,Spain

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:12:14.518389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:12:12.188385Z digest=sha256:ae9407ff28d2a2fb5867fab7972ee16491b8d9cf5a1fc6589d409a57fe51cfa8

Observation f485b723-0a93-459b-837d-0e5adc672e0f · outbound

This paper cites Learning vision-and-language navigation from youtube videos, in: CVPR, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Learning vision-and-language navigation from youtube videos, in: CVPR, pp

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T05:12:12.192628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:12:12.192628Z digest=sha256:b14f91da83f1bae07f764a56de03143703a4737edb376fa937eea790f184d99f

Observation 8794a4ed-015e-4f2b-a699-0e8aee49458f · outbound

This paper cites Self-monitoring navigation agent via auxiliary progress estimation, in: ICLR.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Self-monitoring navigation agent via auxiliary progress estimation, in: ICLR

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:12:14.505709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:12:12.200454Z digest=sha256:5591eb10d6c7e8155ce3b0d484d7457b0ea7fea1a0a276e030c9331cf67b76e5

Observation 65b1f3e3-261b-4320-9e2d-0ed0cfd45906 · outbound

This paper cites Theregretful agent: Heuristic-aided navigation through progress estimation, in: CVPR, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Theregretful agent: Heuristic-aided navigation through progress estimation, in: CVPR, pp

Reference 41

Resolution
verified exact
raw_fallback, observed 2026-08-07T05:12:13.181081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:12:12.204524Z digest=sha256:654000bd7df398e6c42845ecf14c097bbba7f1d87c04aec70d675586d1da1893

Observation 0187b010-9da3-4203-82d9-cbc16ca6e155 · outbound

This paper cites Improving vision-and-language navigation with image-text pairs from the web, in: ECCV, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Improving vision-and-language navigation with image-text pairs from the web, in: ECCV, pp

Reference 42

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T05:12:14.492370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:12:12.208375Z digest=sha256:f753c43ef53c0ef47b4c3c3dbb1db20484bb9384ae639676a4afa2bbac52f7d2

Observation d50d4463-a269-43cd-9579-b13975f35f18 · outbound

This paper cites SOAT: A scene- and object-aware transformer for vision-and-language navigation, in: NeurIPS, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations SOAT: A scene- and object-aware transformer for vision-and-language navigation, in: NeurIPS, pp

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:12:14.478509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:12:12.212340Z digest=sha256:1b97e29f84c5ec666f20f628b1a6cc249a769e8461e46bda65548372a646f3b2

Observation 9396e758-f39a-412f-bd9a-ccae3a5eaafa · outbound

This paper cites Bridging the visual semantic gapinVLNviasemanticallyricherinstructions,in:ECCV,pp.54–69.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Bridging the visual semantic gapinVLNviasemanticallyricherinstructions,in:ECCV,pp.54–69

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T05:12:12.220240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:12:12.220240Z digest=sha256:b30c971dac4264bfdff7a2a47563759854b2e46923d9d94eb35f408046f840b6

Observation b2052fa8-2f1e-4b64-8f95-d40c067ac602 · outbound

This paper cites Teach: Task-driven embodied agents that chat, in: AAAI, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Teach: Task-driven embodied agents that chat, in: AAAI, pp

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:12:14.450898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:12:12.224781Z digest=sha256:906d36a8343f56beb0875834b01ab5346467b77eaaa5b9d3ae0457ab0b2cb44e

Observation 88509fdc-81c0-4ebc-bac5-551fa243ac51 · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation, in: Proceedings of the 40th Annual Meeting of the Association for Computational Lin- guistics, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Bleu: a method for automatic evaluation of machine translation, in: Proceedings of the 40th Annual Meeting of the Association for Computational Lin- guistics, pp

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T05:12:12.235094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:12:12.235094Z digest=sha256:c9f1b43897c78d624be36feb7e92898bbecc540122577083a91c1df90c6ae4e4

Observation 661964f5-77b7-4b2e-9318-550233c04d4a · outbound

This paper cites The road to know-where: An object-and-room informed sequential BERT for indoor vision-language navigation, in: ICCV, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations The road to know-where: An object-and-room informed sequential BERT for indoor vision-language navigation, in: ICCV, pp

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:12:14.436917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:12:12.239091Z digest=sha256:5ee83eabeb2af7df7997f2cdb7c8ab72e93d2a32835254f9df745f63b53deb78

Observation 23b9ec66-062c-4738-82ce-742964fb1c6c · outbound

This paper cites Object- and-actionawaremodelforvisuallanguagenavigation,in:ECCV,pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Object- and-actionawaremodelforvisuallanguagenavigation,in:ECCV,pp

Reference 48

Resolution
verified exact
doi, observed 2026-08-07T05:12:12.419946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:12:12.247499Z digest=sha256:0ab9f6377bab71bd34b2fb0dd186d9456fee0d9b6bc59e31399637c94f5b3631

Observation 2c28b650-edb2-45fc-bfd5-f82a69f624fe · outbound

This paper cites HOP+: history-enhanced and order-aware pre-training for vision- and-language navigation.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations HOP+: history-enhanced and order-aware pre-training for vision- and-language navigation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T05:12:12.256573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:12:12.256573Z digest=sha256:b3b3a36f6ae1771b153a3d55060a6a2cb7c7147b9659a9b4ece83d150d21e70d

Observation 615a9ec6-fb64-4bf0-b055-d42276909213 · outbound

This paper cites Learningtransferablevisualmodelsfromnatural language supervision, in: ICML, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Learningtransferablevisualmodelsfromnatural language supervision, in: ICML, pp

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:12:14.422608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:12:12.260608Z digest=sha256:b7e9dd9e4db3bb3561209de34fb8306903548ec4aaef0004a64d2c12849f3095

Observation 0fea42e3-ee92-4f61-9f2a-ff3cf639a98c · outbound

This paper cites Towards long-horizon vision-language navigation: Platform, benchmark and method, in: CVPR.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Towards long-horizon vision-language navigation: Platform, benchmark and method, in: CVPR

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:12:14.407526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:12:12.268752Z digest=sha256:c03b3ffefe89911326d0cbd2b6fecf41a6592217d43a869d0789a12f961034fe

Observation bb77a352-9934-4565-ae53-20ebda4ace23 · outbound

This paper cites Contrastive search is what you need for neuraltextgeneration.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Contrastive search is what you need for neuraltextgeneration

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:12:14.393190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:12:12.272856Z digest=sha256:a05377f5989ef83b8f949073c180a1d95427b54ce8ed18dcd0d91b82d8108f55

Observation aecb1bbf-1e92-4d93-933c-7851a130c470 · outbound

This paper cites Learning to navigate unseen envi- ronments:Backtranslationwithenvironmentaldropout,in:NAACL- HLT, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Learning to navigate unseen envi- ronments:Backtranslationwithenvironmentaldropout,in:NAACL- HLT, pp

Reference 55

Resolution
verified exact
doi, observed 2026-08-07T05:12:12.405485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:12:12.277578Z digest=sha256:7fce0af8d2246ce4b3797ef720e9d705305104d7ee23ec34a49b68b3d79857a8

Observation 813c2439-0122-4ab6-935c-edc2581e9a80 · outbound

This paper cites Vision-and-dialog navigation, in: CoRL, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Vision-and-dialog navigation, in: CoRL, pp

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:12:14.378885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:12:12.281829Z digest=sha256:62e0845c43fa869358d8fb2046a11133a7f992acd9297144ba141cbabeedd644

Observation fa53d099-7623-4a10-9e7a-de62d4e3d867 · outbound

This paper cites Cider: Consensus- based image description evaluation, in: CVPR, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Cider: Consensus- based image description evaluation, in: CVPR, pp

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T05:12:12.285974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:12:12.285974Z digest=sha256:c51233331ce7a08e143842d09d3c9bd430a7f8f2a4caf7fe92adccef0ec0f86f

Observation 04a3101a-bf63-467a-b49b-d5e2dec4bf35 · outbound

This paper cites Show and tell: A neural image caption generator, in: CVPR, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Show and tell: A neural image caption generator, in: CVPR, pp

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T05:12:12.290006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:12:12.290006Z digest=sha256:6967248bc2a28b25d77318ed5b306dabc674d9a7f9793cc9bd1112da73917596

Observation 3ab685b3-0fb2-443e-b9da-57a4ddd0134d · outbound

This paper cites Soft expert reward learning for vision-and-language navigation, in: ECCV, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Soft expert reward learning for vision-and-language navigation, in: ECCV, pp

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:12:14.363551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:12:12.298071Z digest=sha256:45112633edb93dcab39f89242670b08ccac8c4b8a0b1bc860aca79bcb6782288

Observation 225f25e8-056b-4056-9d25-f6a7a05c2bf7 · outbound

This paper cites GIT: A generative image-to-text transformer for vision and language.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations GIT: A generative image-to-text transformer for vision and language

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:12:14.347964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:12:12.303260Z digest=sha256:e1a65ef32cb451f71e8cd2feeb763c142484376eb627b2cdfb5d1cab9622e8b6

Observation 341561be-031f-4fea-af60-01fe927b83a4 · outbound

This paper cites an unresolved cited work.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:12:14.320063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:12:12.311490Z digest=sha256:0dd15cdfe2c9d7de02558de1c347041baf5ba0ba95860bf1730e4e471b02041b

Observation 63c57a3a-3ba8-4d14-8b8f-16d07ccc25f9 · outbound

This paper cites OFA: unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning frame- work, in: ICML, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations OFA: unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning frame- work, in: ICML, pp

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:12:14.306372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:12:12.319606Z digest=sha256:3342bef1e72a821eecdaae6824758d175a489277f5116a3e0a89e63ddcb9b3e4

Observation 0609499b-c5eb-4540-b01f-9e7fb8b15770 · outbound

This paper cites Less is more: Generating grounded navigation instructions from landmarks, in: CVPR, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Less is more: Generating grounded navigation instructions from landmarks, in: CVPR, pp

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T05:12:12.323819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:12:12.323819Z digest=sha256:fa666438879d1f439bf936479b501e58953a3f6154dcca938fe80931e18d84c0

Observation 9edd9f04-e220-4576-822d-f24e52c0f593 · outbound

This paper cites Reinforced cross-modal matching and self-supervised imitation learning for vision-language navigation, in: CVPR, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Reinforced cross-modal matching and self-supervised imitation learning for vision-language navigation, in: CVPR, pp

Reference 65

Resolution
verified exact
raw_fallback, observed 2026-08-07T05:12:12.756522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:12:12.327661Z digest=sha256:3150347fc8ecc3295be61320fc6e97cd433109fab4e5618dabbe50a16ee87878

Observation 34bea04b-a9f1-4154-b860-46b8db0f2882 · outbound

This paper cites LANA: A language- capablenavigatorforinstructionfollowingandgeneration,in:CVPR, IEEE.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations LANA: A language- capablenavigatorforinstructionfollowingandgeneration,in:CVPR, IEEE

Reference 66

Resolution
verified exact
raw_fallback, observed 2026-08-07T05:12:12.677546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:12:12.331590Z digest=sha256:881f1534289dd7e01de1aca466305ebed86ccaa360faad244dd14503976a7a54

Observation 3ac4d94b-438e-41ff-acfd-bbfae8db2b06 · outbound

This paper cites Show, attend and tell: Neural image :Preprint submitted to Elsevier Page 14 of 15 caption generation with visual attention, in: ICML, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Show, attend and tell: Neural image :Preprint submitted to Elsevier Page 14 of 15 caption generation with visual attention, in: ICML, pp

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:12:14.292112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:12:12.336239Z digest=sha256:bf36b87c3ba5bf3cb61922e76a86658f50c7dde04f34d19a92f72f1d82205409

Observation d4d4cf0d-a8ac-4376-9cf1-1a02599e6059 · outbound

This paper cites Unified vision-language pre-training for image captioning and VQA, in: AAAI, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Unified vision-language pre-training for image captioning and VQA, in: AAAI, pp

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T05:12:12.340247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:12:12.340247Z digest=sha256:487cd23005bdb6d1081430b7360997a43a79a84b5b8fc8b79bfe4c36a6c03a3a

Observation 6babf70e-4066-4e21-8f4c-c977d6cd47cd · outbound

This paper cites Vision-language navigation with self-supervised auxiliary reasoning tasks, in: CVPR, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Vision-language navigation with self-supervised auxiliary reasoning tasks, in: CVPR, pp

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T05:12:12.345285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:12:12.345285Z digest=sha256:6f3bce8470af5cf25c96f181224bcb1b48b3e27c937a24b01c3712ba7925d8f7

Observation ca6b9874-9a29-4ed5-86d9-409884f01816 · outbound

This paper cites Babywalk:Goingfartherinvision-and-languagenavigationbytaking babysteps,in:ACL,pp.2539–2556.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Babywalk:Goingfartherinvision-and-languagenavigationbytaking babysteps,in:ACL,pp.2539–2556

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T05:12:12.349149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:12:12.349149Z digest=sha256:859e3248c236ef00e02e3b829a7a8ca45a36f4d8a7416e5dc8b1511ae92079e7

Observation 026645a0-97a9-4a3b-9725-f08d96dcf612 · outbound

This paper cites an unresolved cited work.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Unresolved cited work

Reference 380

Resolution
unresolved
no resolver link, observed 2026-08-07T05:12:12.085959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:12:12.085959Z digest=sha256:15b0294ca3188f6a84891e3caa594a66150c4905641fba3c8e8a936f03a155d5

Observation b910639c-c43c-4e4b-b389-5cbba8e5a53d · outbound

This paper cites an unresolved cited work.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Unresolved cited work

Reference 676

Resolution
unresolved
no resolver link, observed 2026-08-07T05:12:12.045804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:12:12.045804Z digest=sha256:1fbe5c4e291ab540f70c968104ba464bec06b3c23ca51b159b24eb9d4f2c2a1f

Observation cf621fab-ba65-4bc6-a139-d8eb16438d01 · outbound

This paper cites an unresolved cited work.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Unresolved cited work

Reference 1644

Resolution
unresolved
no resolver link, observed 2026-08-07T05:12:12.243135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:12:12.243135Z digest=sha256:ae87b1bd5f82aeda0ba2ab9e0f5dfca5d5e23c1c8071a3cbdc6dfd8b9e6c1eac

Observation 538109e6-4675-4359-a24c-0044a6b7a330 · outbound

This paper cites an unresolved cited work.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Unresolved cited work

Reference 2022

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:12:14.333845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:12:12.307332Z digest=sha256:aa4d792e1f42ff85214def12d7b1db81f8765409a8d09feab26b388892f37607

Observation 160a8a3a-27e1-40c4-bc54-ab34e3c7420b · outbound

This paper cites an unresolved cited work.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Unresolved cited work

Reference 2024

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T05:12:12.830343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:12:12.315636Z digest=sha256:225f7626cfffb0961d458807f751309128b814278563772128a5b3462207ca6d

Observation 56374228-da84-4be0-885b-62f890fe8cef · outbound

This paper cites an unresolved cited work.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Unresolved cited work

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T05:12:12.229720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:12:12.229720Z digest=sha256:1a613ca531e165181ddf5b27a2071cd9351baa3275be8438f26031bf61c22702

Observation 8fd95b62-9716-4c42-8769-55257bce25fb · outbound

This paper cites an unresolved cited work.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Unresolved cited work

Reference 7367

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:12:14.465023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:12:12.216264Z digest=sha256:f006226bb081fda8cabc60e46c12c6bab77dcb7fedc68d7e1a10e34c2604a295

Pith citing papers

Observation 134c4ef6-5447-4a21-bc5c-649dbd55a063 · inbound

GoViG: Goal-Conditioned Visual Navigation Instruction Generation via Multimodal Reasoning cites this paper.

GoViG: Goal-Conditioned Visual Navigation Instruction Generation via Multimodal Reasoning Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-18T23:02:52.670999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T23:02:14.913189Z digest=sha256:289bbe11a23deb2bc54b39b234598bad5c15aef548bc3860b62b22fd2f9bb8bb

Observation bbd2b0e3-7b33-4aa2-bc60-009a975a2a50 · inbound

Goal-oriented Navigation Instruction Generation with Tour Video Priors cites this paper.

Goal-oriented Navigation Instruction Generation with Tour Video Priors Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-14T04:34:35.053699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:34:35.053699Z digest=sha256:783f63dfc6ecd86f4e9fbc96378cffcffea1a09168097d8c37268b0ba5d9a7aa