Pith. sign in

Paper Citation Record · LEDGER

MemVLN: Episodic and Procedural Memory for Vision-and-Language Navigation

As of 22 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 0 inbound Pith citation observations for arXiv:2607.23504.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.23504 v1

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-30T20:42:04.352249Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

48 of 48 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved47
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4249682e-8f12-4b2b-ae4f-320f69361e1c · outbound

This paper cites GPT-4 Technical Report.

MemVLN: Episodic and Procedural Memory for Vision-and-Language Navigation GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-30T20:42:02.541961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T20:42:02.541961Z digest=sha256:e791ee18e1752e048ba4aae22e39feeafa63f0d18e75f2ff56996a98af2dd713

Observation e67b5f55-f346-4833-8788-7def70da434e · outbound

This paper cites Flamingo: a visual language model for few-shot learning.Advances in neural information processing systems, 35: 23716–23736, 2022.

MemVLN: Episodic and Procedural Memory for Vision-and-Language Navigation Flamingo: a visual language model for few-shot learning.Advances in neural information processing systems, 35: 23716–23736, 2022

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-30T20:42:02.628557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T20:42:02.628557Z digest=sha256:204a3003d15ac6f5e152694223c456ed2554c82b889522c08dfddf9baa3e021b

Observation c85c0990-00a8-4e6d-b5d4-ad78ce59a3fe · outbound

This paper cites Etpnav: Evolving topological planning for vision-language navigation in continuous environments.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024.

MemVLN: Episodic and Procedural Memory for Vision-and-Language Navigation Etpnav: Evolving topological planning for vision-language navigation in continuous environments.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-30T20:42:02.675416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T20:42:02.675416Z digest=sha256:395a222ff1b86ce6826e64bd692282fc9038e420b7770d578cee7ca92b31cf14

Observation 166c0485-c004-449f-81b5-7ca6189cc13a · outbound

This paper cites Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments.

MemVLN: Episodic and Procedural Memory for Vision-and-Language Navigation Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-30T20:42:02.715588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T20:42:02.715588Z digest=sha256:e45dbaddf9ff74d338a8cf42faf49df491ea38efca4683a2fbf6096c2e84bbc0

Observation c924cc41-2f60-424d-9c01-ac2f39d972f8 · outbound

This paper cites Sim-to-real transfer for vision-and-language navigation.

MemVLN: Episodic and Procedural Memory for Vision-and-Language Navigation Sim-to-real transfer for vision-and-language navigation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-30T20:42:02.785592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T20:42:02.785592Z digest=sha256:bdc7ef226ceb8043768eda96f645b3e7e8649d3a4ed078039518d7850dfa1b48

Observation 17d349f9-9cfd-4e5f-a6e6-14f848601dc3 · outbound

This paper cites Qwen3-VL Technical Report.

MemVLN: Episodic and Procedural Memory for Vision-and-Language Navigation Qwen3-VL Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-30T20:42:02.854773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T20:42:02.854773Z digest=sha256:9ee5b6da2c33fce1dd19bd61c2890f9db2642b1dfdf9ba895ff9bf6b3158e672

Observation 338138f6-32f2-45ea-a9a7-c25effd74dff · outbound

This paper cites Topological planning with transformers for vision-and-language navigation.

MemVLN: Episodic and Procedural Memory for Vision-and-Language Navigation Topological planning with transformers for vision-and-language navigation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-30T20:42:02.903996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T20:42:02.903996Z digest=sha256:097b79a63f4d6b5c837e64603d1206b40ec25014333c0c48266b037e147883ae

Observation 16cca270-3de1-432f-8d88-495bcdfd5430 · outbound

This paper cites Weakly-supervised multi-granularity map learning for vision-and-language navigation.

MemVLN: Episodic and Procedural Memory for Vision-and-Language Navigation Weakly-supervised multi-granularity map learning for vision-and-language navigation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-30T20:42:02.981520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T20:42:02.981520Z digest=sha256:81e570f01964bd173d1dfea7a5fd4b5270f333a74b371df604aac31b2b3eec3a

Observation 51a0af7a-2ead-4729-b06b-9357ef1cda00 · outbound

This paper cites NaVILA: Legged Robot Vision-Language-Action Model for Navigation.

MemVLN: Episodic and Procedural Memory for Vision-and-Language Navigation NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-30T20:42:03.068111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T20:42:03.068111Z digest=sha256:cb7c98948645ffd094803ef5e6d1639db3b85f594202a0207b69701396e2d805

Observation c86f0159-f8d6-4547-8650-1b13c692fc31 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

MemVLN: Episodic and Procedural Memory for Vision-and-Language Navigation An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-30T20:42:03.137249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T20:42:03.137249Z digest=sha256:8a8b419d00088a71959b277adebedd62e4436b04248d81a6e783363a16e51459

Observation 94c847d0-1e6b-40a3-a78f-550293eed3ad · outbound

This paper cites Speaker- follower models for vision-and-language navigation.Advances in neural information processing systems, 31, 2018.

MemVLN: Episodic and Procedural Memory for Vision-and-Language Navigation Speaker- follower models for vision-and-language navigation.Advances in neural information processing systems, 31, 2018

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-30T20:42:03.197473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T20:42:03.197473Z digest=sha256:d0cd2322b1e9e91d094ac1e413ea76b7ac77d6131f7c212da7dc64df862dbf9f

Observation 4f368302-63c0-4c80-a8ba-595c46d7c586 · outbound

This paper cites Counterfactual vision-and-language navigation via adversarial path sampler.

MemVLN: Episodic and Procedural Memory for Vision-and-Language Navigation Counterfactual vision-and-language navigation via adversarial path sampler

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-30T20:42:03.241149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T20:42:03.241149Z digest=sha256:913d5fde5e3e5ead9b1461103ebb61978429745c2f4fa5a9972d708279b25fed

Observation e55bdd75-8713-4096-8559-266d6033c395 · outbound

This paper cites Cross-modal map learning for vision and language navigation.

MemVLN: Episodic and Procedural Memory for Vision-and-Language Navigation Cross-modal map learning for vision and language navigation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-30T20:42:03.299565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T20:42:03.299565Z digest=sha256:71fdcb813951d48c166dd16776ab75c5f3ab3b03ee3346da89826eb125d92761

Observation 270290de-cb87-4127-9f80-2746f625f5c3 · outbound

This paper cites Language and visual entity relationship graph for agent navigation.Advances in Neural Information Processing Systems, 33:7685–7696, 2020.

MemVLN: Episodic and Procedural Memory for Vision-and-Language Navigation Language and visual entity relationship graph for agent navigation.Advances in Neural Information Processing Systems, 33:7685–7696, 2020

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-30T20:42:03.349403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T20:42:03.349403Z digest=sha256:6a94e0419039fcd1e7b09c3d83ddcde9482a500585b19e5bdd882525a7caa536

Observation a6b9bfce-89b2-4c02-971b-9be105d8bd8f · outbound

This paper cites Bridging the gap between learning in discrete and continuous environments for vision-and-language navigation.

MemVLN: Episodic and Procedural Memory for Vision-and-Language Navigation Bridging the gap between learning in discrete and continuous environments for vision-and-language navigation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-30T20:42:03.414738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T20:42:03.414738Z digest=sha256:c722d4080d10866b5640f48aecadb9b485fbc15f5825e464b0f396aa9281b5b7

Observation f11714af-b5db-4d05-8e1d-824a81b8b085 · outbound

This paper cites Sim-2-sim transfer for vision-and-language navigation in con- tinuous environments.

MemVLN: Episodic and Procedural Memory for Vision-and-Language Navigation Sim-2-sim transfer for vision-and-language navigation in con- tinuous environments

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-30T20:42:03.492588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T20:42:03.492588Z digest=sha256:b8be40a0e1d9f6a97478a4b9f1683dbfbe7e7533e707f460a0f50b35f2016ced

Observation 3ebf72e2-17f2-4326-92d3-cde805daa004 · outbound

This paper cites Beyond the nav- graph: Vision-and-language navigation in continuous environments.

MemVLN: Episodic and Procedural Memory for Vision-and-Language Navigation Beyond the nav- graph: Vision-and-language navigation in continuous environments

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-30T20:42:03.547400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T20:42:03.547400Z digest=sha256:2bc0cf9e4ef2cf9236683014ccf0ed9448b6cc9ce7e66b54975daeb91936572e

Observation b69145a9-ff25-4ff7-a389-99f55bd69f92 · outbound

This paper cites Waypoint models for instruction-guided navigation in continuous environments.

MemVLN: Episodic and Procedural Memory for Vision-and-Language Navigation Waypoint models for instruction-guided navigation in continuous environments

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-30T20:42:03.624402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T20:42:03.624402Z digest=sha256:dfb38e77dbdcb0ece0e4d3800067c470dded40a290e1f521aa3f90ed5b07f52e

Observation 830de76f-17a1-4e7d-9bae-0b9a47022a18 · outbound

This paper cites Room-across- room: Multilingual vision-and-language navigation with dense spatiotemporal grounding.

MemVLN: Episodic and Procedural Memory for Vision-and-Language Navigation Room-across- room: Multilingual vision-and-language navigation with dense spatiotemporal grounding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-30T20:42:03.706075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T20:42:03.706075Z digest=sha256:b822a7b636f98d7ff8a01c9fe7de9c65e8c3a22d11ba6677271b9dd20af698ec

Observation 77161361-f183-4dea-b951-ef5468de47e1 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

MemVLN: Episodic and Procedural Memory for Vision-and-Language Navigation LLaVA-OneVision: Easy Visual Task Transfer

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-30T20:42:03.770114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T20:42:03.770114Z digest=sha256:13396a62f7004c60e3c197de15a69da542fafb877c43614b1f1bc448fd548cd2

Observation 05df4398-8db7-45a9-bdc6-e94aae7c0cc0 · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models.

MemVLN: Episodic and Procedural Memory for Vision-and-Language Navigation Llama-vid: An image is worth 2 tokens in large language models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-30T20:42:03.851911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T20:42:03.851911Z digest=sha256:6f85a293f522247aa0ccdbb8ad8bcb380cb5442d04edc1395146c51402c8121d

Observation a1967f81-244f-4340-a806-6019a8046a72 · outbound

This paper cites Vila: On pre-training for visual language models.

MemVLN: Episodic and Procedural Memory for Vision-and-Language Navigation Vila: On pre-training for visual language models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-30T20:42:03.945113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T20:42:03.945113Z digest=sha256:54bd59a657f8e315b1830469a5c170e8de5f09d1109e299a77a3295baffee3e4

Observation 6753bae8-a674-4081-8b2d-524570ad9582 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023.

MemVLN: Episodic and Procedural Memory for Vision-and-Language Navigation Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-30T20:42:03.969709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T20:42:03.969709Z digest=sha256:575323d5a440c5aead152011b6389e3215c683109edd8be5ee65a8fc80faff5e

Observation bf81f264-832a-4a0b-9f3e-efb10f36af84 · outbound

This paper cites InstructNav: Zero-shot System for Generic Instruction Navigation in Unexplored Environment.

MemVLN: Episodic and Procedural Memory for Vision-and-Language Navigation InstructNav: Zero-shot System for Generic Instruction Navigation in Unexplored Environment

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-30T20:42:04.019111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T20:42:04.019111Z digest=sha256:405178c4e6975de7b06bb25aa9ffcc372a5aec8377363e39b6a0724c970db370

Observation ae4c0ceb-7a80-412a-87d0-b5379d9e0139 · outbound

This paper cites Discuss before moving: Visual language navigation via multi-expert discussions.

MemVLN: Episodic and Procedural Memory for Vision-and-Language Navigation Discuss before moving: Visual language navigation via multi-expert discussions

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-30T20:42:04.124407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T20:42:04.124407Z digest=sha256:a89bc868afbf59c8bd4804195e316423100b3c8450fe350a89d969ce8a0e617a

Observation 80dd1ef7-0309-4364-90ea-9df1c482d955 · outbound

This paper cites Deepstack: Deeply stacking visual tokens is surprisingly simple and effective for lmms.

MemVLN: Episodic and Procedural Memory for Vision-and-Language Navigation Deepstack: Deeply stacking visual tokens is surprisingly simple and effective for lmms

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-30T20:42:04.151044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T20:42:04.151044Z digest=sha256:8ae44ee2ef135f298986dd4aabc6ac1ff06e67faf43b73636983bda60891ddc5

Observation 0d68cc44-7ddd-4136-8f03-60e9ada92ccc · outbound

This paper cites Reverie: Remote embodied visual referring expression in real indoor environments.

MemVLN: Episodic and Procedural Memory for Vision-and-Language Navigation Reverie: Remote embodied visual referring expression in real indoor environments

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-30T20:42:04.159862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T20:42:04.159862Z digest=sha256:d619ca585cd2974d10b86d40ece1dadc04d5343e12baca1578ae6e88e259422c

Observation dfa1ae66-9b8d-43fe-b109-05de296d5261 · outbound

This paper cites Learning transferable visual models from natural language supervision.

MemVLN: Episodic and Procedural Memory for Vision-and-Language Navigation Learning transferable visual models from natural language supervision

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-30T20:42:04.170592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T20:42:04.170592Z digest=sha256:4d56d180c9b8fbe1b99353aff25f6b4abbf8008797cec322185b9b7022f75d76

Observation 7d4bb9c3-83df-4640-be24-8a1c883918ee · outbound

This paper cites Habitat-Matterport 3D Dataset (HM3D): 1000 Large-scale 3D Environments for Embodied AI.

MemVLN: Episodic and Procedural Memory for Vision-and-Language Navigation Habitat-Matterport 3D Dataset (HM3D): 1000 Large-scale 3D Environments for Embodied AI

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-30T20:42:04.177491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T20:42:04.177491Z digest=sha256:e1689aef912152726ae34953bf547fa427cf32992552c88f99e52dc6af02ed3f

Observation 8f7755a6-a814-45d4-b3b8-f6871d6fb957 · outbound

This paper cites Language- aligned waypoint (law) supervision for vision-and-language navigation in continuous envi- ronments.

MemVLN: Episodic and Procedural Memory for Vision-and-Language Navigation Language- aligned waypoint (law) supervision for vision-and-language navigation in continuous envi- ronments

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-30T20:42:04.183795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T20:42:04.183795Z digest=sha256:fc2b63cbdd19cfb8fbad041904322f7235177b1130818861f78cc241d2ca2eb1

Observation 0d9ce42f-e30e-4f61-a6a1-c1de52d2c815 · outbound

This paper cites A reduction of imitation learning and structured prediction to no-regret online learning.

MemVLN: Episodic and Procedural Memory for Vision-and-Language Navigation A reduction of imitation learning and structured prediction to no-regret online learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-30T20:42:04.190707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T20:42:04.190707Z digest=sha256:1d00f3a7c2b2cd1de08f8d7631979c732d4a0c3f470f3aae94ca1f6fb8ed4c4e

Observation 13fd62c1-a987-4850-a83e-dbb1fbd4111b · outbound

This paper cites Habitat: A platform for embodied ai research.

MemVLN: Episodic and Procedural Memory for Vision-and-Language Navigation Habitat: A platform for embodied ai research

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-30T20:42:04.199713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T20:42:04.199713Z digest=sha256:3efb9c9c1beecfb69d0ebe7700d391c71d83778229c3fdc3500b59e41980aa09

Observation 84ea8eb2-66f5-4c3b-a52c-24283fcd70d0 · outbound

This paper cites Learning to navigate unseen environments: Back translation with environmental dropout.

MemVLN: Episodic and Procedural Memory for Vision-and-Language Navigation Learning to navigate unseen environments: Back translation with environmental dropout

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-30T20:42:04.210468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T20:42:04.210468Z digest=sha256:9e5b38c311eb4a4facc76b3e931799f92d987e0b2b6f3941c4bbb8a680a78aa9

Observation 3770f99d-138f-40ca-bceb-704b747b0121 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

MemVLN: Episodic and Procedural Memory for Vision-and-Language Navigation Gemini: A Family of Highly Capable Multimodal Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-30T20:42:04.226291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T20:42:04.226291Z digest=sha256:16fdae887b340b09ea875280ed1b5ffccd3b1c91f7a919f7435f66fbf04712a9

Observation 10b9fe1e-98db-47c8-bd49-2c27f023c4c8 · outbound

This paper cites Vision-and-dialog navigation.

MemVLN: Episodic and Procedural Memory for Vision-and-Language Navigation Vision-and-dialog navigation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-30T20:42:04.238540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T20:42:04.238540Z digest=sha256:337abdc9f33bd8f29afeec6e529bbfe27998cabf2e50467c24d31cca7a398bee

Observation 591f0d4e-cbb9-43e2-a3ed-f029809c1d62 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

MemVLN: Episodic and Procedural Memory for Vision-and-Language Navigation LLaMA: Open and Efficient Foundation Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-30T20:42:04.245014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T20:42:04.245014Z digest=sha256:21a8b04129f97bb3cb9ae149fed799dfded124a0e8c4d286fb00ba737fb23a2e

Observation 08aaa791-0068-497c-8215-4577669021dc · outbound

This paper cites Dreamwalker: Mental planning for continuous vision-language navigation.

MemVLN: Episodic and Procedural Memory for Vision-and-Language Navigation Dreamwalker: Mental planning for continuous vision-language navigation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-30T20:42:04.256233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T20:42:04.256233Z digest=sha256:f28e2a1218714fc215d75926b1e20789d3cb074abe42cbcee54c06bf6533f56b

Observation b07c4916-dd03-4954-9ff0-9a49313f3033 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

MemVLN: Episodic and Procedural Memory for Vision-and-Language Navigation Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-30T20:42:04.265665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T20:42:04.265665Z digest=sha256:35d3c59eba0b1791063df1f4ec0698aef55729c5cd75df280797be7d81e23c50

Observation b9ed50af-392c-4949-a1fb-6d13ddcc4585 · outbound

This paper cites Gridmm: Grid memory map for vision-and-language navigation.

MemVLN: Episodic and Procedural Memory for Vision-and-Language Navigation Gridmm: Grid memory map for vision-and-language navigation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-30T20:42:04.273325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T20:42:04.273325Z digest=sha256:8b0fabd40d1523bf882538614079de69eca99e8ae0393570ddc3b9d9376e6fb1

Observation 96636e53-13ff-47ef-b6f8-59243aed4443 · outbound

This paper cites Lookahead exploration with neural radiance representation for continuous vision-language navigation.

MemVLN: Episodic and Procedural Memory for Vision-and-Language Navigation Lookahead exploration with neural radiance representation for continuous vision-language navigation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-30T20:42:04.280485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T20:42:04.280485Z digest=sha256:276c5d44a6c7c9528360b3c0ee26bd9f92e7969a2cc715a657ba0a410583b03c

Observation 0ff8670e-d957-4cdc-b122-94b3404acb98 · outbound

This paper cites Sim-to-Real Transfer via 3D Feature Fields for Vision-and-Language Navigation.

MemVLN: Episodic and Procedural Memory for Vision-and-Language Navigation Sim-to-Real Transfer via 3D Feature Fields for Vision-and-Language Navigation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-30T20:42:04.289618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T20:42:04.289618Z digest=sha256:e796c5d47573cedcb539e29d7e308d05f435daa21b92e660e1764c9d9a518227

Observation 997c602b-1ff2-4595-b5b4-bfde57e93136 · outbound

This paper cites Scaling data generation in vision-and-language navigation.

MemVLN: Episodic and Procedural Memory for Vision-and-Language Navigation Scaling data generation in vision-and-language navigation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-30T20:42:04.300070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T20:42:04.300070Z digest=sha256:b1473eb4a25a833c6e939e17a0f8d4409067cbc3b6b06f438dd72b8df2e609d2

Observation 5bd1de4b-18ff-4341-9a0f-39851b05a53f · outbound

This paper cites StreamVLN: Streaming Vision-and-Language Navigation via SlowFast Context Modeling.

MemVLN: Episodic and Procedural Memory for Vision-and-Language Navigation StreamVLN: Streaming Vision-and-Language Navigation via SlowFast Context Modeling

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-30T20:42:04.312697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T20:42:04.312697Z digest=sha256:086cde6980504fd9ebafa60db2fbdc92023fe2d74819217400f83c0bee764dad

Observation d85f88cc-a354-4a16-871c-53d9a16153ca · outbound

This paper cites Gibson env: Real-world perception for embodied agents.

MemVLN: Episodic and Procedural Memory for Vision-and-Language Navigation Gibson env: Real-world perception for embodied agents

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-30T20:42:04.320554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T20:42:04.320554Z digest=sha256:7649a70c3c6199c38e84a9bdd28be54e60b9fc9e9357caf5ff2ac2c23301fbcc

Observation 62381cda-58df-4449-8cfe-7b594c6a817f · outbound

This paper cites Uni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks.

MemVLN: Episodic and Procedural Memory for Vision-and-Language Navigation Uni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-30T20:42:04.329176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T20:42:04.329176Z digest=sha256:01c6dcdbe412125fd921c811429a7486129c8a53be4f7de79b14f42e757daee8

Observation 9018df27-be93-47a9-a72f-0f7de330cdc5 · outbound

This paper cites NaVid: Video-based VLM Plans the Next Step for Vision-and-Language Navigation.

MemVLN: Episodic and Procedural Memory for Vision-and-Language Navigation NaVid: Video-based VLM Plans the Next Step for Vision-and-Language Navigation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-07-30T20:42:04.335446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T20:42:04.335446Z digest=sha256:c61c4465fb7a73d792d4b7c4e30a7a56cca9d0abffb344fc5891bb21721b48ab

Observation c87b155e-1dd5-4476-8cec-141e67f181fa · outbound

This paper cites Lyra: An efficient and speech-centric framework for omni-cognition.

MemVLN: Episodic and Procedural Memory for Vision-and-Language Navigation Lyra: An efficient and speech-centric framework for omni-cognition

Reference 47

Resolution
unresolved
no resolver link, observed 2026-07-30T20:42:04.343242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T20:42:04.343242Z digest=sha256:64af196661585b5de3f160283b94412d1ae688b64c3509d3b8c126d4cea10e5e

Observation bd217528-721c-44f4-9ff0-27980426c0c5 · outbound

This paper cites Navgpt: Explicit reasoning in vision-and-language navigation with large language models.

MemVLN: Episodic and Procedural Memory for Vision-and-Language Navigation Navgpt: Explicit reasoning in vision-and-language navigation with large language models

Reference 48

Resolution
malformed identifier
no resolver link, observed 2026-07-30T20:42:04.352249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T20:42:04.352249Z digest=sha256:6262308014ba435cc07889300df611437bc0205523077d8c12c89444d2eecb3c

Pith citing papers

No inbound Pith citation observations are available.