Pith. sign in

Paper Citation Record · LEDGER

HALO: Human Preference Aligned Offline Reward Learning for Robot Navigation

As of 8 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 4 inbound Pith citation observations for arXiv:2508.01539.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.01539 v1

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T05:41:42.393210Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T11:03:51.194382Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T21:38:59.525642Z

Reference resolution

50 of 50 outbound references displayed

  • verified exact3
  • verified fuzzy9
  • unresolved37
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 447b2ff4-7e7c-49d7-a291-9b33bddd2fdd · outbound

This paper cites Paluch, J.

HALO: Human Preference Aligned Offline Reward Learning for Robot Navigation Paluch, J

Reference 1

Resolution
verified exact
doi, observed 2026-08-06T05:41:42.445999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T05:41:41.898677Z digest=sha256:b1ec8e924883bd5259ef02b75f49e81815b5c5c735b3f68bbcede2836486e47f

Observation 2ec8c764-df64-4d11-9d56-359dd62ce709 · outbound

This paper cites an unresolved cited work.

HALO: Human Preference Aligned Offline Reward Learning for Robot Navigation Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-06T05:41:44.268560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T05:41:41.910258Z digest=sha256:63dc575442c3f1ce7cc93dcdc8199ac3fd4b3728e95a3de5e92b3c8b0979f968

Observation a11ebb28-33ba-4a3c-93d7-9451a6be7ba2 · outbound

This paper cites BehAV: Behavioral Rule Guided Autonomy Using VLMs for Robot Navigation in Outdoor Scenes.

HALO: Human Preference Aligned Offline Reward Learning for Robot Navigation BehAV: Behavioral Rule Guided Autonomy Using VLMs for Robot Navigation in Outdoor Scenes

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T05:41:41.922261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:41:41.922261Z digest=sha256:749f8cda18c860789553438535bdcc642ae46803c8acf159dbf3be1c4e0779e8

Observation c4176d68-2e40-4bfe-a753-303db470432b · outbound

This paper cites ViNT: A Foundation Model for Visual Navigation.

HALO: Human Preference Aligned Offline Reward Learning for Robot Navigation ViNT: A Foundation Model for Visual Navigation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T05:41:41.942248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:41:41.942248Z digest=sha256:5e1bb47b3a2da9e5865829982203ad0d0a376be3cd61820c5f4e799939dc7808

Observation 2c164fd9-6367-4b9d-b349-76e0ce34874b · outbound

This paper cites CROSS-GAiT: Cross-Attention-Based Multimodal Representation Fusion for Parametric Gait Adaptation in Complex Terrains.

HALO: Human Preference Aligned Offline Reward Learning for Robot Navigation CROSS-GAiT: Cross-Attention-Based Multimodal Representation Fusion for Parametric Gait Adaptation in Complex Terrains

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T05:41:41.952172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:41:41.952172Z digest=sha256:36d40625fb6ece857272f9251f6473a7c5318b83433a1ef8e58658e21dba9535

Observation 008c4d89-bd89-41cb-aa6b-87abc623945c · outbound

This paper cites Social-LLaVA: Enhancing Robot Navigation through Human-Language Reasoning in Social Spaces.

HALO: Human Preference Aligned Offline Reward Learning for Robot Navigation Social-LLaVA: Enhancing Robot Navigation through Human-Language Reasoning in Social Spaces

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T05:41:41.959671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:41:41.959671Z digest=sha256:de9ceb294f4ce73e74820092ceb8cfd491e7f158ff011976b52e11c74f2fb5d2

Observation 26bce7b1-d0f1-4233-9dc8-683c84c58121 · outbound

This paper cites Offline Reinforcement Learning for Visual Navigation.

HALO: Human Preference Aligned Offline Reward Learning for Robot Navigation Offline Reinforcement Learning for Visual Navigation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T05:41:41.966399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:41:41.966399Z digest=sha256:3280f4a88626788b79611a4b40d35e5be1d7e0cc40a19eb05e4bcaf86aadcdd2

Observation 211d81ce-997e-4b8c-81cd-243c1b8c3166 · outbound

This paper cites an unresolved cited work.

HALO: Human Preference Aligned Offline Reward Learning for Robot Navigation Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-06T05:41:44.241719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T05:41:41.975969Z digest=sha256:4934b6246cb7304828fbd1474b65009289f516e7519c06c1bc40f7b479e81195

Observation 84047cca-00d1-41af-b27e-2ba892b5c6f4 · outbound

This paper cites Improving Generalization in Reinforcement Learning Training Regimes for Social Robot Navigation.

HALO: Human Preference Aligned Offline Reward Learning for Robot Navigation Improving Generalization in Reinforcement Learning Training Regimes for Social Robot Navigation

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-06T05:41:43.138614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T05:41:41.985013Z digest=sha256:49109641117d5f64469e98ccb9f85f3a0a663e1300b81aedcb144e4830ad405a

Observation 5cd5aadd-c4bd-4ac0-988b-5d5c93c4c7e1 · outbound

This paper cites SoNIC: Safe Social Navigation with Adaptive Conformal Inference and Constrained Reinforcement Learning.

HALO: Human Preference Aligned Offline Reward Learning for Robot Navigation SoNIC: Safe Social Navigation with Adaptive Conformal Inference and Constrained Reinforcement Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T05:41:41.991540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:41:41.991540Z digest=sha256:52d91379189ef65ad4f8b7f6cc7a52ac2399cc2b537dc21b023fe0f2eedd408b

Observation ce30310a-c7da-46ee-824a-a33cf9d33c36 · outbound

This paper cites Jiang, P.

HALO: Human Preference Aligned Offline Reward Learning for Robot Navigation Jiang, P

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:41:44.216441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T05:41:41.998346Z digest=sha256:e358adc7a450e2f3007fe67c92c89e9c4fd13a09501d5d22f6dac81aa5c6d85e

Observation d580991e-99c6-491a-a5a2-e269b95bd9f0 · outbound

This paper cites Karnan, A.

HALO: Human Preference Aligned Offline Reward Learning for Robot Navigation Karnan, A

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:41:44.187969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T05:41:42.004766Z digest=sha256:a55baeae239d887b0a4d59c4ba433803d21117c1e5eabb74794dedf35196a8bb

Observation 7d6cab99-410d-44ea-bd41-66f6294efe59 · outbound

This paper cites VAPOR: Legged Robot Navigation in Outdoor Vegetation Using Offline Reinforcement Learning.

HALO: Human Preference Aligned Offline Reward Learning for Robot Navigation VAPOR: Legged Robot Navigation in Outdoor Vegetation Using Offline Reinforcement Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T05:41:42.012865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:41:42.012865Z digest=sha256:e6e02d7030950f52763574f4738a84ac1239b2e68541054a7a213827d48a5d52

Observation 769af3b3-c51b-4e79-9bf4-39931cc1d620 · outbound

This paper cites Caesar, V.

HALO: Human Preference Aligned Offline Reward Learning for Robot Navigation Caesar, V

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T05:41:42.023020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:41:42.023020Z digest=sha256:9d00d6e8a1abe6d0396ff1e222d2dc2c5fa65675981da4bb7fab268e9b8e20cf

Observation 6dcbe0c9-1d50-4df9-ac46-93983f6cb10e · outbound

This paper cites A2D2: Audi Autonomous Driving Dataset.

HALO: Human Preference Aligned Offline Reward Learning for Robot Navigation A2D2: Audi Autonomous Driving Dataset

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T05:41:42.035762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:41:42.035762Z digest=sha256:20510e66324b2d8a5743be652e86d29588d5c2db8053c4cca938b0c181efc558

Observation f62aec65-6ef1-4b97-a5af-745d3c83d26e · outbound

This paper cites Kumar, A.

HALO: Human Preference Aligned Offline Reward Learning for Robot Navigation Kumar, A

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T05:41:42.066004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:41:42.066004Z digest=sha256:caa2a53c669a5b108f6d14b183180ad39696607f92a205b2c500fc3b13596c04

Observation 514ea7c8-6cc3-4cf4-a0ea-3e4959d22580 · outbound

This paper cites Kapoor, S.

HALO: Human Preference Aligned Offline Reward Learning for Robot Navigation Kapoor, S

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:41:44.102127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T05:41:42.072754Z digest=sha256:ae645666463f57c2b8be1e54648f6c4fcf67385f00c31e610fd309e64488d5b0

Observation 1f417724-f7b0-4cad-82a6-103914ca6f3d · outbound

This paper cites Liang, U.

HALO: Human Preference Aligned Offline Reward Learning for Robot Navigation Liang, U

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:41:44.065291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T05:41:42.082464Z digest=sha256:33589d69df11454464791cf4cb4b6c7a13917b41efd4c8496d2fdcb1d0c44214

Observation 8aa917d6-18e0-4a3d-a73e-dfc105c476ac · outbound

This paper cites Patel, N.

HALO: Human Preference Aligned Offline Reward Learning for Robot Navigation Patel, N

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:41:44.018566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T05:41:42.088954Z digest=sha256:ba8bcec29646f9af0acc31bd3abf93936518b2ea2727c14e1d54650d2a3fa8d8

Observation a9029285-c2f9-4d4c-82fa-3c3c26941291 · outbound

This paper cites Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey.

HALO: Human Preference Aligned Offline Reward Learning for Robot Navigation Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T05:41:42.094226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:41:42.094226Z digest=sha256:58d851e8b98dc2b3da96c38c0c58f858ebc0e59b1d3207fad62e7cfc4159b48b

Observation db93ab8e-02c2-44f8-960c-e07eed811093 · outbound

This paper cites Towards Reasoning in Large Language Models: A Survey.

HALO: Human Preference Aligned Offline Reward Learning for Robot Navigation Towards Reasoning in Large Language Models: A Survey

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T05:41:42.108363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:41:42.108363Z digest=sha256:7dcc0a8a1c016f7fc2a4d0323a2e99455aac1cde7c48ef098a8c2eac19080962

Observation a5dbacfa-40e5-4177-9235-3f517a90ed75 · outbound

This paper cites LogicBench: Towards Systematic Evaluation of Logical Reasoning Ability of Large Language Models.

HALO: Human Preference Aligned Offline Reward Learning for Robot Navigation LogicBench: Towards Systematic Evaluation of Logical Reasoning Ability of Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T05:41:42.117693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:41:42.117693Z digest=sha256:30760a8f3906201341ec6c697b6d52cbcdaf2cb96dff04d9e622b2f989e8cf1c

Observation 5d3615b1-23d6-47be-8485-1f2fc8897293 · outbound

This paper cites an unresolved cited work.

HALO: Human Preference Aligned Offline Reward Learning for Robot Navigation Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-06T05:41:43.986811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T05:41:42.128414Z digest=sha256:1924230043438a4463d81afc1d12639e1884a56fca4c88498b7a1ed7d805df41

Observation 73adf8aa-a8c5-42b2-8ad9-bc7b26c7ca4f · outbound

This paper cites Huang, O.

HALO: Human Preference Aligned Offline Reward Learning for Robot Navigation Huang, O

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:41:43.960410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T05:41:42.137295Z digest=sha256:216834b91aad48729340ee80005da75c4b156df036a7b8af27f45a3583edd07c

Observation 972c6a70-f895-47ea-9862-77cb281e1147 · outbound

This paper cites an unresolved cited work.

HALO: Human Preference Aligned Offline Reward Learning for Robot Navigation Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-06T05:41:43.904495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T05:41:42.152092Z digest=sha256:52877b46edf4c29f49f8409c4bcc46b34e5f466b0f53d9f4ea87647a1c3d9249

Observation 11d7e473-5c36-4862-937d-a2013cca1700 · outbound

This paper cites Robot Navigation Using Physically Grounded Vision-Language Models in Outdoor Environments.

HALO: Human Preference Aligned Offline Reward Learning for Robot Navigation Robot Navigation Using Physically Grounded Vision-Language Models in Outdoor Environments

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T05:41:42.161527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:41:42.161527Z digest=sha256:59d0ca07884d49ed381b5c11a263dcefe991dc4fb048c335ad51375f4ff5dca5

Observation 2895e9fc-018e-4b83-9b35-57610a6036af · outbound

This paper cites an unresolved cited work.

HALO: Human Preference Aligned Offline Reward Learning for Robot Navigation Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-06T05:41:43.876693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T05:41:42.169468Z digest=sha256:a993202d4c977d7a8a412e82eb428fa0d6b4c7cee8ebf2134404f20c4e1ac98a

Observation cf4ee867-a2ce-44d8-81f7-cb2626279152 · outbound

This paper cites Resilient Timed Elastic Band Planner for Collision-Free Navigation in Unknown Environments.

HALO: Human Preference Aligned Offline Reward Learning for Robot Navigation Resilient Timed Elastic Band Planner for Collision-Free Navigation in Unknown Environments

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-08-06T05:41:42.818139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T05:41:42.180688Z digest=sha256:46f6cb133e4dd8ababe23619e6121afe30019d309c6220267daa70b8fd821a86

Observation 2e6bb728-0397-46b0-842a-74c74711489b · outbound

This paper cites Dosovitskiy, G.

HALO: Human Preference Aligned Offline Reward Learning for Robot Navigation Dosovitskiy, G

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T05:41:42.191176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:41:42.191176Z digest=sha256:5b2dd6f9bb5908a528666a18c3206876cc8f079c1d75aedd1becd6d0b03f3baf

Observation 7bd0c8b2-bb76-49eb-9889-9f587c83d6ab · outbound

This paper cites an unresolved cited work.

HALO: Human Preference Aligned Offline Reward Learning for Robot Navigation Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-06T05:41:43.819754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T05:41:42.197893Z digest=sha256:4c90940320ceaa249f26eaff0e04de6dd3a0a1a71bf8fc8ebaa02ecf1946d1fe

Observation c991c3d9-3363-4ab2-af98-b60d1ebf6e51 · outbound

This paper cites an unresolved cited work.

HALO: Human Preference Aligned Offline Reward Learning for Robot Navigation Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T05:41:42.203683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:41:42.203683Z digest=sha256:f07a1ea95a4c7168115c32036a207833720f0af698149a365f3ffe64189b5b72

Observation bf990ad4-6fe6-4e04-848f-c4f42f575879 · outbound

This paper cites PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-training.

HALO: Human Preference Aligned Offline Reward Learning for Robot Navigation PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-training

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T05:41:42.210585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:41:42.210585Z digest=sha256:ddfd1bc4ea8e75522cbece8f318b0abf4f5b5323720e462f28b6fb680eb17fd7

Observation 1debe993-d244-4a08-84cc-9bc6ccd864dc · outbound

This paper cites RL-VLM-F: Reinforcement Learning from Vision Language Foundation Model Feedback.

HALO: Human Preference Aligned Offline Reward Learning for Robot Navigation RL-VLM-F: Reinforcement Learning from Vision Language Foundation Model Feedback

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T05:41:42.219536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:41:42.219536Z digest=sha256:58b91e26ada219b55cfe5e03c4858341eca1fa27d62995180a34836e5b748a16

Observation f8be3d71-d44e-4dd6-a8b7-807fd669d6ff · outbound

This paper cites an unresolved cited work.

HALO: Human Preference Aligned Offline Reward Learning for Robot Navigation Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-06T05:41:43.733022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T05:41:42.244546Z digest=sha256:75df58b47a948b8638cb30536dd22ec5227959276559c12d26167bb90fbdae81

Observation a4ff4c6b-9d2e-4e44-9182-ff5299d37b8d · outbound

This paper cites Fine-Tuning Large Vision-Language Models as Decision-Making Agents via Reinforcement Learning.

HALO: Human Preference Aligned Offline Reward Learning for Robot Navigation Fine-Tuning Large Vision-Language Models as Decision-Making Agents via Reinforcement Learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T05:41:42.254081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:41:42.254081Z digest=sha256:7e60408b9e8a74c9939b08cc5ce89dd9621d6fb8ecf6040feb9ca1868910be1c

Observation 9187e42d-90ed-4e90-87c8-1db10a76ff20 · outbound

This paper cites NaVid: Video-based VLM Plans the Next Step for Vision-and-Language Navigation.

HALO: Human Preference Aligned Offline Reward Learning for Robot Navigation NaVid: Video-based VLM Plans the Next Step for Vision-and-Language Navigation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T05:41:42.263684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:41:42.263684Z digest=sha256:540ad7ef90f62dc241870e7b8f4f792f74e9ae736e95e7979504207f27a31d54

Observation 5b0c29b2-a42e-48d3-a0f6-0bab735b5f1c · outbound

This paper cites VL-TGS: Trajectory Generation and Selection using Vision Language Models in Mapless Outdoor Environments.

HALO: Human Preference Aligned Offline Reward Learning for Robot Navigation VL-TGS: Trajectory Generation and Selection using Vision Language Models in Mapless Outdoor Environments

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T05:41:42.275732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:41:42.275732Z digest=sha256:e0a5c8e07c3bcb362afd880d67c1ea5b806d6ba54cb6bf4d2106e813192d856b

Observation 889620ba-285b-4180-a4bd-b176174195c5 · outbound

This paper cites Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems.

HALO: Human Preference Aligned Offline Reward Learning for Robot Navigation Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T05:41:42.284447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:41:42.284447Z digest=sha256:b3dfd13bca40c11ed3fe95ef2e21111c7956c0d6162d622601197e323b25b657

Observation 7bc873a4-4dbe-4ee1-8c79-250333343146 · outbound

This paper cites an unresolved cited work.

HALO: Human Preference Aligned Offline Reward Learning for Robot Navigation Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-06T05:41:43.689975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T05:41:42.292391Z digest=sha256:9e4838ba2f16299f203a9790b62452dafc9e8d865d4d0457c0ffb5ba78087271

Observation 4639882f-adba-4a59-be20-18e10c678589 · outbound

This paper cites Oquab, T.

HALO: Human Preference Aligned Offline Reward Learning for Robot Navigation Oquab, T

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T05:41:42.299094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:41:42.299094Z digest=sha256:8d07acfdcd4b68214f2e35da44182bc47306998933630c6c039865ba2fa5f706

Observation 188869df-076a-41a1-97c4-638f36f7f3db · outbound

This paper cites Offline Reinforcement Learning with Implicit Q-Learning.

HALO: Human Preference Aligned Offline Reward Learning for Robot Navigation Offline Reinforcement Learning with Implicit Q-Learning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T05:41:42.308957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:41:42.308957Z digest=sha256:92cac15fc3dba6010a8005d6fc9e6dd92696b3157bc0c5ed2e096e9e99786e2d

Observation 49c3d291-48b3-475d-a818-76b8f1702a53 · outbound

This paper cites Fujimoto, H.

HALO: Human Preference Aligned Offline Reward Learning for Robot Navigation Fujimoto, H

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T05:41:42.315578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:41:42.315578Z digest=sha256:06e7b9032b56b46f63f486526a4a6a9e295752454d014ea0f0f71cb7bc929965

Observation 3d25c821-87d7-4992-9060-40219e793cba · outbound

This paper cites an unresolved cited work.

HALO: Human Preference Aligned Offline Reward Learning for Robot Navigation Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-06T05:41:43.605492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T05:41:42.326150Z digest=sha256:4fccbfa7f118d44725f348cab790b46ca574da0fb1d3643fb3d88ac3cdebfcc5

Observation bb7ede81-6f6f-4105-900d-e51358250c5b · outbound

This paper cites Nazeri, J.

HALO: Human Preference Aligned Offline Reward Learning for Robot Navigation Nazeri, J

Reference 45

Resolution
metadata mismatch
raw_fallback, observed 2026-08-06T05:41:42.589513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T05:41:42.334510Z digest=sha256:5bddd182cdc9450b41068c7a81841c5111f44b36ecbc0fa1375fe369d456bcd9

Observation 379f94f0-fa1b-4266-96ec-9aba6588e650 · outbound

This paper cites Alt and M.

HALO: Human Preference Aligned Offline Reward Learning for Robot Navigation Alt and M

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:41:43.576104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T05:41:42.341285Z digest=sha256:dd3a58bcd402c486fdd7f36427852e62068d8de892054f44cbb79b02029dc35e

Observation 84f69c58-bd53-4431-889f-5e35c3b68a4b · outbound

This paper cites Fujimoto and S.

HALO: Human Preference Aligned Offline Reward Learning for Robot Navigation Fujimoto and S

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:41:43.550165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T05:41:42.349958Z digest=sha256:854c8c189a4e224d1cec77bf87e215e1b9d8a71d978dc98a209313468282fde6

Observation b3622393-bca0-4325-b54e-97a502504c47 · outbound

This paper cites Tarasov, A.

HALO: Human Preference Aligned Offline Reward Learning for Robot Navigation Tarasov, A

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:41:43.514991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T05:41:42.387708Z digest=sha256:21a4b838b23ec73294f549fed3f3121b508889b9161e05c93521a7481d7258c8

Observation 7deac5a7-971a-4e7e-9fcb-c0f8a8508c9e · outbound

This paper cites an unresolved cited work.

HALO: Human Preference Aligned Offline Reward Learning for Robot Navigation Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-06T05:41:43.472786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T05:41:42.393210Z digest=sha256:b863f878d1c76a99c0ce5bc375aa20de87815e975da72c7d54f827ba849ca535

Observation 82fe68fc-0203-4da4-8ec9-cc343b14cfca · outbound

This paper cites an unresolved cited work.

HALO: Human Preference Aligned Offline Reward Learning for Robot Navigation Unresolved cited work

Reference 2022

Resolution
unresolved
raw_fallback, observed 2026-08-06T05:41:43.927628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T05:41:42.144265Z digest=sha256:3b7b37cfe0978a0b3b351e18fdbd7f2dc3d3ec7f1a4f738e4953acc56d31787d

Observation 6cda450d-7312-4cda-8761-650da2fc7941 · outbound

This paper cites an unresolved cited work.

HALO: Human Preference Aligned Offline Reward Learning for Robot Navigation Unresolved cited work

Reference 2024

Resolution
unresolved
raw_fallback, observed 2026-08-06T05:41:43.773742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T05:41:42.234906Z digest=sha256:6a17a1faf2d6ac154d322a140c9a0782e96df9c7fe3777d36a6b8bb8bd2b44af

Pith citing papers

Observation eee1caca-fdb9-4b91-92f1-fe5dd3906958 · inbound

Interpreting Context-Aware Human Preferences for Multi-Objective Robot Navigation cites this paper.

Interpreting Context-Aware Human Preferences for Multi-Objective Robot Navigation HALO: Human Preference Aligned Offline Reward Learning for Robot Navigation

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-15T09:25:22.284010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T09:20:24.694980Z digest=sha256:9ad22790f81c4a7270f46f4932675e8a1504d6f71d263272c63d8b5a22438aed

Observation c3309775-2f63-4448-8057-d8ba1c33babe · inbound

From Imitation to Alignment: Human-Preference Flow Policies for Long-Horizon Sidewalk Navigation cites this paper.

From Imitation to Alignment: Human-Preference Flow Policies for Long-Horizon Sidewalk Navigation HALO: Human Preference Aligned Offline Reward Learning for Robot Navigation

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-03T11:38:04.680186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T09:29:43.030058Z digest=sha256:42b482fc73714da04573570c174ebe18ee90224066f61901b30b50dcbc48a9d3

Observation 573bdfba-b5b0-44f2-ac81-57da7299b166 · inbound

VEGA: Learning Navigation VLAs from In-the-Wild Egocentric Video with Geometric Trajectory Supervision cites this paper.

VEGA: Learning Navigation VLAs from In-the-Wild Egocentric Video with Geometric Trajectory Supervision HALO: Human Preference Aligned Offline Reward Learning for Robot Navigation

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-03T21:38:59.528002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T00:15:26.619238Z digest=sha256:2c47e6b34111e4d69c9d94be7dc1042a941301307f144f01095ea08fec8bf6de

Observation a1802dc2-07e2-429d-8bc2-5355b01cac3f · inbound

VEGA: Learning Navigation VLAs from In-the-Wild Egocentric Video with Geometric Trajectory Supervision cites this paper.

VEGA: Learning Navigation VLAs from In-the-Wild Egocentric Video with Geometric Trajectory Supervision HALO: Human Preference Aligned Offline Reward Learning for Robot Navigation

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-02T11:03:51.194382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:03:51.194382Z digest=sha256:1b0dbe3d234d2628f4b81d285fc4465de07ba6353943d44d0e3781a47573591b