Pith. sign in

Paper Citation Record · LEDGER

Discounted Beta-Bernoulli Reward Estimation for Sample-Efficient Reinforcement Learning with Verifiable Rewards

As of 20 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 0 inbound Pith citation observations for arXiv:2603.18444.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2603.18444 v2

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-13T22:38:02.945576Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

54 of 54 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved54
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ee13414a-61dc-4dc7-8021-5d2f7949d656 · outbound

This paper cites ObjectNav Revisited: On Evaluation of Embodied Agents Navigating to Objects.

Discounted Beta-Bernoulli Reward Estimation for Sample-Efficient Reinforcement Learning with Verifiable Rewards ObjectNav Revisited: On Evaluation of Embodied Agents Navigating to Objects

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-13T22:38:02.945576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:38:02.945576Z digest=sha256:f198b74dfd823fbb235358b64907a404be216f38bef4ffc24ef1d10e6dee71a8

Observation bab5cdb9-12ab-4224-913b-44d72fe7cb78 · outbound

This paper cites Bridging zero-shot object navigation and foundation models through pixel-guided navigation skill,.

Discounted Beta-Bernoulli Reward Estimation for Sample-Efficient Reinforcement Learning with Verifiable Rewards Bridging zero-shot object navigation and foundation models through pixel-guided navigation skill,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-13T22:38:02.945576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:38:02.945576Z digest=sha256:9c5b277e6b417ad0d073b4ce78e458616fb202e7300886c23e4aff9a66d737d7

Observation b88404a6-763b-4eed-8e85-21e4a1cc880e · outbound

This paper cites V oila: Visual- observation-only imitation learning for autonomous navigation,.

Discounted Beta-Bernoulli Reward Estimation for Sample-Efficient Reinforcement Learning with Verifiable Rewards V oila: Visual- observation-only imitation learning for autonomous navigation,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-13T22:38:02.945576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:38:02.945576Z digest=sha256:d746dfa5490c87984080140515ee5f17f9b83d0e994e5d2254ddf519bcb8684c

Observation 58a525dd-6cc7-4ab9-a1b1-e3980464365a · outbound

This paper cites Zson: Zero-shot object-goal navigation using multimodal goal embed- dings,.

Discounted Beta-Bernoulli Reward Estimation for Sample-Efficient Reinforcement Learning with Verifiable Rewards Zson: Zero-shot object-goal navigation using multimodal goal embed- dings,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-13T22:38:02.945576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:38:02.945576Z digest=sha256:4528f856071e87f3ffd9c418dd0d5a2226b191b5a310d0c9d40477644ea64897

Observation 714bca8c-96b0-49a3-a054-9c28dc97cfad · outbound

This paper cites Zero-shot active visual search (zavis): Intelligent object search for robotic assistants,.

Discounted Beta-Bernoulli Reward Estimation for Sample-Efficient Reinforcement Learning with Verifiable Rewards Zero-shot active visual search (zavis): Intelligent object search for robotic assistants,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-13T22:38:02.945576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:38:02.945576Z digest=sha256:06e44d6be7500c57e1fc8d973fc8661a73d0036d6c88dd1b381917c0a0fcc9ee

Observation 8d31a1de-b4a2-47d7-9873-0616e56f2d04 · outbound

This paper cites Zero-shot object goal visual navigation,.

Discounted Beta-Bernoulli Reward Estimation for Sample-Efficient Reinforcement Learning with Verifiable Rewards Zero-shot object goal visual navigation,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-13T22:38:02.945576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:38:02.945576Z digest=sha256:ceb72b7ba6202903b7808c1609d7b41b8eeca19c8211f4d7dbd7067d595f8854

Observation 3c8d42a3-5f05-4a76-bfee-4d1884f8e15f · outbound

This paper cites Hierarchies of planning and reinforcement learning for robot navigation,.

Discounted Beta-Bernoulli Reward Estimation for Sample-Efficient Reinforcement Learning with Verifiable Rewards Hierarchies of planning and reinforcement learning for robot navigation,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-13T22:38:02.945576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:38:02.945576Z digest=sha256:9a976ece12fa5707a26da91c39643edab85c233fbf34f5e553be604bc8153533

Observation ab6858b9-e292-4437-a020-e20ccafd8cc0 · outbound

This paper cites Esc: Exploration with soft commonsense constraints for zero- shot object navigation,.

Discounted Beta-Bernoulli Reward Estimation for Sample-Efficient Reinforcement Learning with Verifiable Rewards Esc: Exploration with soft commonsense constraints for zero- shot object navigation,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-13T22:38:02.945576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:38:02.945576Z digest=sha256:720c724069cb2eb35daac3a2f4ba245e739b4e0cd3bcbe8d67d2acc1f398831e

Observation d577743e-939c-408b-9d4f-db2b5c29def5 · outbound

This paper cites Sg-nav: Online 3d scene graph prompting for llm-based zero-shot object navigation,.

Discounted Beta-Bernoulli Reward Estimation for Sample-Efficient Reinforcement Learning with Verifiable Rewards Sg-nav: Online 3d scene graph prompting for llm-based zero-shot object navigation,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-13T22:38:02.945576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:38:02.945576Z digest=sha256:4ca3f38efb54b3d90b2625a94c45c8608193eb3044c8dc193bbd7a9ddc6dea1e

Observation cdf9663e-e929-4c8d-849f-e69b0e9f0c6a · outbound

This paper cites Vlfm: Vision- language frontier maps for zero-shot semantic navigation,.

Discounted Beta-Bernoulli Reward Estimation for Sample-Efficient Reinforcement Learning with Verifiable Rewards Vlfm: Vision- language frontier maps for zero-shot semantic navigation,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-13T22:38:02.945576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:38:02.945576Z digest=sha256:130a0eb88e052e08d2fea495bd46ad4f08a2db41bbbb03e7305df74b10f67abd

Observation 4c53bf7a-33c8-4c47-b896-a78f15fd179d · outbound

This paper cites V oronav: voronoi-based zero-shot object navigation with large language model,.

Discounted Beta-Bernoulli Reward Estimation for Sample-Efficient Reinforcement Learning with Verifiable Rewards V oronav: voronoi-based zero-shot object navigation with large language model,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-13T22:38:02.945576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:38:02.945576Z digest=sha256:15748558b155036e21418ce35162eba0614390e364081c10aab1667b6ed081c1

Observation 35611784-b131-465b-9273-15d5173dfbb2 · outbound

This paper cites TopV-Nav: Unlocking the Top-View Spatial Reasoning Potential of MLLM for Zero-shot Object Navigation.

Discounted Beta-Bernoulli Reward Estimation for Sample-Efficient Reinforcement Learning with Verifiable Rewards TopV-Nav: Unlocking the Top-View Spatial Reasoning Potential of MLLM for Zero-shot Object Navigation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-13T22:38:02.945576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:38:02.945576Z digest=sha256:ce1536947def45ab0c05e5bc3e2e813048d8de645b2275302d64631e16aa8894

Observation d9b4a430-86be-40b7-83e2-0d1fa2aa0616 · outbound

This paper cites Cows on pasture: Baselines and benchmarks for language-driven zero-shot object navigation,.

Discounted Beta-Bernoulli Reward Estimation for Sample-Efficient Reinforcement Learning with Verifiable Rewards Cows on pasture: Baselines and benchmarks for language-driven zero-shot object navigation,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-13T22:38:02.945576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:38:02.945576Z digest=sha256:672fd3c87cf49168ead0f4d242610e3f54fdd71258a00108af59f99c4ebb3ab7

Observation 55c31205-0279-4a85-98aa-d00e603c4030 · outbound

This paper cites Star-searcher: A complete and efficient aerial system for autonomous target search in complex unknown environments,.

Discounted Beta-Bernoulli Reward Estimation for Sample-Efficient Reinforcement Learning with Verifiable Rewards Star-searcher: A complete and efficient aerial system for autonomous target search in complex unknown environments,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-13T22:38:02.945576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:38:02.945576Z digest=sha256:ab498b4176788c4c3e44ea1560dbcee1ff5b07f73f2bc26bf64b0de91430e76f

Observation 2a017583-2025-415a-a5cb-6a88614ab482 · outbound

This paper cites Occupancy Grids: A Stochastic Spatial Representation for Active Robot Perception.

Discounted Beta-Bernoulli Reward Estimation for Sample-Efficient Reinforcement Learning with Verifiable Rewards Occupancy Grids: A Stochastic Spatial Representation for Active Robot Perception

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-13T22:38:02.945576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:38:02.945576Z digest=sha256:2f5740612edad1587fa71bb05e7343abce5d2ddd6cf01b10a6aa9447bd75b18a

Observation 422750fd-de7f-483e-8577-edb12fd81549 · outbound

This paper cites A frontier-based approach for autonomous exploration,.

Discounted Beta-Bernoulli Reward Estimation for Sample-Efficient Reinforcement Learning with Verifiable Rewards A frontier-based approach for autonomous exploration,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-13T22:38:02.945576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:38:02.945576Z digest=sha256:062fa71e5f5b00a5b97f29309737599e020481308bb0b9f97cfcb3ed4e4aaa39

Observation 4a29c07d-e87f-4229-ba10-40fa2b9788c6 · outbound

This paper cites Information gain-based exploration using rao-blackwellized particle filters,.

Discounted Beta-Bernoulli Reward Estimation for Sample-Efficient Reinforcement Learning with Verifiable Rewards Information gain-based exploration using rao-blackwellized particle filters,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-13T22:38:02.945576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:38:02.945576Z digest=sha256:a97fd77a52b5814cac4f3d9f0466fdfbca2aa2e947cd896218ba13a43837c75c

Observation 5f96670a-d8de-4044-9e9d-eea88f0a39c0 · outbound

This paper cites Gamap: Zero-shot object goal navigation with multi-scale geometric-affordance guidance,.

Discounted Beta-Bernoulli Reward Estimation for Sample-Efficient Reinforcement Learning with Verifiable Rewards Gamap: Zero-shot object goal navigation with multi-scale geometric-affordance guidance,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-13T22:38:02.945576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:38:02.945576Z digest=sha256:78aea130a3c60f07c70970b7eed348bd4369362044cc4711e69e2069fe6f6c1b

Observation a193f278-1470-4c96-bb39-37a1850a4e56 · outbound

This paper cites L3mvn: Leveraging large language models for visual target navigation,.

Discounted Beta-Bernoulli Reward Estimation for Sample-Efficient Reinforcement Learning with Verifiable Rewards L3mvn: Leveraging large language models for visual target navigation,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-13T22:38:02.945576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:38:02.945576Z digest=sha256:c370183eaf00536912b15273dff99ebf77b10292bf5ddf3d1019264469436741

Observation a2a3ee19-2892-410c-8e51-7d8a77e153c0 · outbound

This paper cites Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,.

Discounted Beta-Bernoulli Reward Estimation for Sample-Efficient Reinforcement Learning with Verifiable Rewards Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-13T22:38:02.945576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:38:02.945576Z digest=sha256:eb787ee15e5d24b2a0ba4b957910c1746fd301c8598b04707c1d484c69070a9b

Observation 22462d54-9508-4ef8-a93f-6e0834f090bb · outbound

This paper cites Qwen2.5-VL Technical Report.

Discounted Beta-Bernoulli Reward Estimation for Sample-Efficient Reinforcement Learning with Verifiable Rewards Qwen2.5-VL Technical Report

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-13T22:38:02.945576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:38:02.945576Z digest=sha256:f3e3b8578c0e653d556938bf315bfa50edf85b089c62ad8627372373825d03eb

Observation ed768df3-022a-48c8-adae-1b0ba64dc7a9 · outbound

This paper cites GPT-4 Technical Report.

Discounted Beta-Bernoulli Reward Estimation for Sample-Efficient Reinforcement Learning with Verifiable Rewards GPT-4 Technical Report

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-13T22:38:02.945576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:38:02.945576Z digest=sha256:a163e48e5b89c5df74e22409577009c2a7573da955847ae76141d2b4ece4a801

Observation 9507e9ee-f4c0-4ff2-9de5-f963df4d93df · outbound

This paper cites OpenFMNav: Towards open-set zero-shot object navigation via vision-language foundation models,.

Discounted Beta-Bernoulli Reward Estimation for Sample-Efficient Reinforcement Learning with Verifiable Rewards OpenFMNav: Towards open-set zero-shot object navigation via vision-language foundation models,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-13T22:38:02.945576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:38:02.945576Z digest=sha256:ad83599b4177b4e2956ab08dc95f7f3bae87045e5b4a367ec69d9d578ad8611a

Observation 74175aed-4548-43bc-ae8e-b3906f67fe21 · outbound

This paper cites Egtr: Extracting graph from transformer for scene graph generation,.

Discounted Beta-Bernoulli Reward Estimation for Sample-Efficient Reinforcement Learning with Verifiable Rewards Egtr: Extracting graph from transformer for scene graph generation,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-13T22:38:02.945576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:38:02.945576Z digest=sha256:53446066e87a8e38d73f4a01de28c4ea05a7fabaae9d7215b3d863b4927e4b4f

Observation 5389475a-3a47-4205-a3db-5ae648be79e9 · outbound

This paper cites Hi- erarchical Open-V ocabulary 3D Scene Graphs for Language-Grounded Robot Navigation,.

Discounted Beta-Bernoulli Reward Estimation for Sample-Efficient Reinforcement Learning with Verifiable Rewards Hi- erarchical Open-V ocabulary 3D Scene Graphs for Language-Grounded Robot Navigation,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-13T22:38:02.945576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:38:02.945576Z digest=sha256:4f2ff0c7f2c1e64f8a16e1c7928ddd77f6d1942aec20e70c00a53132337b9434

Observation 9fd13e79-6973-4696-92ce-2953a7facb63 · outbound

This paper cites Expanding scene graph boundaries: Fully open-vocabulary scene graph generation via visual- concept alignment and retention,.

Discounted Beta-Bernoulli Reward Estimation for Sample-Efficient Reinforcement Learning with Verifiable Rewards Expanding scene graph boundaries: Fully open-vocabulary scene graph generation via visual- concept alignment and retention,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-13T22:38:02.945576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:38:02.945576Z digest=sha256:294ce943caba20e77d2e238344d6fd5db3ba6e6432319580370736911d178d18

Observation f4e1b900-4d03-457e-92b4-c0aadcf8c549 · outbound

This paper cites Cognav: Cognitive process modeling for object goal navigation with llms,.

Discounted Beta-Bernoulli Reward Estimation for Sample-Efficient Reinforcement Learning with Verifiable Rewards Cognav: Cognitive process modeling for object goal navigation with llms,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-13T22:38:02.945576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:38:02.945576Z digest=sha256:ed485bd12840442f980446dca71c03a8ffc268ccf07d360176375abe7abcd851

Observation e50425c1-1a04-4463-9ac3-e78fabd282b4 · outbound

This paper cites Beliefmapnav: 3d voxel- based belief map for zero-shot object navigation,.

Discounted Beta-Bernoulli Reward Estimation for Sample-Efficient Reinforcement Learning with Verifiable Rewards Beliefmapnav: 3d voxel- based belief map for zero-shot object navigation,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-13T22:38:02.945576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:38:02.945576Z digest=sha256:0e9e3b50a33db67d585b1a1260f0e7481c50c1a4608962752034e9a7fb5890ad

Observation dfbd27a7-057d-4c53-a748-9e96f56effc5 · outbound

This paper cites Object goal navigation using goal-oriented semantic exploration,.

Discounted Beta-Bernoulli Reward Estimation for Sample-Efficient Reinforcement Learning with Verifiable Rewards Object goal navigation using goal-oriented semantic exploration,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-13T22:38:02.945576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:38:02.945576Z digest=sha256:7ce2e0d63c211c70a22aa6bedce20ddcf675e8cc3e1763564f3d1098d8c42f84

Observation d929c26c-16d1-43c2-a77f-57a031226a53 · outbound

This paper cites Habitat-Matterport 3D Dataset (HM3D): 1000 Large-scale 3D Environments for Embodied AI.

Discounted Beta-Bernoulli Reward Estimation for Sample-Efficient Reinforcement Learning with Verifiable Rewards Habitat-Matterport 3D Dataset (HM3D): 1000 Large-scale 3D Environments for Embodied AI

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-13T22:38:02.945576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:38:02.945576Z digest=sha256:4bc4128f6d8a01f49d94b4fb6d2066ffd75e98f194c111de4bb621b71fbda58c

Observation 9f9942e5-bbd3-4c18-8de3-a0ef9ac544b4 · outbound

This paper cites Matterport3d: Learning from rgb-d data in indoor environments,.

Discounted Beta-Bernoulli Reward Estimation for Sample-Efficient Reinforcement Learning with Verifiable Rewards Matterport3d: Learning from rgb-d data in indoor environments,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-13T22:38:02.945576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:38:02.945576Z digest=sha256:c839db8b5de95a21b0604757eb803a9af81960292ae86c409f8fcc526bc84c94

Observation 5d0cc65b-40b5-424c-90e2-fb4bfedd6878 · outbound

This paper cites Trihelper: Zero-shot object navigation with dynamic assistance,.

Discounted Beta-Bernoulli Reward Estimation for Sample-Efficient Reinforcement Learning with Verifiable Rewards Trihelper: Zero-shot object navigation with dynamic assistance,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-13T22:38:02.945576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:38:02.945576Z digest=sha256:1cd702a3cd13991beecb5fbc96284c76750d3327d00b81e855e94e3685a7bed0

Observation bf0da835-7476-4d63-811f-5304dcff8de9 · outbound

This paper cites How to not train your dragon: Training-free embodied object goal navigation with semantic frontiers,.

Discounted Beta-Bernoulli Reward Estimation for Sample-Efficient Reinforcement Learning with Verifiable Rewards How to not train your dragon: Training-free embodied object goal navigation with semantic frontiers,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-13T22:38:02.945576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:38:02.945576Z digest=sha256:53a5c6a4fecd012b69a26e9401f45bebbd023790cfaefa1ab931fa099673c4de

Observation 9e5c6740-0425-4705-be05-4dd17b64711b · outbound

This paper cites Fuel: Fast uav exploration using incremental frontier structure and hierarchical planning,.

Discounted Beta-Bernoulli Reward Estimation for Sample-Efficient Reinforcement Learning with Verifiable Rewards Fuel: Fast uav exploration using incremental frontier structure and hierarchical planning,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-13T22:38:02.945576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:38:02.945576Z digest=sha256:dfc28e441d7f0d2012ea7897477472614fb6501fb61e20905e2fb678e7cc8d3e

Observation a8a622e0-f555-4d40-a1b3-684ba0591a86 · outbound

This paper cites Think holistically, act down-to-earth: A semantic navigation strategy with continuous environmental representation and multi-step forward planning,.

Discounted Beta-Bernoulli Reward Estimation for Sample-Efficient Reinforcement Learning with Verifiable Rewards Think holistically, act down-to-earth: A semantic navigation strategy with continuous environmental representation and multi-step forward planning,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-13T22:38:02.945576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:38:02.945576Z digest=sha256:b79a0c9c69f8fb85a229c9cd8174fb05ab42361b54c782fe61965db2226313d9

Observation cd5a7e54-6a4a-4643-84a6-91ee7833e11f · outbound

This paper cites Open scene graphs for open-world object- goal navigation,.

Discounted Beta-Bernoulli Reward Estimation for Sample-Efficient Reinforcement Learning with Verifiable Rewards Open scene graphs for open-world object- goal navigation,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-13T22:38:02.945576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:38:02.945576Z digest=sha256:caca0e180f5bd3990a4d94c61f4f87c2a9a06f601d36b3907381e4ed5ab44e02

Observation 61156cb8-76b5-4303-9e8e-532b688f8a68 · outbound

This paper cites Agent-centric relation graph for object visual navigation,.

Discounted Beta-Bernoulli Reward Estimation for Sample-Efficient Reinforcement Learning with Verifiable Rewards Agent-centric relation graph for object visual navigation,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-13T22:38:02.945576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:38:02.945576Z digest=sha256:d3ca1e80b2cb95afdbd582f5d9b04c6d672ff1a2b9907a5b2aac88288d313887

Observation beb2cef6-86b4-4c61-a403-8ee98f9aae9b · outbound

This paper cites Imaginenav: Prompting vision- language models as embodied navigator through scene imagination,.

Discounted Beta-Bernoulli Reward Estimation for Sample-Efficient Reinforcement Learning with Verifiable Rewards Imaginenav: Prompting vision- language models as embodied navigator through scene imagination,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-13T22:38:02.945576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:38:02.945576Z digest=sha256:7c7a4e421852cd839316054310aaf73ceb21f80c47f5d671e6437a89593df993

Observation e3dfd92a-bb50-42e0-a110-bb69bb6477b8 · outbound

This paper cites BERT: Pre- training of deep bidirectional transformers for language understanding,.

Discounted Beta-Bernoulli Reward Estimation for Sample-Efficient Reinforcement Learning with Verifiable Rewards BERT: Pre- training of deep bidirectional transformers for language understanding,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-13T22:38:02.945576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:38:02.945576Z digest=sha256:27e9806c956dd8c98a6e8668ec99e06efdefa193a038b3b423a55517d6cc8b33

Observation 05eb2658-667e-4015-8179-7f85b963f23f · outbound

This paper cites Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning,.

Discounted Beta-Bernoulli Reward Estimation for Sample-Efficient Reinforcement Learning with Verifiable Rewards Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-13T22:38:02.945576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:38:02.945576Z digest=sha256:bd5dee6bb4b59aecafb56e1eb436cc87e763abbf5f6735670e75a1e8a5e3f041

Observation f2632f36-53a3-49c2-858f-6f2d1d164a4f · outbound

This paper cites Chatnav: Leveraging llm to zero-shot semantic reasoning in object navigation,.

Discounted Beta-Bernoulli Reward Estimation for Sample-Efficient Reinforcement Learning with Verifiable Rewards Chatnav: Leveraging llm to zero-shot semantic reasoning in object navigation,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-13T22:38:02.945576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:38:02.945576Z digest=sha256:55f6154c5509e16c47e8ea2965b2ef92b62d4ee7b9000df0ac20a41f6efbbd41

Observation 3649dc97-8331-421d-8d70-ddaa592d0f1b · outbound

This paper cites E²ba: Environment exploration and backtracking agent for visual language object navigation,.

Discounted Beta-Bernoulli Reward Estimation for Sample-Efficient Reinforcement Learning with Verifiable Rewards E²ba: Environment exploration and backtracking agent for visual language object navigation,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-13T22:38:02.945576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:38:02.945576Z digest=sha256:77f8e592eabfaa86fa8698fc09c2cba5ec835bf53f7fbb4e6601eee0d3a5aa4b

Observation b0ba9d8e-8ec6-4fc3-86a9-149c4d3179a5 · outbound

This paper cites Strive: Structured representation integrating vlm reasoning for efficient object navigation,.

Discounted Beta-Bernoulli Reward Estimation for Sample-Efficient Reinforcement Learning with Verifiable Rewards Strive: Structured representation integrating vlm reasoning for efficient object navigation,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-13T22:38:02.945576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:38:02.945576Z digest=sha256:3c177c34d2937ffedb4f646ab887d5037d082012b5da4382ac10b6cff32c92b1

Observation ffaeb69b-ceef-4fb5-b252-f5004682a512 · outbound

This paper cites Visual semantic navigation using scene priors,.

Discounted Beta-Bernoulli Reward Estimation for Sample-Efficient Reinforcement Learning with Verifiable Rewards Visual semantic navigation using scene priors,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-13T22:38:02.945576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:38:02.945576Z digest=sha256:b4f220aa368c59b2991b2a0cf81eeb4ee250cfc44203591a3c85b5f666a0cac2

Observation 48e4afe0-a828-4c03-bf06-5a2a94f15a2a · outbound

This paper cites Imagine before go: Self-supervised generative map for object goal navigation,.

Discounted Beta-Bernoulli Reward Estimation for Sample-Efficient Reinforcement Learning with Verifiable Rewards Imagine before go: Self-supervised generative map for object goal navigation,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-13T22:38:02.945576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:38:02.945576Z digest=sha256:caabaf14a0390d4f2541ddbfff6cecaa5784a6993fb5f7fad2be6b675bb1aaee

Observation ec82943f-9036-4602-94bf-d260b783dda4 · outbound

This paper cites Hierarchical object-to-zone graph for object navigation,.

Discounted Beta-Bernoulli Reward Estimation for Sample-Efficient Reinforcement Learning with Verifiable Rewards Hierarchical object-to-zone graph for object navigation,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-07-13T22:38:02.945576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:38:02.945576Z digest=sha256:83796e1bb16f077cf2f3f00adaa7d3f20797dc0d237f28b4c4c48f7aaa6026e4

Observation 0f547d67-5653-4d00-b199-d77e75053ac1 · outbound

This paper cites Search for or navigate to? dual adaptive thinking for object navigation,.

Discounted Beta-Bernoulli Reward Estimation for Sample-Efficient Reinforcement Learning with Verifiable Rewards Search for or navigate to? dual adaptive thinking for object navigation,

Reference 47

Resolution
unresolved
no resolver link, observed 2026-07-13T22:38:02.945576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:38:02.945576Z digest=sha256:6f239aa0bb1dac7be745d2f30491ad46fa41b06565cd91d7be9bf1ee893ed930

Observation a57d928d-7ec9-4039-91c1-f435bcaacdc0 · outbound

This paper cites Layout-based causal inference for object navigation,.

Discounted Beta-Bernoulli Reward Estimation for Sample-Efficient Reinforcement Learning with Verifiable Rewards Layout-based causal inference for object navigation,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-07-13T22:38:02.945576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:38:02.945576Z digest=sha256:0f4f772f38ade414b02f884441cf39827d69b33a0cfca875527cf467ca2c0c0e

Observation 985abf22-8817-4706-bebe-607ce89eaf30 · outbound

This paper cites Poni: Potential functions for objectgoal navigation with interaction-free learning,.

Discounted Beta-Bernoulli Reward Estimation for Sample-Efficient Reinforcement Learning with Verifiable Rewards Poni: Potential functions for objectgoal navigation with interaction-free learning,

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-13T22:38:02.945576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:38:02.945576Z digest=sha256:9d0f1c3351a415c17967a8f0594383904e45c4f91cb602d4c1e554040cbfd5cc

Observation 3fcb54d5-3554-4f55-bb2c-dfef01caf4aa · outbound

This paper cites Learning object relation graph and tentative policy for visual navigation,.

Discounted Beta-Bernoulli Reward Estimation for Sample-Efficient Reinforcement Learning with Verifiable Rewards Learning object relation graph and tentative policy for visual navigation,

Reference 50

Resolution
unresolved
no resolver link, observed 2026-07-13T22:38:02.945576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:38:02.945576Z digest=sha256:faba161fdea14bfd8bb25e21efad7983b2e5d68493a10e8a639d305db8bea072

Observation b25dea6e-8860-4165-98be-2444f0c34dbc · outbound

This paper cites Yolov7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors,.

Discounted Beta-Bernoulli Reward Estimation for Sample-Efficient Reinforcement Learning with Verifiable Rewards Yolov7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors,

Reference 51

Resolution
unresolved
no resolver link, observed 2026-07-13T22:38:02.945576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:38:02.945576Z digest=sha256:0d4853046c29e968e2f98c9dd4781276d7f02c8baacd1da85ba02cced367bf27

Observation 613d9442-cc37-4e34-ba49-1f29faefd448 · outbound

This paper cites Grounding dino: Marrying dino with grounded pre-training for open-set object detection,.

Discounted Beta-Bernoulli Reward Estimation for Sample-Efficient Reinforcement Learning with Verifiable Rewards Grounding dino: Marrying dino with grounded pre-training for open-set object detection,

Reference 52

Resolution
unresolved
no resolver link, observed 2026-07-13T22:38:02.945576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:38:02.945576Z digest=sha256:7a84dbe11d0537cd98768ff7bbbfd3d366939c06ffe2d5b25ffac8fb4909b424

Observation 6c30f039-8c2d-42e5-a8bf-61ad55d85adb · outbound

This paper cites Faster Segment Anything: Towards Lightweight SAM for Mobile Applications.

Discounted Beta-Bernoulli Reward Estimation for Sample-Efficient Reinforcement Learning with Verifiable Rewards Faster Segment Anything: Towards Lightweight SAM for Mobile Applications

Reference 53

Resolution
unresolved
no resolver link, observed 2026-07-13T22:38:02.945576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:38:02.945576Z digest=sha256:6349a9d53e18dda0617cc4a2caf50aaa0e65205da8b2d643aff96a874f00c977

Observation 07179b9a-ed6c-4973-bb92-8507893ebaf6 · outbound

This paper cites Unigoal: Towards universal zero-shot goal-oriented navigation,.

Discounted Beta-Bernoulli Reward Estimation for Sample-Efficient Reinforcement Learning with Verifiable Rewards Unigoal: Towards universal zero-shot goal-oriented navigation,

Reference 54

Resolution
unresolved
no resolver link, observed 2026-07-13T22:38:02.945576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:38:02.945576Z digest=sha256:69e94fb374706c41b7704de877575897ed4afab2efab7d71cb1d0453768aca8c

Pith citing papers

No inbound Pith citation observations are available.