Pith. sign in

Paper Citation Record · LEDGER

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models

As of 15 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 0 inbound Pith citation observations for arXiv:2505.20718.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.20718 v2

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:51:46.426406Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

32 of 32 outbound references displayed

  • verified exact1
  • verified fuzzy24
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 54c4669a-2f11-403f-baaa-d4c2a298f02a · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:43.474536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:51:43.474536Z digest=sha256:7ddddae5c844ef45804f66e101b383aa75ffad3fcdb21903ff9c801157d768aa

Observation 6ef6bfa1-9b6e-4479-8d86-5dd33924f34b · outbound

This paper cites Spatialvlm: Endowing vision-language models with spatial reasoning capabilities.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models Spatialvlm: Endowing vision-language models with spatial reasoning capabilities

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:43.540009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:51:43.540009Z digest=sha256:cd95b022dbf64acb65b1cf93e864415d820b400a809db85432c285eea6209b9d

Observation 31cb6773-2e01-4d29-8031-b6364f4d8287 · outbound

This paper cites Tracking anything with decoupled video segmenta- tion.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models Tracking anything with decoupled video segmenta- tion

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:50.969592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:51:43.644671Z digest=sha256:b09e690c17570a990e23c6d3fbb5a46f7e8a106b528054f3678acbb5420184f0

Observation 5b26f5f9-35cb-44d4-9ce4-4137ccf72add · outbound

This paper cites an unresolved cited work.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:51:50.855505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:51:43.761418Z digest=sha256:ca35575b0231ddc1822323fc5fa47d68d550fef7ba88cf261d82269f306eb160

Observation 704b213d-417b-4f9a-bad5-8c4c6730ff8b · outbound

This paper cites Proactive multi-camera collaboration for 3d human pose estimation.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models Proactive multi-camera collaboration for 3d human pose estimation

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:50.692694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:51:43.901893Z digest=sha256:6340821578008394770f33fe95f3cc11c0691180aced68efdee57e50a8401ce2

Observation b18b71a2-8040-4c01-aef6-96f5211b6d6d · outbound

This paper cites Enhancing continuous control of mobile robots for end-to-end visual active tracking.Robotics and Autonomous Systems, 142:103799, 2021.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models Enhancing continuous control of mobile robots for end-to-end visual active tracking.Robotics and Autonomous Systems, 142:103799, 2021

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:50.605703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:51:44.059949Z digest=sha256:d5ad280063787855a3bab1cafde7e72264e237a909e37cf083ce687b07c8a4c0

Observation f6d464b0-ed64-428e-939c-1a569fac0b7d · outbound

This paper cites E-vat: An asymmetric end-to-end approach to visual active exploration and tracking.IEEE Robotics and Automation Letters, 7(2):4259–4266, 2022.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models E-vat: An asymmetric end-to-end approach to visual active exploration and tracking.IEEE Robotics and Automation Letters, 7(2):4259–4266, 2022

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:50.455798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:51:44.177272Z digest=sha256:dc15efbdaed3a216c82c63ec8ac71148fa721e139a6944bdb406dfda1f1c3ac8

Observation cbf17c7f-be72-4e1e-949f-5826f5d84ac8 · outbound

This paper cites D-vat: End-to-end visual active tracking for micro aerial vehicles.IEEE Robotics and Automation Letters, 2024.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models D-vat: End-to-end visual active tracking for micro aerial vehicles.IEEE Robotics and Automation Letters, 2024

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:50.324369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:51:44.299789Z digest=sha256:20f50ff97df33009e6ab3005b2401ea8ef3295addeb696cd004186fe89c3d33e

Observation 16a841a7-1f69-462c-b24b-69874cc65de2 · outbound

This paper cites Memory sharing for large language model based agents.Arxiv Preprint Arxiv:2404.09982, 2024.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models Memory sharing for large language model based agents.Arxiv Preprint Arxiv:2404.09982, 2024

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:44.418085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:51:44.418085Z digest=sha256:97ca441e497cede1cadc9f68bfb7c690ea6cea4a64e18ad27098262ad082ccf7

Observation 93964576-1afd-41f8-b39a-fd8b7be5de22 · outbound

This paper cites An embodied generalist agent in 3d world.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models An embodied generalist agent in 3d world

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:50.182089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:51:44.500183Z digest=sha256:44b88d8a6ec41b454344d1a4fc03520a69936be5d23f172544171e2c20016817

Observation 6abaf5df-2d3c-4371-b70e-d389b36ce974 · outbound

This paper cites Conquering Ghosts: Relation Learning for Information Reliability Representation and End-to-End Robust Navigation.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models Conquering Ghosts: Relation Learning for Information Reliability Representation and End-to-End Robust Navigation

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:51:46.630634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:51:44.571948Z digest=sha256:af055a00454af3bff4b37fb94da56c0efa6152069054f7383367dfc3bfb8945b

Observation 55664fc4-dc56-47dd-83c1-d2bd62cba002 · outbound

This paper cites OpenVLA: An open-source vision-language-action model.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models OpenVLA: An open-source vision-language-action model

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:50.001978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:51:44.643669Z digest=sha256:73831ac24ef9688f7df2677a874f7b5cd34c2f5f2363aa88ec14edfa4851dc77

Observation 7f395053-766f-41b7-ad25-7da0c9e89a57 · outbound

This paper cites A novel performance evaluation methodology for single-target trackers.IEEE Transactions on Pattern Analysis and Machine Intelligence, 38(11):2137–2155, Nov 2016.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models A novel performance evaluation methodology for single-target trackers.IEEE Transactions on Pattern Analysis and Machine Intelligence, 38(11):2137–2155, Nov 2016

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:49.824518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:51:44.737506Z digest=sha256:4a1dad6aedd2c730410c89843e13a71b6d2bd03ddb214c8d6262472aa0c2eeb7

Observation 97273674-23de-4b91-ae61-801372c2e4ee · outbound

This paper cites Person following robot based on real time single object tracking and rgb-d image.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models Person following robot based on real time single object tracking and rgb-d image

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:49.727647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:51:44.829613Z digest=sha256:6288c256cda117ec22728211da76655af8dc193a027552766e77b2a40c34e8f7

Observation 225dd818-5576-43b6-8d35-e74c33fcbff8 · outbound

This paper cites Blip-2: Boot- strapping language-image pre-training with frozen image encoders and large language models.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models Blip-2: Boot- strapping language-image pre-training with frozen image encoders and large language models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:49.536692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:51:44.927307Z digest=sha256:c1ec4f6d6aff8956f9e6bc26147bd47a4ee331facbd51cf03e6bd91f88791069

Observation 4417db72-92aa-45cc-8bf4-bc54c8dfe350 · outbound

This paper cites Vi- sual instruction tuning.Advances in Neural Information Processing Systems, 36:34892–34916, 2023.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models Vi- sual instruction tuning.Advances in Neural Information Processing Systems, 36:34892–34916, 2023

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:49.380606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:51:45.026284Z digest=sha256:60b2620fc177a11cd7c7339c85d4189166dd7be8d9b3c688928fb2d317141f92

Observation 6f02cfff-2347-4e62-8510-e8c41d2acedc · outbound

This paper cites End-to-end active object tracking via reinforcement learning.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models End-to-end active object tracking via reinforcement learning

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:49.120498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:51:45.123410Z digest=sha256:4ee5822d4d38f94f92b3b5606e4665711fc420e33e310aff005cb78252d1ab27

Observation 093d7d60-5442-4e70-97ac-08a00379108d · outbound

This paper cites Curious george: An attentive semantic robot.Robotics and Autonomous Systems, 56(6):503–511, 2008.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models Curious george: An attentive semantic robot.Robotics and Autonomous Systems, 56(6):503–511, 2008

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:48.945402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:51:45.249150Z digest=sha256:590070f591d1b893528571dbbe53f621f64f828f90e50cff10b4de03a2465f07

Observation 2ccb200f-9763-4a12-b6f4-52931b22ea36 · outbound

This paper cites The hands-free push-cart: Autonomous following in front by predicting user trajectory around obstacles.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models The hands-free push-cart: Autonomous following in front by predicting user trajectory around obstacles

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:48.787461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:51:45.348257Z digest=sha256:6cfaf658602319ee80f0755d8ba65bebdc81f2620adeab328a81d038554a94be

Observation fa7968d5-f715-4a43-bfc7-33a0b93f4ed8 · outbound

This paper cites Unrealcv: Virtual worlds for computer vision.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models Unrealcv: Virtual worlds for computer vision

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:48.557921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:51:45.423214Z digest=sha256:02455130c64e83c3d1d33183de08ddd31005a2009b2f61f2ae07bbdf77308e9c

Observation a0c6ebb1-b3d7-45f0-a4e9-0f8eadb2f2ef · outbound

This paper cites Tracking multiple moving targets with a mobile robot using particle filters and statistical data association.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models Tracking multiple moving targets with a mobile robot using particle filters and statistical data association

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:48.406896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:51:45.501658Z digest=sha256:e833553079c2372e717cc8319d4b6a90c4c0385e4d1b498f782e80c9cf9c3c8c

Observation f929f2e8-5a5c-43f3-b8ed-543c40b5b99e · outbound

This paper cites Accurate and real-time 3-d tracking for the following robots by fusing vision and ultrasonar information.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models Accurate and real-time 3-d tracking for the following robots by fusing vision and ultrasonar information

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:48.185872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:51:45.573732Z digest=sha256:308370bdda0470cc0b384a2f930ccb41e912801aab67021b044c30b15df40d58

Observation f698e522-b59a-4366-81c7-c3897d833e31 · outbound

This paper cites Vlfm: Vision-language frontier maps for zero-shot semantic navigation.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models Vlfm: Vision-language frontier maps for zero-shot semantic navigation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:47.971967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:51:45.633871Z digest=sha256:53b9f77c4d5d0076541e0291ee2714ef99afac1d040b35f06f40a4c4f6394bcb

Observation bf9ea878-827c-4e9a-801d-5334361f5f9f · outbound

This paper cites NaVid: Video-based VLM Plans the Next Step for Vision-and-Language Navigation.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models NaVid: Video-based VLM Plans the Next Step for Vision-and-Language Navigation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:45.712882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:51:45.712882Z digest=sha256:b9f07d8d7de7e0ee7d185b65a150fbbd36028312e083fa87bc83d85c56fe8420

Observation ae97a3f0-c3f1-4b96-9694-406729479488 · outbound

This paper cites Vision- language models for vision tasks: A survey.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models Vision- language models for vision tasks: A survey.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:45.745565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:51:45.745565Z digest=sha256:ae034822b801e481f8d758d6005cbabe7f558ec07f178b6b7caf150914317b83

Observation b8c1f7a8-18fb-4977-97da-77328d528d46 · outbound

This paper cites Ad-vat+: An asymmetric dueling mechanism for learning and understanding visual active tracking.IEEE Transactions on Pattern Analysis and Machine Intelligence, 43(5):1467–1482, 2019.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models Ad-vat+: An asymmetric dueling mechanism for learning and understanding visual active tracking.IEEE Transactions on Pattern Analysis and Machine Intelligence, 43(5):1467–1482, 2019

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:47.772963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:51:45.848692Z digest=sha256:deb22354791450b76b5b23ca05bcb2898499670861ba8871f17e89588bd82530

Observation 03053d95-a475-4b71-9abf-f544c9f0f1a0 · outbound

This paper cites AD-V AT: An asymmetric dueling mechanism for learning visual active tracking.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models AD-V AT: An asymmetric dueling mechanism for learning visual active tracking

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:47.591371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:51:45.913847Z digest=sha256:3cfb161a23e86e98ee890b8464bbe3eae549436e895e431eb150574d5f465fc7

Observation e7dcbe83-8bb0-40c4-ad19-73d0dcef778a · outbound

This paper cites Towards distraction-robust active visual tracking.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models Towards distraction-robust active visual tracking

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:47.403448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:51:45.993230Z digest=sha256:7a08b5ac81fb4b40d6ad6eb6a604fbaa171111e3d4252402fe8bc617d28e1987

Observation 4ecb89fa-7d86-472b-9434-d00fa02bbbab · outbound

This paper cites Empowering embodied visual tracking with visual foundation models and offline rl.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models Empowering embodied visual tracking with visual foundation models and offline rl

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:47.194940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:51:46.111382Z digest=sha256:30ae32bb0651d7f78be1b9c07403246461446d03a33dacb7102e7788dd1f3a68

Observation c0649b78-027a-409b-add2-5304a117d34a · outbound

This paper cites UnrealZoo: Enriching Photo-realistic Virtual Worlds for Embodied AI.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models UnrealZoo: Enriching Photo-realistic Virtual Worlds for Embodied AI

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:46.216740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:51:46.216740Z digest=sha256:56ec956729967272e3e47bd3f37e40c5c8fd6d4b2be9ed269968c7f47fd0e6f3

Observation 7286d023-c9cd-4e30-b802-7eceb1b93786 · outbound

This paper cites On deep recurrent reinforcement learning for active visual tracking of space noncooperative objects.IEEE Robotics and Automation Letters, 8(8):4418–4425, 2023.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models On deep recurrent reinforcement learning for active visual tracking of space noncooperative objects.IEEE Robotics and Automation Letters, 8(8):4418–4425, 2023

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:47.054337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:51:46.295861Z digest=sha256:282ce5fd0a4681dd2473c370888ad26dca4acab0a0da11790f3aa2acdbd537b7

Observation 571ac838-6b22-4f3d-9760-004a84501875 · outbound

This paper cites Navgpt-2: Unleashing navigational reasoning capability for large vision-language models.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models Navgpt-2: Unleashing navigational reasoning capability for large vision-language models

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:46.915857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:51:46.426406Z digest=sha256:3bcb81a445d92d381c96c71f74974f717759030492abb5e7f0ce8a4d5256de9e

Pith citing papers

No inbound Pith citation observations are available.