Pith. sign in

Paper Citation Record · LEDGER

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models

As of 15 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 0 inbound Pith citation observations for arXiv:2505.20718.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.20718 v2

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:51:46.426406Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

32 of 32 outbound references displayed

  • verified exact1
  • verified fuzzy24
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 54c4669a-2f11-403f-baaa-d4c2a298f02a · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:43.474536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:51:43.474536Z digest=sha256:7ddddae5c844ef45804f66e101b383aa75ffad3fcdb21903ff9c801157d768aa

Observation 6ef6bfa1-9b6e-4479-8d86-5dd33924f34b · outbound

This paper cites Spatialvlm: Endowing vision-language models with spatial reasoning capabilities.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models Spatialvlm: Endowing vision-language models with spatial reasoning capabilities

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:43.540009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:51:43.540009Z digest=sha256:cd95b022dbf64acb65b1cf93e864415d820b400a809db85432c285eea6209b9d

Observation 31cb6773-2e01-4d29-8031-b6364f4d8287 · outbound

This paper cites Tracking anything with decoupled video segmenta- tion.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models Tracking anything with decoupled video segmenta- tion

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:50.969592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:51:43.644671Z digest=sha256:5e37cebff5f6ad299d37957a5cefc0700de5bd5592e74cfe7dd5969f86f3d82e

Observation 5b26f5f9-35cb-44d4-9ce4-4137ccf72add · outbound

This paper cites an unresolved cited work.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:51:50.855505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:51:43.761418Z digest=sha256:861e722deaf83a260e6740b08f384c992a07edf2d5675c930a131775718f33f6

Observation 704b213d-417b-4f9a-bad5-8c4c6730ff8b · outbound

This paper cites Proactive multi-camera collaboration for 3d human pose estimation.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models Proactive multi-camera collaboration for 3d human pose estimation

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:50.692694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:51:43.901893Z digest=sha256:84194cc1200e4ad76cb047d2e5b85bf2a946b8de61b6b0c7b7f790c802f90625

Observation b18b71a2-8040-4c01-aef6-96f5211b6d6d · outbound

This paper cites Enhancing continuous control of mobile robots for end-to-end visual active tracking.Robotics and Autonomous Systems, 142:103799, 2021.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models Enhancing continuous control of mobile robots for end-to-end visual active tracking.Robotics and Autonomous Systems, 142:103799, 2021

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:50.605703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:51:44.059949Z digest=sha256:b06993bce917006bcfc4013fe85f5672b1a25e9f84735d5852fab737b7fd2740

Observation f6d464b0-ed64-428e-939c-1a569fac0b7d · outbound

This paper cites E-vat: An asymmetric end-to-end approach to visual active exploration and tracking.IEEE Robotics and Automation Letters, 7(2):4259–4266, 2022.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models E-vat: An asymmetric end-to-end approach to visual active exploration and tracking.IEEE Robotics and Automation Letters, 7(2):4259–4266, 2022

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:50.455798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:51:44.177272Z digest=sha256:8db4531d4bf2d7d20b90c4f4a8dd78eb60d6fb2967037d5749620bd6b8e1cc4a

Observation cbf17c7f-be72-4e1e-949f-5826f5d84ac8 · outbound

This paper cites D-vat: End-to-end visual active tracking for micro aerial vehicles.IEEE Robotics and Automation Letters, 2024.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models D-vat: End-to-end visual active tracking for micro aerial vehicles.IEEE Robotics and Automation Letters, 2024

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:50.324369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:51:44.299789Z digest=sha256:497a32ee8bafb412fc26445f34c0caa4a25749080020c11b819215e308ebc57e

Observation 16a841a7-1f69-462c-b24b-69874cc65de2 · outbound

This paper cites Memory sharing for large language model based agents.Arxiv Preprint Arxiv:2404.09982, 2024.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models Memory sharing for large language model based agents.Arxiv Preprint Arxiv:2404.09982, 2024

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:44.418085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:51:44.418085Z digest=sha256:97ca441e497cede1cadc9f68bfb7c690ea6cea4a64e18ad27098262ad082ccf7

Observation 93964576-1afd-41f8-b39a-fd8b7be5de22 · outbound

This paper cites An embodied generalist agent in 3d world.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models An embodied generalist agent in 3d world

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:50.182089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:51:44.500183Z digest=sha256:47e97b203bd36dc8ac82ad7f3f2284c7e86c5e6eabe5d928bc7b9b113805651c

Observation 6abaf5df-2d3c-4371-b70e-d389b36ce974 · outbound

This paper cites Conquering Ghosts: Relation Learning for Information Reliability Representation and End-to-End Robust Navigation.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models Conquering Ghosts: Relation Learning for Information Reliability Representation and End-to-End Robust Navigation

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:51:46.630634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:51:44.571948Z digest=sha256:de14a72b7861e48b6c73e2c2141785e30063f21839bb99e323da94507509334f

Observation 55664fc4-dc56-47dd-83c1-d2bd62cba002 · outbound

This paper cites OpenVLA: An open-source vision-language-action model.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models OpenVLA: An open-source vision-language-action model

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:50.001978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:51:44.643669Z digest=sha256:492b1518cc25c549f9a10e9f3fbe648113e09e6b4d133e3dd212951018ceb6d6

Observation 7f395053-766f-41b7-ad25-7da0c9e89a57 · outbound

This paper cites A novel performance evaluation methodology for single-target trackers.IEEE Transactions on Pattern Analysis and Machine Intelligence, 38(11):2137–2155, Nov 2016.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models A novel performance evaluation methodology for single-target trackers.IEEE Transactions on Pattern Analysis and Machine Intelligence, 38(11):2137–2155, Nov 2016

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:49.824518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:51:44.737506Z digest=sha256:918762bdca708000109758f625a057a216ff602874bf64e79a7f98936191b6c5

Observation 97273674-23de-4b91-ae61-801372c2e4ee · outbound

This paper cites Person following robot based on real time single object tracking and rgb-d image.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models Person following robot based on real time single object tracking and rgb-d image

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:49.727647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:51:44.829613Z digest=sha256:6e25c53e3ab64b7e9b736aa37aa7e29269a728df9b7b6cff37ce8f7708ead344

Observation 225dd818-5576-43b6-8d35-e74c33fcbff8 · outbound

This paper cites Blip-2: Boot- strapping language-image pre-training with frozen image encoders and large language models.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models Blip-2: Boot- strapping language-image pre-training with frozen image encoders and large language models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:49.536692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:51:44.927307Z digest=sha256:fe95a0061f4c372c88ec4ce513416e8b5fd7d339ec289ec944ff4d77962a0cb5

Observation 4417db72-92aa-45cc-8bf4-bc54c8dfe350 · outbound

This paper cites Vi- sual instruction tuning.Advances in Neural Information Processing Systems, 36:34892–34916, 2023.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models Vi- sual instruction tuning.Advances in Neural Information Processing Systems, 36:34892–34916, 2023

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:49.380606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:51:45.026284Z digest=sha256:8a2c1551122a269c83fbaebe3a38b109d7db4c9ca9a12ac5b0f48a6d12c074cc

Observation 6f02cfff-2347-4e62-8510-e8c41d2acedc · outbound

This paper cites End-to-end active object tracking via reinforcement learning.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models End-to-end active object tracking via reinforcement learning

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:49.120498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:51:45.123410Z digest=sha256:287ec67020753a5ffb271807e7ce67ce41cf8950d00161c8edefca4a51c7a1ef

Observation 093d7d60-5442-4e70-97ac-08a00379108d · outbound

This paper cites Curious george: An attentive semantic robot.Robotics and Autonomous Systems, 56(6):503–511, 2008.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models Curious george: An attentive semantic robot.Robotics and Autonomous Systems, 56(6):503–511, 2008

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:48.945402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:51:45.249150Z digest=sha256:40c857e203f67be0eb85cd7fd7720374ba4757d236b4fb66cb1ebbc7e9e312c5

Observation 2ccb200f-9763-4a12-b6f4-52931b22ea36 · outbound

This paper cites The hands-free push-cart: Autonomous following in front by predicting user trajectory around obstacles.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models The hands-free push-cart: Autonomous following in front by predicting user trajectory around obstacles

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:48.787461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:51:45.348257Z digest=sha256:99422d6d872feebd06c7414930b6d77c649b64023017f8e704298c7fc36dfc9b

Observation fa7968d5-f715-4a43-bfc7-33a0b93f4ed8 · outbound

This paper cites Unrealcv: Virtual worlds for computer vision.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models Unrealcv: Virtual worlds for computer vision

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:48.557921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:51:45.423214Z digest=sha256:fec550a312d5636ad76ecd4e89fa2fdd3582ffb6fab93ec5b815486bd785e6ae

Observation a0c6ebb1-b3d7-45f0-a4e9-0f8eadb2f2ef · outbound

This paper cites Tracking multiple moving targets with a mobile robot using particle filters and statistical data association.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models Tracking multiple moving targets with a mobile robot using particle filters and statistical data association

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:48.406896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:51:45.501658Z digest=sha256:6df03041ab22738eb35c6f0905f4524e1cf3789895090a66ba9d6b2cb92079e4

Observation f929f2e8-5a5c-43f3-b8ed-543c40b5b99e · outbound

This paper cites Accurate and real-time 3-d tracking for the following robots by fusing vision and ultrasonar information.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models Accurate and real-time 3-d tracking for the following robots by fusing vision and ultrasonar information

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:48.185872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:51:45.573732Z digest=sha256:9f67fd7df07f970d3f90c15544d082ae4941908695de32ba998fbdac5f502e31

Observation f698e522-b59a-4366-81c7-c3897d833e31 · outbound

This paper cites Vlfm: Vision-language frontier maps for zero-shot semantic navigation.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models Vlfm: Vision-language frontier maps for zero-shot semantic navigation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:47.971967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:51:45.633871Z digest=sha256:34d955a3ef9090e733f9bf349826cae1120abf94b359120a1ad1227678449500

Observation bf9ea878-827c-4e9a-801d-5334361f5f9f · outbound

This paper cites NaVid: Video-based VLM Plans the Next Step for Vision-and-Language Navigation.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models NaVid: Video-based VLM Plans the Next Step for Vision-and-Language Navigation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:45.712882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:51:45.712882Z digest=sha256:b9f07d8d7de7e0ee7d185b65a150fbbd36028312e083fa87bc83d85c56fe8420

Observation ae97a3f0-c3f1-4b96-9694-406729479488 · outbound

This paper cites Vision- language models for vision tasks: A survey.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models Vision- language models for vision tasks: A survey.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:45.745565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:51:45.745565Z digest=sha256:ae034822b801e481f8d758d6005cbabe7f558ec07f178b6b7caf150914317b83

Observation b8c1f7a8-18fb-4977-97da-77328d528d46 · outbound

This paper cites Ad-vat+: An asymmetric dueling mechanism for learning and understanding visual active tracking.IEEE Transactions on Pattern Analysis and Machine Intelligence, 43(5):1467–1482, 2019.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models Ad-vat+: An asymmetric dueling mechanism for learning and understanding visual active tracking.IEEE Transactions on Pattern Analysis and Machine Intelligence, 43(5):1467–1482, 2019

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:47.772963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:51:45.848692Z digest=sha256:33aed594cee6490d31806caa0a044f24eb3ca7d022e044e3116a5495e96b2b4b

Observation 03053d95-a475-4b71-9abf-f544c9f0f1a0 · outbound

This paper cites AD-V AT: An asymmetric dueling mechanism for learning visual active tracking.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models AD-V AT: An asymmetric dueling mechanism for learning visual active tracking

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:47.591371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:51:45.913847Z digest=sha256:fd2cfcde89463fdcbb4999b410ff41960fc8cb890b6d77c4f0a46c3603396066

Observation e7dcbe83-8bb0-40c4-ad19-73d0dcef778a · outbound

This paper cites Towards distraction-robust active visual tracking.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models Towards distraction-robust active visual tracking

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:47.403448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:51:45.993230Z digest=sha256:192c75005682d03c9dc160da9ce5036b55c5521e70954f1fa23f03cb6a9301bd

Observation 4ecb89fa-7d86-472b-9434-d00fa02bbbab · outbound

This paper cites Empowering embodied visual tracking with visual foundation models and offline rl.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models Empowering embodied visual tracking with visual foundation models and offline rl

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:47.194940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:51:46.111382Z digest=sha256:bfd938096ac2ed546312f280c3e0807410e5132d1ab692853e53f54006b13f11

Observation c0649b78-027a-409b-add2-5304a117d34a · outbound

This paper cites UnrealZoo: Enriching Photo-realistic Virtual Worlds for Embodied AI.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models UnrealZoo: Enriching Photo-realistic Virtual Worlds for Embodied AI

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:46.216740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:51:46.216740Z digest=sha256:56ec956729967272e3e47bd3f37e40c5c8fd6d4b2be9ed269968c7f47fd0e6f3

Observation 7286d023-c9cd-4e30-b802-7eceb1b93786 · outbound

This paper cites On deep recurrent reinforcement learning for active visual tracking of space noncooperative objects.IEEE Robotics and Automation Letters, 8(8):4418–4425, 2023.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models On deep recurrent reinforcement learning for active visual tracking of space noncooperative objects.IEEE Robotics and Automation Letters, 8(8):4418–4425, 2023

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:47.054337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:51:46.295861Z digest=sha256:c587f2816c86e88a441d5c9ef1c3537bb8ef7bc10a7519be1d66d8df8828d7f9

Observation 571ac838-6b22-4f3d-9760-004a84501875 · outbound

This paper cites Navgpt-2: Unleashing navigational reasoning capability for large vision-language models.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models Navgpt-2: Unleashing navigational reasoning capability for large vision-language models

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:46.915857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:51:46.426406Z digest=sha256:9e76d893a4fdbd031f2359d5e9890d31bfc29d0ffae5e24ec64fe26f686eb618

Pith citing papers

No inbound Pith citation observations are available.