Pith. sign in

Paper Citation Record · LEDGER

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models

As of 10 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 0 inbound Pith citation observations for arXiv:2505.20718.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.20718 v2

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:51:46.426406Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

32 of 32 outbound references displayed

  • verified exact1
  • verified fuzzy24
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 54c4669a-2f11-403f-baaa-d4c2a298f02a · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:43.474536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:51:43.474536Z digest=sha256:e885da5c3df047c2b7f332167dff810e6e7d16a0c4cb6539ec0b92f8a85ecc11

Observation 6ef6bfa1-9b6e-4479-8d86-5dd33924f34b · outbound

This paper cites Spatialvlm: Endowing vision-language models with spatial reasoning capabilities.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models Spatialvlm: Endowing vision-language models with spatial reasoning capabilities

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:43.540009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:51:43.540009Z digest=sha256:9484a3ea24aed78a4e7152d7cb228a000258f19cf7595784f8dd7db4e2b3da14

Observation 31cb6773-2e01-4d29-8031-b6364f4d8287 · outbound

This paper cites Tracking anything with decoupled video segmenta- tion.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models Tracking anything with decoupled video segmenta- tion

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:50.969592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:51:43.644671Z digest=sha256:1ee94f3234a1d840721336924e359611816104d1a006c92e5fc6271e70021117

Observation 5b26f5f9-35cb-44d4-9ce4-4137ccf72add · outbound

This paper cites an unresolved cited work.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:51:50.855505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:51:43.761418Z digest=sha256:db4a34a5220a42c88b0ffc1cddec5a6141dc27b1bc978130d5c8b29bbf00f660

Observation 704b213d-417b-4f9a-bad5-8c4c6730ff8b · outbound

This paper cites Proactive multi-camera collaboration for 3d human pose estimation.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models Proactive multi-camera collaboration for 3d human pose estimation

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:50.692694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:51:43.901893Z digest=sha256:383c767e1ea4df9d85bdca15151724a0f2591130ade807306fd82baebc1a0813

Observation b18b71a2-8040-4c01-aef6-96f5211b6d6d · outbound

This paper cites Enhancing continuous control of mobile robots for end-to-end visual active tracking.Robotics and Autonomous Systems, 142:103799, 2021.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models Enhancing continuous control of mobile robots for end-to-end visual active tracking.Robotics and Autonomous Systems, 142:103799, 2021

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:50.605703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:51:44.059949Z digest=sha256:4c9d51f8f6635af7f1eb1f5ef6ec0978346503b0eb94c3ad6aa8a316d94bf70d

Observation f6d464b0-ed64-428e-939c-1a569fac0b7d · outbound

This paper cites E-vat: An asymmetric end-to-end approach to visual active exploration and tracking.IEEE Robotics and Automation Letters, 7(2):4259–4266, 2022.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models E-vat: An asymmetric end-to-end approach to visual active exploration and tracking.IEEE Robotics and Automation Letters, 7(2):4259–4266, 2022

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:50.455798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:51:44.177272Z digest=sha256:158e0dafca73877b4bb181457852829fc163b942b60e12b92708e79440308bf2

Observation cbf17c7f-be72-4e1e-949f-5826f5d84ac8 · outbound

This paper cites D-vat: End-to-end visual active tracking for micro aerial vehicles.IEEE Robotics and Automation Letters, 2024.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models D-vat: End-to-end visual active tracking for micro aerial vehicles.IEEE Robotics and Automation Letters, 2024

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:50.324369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:51:44.299789Z digest=sha256:51bf28daca59069ab621ad4e29f4020c0007cf8fcbd2eb297649cb271f7f99a0

Observation 16a841a7-1f69-462c-b24b-69874cc65de2 · outbound

This paper cites Memory sharing for large language model based agents.Arxiv Preprint Arxiv:2404.09982, 2024.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models Memory sharing for large language model based agents.Arxiv Preprint Arxiv:2404.09982, 2024

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:44.418085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:51:44.418085Z digest=sha256:b8789d27054759b5410c533d21bf43e49f3f413113a85ef93ab6bb077bc8e50c

Observation 93964576-1afd-41f8-b39a-fd8b7be5de22 · outbound

This paper cites An embodied generalist agent in 3d world.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models An embodied generalist agent in 3d world

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:50.182089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:51:44.500183Z digest=sha256:e425dbc28ce7cc0b6626274f3410eeef82916d06f0c162f8ae3a901dadaa537a

Observation 6abaf5df-2d3c-4371-b70e-d389b36ce974 · outbound

This paper cites Conquering Ghosts: Relation Learning for Information Reliability Representation and End-to-End Robust Navigation.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models Conquering Ghosts: Relation Learning for Information Reliability Representation and End-to-End Robust Navigation

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:51:46.630634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:51:44.571948Z digest=sha256:5a680858be9260dc399eb5cc952234aceaec71a343e9e82955034f2307cd06bc

Observation 55664fc4-dc56-47dd-83c1-d2bd62cba002 · outbound

This paper cites OpenVLA: An open-source vision-language-action model.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models OpenVLA: An open-source vision-language-action model

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:50.001978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:51:44.643669Z digest=sha256:b4d7f0595f16abd5825beb32f845778b3d75a4d36ea48c446f3ef5a8bbc4d2a9

Observation 7f395053-766f-41b7-ad25-7da0c9e89a57 · outbound

This paper cites A novel performance evaluation methodology for single-target trackers.IEEE Transactions on Pattern Analysis and Machine Intelligence, 38(11):2137–2155, Nov 2016.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models A novel performance evaluation methodology for single-target trackers.IEEE Transactions on Pattern Analysis and Machine Intelligence, 38(11):2137–2155, Nov 2016

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:49.824518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:51:44.737506Z digest=sha256:880cc71cf54583d68e4ff038df38a0a89af61504d4dd113619e8e4894df8193b

Observation 97273674-23de-4b91-ae61-801372c2e4ee · outbound

This paper cites Person following robot based on real time single object tracking and rgb-d image.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models Person following robot based on real time single object tracking and rgb-d image

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:49.727647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:51:44.829613Z digest=sha256:f1d0948c7e444d95490e0629f88b6095da5462c27be8ecb0d7e3d416470496ac

Observation 225dd818-5576-43b6-8d35-e74c33fcbff8 · outbound

This paper cites Blip-2: Boot- strapping language-image pre-training with frozen image encoders and large language models.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models Blip-2: Boot- strapping language-image pre-training with frozen image encoders and large language models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:49.536692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:51:44.927307Z digest=sha256:d93881388407d23f0d85010acac5bbd7a0f0b8e3e5864606876ca9b5af0f4b97

Observation 4417db72-92aa-45cc-8bf4-bc54c8dfe350 · outbound

This paper cites Vi- sual instruction tuning.Advances in Neural Information Processing Systems, 36:34892–34916, 2023.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models Vi- sual instruction tuning.Advances in Neural Information Processing Systems, 36:34892–34916, 2023

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:49.380606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:51:45.026284Z digest=sha256:833f3736f41cc0f46e2dabc699cd04387b4443fc54f2d67ac251717d559627a9

Observation 6f02cfff-2347-4e62-8510-e8c41d2acedc · outbound

This paper cites End-to-end active object tracking via reinforcement learning.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models End-to-end active object tracking via reinforcement learning

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:49.120498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:51:45.123410Z digest=sha256:6a0bda26f6a1ef285f7ccb1829c858a4732753b2d14ac4ee0a356fe35aa24370

Observation 093d7d60-5442-4e70-97ac-08a00379108d · outbound

This paper cites Curious george: An attentive semantic robot.Robotics and Autonomous Systems, 56(6):503–511, 2008.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models Curious george: An attentive semantic robot.Robotics and Autonomous Systems, 56(6):503–511, 2008

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:48.945402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:51:45.249150Z digest=sha256:627740a10f1f05e4a575dbeeb4f2ca8b6111976e70989214db5da29450f51835

Observation 2ccb200f-9763-4a12-b6f4-52931b22ea36 · outbound

This paper cites The hands-free push-cart: Autonomous following in front by predicting user trajectory around obstacles.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models The hands-free push-cart: Autonomous following in front by predicting user trajectory around obstacles

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:48.787461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:51:45.348257Z digest=sha256:2ce46292b0902faa8652bd7939e04a3188c505bc324a2459066007cf4d1c6cd9

Observation fa7968d5-f715-4a43-bfc7-33a0b93f4ed8 · outbound

This paper cites Unrealcv: Virtual worlds for computer vision.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models Unrealcv: Virtual worlds for computer vision

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:48.557921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:51:45.423214Z digest=sha256:277b075a4a8a678d32bc877ab9d7f583c44f328f6f93173e2d65e23af3d56c3f

Observation a0c6ebb1-b3d7-45f0-a4e9-0f8eadb2f2ef · outbound

This paper cites Tracking multiple moving targets with a mobile robot using particle filters and statistical data association.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models Tracking multiple moving targets with a mobile robot using particle filters and statistical data association

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:48.406896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:51:45.501658Z digest=sha256:f1fb5bd19c8d8b7cb4a76cddcc3f73ae5c8df49747dc22948eea64044dc57e60

Observation f929f2e8-5a5c-43f3-b8ed-543c40b5b99e · outbound

This paper cites Accurate and real-time 3-d tracking for the following robots by fusing vision and ultrasonar information.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models Accurate and real-time 3-d tracking for the following robots by fusing vision and ultrasonar information

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:48.185872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:51:45.573732Z digest=sha256:c31ba9a99fdf94463af50fe2d0ac9edcd1ce4e1e140080475006e3b9873a9159

Observation f698e522-b59a-4366-81c7-c3897d833e31 · outbound

This paper cites Vlfm: Vision-language frontier maps for zero-shot semantic navigation.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models Vlfm: Vision-language frontier maps for zero-shot semantic navigation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:47.971967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:51:45.633871Z digest=sha256:fdc1dbed0a889764a2f9fb29c7924163229cf9fe5aba53d8db36532449ffb10f

Observation bf9ea878-827c-4e9a-801d-5334361f5f9f · outbound

This paper cites NaVid: Video-based VLM Plans the Next Step for Vision-and-Language Navigation.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models NaVid: Video-based VLM Plans the Next Step for Vision-and-Language Navigation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:45.712882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:51:45.712882Z digest=sha256:73ec54133eb1b142ae33ee0f609a5799f46b6c4d6c2cab4d0769dc023ad4f40e

Observation ae97a3f0-c3f1-4b96-9694-406729479488 · outbound

This paper cites Vision- language models for vision tasks: A survey.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models Vision- language models for vision tasks: A survey.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:45.745565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:51:45.745565Z digest=sha256:997958f5e3861304e0583b790829af6c63e2dca97491b685910fc97001c52bd6

Observation b8c1f7a8-18fb-4977-97da-77328d528d46 · outbound

This paper cites Ad-vat+: An asymmetric dueling mechanism for learning and understanding visual active tracking.IEEE Transactions on Pattern Analysis and Machine Intelligence, 43(5):1467–1482, 2019.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models Ad-vat+: An asymmetric dueling mechanism for learning and understanding visual active tracking.IEEE Transactions on Pattern Analysis and Machine Intelligence, 43(5):1467–1482, 2019

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:47.772963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:51:45.848692Z digest=sha256:a03f5cb0ce34ff3ff66c8bcf5e7db730e5f35ae9f94cd2db15926b0b80dcad6f

Observation 03053d95-a475-4b71-9abf-f544c9f0f1a0 · outbound

This paper cites AD-V AT: An asymmetric dueling mechanism for learning visual active tracking.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models AD-V AT: An asymmetric dueling mechanism for learning visual active tracking

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:47.591371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:51:45.913847Z digest=sha256:48f9325d1b1770c1380379cc342d83e92168e7469270aada408bb8d0706ea970

Observation e7dcbe83-8bb0-40c4-ad19-73d0dcef778a · outbound

This paper cites Towards distraction-robust active visual tracking.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models Towards distraction-robust active visual tracking

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:47.403448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:51:45.993230Z digest=sha256:123a80226f3a5b422e54f2cd969ae2d37a8d4a3172194f9907c0221a5f2bdc45

Observation 4ecb89fa-7d86-472b-9434-d00fa02bbbab · outbound

This paper cites Empowering embodied visual tracking with visual foundation models and offline rl.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models Empowering embodied visual tracking with visual foundation models and offline rl

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:47.194940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:51:46.111382Z digest=sha256:7cc2a7b32dec38bfab7c5a4f52d4a7dc8cd9a22697ead7e3ee8cdf348b0c72ac

Observation c0649b78-027a-409b-add2-5304a117d34a · outbound

This paper cites UnrealZoo: Enriching Photo-realistic Virtual Worlds for Embodied AI.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models UnrealZoo: Enriching Photo-realistic Virtual Worlds for Embodied AI

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:46.216740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:51:46.216740Z digest=sha256:50aba12e7c63a76827c9bee473f88a02748646f72a6a6b5124f7cf87c7d51472

Observation 7286d023-c9cd-4e30-b802-7eceb1b93786 · outbound

This paper cites On deep recurrent reinforcement learning for active visual tracking of space noncooperative objects.IEEE Robotics and Automation Letters, 8(8):4418–4425, 2023.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models On deep recurrent reinforcement learning for active visual tracking of space noncooperative objects.IEEE Robotics and Automation Letters, 8(8):4418–4425, 2023

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:47.054337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:51:46.295861Z digest=sha256:5b4bdb05b02c98b80600f9e926e3c9f7d5daeee2384df5172ce3f167cd5f4aba

Observation 571ac838-6b22-4f3d-9760-004a84501875 · outbound

This paper cites Navgpt-2: Unleashing navigational reasoning capability for large vision-language models.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models Navgpt-2: Unleashing navigational reasoning capability for large vision-language models

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:46.915857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:51:46.426406Z digest=sha256:d4d0de70bfc68b93acf39fab8f6b8a140fa8cf7ae5b140cbd3def2994d82da85

Pith citing papers

No inbound Pith citation observations are available.