Pith. sign in

Paper Citation Record · LEDGER

Visual Grounding in Zero-Shot Vision-Language Control

As of 8 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 0 inbound Pith citation observations for arXiv:2608.06154.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.06154 v1

Coverage vector

measured 44 of 44 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:57:12.082156Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

44 of 44 outbound references displayed

  • verified exact0
  • verified fuzzy25
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a603f8c3-022b-4147-8169-63fe57346953 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Visual Grounding in Zero-Shot Vision-Language Control Learning transferable visual models from natural language supervision,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:57:16.790543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:57:08.270889Z digest=sha256:b6afe79a44ddbfe0d9eb60125312f27098d1c9c9129da6e3f07b4c383f2e0dab

Observation 0e43c405-639a-4c5b-b059-ff170912dad6 · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

Visual Grounding in Zero-Shot Vision-Language Control Improved Baselines with Visual Instruction Tuning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:57:08.301524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:57:08.301524Z digest=sha256:bfb4824ade1693ed4583f0604f697aace6093ba678476744abdfaffc7dbe849b

Observation 4603868a-d4e8-4c73-89f7-1d53b5515177 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Visual Grounding in Zero-Shot Vision-Language Control Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:57:08.373123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:57:08.373123Z digest=sha256:94150d1a0fea41a89b54fbcb37a0dcc8976d70dc4a973d1c74b3e097d4bba3ae

Observation a1a210d9-32bb-4b99-8083-5b762cb6b2ef · outbound

This paper cites Qwen2.5-VL Technical Report.

Visual Grounding in Zero-Shot Vision-Language Control Qwen2.5-VL Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:57:08.442796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:57:08.442796Z digest=sha256:9a30369664f36d93957838c4cd877a370461a392d545352c1e3c03cceb5a4214

Observation dbb8b77a-3727-4987-a03a-e8a06282aa5f · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

Visual Grounding in Zero-Shot Vision-Language Control MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:57:08.553499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:57:08.553499Z digest=sha256:853d337d4b65beca8ede340ebae3fc2a79fe74cb457ee79c719ed6b7c1551aaf

Observation fb1a59d1-ed3d-4bad-9ccb-6af984254bd5 · outbound

This paper cites Qwen3-VL Technical Report.

Visual Grounding in Zero-Shot Vision-Language Control Qwen3-VL Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:57:08.643309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:57:08.643309Z digest=sha256:259b7f0b87edf7df60ccbfbb89e0ea2e24628eff083dba4eec98579339221eb6

Observation 4d55ec77-a452-4a5a-becb-390200596896 · outbound

This paper cites Gemma 4 Technical Report.

Visual Grounding in Zero-Shot Vision-Language Control Gemma 4 Technical Report

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:57:08.729258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:57:08.729258Z digest=sha256:66c6f90137b7040a7b15cde2fae36cdc9e3db856a4e8adbb79d9a3c1db977725

Observation ec3159d7-5976-4240-90d2-7389098270cd · outbound

This paper cites Qwen3.5: Towards native multimodal agents,.

Visual Grounding in Zero-Shot Vision-Language Control Qwen3.5: Towards native multimodal agents,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:57:16.579700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:57:08.814585Z digest=sha256:6799e3bc38e8a52babb836ce36542c09cf656f2af2e560dd3f459974142d5dc1

Observation f585939b-9dd9-4c7b-88d2-52e812ad10b1 · outbound

This paper cites Introducing Mistral 3,.

Visual Grounding in Zero-Shot Vision-Language Control Introducing Mistral 3,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:57:16.388390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:57:08.880443Z digest=sha256:c781c9b29546f2711c1752fa211914b3f0792cf517638091dbacd843fd70dde1

Observation 06420827-4300-4bed-81da-c6f12310a09a · outbound

This paper cites MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training Recipe.

Visual Grounding in Zero-Shot Vision-Language Control MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training Recipe

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:57:08.961126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:57:08.961126Z digest=sha256:f9a2842991674a121e5485db281b2416a52270377d40109b44cf1fe9951feb2e

Observation 9ffb7a4b-5bc3-433f-b84e-657efa32534c · outbound

This paper cites SmolVLM: Redefining small and efficient multimodal models.

Visual Grounding in Zero-Shot Vision-Language Control SmolVLM: Redefining small and efficient multimodal models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:57:09.055707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:57:09.055707Z digest=sha256:f7cf676d6b0707ad1020e65dda2f2e38c0f729841e82fa23f652788882f5eba3

Observation 2e96abb8-7e75-49cf-bad1-44b987d96f15 · outbound

This paper cites NavGPT: Explicit reasoning in vision- and-language navigation with large language models,.

Visual Grounding in Zero-Shot Vision-Language Control NavGPT: Explicit reasoning in vision- and-language navigation with large language models,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:57:16.185974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:57:09.150370Z digest=sha256:1aa1d266dbdd2477f903b16f53830c9fc821705203eba6cc245d53f9296db813

Observation 2ac4a8d0-ddbe-4845-83ab-753e13619f30 · outbound

This paper cites VLM-Social-Nav: Socially aware robot navigation through scoring using vision-language models,.

Visual Grounding in Zero-Shot Vision-Language Control VLM-Social-Nav: Socially aware robot navigation through scoring using vision-language models,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:57:15.988861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:57:09.191891Z digest=sha256:2fd2da500f5fc6311a0e5bf231f0c2c2d0dc1b6ec6d88ba61aa5631352d327f2

Observation 1617f2b0-e8b3-4d04-a32d-24e9620e142a · outbound

This paper cites VLM-GroNav: Robot navigation using physically grounded vision-language models in outdoor environments,.

Visual Grounding in Zero-Shot Vision-Language Control VLM-GroNav: Robot navigation using physically grounded vision-language models in outdoor environments,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:57:15.794638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:57:09.271864Z digest=sha256:84212a6c6751bba3f03621f0e5f15911d9fa0491cc57042e7949ba8a2fae000c

Observation f31e68ee-2539-4c67-bb19-0ef0a3340012 · outbound

This paper cites MapNav: A novel memory representation via annotated semantic maps for VLM-based vision-and-language navigation,.

Visual Grounding in Zero-Shot Vision-Language Control MapNav: A novel memory representation via annotated semantic maps for VLM-based vision-and-language navigation,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:57:15.624658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:57:09.376035Z digest=sha256:17c3188b10fd44ec2b081a72e5ed000217dd2f476dfedce7d48b2a45588d1981

Observation 70fc3451-2ab0-43bf-bfdb-0fc46a2249cf · outbound

This paper cites HazardVLM: A video language model for real-time hazard description in automated driving systems,.

Visual Grounding in Zero-Shot Vision-Language Control HazardVLM: A video language model for real-time hazard description in automated driving systems,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:57:15.405609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:57:09.469350Z digest=sha256:48796cd99a1e4176f035911d168504819ebfc5acd42e549b863aa846a75e2f62

Observation 25951574-ea08-42cf-87f1-209bb847b839 · outbound

This paper cites LLM-powered cooperative perception framework for mixed UA V-vehicle platoons,.

Visual Grounding in Zero-Shot Vision-Language Control LLM-powered cooperative perception framework for mixed UA V-vehicle platoons,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:57:15.178944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:57:09.543585Z digest=sha256:e45741712493a61297b3e59b78dbf556d4f81f4dddc96b2c8c7575be0b9e4ea1

Observation c5278006-8a5e-4c69-99f3-2bcd668bf024 · outbound

This paper cites Semantic scene understand- ing with large language models on unmanned aerial vehicles,.

Visual Grounding in Zero-Shot Vision-Language Control Semantic scene understand- ing with large language models on unmanned aerial vehicles,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:57:14.962984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:57:09.642706Z digest=sha256:6fa358fc95b5833e9b07645552268a7c908c3ff5aaccf36340c0d265462bc18d

Observation 91026761-6ef7-4e28-b7b5-ce3798e118e0 · outbound

This paper cites Shortcut learning in deep neural networks,.

Visual Grounding in Zero-Shot Vision-Language Control Shortcut learning in deep neural networks,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:57:09.715223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:57:09.715223Z digest=sha256:4117d40d8034537fbdb6d605f824b26c98117c99bb4435169e976cd2956249a0

Observation f29b433f-6864-482b-8456-418ec2180865 · outbound

This paper cites Beyond accuracy: Behavioral testing of NLP models with CheckList,.

Visual Grounding in Zero-Shot Vision-Language Control Beyond accuracy: Behavioral testing of NLP models with CheckList,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:57:14.751966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:57:09.835083Z digest=sha256:457da5cb26fc966baef083817dd90245ce1c24dae6bbd17f89145fda6cabb999

Observation ca7c9e58-04a9-4a16-9f56-56cadde12161 · outbound

This paper cites Holistic Evaluation of Language Models.

Visual Grounding in Zero-Shot Vision-Language Control Holistic Evaluation of Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T13:57:09.972585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:57:09.972585Z digest=sha256:a0b3480facfa7054db07df30076e7a6ee08dafc1f096b86ab9fd2585119e76b6

Observation e00c28b0-ce6a-4c24-a854-1d9bf318dca4 · outbound

This paper cites Metamorphic testing: A review of challenges and opportunities,.

Visual Grounding in Zero-Shot Vision-Language Control Metamorphic testing: A review of challenges and opportunities,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:57:10.087742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:57:10.087742Z digest=sha256:40dbdc66f6e2ed46f2d8fa2b1178357d62cc0cd1085c598c30c4bef7ab918493

Observation c77f0cb6-c974-420b-b245-66d362744f41 · outbound

This paper cites DeepTest: Automated testing of DNN-driven autonomous cars,.

Visual Grounding in Zero-Shot Vision-Language Control DeepTest: Automated testing of DNN-driven autonomous cars,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:57:14.544924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:57:10.174166Z digest=sha256:94bca079115d7485a68e5f2044af3e2ca1e4f8de07c9af48d448ca45a2b6ddfe

Observation 7e56556c-58b4-4bdd-8297-6dc57e25b612 · outbound

This paper cites DeepRoad: GAN-based metamorphic testing and input validation framework for autonomous driving systems,.

Visual Grounding in Zero-Shot Vision-Language Control DeepRoad: GAN-based metamorphic testing and input validation framework for autonomous driving systems,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:57:14.289522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:57:10.263345Z digest=sha256:0faebd121ca9f99a69c897f6e77f23bdec4cb4611d19cc45a65a850f2522ffd1

Observation 3d13e243-58ed-4600-a09d-910ea348d10e · outbound

This paper cites Metamorphic testing for semantic invariance in large language models,.

Visual Grounding in Zero-Shot Vision-Language Control Metamorphic testing for semantic invariance in large language models,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:57:14.128045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:57:10.360725Z digest=sha256:33523b2fd673597b62b2b4390da93d36091faae391cd1915d89b2d4e4e9c4ecb

Observation b0ec6d09-d4c1-48cf-8a38-2302de9e0bad · outbound

This paper cites Semantic invariance in agentic AI,.

Visual Grounding in Zero-Shot Vision-Language Control Semantic invariance in agentic AI,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:57:13.966883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:57:10.489405Z digest=sha256:ec61238d6c6e8b4d036ada2ed941c25d3dd4b080f1f9bcb4db99fda0142ea1c1

Observation 30cb1895-c93d-4bb4-a201-3ef5de9543b3 · outbound

This paper cites Energy-aware multilingual evaluation of large language models,.

Visual Grounding in Zero-Shot Vision-Language Control Energy-aware multilingual evaluation of large language models,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:57:13.818647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:57:10.601401Z digest=sha256:c633041f18c2b358827187b95be947971501786a12d3074e02cb8ca10cd73181

Observation d6c3c2ca-c79c-444a-b2a7-7605c6b4ab47 · outbound

This paper cites EdgeShard: Efficient LLM inference via collaborative edge computing,.

Visual Grounding in Zero-Shot Vision-Language Control EdgeShard: Efficient LLM inference via collaborative edge computing,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:57:13.611251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:57:10.719094Z digest=sha256:a9afa2679c0434f9269411ca7fc50eccfd5901f5056c0de749c19dde2282ce49

Observation 225b71da-b3d6-4df4-9682-930d474951cf · outbound

This paper cites Power hungry processing: Watts driving the cost of AI deployment?.

Visual Grounding in Zero-Shot Vision-Language Control Power hungry processing: Watts driving the cost of AI deployment?

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:57:13.454238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:57:10.811054Z digest=sha256:81f4a0c8a61a087d5ec7d24175c98430e2300030e5161d65d3db0e23aedf1804

Observation 87e4c2cd-8ee4-431e-a0bb-8a28e650ca2b · outbound

This paper cites An environment for autonomous driving decision- making,.

Visual Grounding in Zero-Shot Vision-Language Control An environment for autonomous driving decision- making,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:57:13.313176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:57:10.906404Z digest=sha256:937bae4aa0933fea360f1472bbf292d9ed46bf83ae330a900895792a8977020a

Observation ff57a3ee-4a1a-4b0c-a83c-a2dc768af9cd · outbound

This paper cites Learning to fly—a Gym environment with PyBullet physics for reinforcement learning of multi-agent quadcopter control,.

Visual Grounding in Zero-Shot Vision-Language Control Learning to fly—a Gym environment with PyBullet physics for reinforcement learning of multi-agent quadcopter control,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:57:13.146672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:57:11.001892Z digest=sha256:95582014cadbcebc028db50897eb46c0d4ba1193a4becd700ee3f3d0a950e667

Observation 74bdaa22-0965-492e-9f57-c6556a94b078 · outbound

This paper cites Gymnasium: A Standard Interface for Reinforcement Learning Environments.

Visual Grounding in Zero-Shot Vision-Language Control Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T13:57:11.105161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:57:11.105161Z digest=sha256:14cd6796c474e22055fa269621031da4aaa241bf8108f642e792270b1dbc9b47

Observation 7b990891-95c6-4224-a76f-fa9af2c2d912 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large lan- guage models,.

Visual Grounding in Zero-Shot Vision-Language Control Chain-of-thought prompting elicits reasoning in large lan- guage models,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:57:11.192482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:57:11.192482Z digest=sha256:c8b8e351c9a2cecbf8a2204ab9c2b4f8e1490c4f1498a857813fe0313224fedb

Observation e81779ee-264f-462c-bdf0-47724d25e710 · outbound

This paper cites Language models don’t always say what they think: Unfaithful explanations in chain- of-thought prompting,.

Visual Grounding in Zero-Shot Vision-Language Control Language models don’t always say what they think: Unfaithful explanations in chain- of-thought prompting,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:57:13.001615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:57:11.303725Z digest=sha256:025bbfd36279d42a72020c42b1900318e15a650c358aa38ed88f603d9667eb38

Observation 4ec7dd12-c123-451e-b0c2-20d72d6cdf58 · outbound

This paper cites Measuring Faithfulness in Chain-of-Thought Reasoning.

Visual Grounding in Zero-Shot Vision-Language Control Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:57:11.374361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:57:11.374361Z digest=sha256:b40ae601105b71ff6b4b88939a4f52e95f223c0df0d0810ad7a76ef36d098c7f

Observation db19dba4-69c2-4470-8fa9-744f8721c355 · outbound

This paper cites Do Prompt-Based Models Really Understand the Meaning of their Prompts?.

Visual Grounding in Zero-Shot Vision-Language Control Do Prompt-Based Models Really Understand the Meaning of their Prompts?

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T13:57:11.478970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:57:11.478970Z digest=sha256:57bc94b67915ce599fb6789006261e27c4450fd04287bd053bafe71299c75362

Observation d7353651-c3b0-43ac-b1a7-9db0d51c80dc · outbound

This paper cites PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts.

Visual Grounding in Zero-Shot Vision-Language Control PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T13:57:11.549507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:57:11.549507Z digest=sha256:e8fe5c6005668dd4005f7aabab6296859f5599d276864c16e16a8759b6424c92

Observation 8ebd7ffe-3c10-49de-9ef6-83dd799032d7 · outbound

This paper cites Measuring and improving consistency in pretrained language models,.

Visual Grounding in Zero-Shot Vision-Language Control Measuring and improving consistency in pretrained language models,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:57:12.854441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:57:11.628585Z digest=sha256:04c3d19965efafe3cec12f0e3bd57f954f5b68f35e6a71e2a9e82d77ee9fe426

Observation 0cfbf370-8663-4870-bbf6-caf7a772d722 · outbound

This paper cites Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order Sensitivity.

Visual Grounding in Zero-Shot Vision-Language Control Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order Sensitivity

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T13:57:11.698424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:57:11.698424Z digest=sha256:c89b0c54266dee7c24ddd37e2dfe1a9b8c031b923eeed34d01060bf585ca38ec

Observation 57408b9b-37d8-41a9-ba60-92e3229aad1c · outbound

This paper cites Probing classifiers: Promises, shortcomings, and ad- vances,.

Visual Grounding in Zero-Shot Vision-Language Control Probing classifiers: Promises, shortcomings, and ad- vances,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:57:12.691304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:57:11.782731Z digest=sha256:813561f475f234b5ee1fbe98b76c82930c9bc39e658321a3f348c6134d9783b1

Observation 3e1d8560-31a1-4c9a-b891-4f3e30f6441a · outbound

This paper cites Con- strained model predictive control: Stability and optimality,.

Visual Grounding in Zero-Shot Vision-Language Control Con- strained model predictive control: Stability and optimality,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T13:57:11.866001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:57:11.866001Z digest=sha256:6f902216251973966f26d64e121ea48d459595880678f989122c8f131d03d823

Observation bc9e8cc5-c59c-41d8-881c-0040b5fd8e28 · outbound

This paper cites A simple sequentially rejective multiple test procedure,.

Visual Grounding in Zero-Shot Vision-Language Control A simple sequentially rejective multiple test procedure,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T13:57:11.927436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:57:11.927436Z digest=sha256:e6d0c3827fd08f618bd21c49a36275f35ba9ff6e10c3c05cc63a0a98c8cb1586

Observation 0a71f3cc-ccf6-462e-aee2-b81059cc820c · outbound

This paper cites SciPy 1.0: Fundamental algorithms for scientific computing in Python,.

Visual Grounding in Zero-Shot Vision-Language Control SciPy 1.0: Fundamental algorithms for scientific computing in Python,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:57:12.541497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:57:12.001667Z digest=sha256:6dc17b53fe1cd676d455fa37f73b50552d1065b1ee18141dcccfcbdd3d70080d

Observation f6350d25-8935-4296-a0fa-a42e0757c57c · outbound

This paper cites Transformers: State-of-the-art natural language pro- cessing,.

Visual Grounding in Zero-Shot Vision-Language Control Transformers: State-of-the-art natural language pro- cessing,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:57:12.424487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:57:12.082156Z digest=sha256:44f67efdd89a0bc0224a9cda44577b63ca2f2ed84ab33bb6a4d4022703b53d32

Pith citing papers

No inbound Pith citation observations are available.