Pith. sign in

Paper Citation Record · LEDGER

Hierarchical Question-Answering for Driving Scene Understanding Using Vision-Language Models

As of 15 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2506.02615.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.02615 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:23:14.558335Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact1
  • verified fuzzy15
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b0cbb2dd-c1e8-4c7c-8218-59c0e5f81b4c · outbound

This paper cites Spherical transformer for lidar-based 3d recognition,.

Hierarchical Question-Answering for Driving Scene Understanding Using Vision-Language Models Spherical transformer for lidar-based 3d recognition,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:23:18.242083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:23:12.510042Z digest=sha256:b89261e8a71704acf0e063ea39f2ff986e88bfcd047fdc695987a1c654a145ae

Observation db2325a5-b133-499e-b29b-043b829e37ba · outbound

This paper cites Rea- son2drive: Towards interpretable and chain-based reasoning for au- tonomous driving,.

Hierarchical Question-Answering for Driving Scene Understanding Using Vision-Language Models Rea- son2drive: Towards interpretable and chain-based reasoning for au- tonomous driving,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:23:17.994358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:23:12.552518Z digest=sha256:060d36dda687cffc9aa08766f284494b2f8bfe8b6d1846faf10808c9c4d20fc4

Observation ebd7634f-7846-439d-a3ec-c488704bf2aa · outbound

This paper cites GPT-4V as Traffic Assistant: An In-depth Look at Vision Language Model on Complex Traffic Events.

Hierarchical Question-Answering for Driving Scene Understanding Using Vision-Language Models GPT-4V as Traffic Assistant: An In-depth Look at Vision Language Model on Complex Traffic Events

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:23:12.643347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:23:12.643347Z digest=sha256:acf9d58f7123e13d80db00eccf532f80a76a76377e5b5599e2fd40f97eac486f

Observation b21e0480-4d1f-4f23-b446-a5dcb535547a · outbound

This paper cites DriveVLM: The convergence of autonomous driving and large vision-language models,.

Hierarchical Question-Answering for Driving Scene Understanding Using Vision-Language Models DriveVLM: The convergence of autonomous driving and large vision-language models,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:23:17.756695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:23:12.700488Z digest=sha256:3bc4394d0263b42a0e3a3cd77870b44952f1893c65a9057223c0f119d1e2806f

Observation a2ca9fb8-e7f4-4487-acec-05afb4933ac4 · outbound

This paper cites BLIP: Bootstrapping language- image pre-training for unified vision-language understanding and generation,.

Hierarchical Question-Answering for Driving Scene Understanding Using Vision-Language Models BLIP: Bootstrapping language- image pre-training for unified vision-language understanding and generation,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:23:17.438610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:23:12.731307Z digest=sha256:f73a345239ef62044da12b3ddecef3703715e938fa7bca196432f0cd50c1277e

Observation 6963edf6-5390-4462-8e3d-3c80c5aa53aa · outbound

This paper cites Clip2scene: Towards label-efficient 3d scene understanding by clip,.

Hierarchical Question-Answering for Driving Scene Understanding Using Vision-Language Models Clip2scene: Towards label-efficient 3d scene understanding by clip,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:23:17.195480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:23:12.811578Z digest=sha256:620431b14de977ed6ad363957afd6cab4ecdaf9777f28c08f625d6da13d6268a

Observation 597f786b-fe0c-447d-9e0e-11ca408ee3c9 · outbound

This paper cites Jiang, A. Tagliasacchi, M. Pollefeys, and T. Funkhouser, “Openscene: 3d scene understanding with open vocab- ularies,.

Hierarchical Question-Answering for Driving Scene Understanding Using Vision-Language Models Jiang, A. Tagliasacchi, M. Pollefeys, and T. Funkhouser, “Openscene: 3d scene understanding with open vocab- ularies,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:23:16.935102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:23:12.896521Z digest=sha256:de611ab012d24139f8ded781c2feb6a0a1bfc93b62b90f51c2842ce9f9e0d1f8

Observation e316f44c-cf65-42b6-8a72-d5771e19f43d · outbound

This paper cites Pla: Language-driven open-vocabulary 3d scene understanding,.

Hierarchical Question-Answering for Driving Scene Understanding Using Vision-Language Models Pla: Language-driven open-vocabulary 3d scene understanding,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:23:16.597432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:23:13.020546Z digest=sha256:7b165c605174231da88103768a779ae7bb3fb6b9b7c89a610d46e69f9925f092

Observation 8e082adb-b2e0-4caf-8991-74438cb51821 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Hierarchical Question-Answering for Driving Scene Understanding Using Vision-Language Models Learning transferable visual models from natural language supervision,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:23:13.119412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:23:13.119412Z digest=sha256:18e53b3015e53f8ae0c8e93186788d75ba59320ba13a6b71fd5985645a69f530

Observation 0c7d06ac-d4f1-4a68-9d5a-24fd23cf89eb · outbound

This paper cites The traffic scene un- derstanding and prediction based on image captioning,.

Hierarchical Question-Answering for Driving Scene Understanding Using Vision-Language Models The traffic scene un- derstanding and prediction based on image captioning,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:23:16.384009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:23:13.240502Z digest=sha256:b075fc1296d0456491f117bb54fed6f7b05d4ab21d4abc86af0928d23e9eecd5

Observation 101ebe9c-9183-4f49-82aa-f422d7fb56a6 · outbound

This paper cites Delving into clip latent space for video anomaly recognition,.

Hierarchical Question-Answering for Driving Scene Understanding Using Vision-Language Models Delving into clip latent space for video anomaly recognition,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:23:16.108961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:23:13.339597Z digest=sha256:a195bfc3f82af8f0d767ada6e41349a555e78065b2a381a817bae8d806c9ef42

Observation 212ca42e-a757-42f2-ac76-e26492b69d59 · outbound

This paper cites Learning to prompt for vision-language models,.

Hierarchical Question-Answering for Driving Scene Understanding Using Vision-Language Models Learning to prompt for vision-language models,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:23:13.482123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:23:13.482123Z digest=sha256:e4446af61ba30186a7cb1a9479ac1246b65b7b0e5ee0ac93bdd4b4ffcd77b2e5

Observation 08761c9a-b025-4ba5-8f76-416a3fc6726a · outbound

This paper cites LiDAR-LLM: Exploring the Potential of Large Language Models for 3D LiDAR Understanding.

Hierarchical Question-Answering for Driving Scene Understanding Using Vision-Language Models LiDAR-LLM: Exploring the Potential of Large Language Models for 3D LiDAR Understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:23:13.595461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:23:13.595461Z digest=sha256:7cc5b1ae63f127605a29de07adcbf3234cc5dd5eccc5faaf0b9d2c02b47ef941

Observation 43ab1c19-b7c8-48aa-aa60-4a48018b4346 · outbound

This paper cites Vectornet: Encoding hd maps and agent dynamics from vectorized representation,.

Hierarchical Question-Answering for Driving Scene Understanding Using Vision-Language Models Vectornet: Encoding hd maps and agent dynamics from vectorized representation,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:23:16.011931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:23:13.727279Z digest=sha256:50c55503c3c7660a67fb1394252435742602b3bba17cf2ef0d18abca0d6ebc5d

Observation bfe27062-e1d8-437a-a555-578abb92c5fb · outbound

This paper cites Reimagining an autonomous vehicle.

Hierarchical Question-Answering for Driving Scene Understanding Using Vision-Language Models Reimagining an autonomous vehicle

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:23:14.984276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:23:13.817011Z digest=sha256:f02a60d50a3d94b8c34b9f17f370425bde6eada63d0a81366736ce34b4f31295

Observation a0a3c807-ec43-42a4-8d44-2a30e46a5e91 · outbound

This paper cites Explainable artificial intelligence (xai): Concepts, taxonomies, opportunities and challenges toward responsible ai,.

Hierarchical Question-Answering for Driving Scene Understanding Using Vision-Language Models Explainable artificial intelligence (xai): Concepts, taxonomies, opportunities and challenges toward responsible ai,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:23:15.867593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:23:13.891608Z digest=sha256:5aeeca5ecff3d2ddd5b22e6015131421dad5c38d790e34a2497b5438c5a7d59f

Observation 26ff000f-8917-4c67-bca6-31f7d4ad6930 · outbound

This paper cites From automation to autonomy and autonomous vehicles: Challenges and opportunities for human-computer interaction,.

Hierarchical Question-Answering for Driving Scene Understanding Using Vision-Language Models From automation to autonomy and autonomous vehicles: Challenges and opportunities for human-computer interaction,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:23:15.758020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:23:13.955573Z digest=sha256:21dd882c9eaa0ac2128c64fbec077343f1a8cc7c617b690a850bc34bb5d1e0a4

Observation fd326759-a542-4496-b394-370e4317e8fb · outbound

This paper cites LanguageMPC: Large Language Models as Decision Makers for Autonomous Driving.

Hierarchical Question-Answering for Driving Scene Understanding Using Vision-Language Models LanguageMPC: Large Language Models as Decision Makers for Autonomous Driving

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:23:14.026613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:23:14.026613Z digest=sha256:97b9e31142e966f3986c4463387222c6df0426eba5597cc8d084f816e09d90cb

Observation 009525c1-7108-4b0f-841b-2898fddf9f13 · outbound

This paper cites Lm-nav: Robotic navigation with large pre-trained models of language, vision, and action,.

Hierarchical Question-Answering for Driving Scene Understanding Using Vision-Language Models Lm-nav: Robotic navigation with large pre-trained models of language, vision, and action,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:23:15.644833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:23:14.068576Z digest=sha256:5d7a55c8ae3521868c657d594fb7b3813b8254bf6bb6e6c78bd62f9469f639d6

Observation 439f4f6a-e764-44cf-bd1c-228f1e3b279a · outbound

This paper cites DiLu: A Knowledge-Driven Approach to Autonomous Driving with Large Language Models.

Hierarchical Question-Answering for Driving Scene Understanding Using Vision-Language Models DiLu: A Knowledge-Driven Approach to Autonomous Driving with Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T11:23:14.113366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:23:14.113366Z digest=sha256:ee927fa9f8fa96858bd3d43902e7e77904b78d190a4ff1ccf4c9de958052ef0f

Observation c9dcfaf7-e65c-424a-92cf-24e977e57cb6 · outbound

This paper cites Driving with llms: Fusing object- level vector modality for explainable autonomous driving,.

Hierarchical Question-Answering for Driving Scene Understanding Using Vision-Language Models Driving with llms: Fusing object- level vector modality for explainable autonomous driving,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:23:15.445095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:23:14.182507Z digest=sha256:c4fde1f86451f951d906576811d42f2372efd7ee4a2ec66c6f30cbe9fa7a5f26

Observation a06e06f1-6544-4233-b29f-5e7294f6bfdf · outbound

This paper cites GPT-Driver: Learning to Drive with GPT.

Hierarchical Question-Answering for Driving Scene Understanding Using Vision-Language Models GPT-Driver: Learning to Drive with GPT

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:23:14.218939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:23:14.218939Z digest=sha256:6e464360d555ee314d14dcb1110c88aa0a2be5e546c5344b9d297339652966e1

Observation 9bc840b0-96eb-4faf-bba9-7f50d9c1cd97 · outbound

This paper cites Drivegpt4: Interpretable end-to-end autonomous driving via large language model,.

Hierarchical Question-Answering for Driving Scene Understanding Using Vision-Language Models Drivegpt4: Interpretable end-to-end autonomous driving via large language model,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:23:14.317834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:23:14.317834Z digest=sha256:fd2a4c38bb1fc637cc525f2156acc13ceae36839b93fcacc0b54474128b9422d

Observation f106eb8c-399d-4736-be2e-3d6bc7a9dd8d · outbound

This paper cites RAG-Driver: Generalisable Driving Explanations with Retrieval-Augmented In-Context Learning in Multi-Modal Large Lan- guage Model,.

Hierarchical Question-Answering for Driving Scene Understanding Using Vision-Language Models RAG-Driver: Generalisable Driving Explanations with Retrieval-Augmented In-Context Learning in Multi-Modal Large Lan- guage Model,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T11:23:14.375874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:23:14.375874Z digest=sha256:43fc0db79ee558236586b0e08963ff88326895ea2e140a00e5db3c91168abbc9

Observation 9962cfb5-f5b4-4242-9d78-35efb8d775fc · outbound

This paper cites DriveMLM: Aligning Multi-Modal Large Language Models with Behavioral Planning States for Autonomous Driving,.

Hierarchical Question-Answering for Driving Scene Understanding Using Vision-Language Models DriveMLM: Aligning Multi-Modal Large Language Models with Behavioral Planning States for Autonomous Driving,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:23:14.454649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:23:14.454649Z digest=sha256:7b9fdcc0ebe3c66f8f2c551d8f076b03c8a2f0959cb3351bce72eb4c175df150

Observation ecc36b9e-663c-4f8c-9e43-53724ef1cebb · outbound

This paper cites Lmdrive: Closed-loop end-to-end driving with large language models,.

Hierarchical Question-Answering for Driving Scene Understanding Using Vision-Language Models Lmdrive: Closed-loop end-to-end driving with large language models,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:23:14.503544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:23:14.503544Z digest=sha256:61f948d273c0816f947cf777d6a925b6061324b5d9d01eda4aa2435cdeb12ae4

Observation 0de58213-75e5-4885-9290-f651cc754242 · outbound

This paper cites Lingoqa: Visual question answering for au- tonomous driving,.

Hierarchical Question-Answering for Driving Scene Understanding Using Vision-Language Models Lingoqa: Visual question answering for au- tonomous driving,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:23:15.227697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:23:14.558335Z digest=sha256:08888d4c5224d7fc98c1cfa3f93f943321fbeff66e1d18f62061262ff6081ab4

Pith citing papers

No inbound Pith citation observations are available.