Pith. sign in

Paper Citation Record · LEDGER

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios

As of 14 August 2026, this Paper Citation Record lists 70 of 70 outbound references and 6 inbound Pith citation observations for arXiv:2501.04671.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.04671 v2

Coverage vector

measured 70 of 70 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T21:30:36.752852Z

measured 76 of 76 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:03:02.400451Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T21:28:58.962434Z

Reference resolution

70 of 70 outbound references displayed

  • verified exact0
  • verified fuzzy40
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6eac0e54-24dc-4f11-a550-b4dfa5e8b48a · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T21:30:36.149229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:30:36.149229Z digest=sha256:9786d28bedcf1d62dd1dbbb501faec0c83395c260466573de4b128c6a3906dfc

Observation 2e26796c-e549-427a-a216-75b008ededc2 · outbound

This paper cites Lawrence Zitnick, and Devi Parikh.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Lawrence Zitnick, and Devi Parikh

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:39.197797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:30:36.162752Z digest=sha256:44a46433b6ecd3f1bec429698c89a3772b34969912585c425670b09ddb7715fd

Observation d8725a61-6142-439d-8372-47b3fd01cea6 · outbound

This paper cites Response Wide Shut: Surprising Observations in Basic Vision Language Model Capabilities.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Response Wide Shut: Surprising Observations in Basic Vision Language Model Capabilities

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T21:30:36.172684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:30:36.172684Z digest=sha256:8dcc1280074d5851fd149470bcf887c2abb898e78ec15741ad6e83e8a42d5486

Observation 170c53ea-a194-4a52-8804-b90bc786f8a9 · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T21:30:36.183356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:30:36.183356Z digest=sha256:fca12e1f6821f93dc7d5364a010f0eeee1802d3a9fa7366a1e159cd5c9699a38

Observation a2bde1ed-0c76-42ea-86fe-afe2792a6806 · outbound

This paper cites PuzzleVQA: Diagnosing Multimodal Reasoning Challenges of Language Models with Abstract Visual Patterns.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios PuzzleVQA: Diagnosing Multimodal Reasoning Challenges of Language Models with Abstract Visual Patterns

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T21:30:36.202132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:30:36.202132Z digest=sha256:002c34fa9613a7c0dc95b4a78030e1890ab8274f24a60b6b370230c46e879952

Observation e9b4910e-bd5c-4a08-8208-9d3b0941154b · outbound

This paper cites Gpt-4 technical report, 2024.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Gpt-4 technical report, 2024

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:39.167769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:30:36.212947Z digest=sha256:4c36385fc04432dff477eb2bd4895440e43a7b8743f7f7f9ca5d56b9a3c4e574

Observation f619a843-7cc5-4e5b-8946-22802fc1eb0f · outbound

This paper cites Making the V in VQA matter: El- evating the role of image understanding in visual question answering.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Making the V in VQA matter: El- evating the role of image understanding in visual question answering

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:39.130806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:30:36.223655Z digest=sha256:993a6a6ffe7399669a8b547d08e68743bd0f45430dd167cad192a73e0b0cfdbf

Observation 1983a855-2c2c-4077-b4d2-5c3557fb5f16 · outbound

This paper cites A Survey on LLM-as-a-Judge.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios A Survey on LLM-as-a-Judge

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T21:30:36.240289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:30:36.240289Z digest=sha256:c7b3f81dd79a60b716213a913dd7b4d055c2ae706d23c7392c0528adb545dff8

Observation 6debd138-9dad-49e2-b3b7-33b1df524c3e · outbound

This paper cites Hal- lusionbench: An advanced diagnostic suite for entangled language hallucination and visual illusion in large vision- language models.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Hal- lusionbench: An advanced diagnostic suite for entangled language hallucination and visual illusion in large vision- language models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:39.103518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:30:36.247879Z digest=sha256:cb85b46e9e592ef6c7065b94d601ebb21f9005c6ca55a7d6beedee9efcee978b

Observation 2481ea86-57d9-4fee-ba29-f746ea242c48 · outbound

This paper cites Visual program- ming: Compositional visual reasoning without training.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Visual program- ming: Compositional visual reasoning without training

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:39.065647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:30:36.255730Z digest=sha256:d054fdda09a49096615f11df00d085c1e50baf50589af99d982a8e3188d75542

Observation 7545632b-f99b-434d-b33e-617ce985ee69 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and composi- tional question answering.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Gqa: A new dataset for real-world visual reasoning and composi- tional question answering

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:39.017237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:30:36.262399Z digest=sha256:72ba28971356133f5f2dc8297ac41470dd8dd03206999b2933406294d0a49414

Observation be5f94ff-42c1-44b5-90b0-58647365c8c8 · outbound

This paper cites EMMA: End-to-End Multimodal Model for Autonomous Driving.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios EMMA: End-to-End Multimodal Model for Autonomous Driving

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T21:30:36.269691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:30:36.269691Z digest=sha256:fd88e4d49ff37000f3a1573dc129b3ab4a1caea30ed1b709550ffee280ae9200

Observation 1c722649-ff0e-479a-b36d-4a84de4c5db3 · outbound

This paper cites an unresolved cited work.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:30:38.988473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:30:36.278634Z digest=sha256:48822f4913af842a340499220570eb6fcfd53e3832e3225cc84f1ad8ed5ab15d

Observation a992410b-0699-4e80-bf10-ebca3b119452 · outbound

This paper cites Hallucination augmented contrastive learn- ing for multimodal large language model.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Hallucination augmented contrastive learn- ing for multimodal large language model

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:38.954899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:30:36.288146Z digest=sha256:8be2aebb00ffe7dfc450a3ac946945c6c92623f908da85184f62de65950f1950

Observation 720b1843-4ba2-4ccf-8c86-785e32544d45 · outbound

This paper cites Textual explanations for self-driving vehicles.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Textual explanations for self-driving vehicles

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:38.924396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:30:36.296054Z digest=sha256:35a4a060bcfa893cd99f51adee56b85b490fba371c4537d2952b12a016c62495

Observation 188272a2-6265-4d64-b66a-783612a9a630 · outbound

This paper cites Scaffolding Coordinates to Promote Vision-Language Coordination in Large Multi-Modal Models.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Scaffolding Coordinates to Promote Vision-Language Coordination in Large Multi-Modal Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T21:30:36.312300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:30:36.312300Z digest=sha256:2571d29396b61247c173ceacdafc4ecd8c2589d9b6b9b1d0388b8173df63ceb4

Observation 7bb1aab0-1a22-477b-876e-037b4f056aca · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios LLaVA-OneVision: Easy Visual Task Transfer

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T21:30:36.324818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:30:36.324818Z digest=sha256:013c743744b740076f3f7ede77620007c83f4a95f2f3eda9ec844e7b962373db

Observation 3be75811-bf98-4772-be0c-09b8a9447a60 · outbound

This paper cites From representa- tion to reasoning: Towards both evidence and commonsense reasoning for video question-answering.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios From representa- tion to reasoning: Towards both evidence and commonsense reasoning for video question-answering

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:38.888105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:30:36.335361Z digest=sha256:c67b2ea04a993073caf802e1641d8669661a46eda95b6efe09d112733ea2bd08

Observation 7cbf36dd-cf39-448b-a6b4-9df9eb0d374b · outbound

This paper cites Chain-of-region: Visual language models need details for diagram analysis.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Chain-of-region: Visual language models need details for diagram analysis

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:38.861322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:30:36.343783Z digest=sha256:f2c905813d0a815c898f9aa67e8d1aade4575ce6006e4c037de5fdc0daf66be8

Observation 3accbeb3-2ebb-4ce3-b660-730752a9e7f4 · outbound

This paper cites Enhancing advanced visual reasoning ability of large language models.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Enhancing advanced visual reasoning ability of large language models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:38.834923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:30:36.350843Z digest=sha256:16ed4fe9b0249ed897b421936daeca5819d28b287ef1252e9105a737e76728d5

Observation 1607bf4c-20b8-4e3d-91fd-0c2b3270945f · outbound

This paper cites Vila: On pre-training for vi- sual language models.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Vila: On pre-training for vi- sual language models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:38.800965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:30:36.358581Z digest=sha256:5987b3632f12146ae0f276622b9efd6d901703e6c156e414b7d0bba517e00b67

Observation ab371db2-69cb-4c8a-9f81-cb3562882ca2 · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T21:30:36.365886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:30:36.365886Z digest=sha256:a233f0c932cbbbc468be803d58e60d06237ccd72c9ef84206559cd9ec7051a57

Observation 93eee8a4-0a34-466c-840f-62cbf698f9cf · outbound

This paper cites Visual instruction tuning.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Visual instruction tuning

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:38.743639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:30:36.371144Z digest=sha256:dcb28aff25bf871e159d17ce45b46355f2f1d8ef645ebd96d51b368ff1940caa

Observation 816f331d-2bf9-433b-8222-4abc381eb137 · outbound

This paper cites A Survey on Hallucination in Large Vision-Language Models.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios A Survey on Hallucination in Large Vision-Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T21:30:36.377405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:30:36.377405Z digest=sha256:ddd1cb1565cbe3b92d2ba27c123abd5df012c9dc6514617c0d96bfec21572e9c

Observation c476e120-ccb2-419d-8774-b1239cbad4c0 · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T21:30:36.385263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:30:36.385263Z digest=sha256:d1e3796c4adfee9faea0d6bb05220b5c445a076397709711b725ecc3775bb88b

Observation 44dc0183-0c48-4173-9e4e-b63d19ed9b6d · outbound

This paper cites Aligning with Human Judgement: The Role of Pairwise Preference in Large Language Model Evaluators.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Aligning with Human Judgement: The Role of Pairwise Preference in Large Language Model Evaluators

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T21:30:36.390831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:30:36.390831Z digest=sha256:9298e785cf394437bf279aa7d7059c9e95b472cda65d310552c984dda94f5864

Observation 18877e83-5800-48e1-ae17-c318030e4a98 · outbound

This paper cites SpatialCoT: Advancing Spatial Reasoning through Coordinate Alignment and Chain-of-Thought for Embodied Task Planning.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios SpatialCoT: Advancing Spatial Reasoning through Coordinate Alignment and Chain-of-Thought for Embodied Task Planning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T21:30:36.398003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:30:36.398003Z digest=sha256:ce2ea491fb008fe5549aecccc6bfc33818a0638612f51c97fd0be33081df7950

Observation 3c0eeb34-a7e3-4ab0-b455-8a067e927ebc · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? In European Conference on Computer Vision, pages 216–233.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Mmbench: Is your multi-modal model an all-around player? In European Conference on Computer Vision, pages 216–233

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:38.717175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:30:36.414293Z digest=sha256:d43c41e6f405ab7509e1c14b96dd44ba25b6ac5e59acc56d3f2a9d74e34327a0

Observation c598d8a0-c409-402c-8336-af120426c895 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T21:30:36.420260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:30:36.420260Z digest=sha256:34aab81e3054c9061626650179eda503a3460d26b36ba20dfab62164fafba6a0

Observation a306cbfd-7391-4d0b-a444-94c2384f3c63 · outbound

This paper cites Lingoqa: Visual question an- swering for autonomous driving.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Lingoqa: Visual question an- swering for autonomous driving

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:38.658998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:30:36.430131Z digest=sha256:48e7c2542e235554c8ea0d156da58e8ea69741ba31807ec2d1e19817f1d2a19e

Observation 7e0d9d59-297e-4d3a-acc9-8af6ae3c31b9 · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Ok-vqa: A visual question answering benchmark requiring external knowledge

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:38.628535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:30:36.437404Z digest=sha256:b55e3f5151fd90f09c3f82f435d8595a9147e55608c6d36677d4fa855b857f8c

Observation e5d10df3-bf1c-42e5-8782-8b15ee4477ce · outbound

This paper cites Compositional chain-of-thought prompting for large multimodal models.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Compositional chain-of-thought prompting for large multimodal models

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:38.604162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:30:36.448556Z digest=sha256:b09e734f5b939997bb58e2432b97229de15c9c8d7b87f22b7b69be1a10ca17c8

Observation b5901ceb-6a78-4511-bfbb-ebf7cd7bcf2e · outbound

This paper cites Gpt-4o system card, 2024.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Gpt-4o system card, 2024

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:38.571257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:30:36.455451Z digest=sha256:6c2f81f1c04f9c5c36758d82050140bf553ab0c630bb3bc807f11abb0df869bd

Observation 2cdf85ba-8a17-45f4-a2ed-e9b400298c5a · outbound

This paper cites Openai o1 system card, 2024.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Openai o1 system card, 2024

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:38.536737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:30:36.461161Z digest=sha256:61b8607e94e68b430f6d89b83af3c0abe0d9294f3ff7861ac34e7c9f71339fdb

Observation 315e8c15-ac7a-4028-aea9-04a7aaa22793 · outbound

This paper cites Enhancing visual question answering through question-driven image captions as prompts.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Enhancing visual question answering through question-driven image captions as prompts

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:38.496766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:30:36.466714Z digest=sha256:efbd3702a5a25a0e1ebe9dacbff8162dff4d17517669ef26d70e6ff397dcf05f

Observation ab18df49-45f2-4dd3-b3a7-a91f34b91cc5 · outbound

This paper cites Piaget’s theory of intelligence.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Piaget’s theory of intelligence

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:38.459157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:30:36.472476Z digest=sha256:db2cb127b41ea3123b023b9f58e5df4c6183838f640529e9dd167fc8a725b4f2

Observation cbbb9c9f-49d9-4292-95c3-dbc310f9ad4d · outbound

This paper cites CogCoM: A Visual Language Model with Chain-of-Manipulations Reasoning.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios CogCoM: A Visual Language Model with Chain-of-Manipulations Reasoning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T21:30:36.478381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:30:36.478381Z digest=sha256:debee164a57fb8bb48dd07fae3bd165c8058feea4c1a0103dd36c3a048f76435

Observation afd6f303-18b7-408f-b4ec-c257d5eab6d8 · outbound

This paper cites Nuscenes-qa: A multi-modal visual ques- tion answering benchmark for autonomous driving scenario,.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Nuscenes-qa: A multi-modal visual ques- tion answering benchmark for autonomous driving scenario,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:38.435527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:30:36.485056Z digest=sha256:36fd75b7c4ff8841f8b28ecfea5fc1e4fb32a63d059c93bb20a254ff9aafaea9

Observation e49f37cf-2b07-4896-aaa2-02cc98e11da3 · outbound

This paper cites Prism: A framework for decoupling and assessing the capabilities of vlms.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Prism: A framework for decoupling and assessing the capabilities of vlms

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:38.393214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:30:36.491409Z digest=sha256:93a6102f3f8447711c4c69b1d750268a0fd1313ab53ac734a2c0bb5ace63f5cb

Observation 93b07525-29fd-4455-a296-13be0a3f35d3 · outbound

This paper cites A-okvqa: A benchmark for visual question answering using world knowl- edge.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios A-okvqa: A benchmark for visual question answering using world knowl- edge

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:38.365510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:30:36.497058Z digest=sha256:11f481c74105d09003102521ed4a241ab111bc687acd55d2e17afb3364fdcdfe

Observation 4e7aab82-e0b8-4f80-9015-05064d369046 · outbound

This paper cites Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T21:30:36.506782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:30:36.506782Z digest=sha256:fcc44ca6ab8ba6524c02dcbc5e3e6feec89cbb7d0cfa2a1bc3be50d09c0ad985

Observation 90af283d-834c-48a2-a291-677c09ccb2a9 · outbound

This paper cites Drivelm: Driving with graph visual ques- tion answering.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Drivelm: Driving with graph visual ques- tion answering

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:38.336952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:30:36.514338Z digest=sha256:33758313f6df3b0598a02dd546f726ef314a4c6afbc0e54d95758bbef732ba0d

Observation 7f0ea722-c5fd-4956-939f-032f56899a7d · outbound

This paper cites Towards vqa models that can read.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Towards vqa models that can read

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:38.314568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:30:36.522235Z digest=sha256:4bcc49ebd893e951f999f36015c5946d8902ae4649c488f2739693e51cfacb95

Observation 5b5c25b4-0b85-440f-b652-98f777adcca8 · outbound

This paper cites To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T21:30:36.527626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:30:36.527626Z digest=sha256:a7b43d242c861ee86f49fd589d4319080b3b9b05a1a73f9b60abfc504386921e

Observation e14b3747-d2fb-430b-b096-0899bfa22de4 · outbound

This paper cites Vipergpt: Visual inference via python execution for reasoning.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Vipergpt: Visual inference via python execution for reasoning

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:38.289634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:30:36.534132Z digest=sha256:3b11c4e39785679691610c5e5fa024c141281729ec7dace2ddab52f1f3a80a41

Observation 15e39f8e-855c-481f-9923-0c43491f3ba6 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios LLaMA: Open and Efficient Foundation Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T21:30:36.540535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:30:36.540535Z digest=sha256:25d9cf55a10992e1476d1275fb1f548572d7f9df01dc11e1afe415096d2e3c46

Observation 49cf1156-a5f3-4770-952d-aafffb4ef24c · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T21:30:36.556236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:30:36.556236Z digest=sha256:c2300e1eaff433bf190b0735afe0bccab8a3ea948cdaae1954586eb60777a952

Observation 6f5ddf51-ea5d-4951-8760-992ffa7f679f · outbound

This paper cites an unresolved cited work.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:30:38.251058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:30:36.561374Z digest=sha256:65f25408bf03a9e2e4d8f70a88c8b52eda946d13843a18844554e8ebf06e986e

Observation 98c8e57b-12de-4f7f-9ad7-3b42b7c86340 · outbound

This paper cites Drive anywhere: Generalizable end-to-end autonomous driving with multi-modal foundation models,.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Drive anywhere: Generalizable end-to-end autonomous driving with multi-modal foundation models,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:38.222761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:30:36.570857Z digest=sha256:67679d73bb397c08d8c50ba2a32eb0f6284fa373faae237b6110767df1fe6d37

Observation d7181fe8-ce77-4cc8-9b33-66d1e145248b · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Chain-of-thought prompting elicits reasoning in large language models

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:38.181322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:30:36.577033Z digest=sha256:c9108eccb7e744f0a2cfb7ef116c05ab98869f9c2fcec70135d0022ab2ff8840

Observation d63dfdb6-12da-46c9-a446-3686f6de3944 · outbound

This paper cites Drivegpt4: Interpretable end-to-end autonomous driving via large language model.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Drivegpt4: Interpretable end-to-end autonomous driving via large language model

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T21:30:36.583053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:30:36.583053Z digest=sha256:a957ac2e18376cdaf32ca9a1977ec08a5ca46db108baea53fe03325d62c04890

Observation 25ccc62b-e148-4520-b3d1-7a631ec14ada · outbound

This paper cites List Items One by One: A New Data Source and Learning Paradigm for Multimodal LLMs.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios List Items One by One: A New Data Source and Learning Paradigm for Multimodal LLMs

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T21:30:36.589188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:30:36.589188Z digest=sha256:7d02513da19351d32690be54d189a27e55df1b9f770fc2f261e581009b77b59c

Observation 16276f81-c2a7-4957-a3b5-1284f8ccddc9 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understand- ing and reasoning benchmark for expert agi.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Mmmu: A massive multi-discipline multimodal understand- ing and reasoning benchmark for expert agi

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:38.118245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:30:36.597676Z digest=sha256:3109da125c3bbe90f0af00526c1b5acfa613cb99d0285cab66595c7e24218978

Observation 107f1902-3405-4a7f-80e0-d0f480279289 · outbound

This paper cites Sigmoid loss for language image pre-training.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Sigmoid loss for language image pre-training

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-10T21:30:36.606307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:30:36.606307Z digest=sha256:a44e66bbe81b70687eab716ea2534420d8de4f6479f4b16c7ce50965d409dcb6

Observation 9105b1a2-9d67-4191-9bda-b93fcbc9e225 · outbound

This paper cites Multimodal chain-of-thought rea- soning in language models.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Multimodal chain-of-thought rea- soning in language models

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:38.056246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:30:36.617223Z digest=sha256:d05c1c854fe3252dc008e22b39e4a01006ca254de2d2a61cac9e3e31cad510e7

Observation a497be7e-228a-40db-8662-6f16761e7bd8 · outbound

This paper cites Multimodal chain-of-thought reasoning in language models.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Multimodal chain-of-thought reasoning in language models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-10T21:30:36.626364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:30:36.626364Z digest=sha256:30ebab3641e4c580599302ec044d10d9d73e01ffcb75060c79528ade78f9f283

Observation d8dfc508-a7b9-4b03-b700-9c1c3c72fc18 · outbound

This paper cites Ddcot: Duty-distinct chain-of-thought prompting for multimodal reasoning in language models.Advances in Neu- ral Information Processing Systems, 36:5168–5191, 2023.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Ddcot: Duty-distinct chain-of-thought prompting for multimodal reasoning in language models.Advances in Neu- ral Information Processing Systems, 36:5168–5191, 2023

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:38.011547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:30:36.633857Z digest=sha256:b183c7060c06511fcf72734223a55ebd68dd7d1d82d5a4c6d517a15c7df369be

Observation 4f8b75b4-414a-4196-8d22-946366919bff · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in Neural Information Processing Systems, 36:46595–46623, 2023.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in Neural Information Processing Systems, 36:46595–46623, 2023

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:37.979560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:30:36.640347Z digest=sha256:a7e6a52c55917688fae0abf5e347c97431b2d6adf4708b55134eb5c9ace887d3

Observation 88c8e4ed-9e2b-4904-ae87-bafb6fde5c37 · outbound

This paper cites an unresolved cited work.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:30:37.929860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:30:36.647145Z digest=sha256:506b98c319827755fb8862923c4f18f954d569de202040b9cb1f395b27f85e3e

Observation 03099f83-4d89-4e61-b2e7-621853fab1f8 · outbound

This paper cites an unresolved cited work.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:30:37.896710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:30:36.653178Z digest=sha256:02b27a40acd6f5ddf5711284071fddab35f532862c36f0559812808dc5080955

Observation dbf4994c-883c-449e-a3eb-f488b0e35fe9 · outbound

This paper cites Replicate bound- ing box coordinates exactly as provided in the list.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Replicate bound- ing box coordinates exactly as provided in the list

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:37.861424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:30:36.663540Z digest=sha256:05c9ec812f798ab09ea3b70ab0ed0863abc343182bd74f52a041fa207f4946c0

Observation b6dc4b4f-c5ef-44d4-b990-29549955602d · outbound

This paper cites an unresolved cited work.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:30:37.822666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:30:36.672684Z digest=sha256:61a5de4f5f0bb20f7907afa212b9d6f017b2e93d8e19510b0c99f01fdbefecf0

Observation e89bb2bf-b40d-4a7a-a170-d73304b334e2 · outbound

This paper cites 1 Demonstration 1 Question: [”I am turning right at the next intersec- tion.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios 1 Demonstration 1 Question: [”I am turning right at the next intersec- tion

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:37.786060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:30:36.684686Z digest=sha256:c69d21f12e5c87641ca90d9f2d779f39201726ac07556db5bce7eefdf621d96f

Observation a29777f7-d99d-46cd-b5e7-68a6d40df5fc · outbound

This paper cites an unresolved cited work.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Unresolved cited work

Reference 65

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:30:37.748045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:30:36.692496Z digest=sha256:a229f4df66701d3f16195d786a35610d26e1e5de6b579f9ad940b5c2c68e5c3c

Observation ef6e2bd7-dc18-4447-82e3-2d448fb9565f · outbound

This paper cites an unresolved cited work.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Unresolved cited work

Reference 66

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:30:37.711537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:30:36.699972Z digest=sha256:1c6e88b690486916ff7792a3a9857074e2f08962f51deb0235d70cfa46028fcf

Observation f9faba4c-7921-4df5-8c31-85da6192b6f1 · outbound

This paper cites correct reasoning.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios correct reasoning

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:37.661846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:30:36.710086Z digest=sha256:56c5bd651b6813dda2df8f914c102617b3740d0f2768b5feffb88525832ce9d6

Observation a42d84ca-e754-4ff3-aa01-48bbdf57abf3 · outbound

This paper cites Step-by-Step Instructions:.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Step-by-Step Instructions:

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:37.634427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:30:36.718007Z digest=sha256:dde9cd702694814bf345bea4a62e408229923eaa0b690177e2c31173fb1c0e62

Observation 0a4c0857-cd4a-4e09-9741-40eccefa8d5c · outbound

This paper cites • For each argument, briefly state whether it is correct or not, given the provided correct reasoning.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios • For each argument, briefly state whether it is correct or not, given the provided correct reasoning

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:37.595726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:30:36.733599Z digest=sha256:dc031c09cbb97954d700a25ea476f12a250592a479b76a566744f1ea62cc3696

Observation 64ee6a68-f966-40df-82dc-1caaf70048b5 · outbound

This paper cites • List important points or steps from the correct reasoning that the student omits or directly contradicts.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios • List important points or steps from the correct reasoning that the student omits or directly contradicts

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:37.569194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:30:36.741666Z digest=sha256:c72598288f2f38b1c31c0fe8a9007de9e15b74880d5a1f7e17f46e96f5e9fb2f

Observation 7e0f1922-d574-4a0d-b64e-7aa12ba5d9f5 · outbound

This paper cites 1” if you judge the student’s reasoning is overall correct, “0.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios 1” if you judge the student’s reasoning is overall correct, “0

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:37.528655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:30:36.752852Z digest=sha256:79d032720c03ec94b6241dcdd3443d062df2312d28e2d0d2782e77875bde9644

Pith citing papers

Observation 9c7da779-1d85-4c52-8184-f79c7db28ea5 · inbound

Chain-of-Thought for Autonomous Driving: A Comprehensive Survey and Future Prospects cites this paper.

Chain-of-Thought for Autonomous Driving: A Comprehensive Survey and Future Prospects Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios

Reference 127

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:02.400451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:02.400451Z digest=sha256:f94da181aa0f570609a15fab627b93548a1706ec37cc954a379b5ce79343cb14

Observation 20675e45-9e4a-438a-9623-f38d744f27ab · inbound

Alpamayo-R1: Bridging Reasoning and Action Prediction for Generalizable Autonomous Driving in the Long Tail cites this paper.

Alpamayo-R1: Bridging Reasoning and Action Prediction for Generalizable Autonomous Driving in the Long Tail Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-18T02:35:13.306854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T02:35:13.126171Z digest=sha256:dcf12336dad20c579fbdeaa069da50f4b8bc0eeb2663592eccad39a5f61b9ad1

Observation f7b6ac57-9991-47a6-a0a8-920044a279fd · inbound

OmniDrive-R1: Reinforcement-driven Interleaved Multi-modal Chain-of-Thought for Trustworthy Vision-Language Autonomous Driving cites this paper.

OmniDrive-R1: Reinforcement-driven Interleaved Multi-modal Chain-of-Thought for Trustworthy Vision-Language Autonomous Driving Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:38:37.815554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-16T22:34:00.895252Z digest=sha256:133dd000ffbca669adb27141a21885b4405ba52e76402617ad0123944251f200

Observation 9e10aa20-2566-4a03-94a5-b7daf20abbb7 · inbound

RailVQA: A Benchmark and Framework for Efficient Interpretable Visual Cognition in Automatic Train Operation cites this paper.

RailVQA: A Benchmark and Framework for Efficient Interpretable Visual Cognition in Automatic Train Operation Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:28:04.771507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-14T22:24:13.897439Z digest=sha256:2f37d4e08fd852313a32b63f3b65d527739e345282a0173b52d9572809271664

Observation bec3ebc2-c327-4c0a-97e8-589271f2df65 · inbound

What to Test Next: Interpretable Coverage Gap Discovery in Driving VLMs cites this paper.

What to Test Next: Interpretable Coverage Gap Discovery in Driving VLMs Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:16:16.242167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T15:36:25.219607Z digest=sha256:a6410dfc9a974d6ab94764a399a08496e5369c772fefa81f1f7b646238539d15

Observation 2ae9e646-ecb7-41cd-a9d2-b45ce7697733 · inbound

CaVe-VLM-CoT: An Interpretable Vision-Language Model Framework cites this paper.

CaVe-VLM-CoT: An Interpretable Vision-Language Model Framework Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T21:28:58.963898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-27T00:31:01.692743Z digest=sha256:b9e4589665c618ed535cd16ffd6637ea44843d122dbb9c635e437c51bb3110fb