Pith. sign in

Paper Citation Record · LEDGER

Pedestrian Intention Prediction via Vision-Language Foundation Models

As of 9 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 1 inbound Pith citation observation for arXiv:2507.04141.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.04141 v1

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:59:56.791343Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-30T13:25:04.053283Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T13:34:40.932917Z

Reference resolution

36 of 36 outbound references displayed

  • verified exact2
  • verified fuzzy25
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 89755a4e-910f-4fd8-b046-ae3d7a248e5b · outbound

This paper cites Autonomous vehicles that interact with pedestrians: A survey of theory and practice,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Autonomous vehicles that interact with pedestrians: A survey of theory and practice,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T19:59:53.815180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:59:53.815180Z digest=sha256:b1a323dce897cf6352277adc2a88523fb22c049e4d264ed30d5c5dda64762902

Observation 6cc181ab-3fb6-4ce6-86b4-dfc0be51a1ff · outbound

This paper cites Pedestrian intention prediction: A convolutional bottom-up multi-task approach,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Pedestrian intention prediction: A convolutional bottom-up multi-task approach,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:00:01.306110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:59:53.876159Z digest=sha256:2518991a7dd06c39b4f2a20400397b9007eeb51dd190caa0103bcd8d7fdc706e

Observation 64c0a2e8-e139-4b94-9995-dd787f946b28 · outbound

This paper cites St crossingpose: A spatial- temporal graph convolutional network for skeleton-based pedestrian crossing intention prediction,.

Pedestrian Intention Prediction via Vision-Language Foundation Models St crossingpose: A spatial- temporal graph convolutional network for skeleton-based pedestrian crossing intention prediction,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:00:01.180728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:59:53.941914Z digest=sha256:9c94700ec81a7b1459562830841efc8612a7247defb12e5f19899ec2383b686c

Observation 8a87e33a-a222-4467-b198-55377e59f960 · outbound

This paper cites PIP-Net: Pedestrian Intention Prediction in the Wild.

Pedestrian Intention Prediction via Vision-Language Foundation Models PIP-Net: Pedestrian Intention Prediction in the Wild

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:59:57.232162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:59:54.015290Z digest=sha256:5ffc5ab7fcd6e34ad72bf97867cf96fde0aaa21b08d47470ca72807a122a4736

Observation 0757d55c-1420-4cf7-b0c6-b743acbad52b · outbound

This paper cites Pedestrian action an- ticipation using contextual feature fusion in stacked rnns,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Pedestrian action an- ticipation using contextual feature fusion in stacked rnns,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:00:01.062175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:59:54.088410Z digest=sha256:cf4dbd5a5518b47dc5334733b702dbd363194abfba629db96a8e67b8914e301c

Observation 1fb433de-cfe6-4c8a-aea3-b8e46effe73e · outbound

This paper cites Do they want to cross? understanding pedestrian intention for behavior prediction,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Do they want to cross? understanding pedestrian intention for behavior prediction,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:00:00.945196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:59:54.171753Z digest=sha256:6c3e7e31fbc5f928dc14ceed553a2150174efdc83582db7e3ead3ad7eefb2e9a

Observation 096beed4-6d41-4232-b594-91f5044818fc · outbound

This paper cites Multi-modal hybrid architecture for pedestrian action prediction,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Multi-modal hybrid architecture for pedestrian action prediction,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:00:00.783093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:59:54.251824Z digest=sha256:5a87f991e223f70e07862e089d3222c7bb5a9980d84d232ae94bbd3d6ae08e25

Observation 57272dc5-5f20-416e-9830-2ae0172d7b8e · outbound

This paper cites Pedestrian graph+: A fast pedestrian crossing prediction model based on graph convo- lutional networks,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Pedestrian graph+: A fast pedestrian crossing prediction model based on graph convo- lutional networks,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:00:00.594842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:59:54.354966Z digest=sha256:cb7a9e99183220e66d0dcaa5d27214805b40bc6d715b3be08c4011ae32b5164f

Observation 6d431e29-b0a2-484d-924a-97647a9e5bcc · outbound

This paper cites Visual reasoning using graph con- volutional networks for predicting pedestrian crossing intention,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Visual reasoning using graph con- volutional networks for predicting pedestrian crossing intention,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:00:00.447152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:59:54.433966Z digest=sha256:06605cb786bc38102821e5ca5701979ec240eb7285a3fed4d0dc8baa923a9663

Observation e4e1f283-16ef-4438-84c3-4698cef4f6e0 · outbound

This paper cites CAPformer: Pedestrian crossing action prediction using transformer,.

Pedestrian Intention Prediction via Vision-Language Foundation Models CAPformer: Pedestrian crossing action prediction using transformer,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:00:00.297573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:59:54.497886Z digest=sha256:223d1b0ecda2b28722a14a73a0ce9d2a031792ef7863f641abeee30822325eac

Observation 7521fcd7-f346-42b0-82ff-37c9cd83ea1b · outbound

This paper cites Pit: Progressive interaction transformer for pedestrian crossing intention prediction,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Pit: Progressive interaction transformer for pedestrian crossing intention prediction,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:00:00.122113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:59:54.565677Z digest=sha256:145debe71180a4e788903f3843720fbb95e6ef1a35a26ab04bf81edb505cbb35

Observation 841f712e-7dfd-4388-9ebb-c567f59da6c3 · outbound

This paper cites Predicting pedestrian inten- tions with multimodal intentformer: A co-learning approach,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Predicting pedestrian inten- tions with multimodal intentformer: A co-learning approach,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:59:59.936485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:59:54.646030Z digest=sha256:8576a4703ea849cee434e6b78d3a5d4438bc563997f59cef4f7c17f2bb976001

Observation cbc0cc2c-ff16-477a-aa65-247a29a4ad34 · outbound

This paper cites Pedestrian behavior inter- pretation from pose estimation,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Pedestrian behavior inter- pretation from pose estimation,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:59:59.789304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:59:54.713024Z digest=sha256:6887f6c0b280e6eea70efe35da2d9aaaf9d0488972ac923c28d691af78be29c5

Observation e7e25164-28e8-4cf8-9723-5db6a1c92a99 · outbound

This paper cites Multi-scale pedestrian intent prediction using 3d joint information as spatio-temporal representation,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Multi-scale pedestrian intent prediction using 3d joint information as spatio-temporal representation,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:59:59.600266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:59:54.773984Z digest=sha256:71b9222c50d1c50536652ffacb160fef2989412b3b7463fe8864290b7343d130

Observation d72d0405-b7da-412a-ab62-1046703f990a · outbound

This paper cites Spatiotemporal relationship reasoning for pedestrian intent prediction,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Spatiotemporal relationship reasoning for pedestrian intent prediction,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T19:59:54.861620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:59:54.861620Z digest=sha256:8c959ea04168ce27d949a1ab52d325d6781dbee6dfe684d82e9f010452b7b0f1

Observation 42de927b-2ff4-4089-ad57-1ffa7ba398ac · outbound

This paper cites Real-time intent prediction of pedestrians for autonomous ground vehicles via spatio-temporal densenet,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Real-time intent prediction of pedestrians for autonomous ground vehicles via spatio-temporal densenet,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:59:59.442338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:59:54.933786Z digest=sha256:d39ddaf55d7b6cd3c15b07a7aab374bcfc1b64d3920eb62b3cec39ec4f179e55

Observation aa461dfd-b696-4dcc-9b73-896c835a4f53 · outbound

This paper cites Pedestrian-vehicle information modulation for pedestrian crossing intention prediction,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Pedestrian-vehicle information modulation for pedestrian crossing intention prediction,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:59:59.265412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:59:54.998685Z digest=sha256:d92b158077112da9670526dcae0308cc213e0090694a4571b11d1117054cb29c

Observation 615b7547-8719-485d-bfb7-e2ac534abdcf · outbound

This paper cites Causal reasoning in typical computer vision tasks,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Causal reasoning in typical computer vision tasks,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:59:59.071147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:59:55.083231Z digest=sha256:2ff416d3508ec50439c7c87d6c89d68ba2a0009aed69f2c7860c5cdeab410904

Observation 18069a05-b99a-4246-9453-288c0dc92147 · outbound

This paper cites Vision language models in autonomous driving: A survey and outlook,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Vision language models in autonomous driving: A survey and outlook,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:59:58.920039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:59:55.171703Z digest=sha256:0748df506dabcec9d95c79df9daa53cddd366a6620c317bf953e3fd3bb6a21d2

Observation 35d3e16b-b52f-4d17-b32c-0190c0006e9a · outbound

This paper cites Gpt-4v takes the wheel: Promises and challenges for pedestrian behavior prediction,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Gpt-4v takes the wheel: Promises and challenges for pedestrian behavior prediction,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:59:58.761794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:59:55.267162Z digest=sha256:84f07f52b5286ad9ad70146e22c41a5bea6917e8dc997a3d6ba87944888365b8

Observation ef2988a4-73fc-4e49-8be5-1a12e03fc931 · outbound

This paper cites Omnipredict: Gpt-4o enhanced multi-modal pedestrian crossing intention prediction.

Pedestrian Intention Prediction via Vision-Language Foundation Models Omnipredict: Gpt-4o enhanced multi-modal pedestrian crossing intention prediction

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:59:58.575555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:59:55.367452Z digest=sha256:789eb919fe673325c5ab54778820afad0e696f9467751f815cfdf2c6edade323

Observation a05142d2-747f-4bde-8387-19af28ae0283 · outbound

This paper cites Pedvlm: Pedestrian vision language model for intentions prediction,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Pedvlm: Pedestrian vision language model for intentions prediction,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:59:58.398927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:59:55.459977Z digest=sha256:aeaa087b6704daa56dc0b7349f4e88561f278f0851242efdcf5dc084050330f3

Observation bf34fca6-b13d-4b6f-8afd-f1890ef41e1f · outbound

This paper cites Cross or wait? predicting pedestrian interaction outcomes at unsignalized crossings,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Cross or wait? predicting pedestrian interaction outcomes at unsignalized crossings,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:59:58.222821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:59:55.569120Z digest=sha256:ff52eed9c493823f4621d34e1a0388c154f843d84da29aaffd7100efa0b79911

Observation 5fdb116b-8981-422a-8255-343c2ee44d16 · outbound

This paper cites Feature Importance in Pedestrian Intention Prediction: A Context-Aware Review.

Pedestrian Intention Prediction via Vision-Language Foundation Models Feature Importance in Pedestrian Intention Prediction: A Context-Aware Review

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:59:57.055566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:59:55.663331Z digest=sha256:00433996053d45f092bc1e2ff52c370a11d06d85b70f22cbb8da6a210b561816

Observation ca7ff6da-26b8-4be8-b4bd-6873e3b2a77b · outbound

This paper cites Hierarchical Prompting Taxonomy: A Universal Evaluation Framework for Large Language Models Aligned with Human Cognitive Principles.

Pedestrian Intention Prediction via Vision-Language Foundation Models Hierarchical Prompting Taxonomy: A Universal Evaluation Framework for Large Language Models Aligned with Human Cognitive Principles

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T19:59:55.734120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:59:55.734120Z digest=sha256:b053fda79077411ac77940cdb9af664a48544d8a52e300efff5822b830a6bd6c

Observation fc62114f-4913-4812-9b21-2bc71f52489a · outbound

This paper cites Large Language Models Are Human-Level Prompt Engineers.

Pedestrian Intention Prediction via Vision-Language Foundation Models Large Language Models Are Human-Level Prompt Engineers

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T19:59:55.815565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:59:55.815565Z digest=sha256:448eea504ec7df04175d32e054fa752ab7824748f609a88a5cd6248a272dace9

Observation 9ceed48e-7f92-4f3b-9b9f-8c3bd66adac8 · outbound

This paper cites Benchmark for evaluating pedestrian action prediction,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Benchmark for evaluating pedestrian action prediction,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:59:58.017275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:59:55.901830Z digest=sha256:06391ba86668fbba00aeb1e64944f817285d93547b1657d4a7ee5b094468ac4c

Observation 99d7b5b0-b85f-469a-8584-a08fa55707f7 · outbound

This paper cites Better Zero-Shot Reasoning with Role-Play Prompting.

Pedestrian Intention Prediction via Vision-Language Foundation Models Better Zero-Shot Reasoning with Role-Play Prompting

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T19:59:55.992136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:59:55.992136Z digest=sha256:3092d9bdd21eef25d9687da58496c1933ff2425fcc12a18beb964112b292123f

Observation 699aaa9e-581e-4a0c-afd9-1752e25544cb · outbound

This paper cites Good at captioning, bad at counting: Benchmarking GPT-4V on Earth observation data.

Pedestrian Intention Prediction via Vision-Language Foundation Models Good at captioning, bad at counting: Benchmarking GPT-4V on Earth observation data

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T19:59:56.092349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:59:56.092349Z digest=sha256:9f40b89275ed4ad449a7c8683866d0af5e5ae88f8c8ba64b91d2bcdf244bdcb2

Observation 1cd8d6af-bea9-49c0-8646-1cf809825b86 · outbound

This paper cites Is the pedestrian going to cross? answering by 2d pose estimation,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Is the pedestrian going to cross? answering by 2d pose estimation,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:59:57.884083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:59:56.186360Z digest=sha256:d9ac3c1b717d9d0db50b9c2a40b7b5104a33264230ad10fe2f7bf41c06502e96

Observation 4e747a18-0a1f-454e-8ab8-2bf16ea658da · outbound

This paper cites Chatgpt: Generative pre-trained transformer,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Chatgpt: Generative pre-trained transformer,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:59:57.715335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:59:56.286479Z digest=sha256:d8e89a6caca55e98bff7165fde033c5de3e9f143f1e383e62649bc4d6596ab02

Observation 35420cca-34d7-4473-94d1-ff1a8173274b · outbound

This paper cites Agreeing to cross: How drivers and pedestrians communicate,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Agreeing to cross: How drivers and pedestrians communicate,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:59:57.594088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:59:56.376470Z digest=sha256:5bc6d6eae828c537edf4858168c2d94d99c298132790268255c8a82f7670cb9e

Observation c5cf1eb8-ead7-4137-abbd-bfce2e73541e · outbound

This paper cites PIE: A large-scale dataset and models for pedestrian intention estimation and trajectory prediction,.

Pedestrian Intention Prediction via Vision-Language Foundation Models PIE: A large-scale dataset and models for pedestrian intention estimation and trajectory prediction,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:59:57.417787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:59:56.479906Z digest=sha256:4b8f9f5ceb4a6937dd8890950e9495387c4f12e33e946efa160e8d97e021ef5e

Observation 398caa99-291c-4851-bb0a-8b63ce9b5728 · outbound

This paper cites GPT-4 Technical Report.

Pedestrian Intention Prediction via Vision-Language Foundation Models GPT-4 Technical Report

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T19:59:56.585090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:59:56.585090Z digest=sha256:000ace21f669d90fd4f5d71066a2478dfcb6d61b28667312003c63a490a09759

Observation 768bfb91-1616-4cd3-82f5-8b553a37cf52 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Pedestrian Intention Prediction via Vision-Language Foundation Models MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T19:59:56.685176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:59:56.685176Z digest=sha256:b26793a7e22997c7bba885a7257f07372eaa5016e7c6466cf3b744ac4b5a894a

Observation f989038e-d8a4-4f79-a1a3-12bf90ba02b3 · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

Pedestrian Intention Prediction via Vision-Language Foundation Models LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T19:59:56.791343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:59:56.791343Z digest=sha256:df0952454dbe5774a331fc37bcb27f03bbfc52d93ffc2375b1264392e4255d7e

Pith citing papers

Observation 0b661ac7-40da-44c8-b177-06d0579874c9 · inbound

PEDESTRIANQA: A Benchmark for Vision-Language Models on Pedestrian Intention and Trajectory Prediction cites this paper.

PEDESTRIANQA: A Benchmark for Vision-Language Models on Pedestrian Intention and Trajectory Prediction Pedestrian Intention Prediction via Vision-Language Foundation Models

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:34:40.934373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T13:25:04.053283Z digest=sha256:36450cb0104980643c2321d6bcab19e210bc62a9f2cee08bca5882b27a577e35