Pith. sign in

Paper Citation Record · LEDGER

Pedestrian Intention Prediction via Vision-Language Foundation Models

As of 13 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 1 inbound Pith citation observation for arXiv:2507.04141.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.04141 v1

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:59:56.791343Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-30T13:25:04.053283Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T13:34:40.932917Z

Reference resolution

36 of 36 outbound references displayed

  • verified exact2
  • verified fuzzy25
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 89755a4e-910f-4fd8-b046-ae3d7a248e5b · outbound

This paper cites Autonomous vehicles that interact with pedestrians: A survey of theory and practice,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Autonomous vehicles that interact with pedestrians: A survey of theory and practice,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T19:59:53.815180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:59:53.815180Z digest=sha256:e9d1839bcb641d0b54aba6d10fbf34945385d5e03102e4821c9d602b81f0b3b0

Observation 6cc181ab-3fb6-4ce6-86b4-dfc0be51a1ff · outbound

This paper cites Pedestrian intention prediction: A convolutional bottom-up multi-task approach,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Pedestrian intention prediction: A convolutional bottom-up multi-task approach,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:00:01.306110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:59:53.876159Z digest=sha256:a8c6c30124eed0575a7b3e209f7892ef687b9e6772e284e91675297a58fea167

Observation 64c0a2e8-e139-4b94-9995-dd787f946b28 · outbound

This paper cites St crossingpose: A spatial- temporal graph convolutional network for skeleton-based pedestrian crossing intention prediction,.

Pedestrian Intention Prediction via Vision-Language Foundation Models St crossingpose: A spatial- temporal graph convolutional network for skeleton-based pedestrian crossing intention prediction,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:00:01.180728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:59:53.941914Z digest=sha256:d172920b2ec334b9a388a48caa35f0a54979bef03e7cd3f96f4c734999fa6253

Observation 8a87e33a-a222-4467-b198-55377e59f960 · outbound

This paper cites PIP-Net: Pedestrian Intention Prediction in the Wild.

Pedestrian Intention Prediction via Vision-Language Foundation Models PIP-Net: Pedestrian Intention Prediction in the Wild

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:59:57.232162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:59:54.015290Z digest=sha256:0011288bfe7761152a02041a3c656af1c9b3735fd1e8e3473df2e69702054778

Observation 0757d55c-1420-4cf7-b0c6-b743acbad52b · outbound

This paper cites Pedestrian action an- ticipation using contextual feature fusion in stacked rnns,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Pedestrian action an- ticipation using contextual feature fusion in stacked rnns,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:00:01.062175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:59:54.088410Z digest=sha256:91df8e411885a8b7903f3c9bc197b4a8fb3320ec43d55611e6011f5ad8c4ed11

Observation 1fb433de-cfe6-4c8a-aea3-b8e46effe73e · outbound

This paper cites Do they want to cross? understanding pedestrian intention for behavior prediction,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Do they want to cross? understanding pedestrian intention for behavior prediction,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:00:00.945196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:59:54.171753Z digest=sha256:b939aeb58ecb08aec42bbd72647d3deed7bfc43f953f8b3daf4328f8b7dea0cc

Observation 096beed4-6d41-4232-b594-91f5044818fc · outbound

This paper cites Multi-modal hybrid architecture for pedestrian action prediction,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Multi-modal hybrid architecture for pedestrian action prediction,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:00:00.783093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:59:54.251824Z digest=sha256:7df163d032a3ce1b4642937b88fb219c75cab466979aa0cb0f96563c9c634235

Observation 57272dc5-5f20-416e-9830-2ae0172d7b8e · outbound

This paper cites Pedestrian graph+: A fast pedestrian crossing prediction model based on graph convo- lutional networks,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Pedestrian graph+: A fast pedestrian crossing prediction model based on graph convo- lutional networks,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:00:00.594842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:59:54.354966Z digest=sha256:e054e259a60f4041c6bfb76ce1e2921d8b278b40d89140f2fb1312185bbaa2e5

Observation 6d431e29-b0a2-484d-924a-97647a9e5bcc · outbound

This paper cites Visual reasoning using graph con- volutional networks for predicting pedestrian crossing intention,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Visual reasoning using graph con- volutional networks for predicting pedestrian crossing intention,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:00:00.447152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:59:54.433966Z digest=sha256:ec94d354d491931d3d8e3bed28bb00ce0951cd405ed27be06c0780db00d95a38

Observation e4e1f283-16ef-4438-84c3-4698cef4f6e0 · outbound

This paper cites CAPformer: Pedestrian crossing action prediction using transformer,.

Pedestrian Intention Prediction via Vision-Language Foundation Models CAPformer: Pedestrian crossing action prediction using transformer,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:00:00.297573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:59:54.497886Z digest=sha256:49e5ece2bb4e92242df28bb6e7df4d8532278638f05278b419739a8578215df4

Observation 7521fcd7-f346-42b0-82ff-37c9cd83ea1b · outbound

This paper cites Pit: Progressive interaction transformer for pedestrian crossing intention prediction,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Pit: Progressive interaction transformer for pedestrian crossing intention prediction,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:00:00.122113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:59:54.565677Z digest=sha256:af8b03111c2cb1090c31b0ab031f56e49db6e3822a238307fe6a29bb7c565668

Observation 841f712e-7dfd-4388-9ebb-c567f59da6c3 · outbound

This paper cites Predicting pedestrian inten- tions with multimodal intentformer: A co-learning approach,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Predicting pedestrian inten- tions with multimodal intentformer: A co-learning approach,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:59:59.936485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:59:54.646030Z digest=sha256:589a93da9be0fa9b21de638e2a97952557c4ace0bba48f16173928eec9f554dc

Observation cbc0cc2c-ff16-477a-aa65-247a29a4ad34 · outbound

This paper cites Pedestrian behavior inter- pretation from pose estimation,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Pedestrian behavior inter- pretation from pose estimation,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:59:59.789304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:59:54.713024Z digest=sha256:ef5c85025dd5c05c812a3fc11975fc98d0074273dbe8f7d0ecf3a04fb7b93ba4

Observation e7e25164-28e8-4cf8-9723-5db6a1c92a99 · outbound

This paper cites Multi-scale pedestrian intent prediction using 3d joint information as spatio-temporal representation,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Multi-scale pedestrian intent prediction using 3d joint information as spatio-temporal representation,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:59:59.600266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:59:54.773984Z digest=sha256:7b18107179938acd1e2ff2ada7da471978ad7fb4be0d8ef033d0c8b7bb3e4f76

Observation d72d0405-b7da-412a-ab62-1046703f990a · outbound

This paper cites Spatiotemporal relationship reasoning for pedestrian intent prediction,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Spatiotemporal relationship reasoning for pedestrian intent prediction,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T19:59:54.861620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:59:54.861620Z digest=sha256:68a45b89b08c2bca8e5fb18f98ba460a89aa003cecf1875b672b44006f9120af

Observation 42de927b-2ff4-4089-ad57-1ffa7ba398ac · outbound

This paper cites Real-time intent prediction of pedestrians for autonomous ground vehicles via spatio-temporal densenet,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Real-time intent prediction of pedestrians for autonomous ground vehicles via spatio-temporal densenet,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:59:59.442338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:59:54.933786Z digest=sha256:a6b46dd79dcb9497edfe4a14c78f27b1dc36f9a0cd41e0b782877b99d7e13d41

Observation aa461dfd-b696-4dcc-9b73-896c835a4f53 · outbound

This paper cites Pedestrian-vehicle information modulation for pedestrian crossing intention prediction,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Pedestrian-vehicle information modulation for pedestrian crossing intention prediction,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:59:59.265412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:59:54.998685Z digest=sha256:7b6914e4a01f5327e0176c714d83b2af9f02e697a2f19aa825eb1a87c73b90cc

Observation 615b7547-8719-485d-bfb7-e2ac534abdcf · outbound

This paper cites Causal reasoning in typical computer vision tasks,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Causal reasoning in typical computer vision tasks,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:59:59.071147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:59:55.083231Z digest=sha256:b429111c759ebc8ac5fb17045b8db36394f4ce6d0dc79bac653b7d87446590e5

Observation 18069a05-b99a-4246-9453-288c0dc92147 · outbound

This paper cites Vision language models in autonomous driving: A survey and outlook,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Vision language models in autonomous driving: A survey and outlook,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:59:58.920039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:59:55.171703Z digest=sha256:51b3af2db7120c35186a3fbee99300311b6a11d681d6f24c9375b3b34c5f05f0

Observation 35d3e16b-b52f-4d17-b32c-0190c0006e9a · outbound

This paper cites Gpt-4v takes the wheel: Promises and challenges for pedestrian behavior prediction,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Gpt-4v takes the wheel: Promises and challenges for pedestrian behavior prediction,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:59:58.761794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:59:55.267162Z digest=sha256:efcd7c4fe1edb5602eb0ad100a8a61364821e026917843af661970a8149c5e77

Observation ef2988a4-73fc-4e49-8be5-1a12e03fc931 · outbound

This paper cites Omnipredict: Gpt-4o enhanced multi-modal pedestrian crossing intention prediction.

Pedestrian Intention Prediction via Vision-Language Foundation Models Omnipredict: Gpt-4o enhanced multi-modal pedestrian crossing intention prediction

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:59:58.575555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:59:55.367452Z digest=sha256:0201c434520b228406e986f1a8f21b8638d33f028e33ba09e01933906075d493

Observation a05142d2-747f-4bde-8387-19af28ae0283 · outbound

This paper cites Pedvlm: Pedestrian vision language model for intentions prediction,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Pedvlm: Pedestrian vision language model for intentions prediction,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:59:58.398927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:59:55.459977Z digest=sha256:fd5f21a8a5c15abc6a43e6ecf0aa1c32d8096dffaf126781f960c5a0effe70c0

Observation bf34fca6-b13d-4b6f-8afd-f1890ef41e1f · outbound

This paper cites Cross or wait? predicting pedestrian interaction outcomes at unsignalized crossings,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Cross or wait? predicting pedestrian interaction outcomes at unsignalized crossings,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:59:58.222821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:59:55.569120Z digest=sha256:7f6ec36798239d878114ad684e3493ffa948e98a27ab395c64e886adf45f1827

Observation 5fdb116b-8981-422a-8255-343c2ee44d16 · outbound

This paper cites Feature Importance in Pedestrian Intention Prediction: A Context-Aware Review.

Pedestrian Intention Prediction via Vision-Language Foundation Models Feature Importance in Pedestrian Intention Prediction: A Context-Aware Review

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:59:57.055566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:59:55.663331Z digest=sha256:cf7eaad98051bf02d380a48a0aaa768a03d4ff6add7e4cf6451e93ff7af92289

Observation ca7ff6da-26b8-4be8-b4bd-6873e3b2a77b · outbound

This paper cites Hierarchical Prompting Taxonomy: A Universal Evaluation Framework for Large Language Models Aligned with Human Cognitive Principles.

Pedestrian Intention Prediction via Vision-Language Foundation Models Hierarchical Prompting Taxonomy: A Universal Evaluation Framework for Large Language Models Aligned with Human Cognitive Principles

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T19:59:55.734120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:59:55.734120Z digest=sha256:56f7a69022132a7b2c2d29db25adcdff9f372957cc74efd04c8464c35518e7e3

Observation fc62114f-4913-4812-9b21-2bc71f52489a · outbound

This paper cites Large Language Models Are Human-Level Prompt Engineers.

Pedestrian Intention Prediction via Vision-Language Foundation Models Large Language Models Are Human-Level Prompt Engineers

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T19:59:55.815565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:59:55.815565Z digest=sha256:e8c24a7b8740b759078fa5be88be3902ae8ca368134f4ca7d737db824bbfaf89

Observation 9ceed48e-7f92-4f3b-9b9f-8c3bd66adac8 · outbound

This paper cites Benchmark for evaluating pedestrian action prediction,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Benchmark for evaluating pedestrian action prediction,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:59:58.017275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:59:55.901830Z digest=sha256:0c8f7daa73cd2b8694dae904c2acfc653600b9ea11beedb2727b154e9ed15a98

Observation 99d7b5b0-b85f-469a-8584-a08fa55707f7 · outbound

This paper cites Better Zero-Shot Reasoning with Role-Play Prompting.

Pedestrian Intention Prediction via Vision-Language Foundation Models Better Zero-Shot Reasoning with Role-Play Prompting

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T19:59:55.992136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:59:55.992136Z digest=sha256:e9e2018f122e3f3b590f182f082e321b5ff191b662d77eea04e03065be6babf5

Observation 699aaa9e-581e-4a0c-afd9-1752e25544cb · outbound

This paper cites Good at captioning, bad at counting: Benchmarking GPT-4V on Earth observation data.

Pedestrian Intention Prediction via Vision-Language Foundation Models Good at captioning, bad at counting: Benchmarking GPT-4V on Earth observation data

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T19:59:56.092349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:59:56.092349Z digest=sha256:bc7b8101ceaf4345151ac177c3e56a268e3736b0881aec2fe3d2a84e07a8c352

Observation 1cd8d6af-bea9-49c0-8646-1cf809825b86 · outbound

This paper cites Is the pedestrian going to cross? answering by 2d pose estimation,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Is the pedestrian going to cross? answering by 2d pose estimation,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:59:57.884083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:59:56.186360Z digest=sha256:d6315d893306c1d2585be9f11a5c8d6794cf3b96a351149f9a0ea543b20fd11c

Observation 4e747a18-0a1f-454e-8ab8-2bf16ea658da · outbound

This paper cites Chatgpt: Generative pre-trained transformer,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Chatgpt: Generative pre-trained transformer,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:59:57.715335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:59:56.286479Z digest=sha256:485e560b7a452e949736281557036504b3e32bd6bdb62da7bdcb48706579d300

Observation 35420cca-34d7-4473-94d1-ff1a8173274b · outbound

This paper cites Agreeing to cross: How drivers and pedestrians communicate,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Agreeing to cross: How drivers and pedestrians communicate,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:59:57.594088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:59:56.376470Z digest=sha256:6d1785ad298141eb9d9617fbd0cc9db03e8c7d3eabbd595d6d4f332c9ca9297d

Observation c5cf1eb8-ead7-4137-abbd-bfce2e73541e · outbound

This paper cites PIE: A large-scale dataset and models for pedestrian intention estimation and trajectory prediction,.

Pedestrian Intention Prediction via Vision-Language Foundation Models PIE: A large-scale dataset and models for pedestrian intention estimation and trajectory prediction,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:59:57.417787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:59:56.479906Z digest=sha256:6187f5308ce640953c425865e905f9927fcf0c50e81da8242457b890a284068f

Observation 398caa99-291c-4851-bb0a-8b63ce9b5728 · outbound

This paper cites GPT-4 Technical Report.

Pedestrian Intention Prediction via Vision-Language Foundation Models GPT-4 Technical Report

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T19:59:56.585090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:59:56.585090Z digest=sha256:3f12ba5c2fe10acbd6f6574e5666f3f6d9bc84f84e15c62dda9e16fb209c4648

Observation 768bfb91-1616-4cd3-82f5-8b553a37cf52 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Pedestrian Intention Prediction via Vision-Language Foundation Models MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T19:59:56.685176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:59:56.685176Z digest=sha256:2d8c709dabd1d5b77fc0a96b70ec88c95382485d9846bed9432d5b348dacbfeb

Observation f989038e-d8a4-4f79-a1a3-12bf90ba02b3 · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

Pedestrian Intention Prediction via Vision-Language Foundation Models LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T19:59:56.791343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:59:56.791343Z digest=sha256:d23b4b13dbfe9caa2de1a7252fca17923f05c3cb4c4648821a374b44170e4483

Pith citing papers

Observation 0b661ac7-40da-44c8-b177-06d0579874c9 · inbound

PEDESTRIANQA: A Benchmark for Vision-Language Models on Pedestrian Intention and Trajectory Prediction cites this paper.

PEDESTRIANQA: A Benchmark for Vision-Language Models on Pedestrian Intention and Trajectory Prediction Pedestrian Intention Prediction via Vision-Language Foundation Models

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:34:40.934373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-30T13:25:04.053283Z digest=sha256:de06cc053b2789ce8f38a8889f0efb1639e1202fe297398548d5839fc78de492