Pith. sign in

Paper Citation Record · LEDGER

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models

As of 9 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 0 inbound Pith citation observations for arXiv:2606.09142.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.09142 v1

Coverage vector

measured 53 of 53 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-27T17:13:04.236686Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

53 of 53 outbound references displayed

  • verified exact10
  • verified fuzzy0
  • unresolved40
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b3b16424-1d11-4892-a331-214cc0b4fcae · outbound

This paper cites Pedestrian Behavior Prediction Using Deep Learning Methods for Urban Scenarios: A Review,.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models Pedestrian Behavior Prediction Using Deep Learning Methods for Urban Scenarios: A Review,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-27T17:13:04.236686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:a36b26380d3c8c1655d128d20ff8d674497a24bf68504f708018645d8670c361

Observation bed714ee-335a-474a-9ccf-df0bec14183d · outbound

This paper cites Predicting Pedestrian Crossing Intention in Autonomous Vehicles: A Review,.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models Predicting Pedestrian Crossing Intention in Autonomous Vehicles: A Review,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-27T17:13:04.236686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:ec58a0a6b0c0bbd94cefe046bbe7dfff65392f9675ef7428d87f27c8450ff1ba

Observation 383f2029-8a44-4880-9512-a11756f0e28d · outbound

This paper cites EgoNav: Egocentric Scene-aware Human Trajectory Prediction.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models EgoNav: Egocentric Scene-aware Human Trajectory Prediction

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T00:27:29.693295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:0ee097d50e729b35118b9887fa1115fa4bc7c15d04fd6a0e13cc6ca2744ba102

Observation 07d7f1f7-3e63-4e8b-b169-94183c628fef · outbound

This paper cites Bridging Perspectives: A Survey on Cross-view Collaborative Intelligence with Egocentric-Exocentric Vision,.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models Bridging Perspectives: A Survey on Cross-view Collaborative Intelligence with Egocentric-Exocentric Vision,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-27T17:13:04.236686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:8abcb96cc305a6765f11f8da4e66066e83575ddee9f5e4a4e2daa7923deb0969

Observation ea260b1b-b7f3-4009-b230-f2259288b332 · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video,.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models Ego4d: Around the world in 3,000 hours of egocentric video,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-27T17:13:04.236686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:24650a3463011f4d36fd87190e182f39ec5a8fa2ea53a04a42b3873111caf3a4

Observation 4726cbaa-423f-4dfa-bd47-047b4cc63bcb · outbound

This paper cites An outlook into the future of egocentric vision,.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models An outlook into the future of egocentric vision,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-27T17:13:04.236686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:4d917c84f19f644c5d9747deea4fa38c7465f242a157b467e807f8bda3931bde

Observation e86667fc-7de0-4676-9c80-3a6280931162 · outbound

This paper cites Egolife: Towards egocentric life assistant,.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models Egolife: Towards egocentric life assistant,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-27T17:13:04.236686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:9581652431fcbd27e88615926a5bb25378475017ccba51eb2c2e8f22284684aa

Observation 045b6d09-7e66-4972-bcb1-c3fa6acd7e14 · outbound

This paper cites EgoCogNav: Cognition-aware Human Egocentric Navigation.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models EgoCogNav: Cognition-aware Human Egocentric Navigation

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T00:27:29.660613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:4b5b0afb80827320f8e485f1a0b3d39bd3512fbbf47a02be709bc72d0f55032a

Observation b6facd0e-cec0-4beb-b701-c4b755743cf4 · outbound

This paper cites Lookout: Real-world humanoid egocentric navigation,.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models Lookout: Real-world humanoid egocentric navigation,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-27T17:13:04.236686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:8250323f92b7e103b9fb2e3a36a683bab0cac6e2ce256bcef004382780f69a82

Observation a8e202bb-8cc7-4911-81f0-d38c8fe1147c · outbound

This paper cites HEADS-UP: Head-Mounted Egocentric Dataset for Trajectory Prediction in Blind Assistance Systems.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models HEADS-UP: Head-Mounted Egocentric Dataset for Trajectory Prediction in Blind Assistance Systems

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:27:29.675869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:40311244775676b4b1071e0eddf74b7f9c43cb97ead94f6dabb978cb1e741542

Observation cb1ebd8b-00a6-4617-8d11-29e556b205ee · outbound

This paper cites Egocentric human trajectory forecasting with a wearable camera and multi-modal fusion,.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models Egocentric human trajectory forecasting with a wearable camera and multi-modal fusion,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-27T17:13:04.236686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:104dbe90d3785c900195827b66a8d2beadc59af82b13ff54271a8201009ad4b1

Observation 85033b4b-6be4-41f6-99b7-c197bf8f19e4 · outbound

This paper cites KrishnaCam: Using a longitudinal, single-person, egocentric dataset for scene understanding tasks,.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models KrishnaCam: Using a longitudinal, single-person, egocentric dataset for scene understanding tasks,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-27T17:13:04.236686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:1ead4062536a510f742b50163704942b043c2a43a9f52f40d5ce6429a87b17c3

Observation 833541b9-b19d-4a4c-8302-bd06132b7b72 · outbound

This paper cites Egocentric future localization,.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models Egocentric future localization,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-27T17:13:04.236686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:08849bceb697a31bdc1b6b9ab9eb4c22e56b45031b1a225d2226671016d553fb

Observation 1078c9ee-9460-4c66-845d-30830139088d · outbound

This paper cites Pedestrian intention prediction for autonomous vehicles: A comprehensive survey,.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models Pedestrian intention prediction for autonomous vehicles: A comprehensive survey,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-27T17:13:04.236686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:e6fe669e0bc146dd1b8e625effdc0f1e01d2bf7596365299e28f186a873ddff8

Observation 0915f032-7cf4-4331-a100-210cd1ed0c95 · outbound

This paper cites GPT-4V Takes the Wheel: Promises and Challenges for Pedestrian Behavior Prediction,.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models GPT-4V Takes the Wheel: Promises and Challenges for Pedestrian Behavior Prediction,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-27T17:13:04.236686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:ccd1d7abc23bdee456bf8fb26d7090fd8dfc46414a398c3ab8d34adb66a8289c

Observation d64df1dd-9b5c-48ef-81d1-53719e388f7b · outbound

This paper cites OmniPredict: GPT-4o Enhanced Multi-modal Pedestrian Crossing Intention Prediction,.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models OmniPredict: GPT-4o Enhanced Multi-modal Pedestrian Crossing Intention Prediction,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-27T17:13:04.236686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:0766e2738eb6b21a4840ff10fecead1d506e9a6d96983c920bf83392ad8f8f4d

Observation 451f2793-aea3-46ec-a694-3985bab79be0 · outbound

This paper cites Seeing beyond frames: Zero-shot pedestrian intention prediction with raw temporal video and multimodal cues,.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models Seeing beyond frames: Zero-shot pedestrian intention prediction with raw temporal video and multimodal cues,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-27T17:13:04.236686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:7c19ecec19084618aad83d8357e8c40eb4166d52e34991531c81c6f3dcf64a0e

Observation f3a739d3-43fe-47ca-83a2-37f147ebb3be · outbound

This paper cites Pedestrian Intention Prediction via Vision-Language Foundation Models,.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models Pedestrian Intention Prediction via Vision-Language Foundation Models,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-27T17:13:04.236686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:5ae47f2b29b8c64a357d49125939af1acc8bdc1ad47f42a7e33f6a8453cd16f9

Observation ba4a1df2-e453-46fe-a49b-631b7d2dba3b · outbound

This paper cites Optimizing Vision-Language Model for Road Crossing Intention Estimation,.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models Optimizing Vision-Language Model for Road Crossing Intention Estimation,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-27T17:13:04.236686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:0345184627f1b2b83965cb9b077be43687fc216ae46b82947fdafb3f2ddcd376

Observation a2c2e360-22a0-4e17-8e67-e9159ec078d0 · outbound

This paper cites Pedestrian Vision Language Model for Intentions Prediction,.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models Pedestrian Vision Language Model for Intentions Prediction,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-27T17:13:04.236686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:70e8e548349af4b98167eb87cdd01724e09c328beedd318f823fcd810bdfc856

Observation aee234ee-929e-4b32-8ae2-c61d2a8c9c63 · outbound

This paper cites Application of Vision-Language Model to Pedestrians Behavior and Scene Understanding in Autonomous Driving.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models Application of Vision-Language Model to Pedestrians Behavior and Scene Understanding in Autonomous Driving

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:27:29.690221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:43ab31ba44fe9a423ffcbc1c34778b7778d03b395cf6c93f0ce71852fa72291a

Observation 30a7d0f3-9a98-4bf9-84c2-d60d19d40505 · outbound

This paper cites Vlmped-cot: A large vision-language model with chain-of-thought mechanism for pedestrian crossing intention prediction,.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models Vlmped-cot: A large vision-language model with chain-of-thought mechanism for pedestrian crossing intention prediction,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-27T17:13:04.236686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:b1486cb1af6cfd4060f00ccb8480620cd294b5f7d5b01e92fda0907bfb340232

Observation 3f78746e-079d-4487-a15a-6bf6ffe3da56 · outbound

This paper cites Scaling Egocentric Vision: The EPIC-KITCHENS Dataset.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models Scaling Egocentric Vision: The EPIC-KITCHENS Dataset

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-07-03T00:27:29.665460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:c277a632e1b0e0353aad82cae6edaec64cbfd3a958f28b15333bfe9de9d554ec

Observation d07f28ea-020a-470b-a181-e3513eaab16b · outbound

This paper cites Actor and observer: Joint modeling of first and third-person videos,.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models Actor and observer: Joint modeling of first and third-person videos,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-27T17:13:04.236686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:eea909bb75e3a49aeee99278187b0269039d846116c4ea9bd0bba9056b040846

Observation d6876619-31bf-4cd8-9876-0942c10bed97 · outbound

This paper cites Egovlpv2: Egocentric video-language pre-training with fusion in the backbone,.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models Egovlpv2: Egocentric video-language pre-training with fusion in the backbone,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-27T17:13:04.236686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:871dc552b82a2d72417fd726aabdfbd94ebcc1959e4ef3945973448a17d22a53

Observation 238674b4-3214-4c28-90ac-257b3e17703c · outbound

This paper cites EgoVideo: Exploring Egocentric Foundation Model and Downstream Adaptation.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models EgoVideo: Exploring Egocentric Foundation Model and Downstream Adaptation

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:27:29.686182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:9f46125963736b51bb560e45ea59a65e0023f416d912a019d95b8df6e21da303

Observation 08552ab6-46d6-4623-be20-038fdb3328b9 · outbound

This paper cites Video question answering: Datasets, algorithms and challenges,.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models Video question answering: Datasets, algorithms and challenges,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-06-27T17:13:04.236686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:4ef565b5057513db07ffc9d8f311dc8d1f7458f7f27f419e1c451364bb40b2df

Observation 84b83758-0e70-4074-8500-882c4cf276fa · outbound

This paper cites Video question answering via gradually refined attention over ap- pearance and motion,.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models Video question answering via gradually refined attention over ap- pearance and motion,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-27T17:13:04.236686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:280c35af4c16732c739844a1e08e103fbd01c2672f7bde1ad1e21831952a0990

Observation 526e1b70-464a-40e5-b994-342bd2f59207 · outbound

This paper cites Tgif-qa: Toward spatio- temporal reasoning in visual question answering,.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models Tgif-qa: Toward spatio- temporal reasoning in visual question answering,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-06-27T17:13:04.236686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:d3690097c2746a9bbb2e81f4e931a496864d36ca99fe0eb11c624fe743ad8716

Observation 46884596-2fe6-4aca-bb42-45f976ef3b8e · outbound

This paper cites Activitynet-qa: A dataset for understanding complex web videos via question answering,.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models Activitynet-qa: A dataset for understanding complex web videos via question answering,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-06-27T17:13:04.236686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:00f50993e182f5b5290798f3be4af2a8d4c5c75155c411abb85af7b8159c8c28

Observation d53d2105-5713-4960-a549-09c8bdf2b330 · outbound

This paper cites Next-qa: Next phase of question-answering to explaining temporal actions,.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models Next-qa: Next phase of question-answering to explaining temporal actions,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-06-27T17:13:04.236686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:8670ebe71afcf2c24fae2cac8d25ed2f065aaadf7a228fa7b7a62d0b1be19132

Observation 11f21bec-81d8-46b3-a156-6b2ddc5f1c30 · outbound

This paper cites Agqa: A benchmark for compositional spatio-temporal reasoning,.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models Agqa: A benchmark for compositional spatio-temporal reasoning,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-06-27T17:13:04.236686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:e35dc919d8c25fc6bf617f9e70f4fa15e6bfb4b5aa8c506c1cb3279d7437a807

Observation 0583cc37-0857-4cdb-9046-52b96999dc56 · outbound

This paper cites From representation to reasoning: To- wards both evidence and commonsense reasoning for video question- answering,.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models From representation to reasoning: To- wards both evidence and commonsense reasoning for video question- answering,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-27T17:13:04.236686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:72e46ed85f3864b3894d70bfbbf793ac5851c1f5dc44e507f3ea2de7c31e38f1

Observation 8b5648aa-5505-41a4-b356-09ca8883ed37 · outbound

This paper cites Egotaskqa: Understanding human tasks in egocentric videos,.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models Egotaskqa: Understanding human tasks in egocentric videos,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-06-27T17:13:04.236686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:44f09b4c56a73ba7605333ad3622ff4ab436ab6255402ebe756d7fd584df7eca

Observation 394094fd-5615-476e-bf62-a9a7fce3d8fe · outbound

This paper cites Intentqa: Context-aware video intent reasoning,.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models Intentqa: Context-aware video intent reasoning,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-06-27T17:13:04.236686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:92eb45908b5876ae222786df25f59a2e257a5016b645557c25cbaf136ff2c675

Observation 4dab975f-bb5c-49d6-8f2a-dbbb3f5f33b4 · outbound

This paper cites In the eye of mllm: Benchmarking egocentric video intent understanding with gaze-guided prompting.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models In the eye of mllm: Benchmarking egocentric video intent understanding with gaze-guided prompting

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:27:29.670686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:92d42d4e82245559adf06486eaf657cec5da0b15db5bcda4c8d7870ad7ebcfa6

Observation 531b9dbd-e5fa-4fc7-9901-f87aa630568e · outbound

This paper cites Qwen3 Technical Report.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models Qwen3 Technical Report

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-07-03T00:27:29.667996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:e40d3fd779355bc2cad0e1583fd4112661c3ffc6431c08b5af5ffd0d275021f4

Observation 73064eea-3446-4818-a45d-46df9badeba7 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-07-03T00:27:29.680667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:aa9dbe07a5f9a7bc9ca13976583c0c03db04b4f771c37e2ad547257c2f4ca990

Observation 791ddcdc-faf3-4be6-b06e-8094bc62c4c1 · outbound

This paper cites Grounded question-answering in long egocentric videos,.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models Grounded question-answering in long egocentric videos,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-06-27T17:13:04.236686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:403af3a3fe16414284efaef57b7f5f05385faa29ba46d85a7fb9b78dfcc7c2ff

Observation 6007c9d7-e2e1-48fc-8187-659d98fae112 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models,.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models Chain-of-thought prompting elicits reasoning in large language models,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-06-27T17:13:04.236686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:def2ed446abe32f1bc00b25b88012bf1b2bf608c78c1e767302545915fbd4a44

Observation 5875e811-4bf9-452d-91b9-2b0981eff465 · outbound

This paper cites Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-07-03T00:27:29.678342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:dfa9a2d222f3c48c3dda42444ea6434902ac5eaac2f921aa9d56120923bcc803

Observation f58249e3-3ca7-4fa8-84bf-4296dfe282d1 · outbound

This paper cites Fine-grained visual prompting,.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models Fine-grained visual prompting,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-06-27T17:13:04.236686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:f3737744cfb61e785700f858ba65ef097fa356a7f12ab9385ff7581d51a8df0b

Observation 384b6566-a00f-402e-abaf-f675896a50aa · outbound

This paper cites Large language models are zero-shot reasoners,.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models Large language models are zero-shot reasoners,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-06-27T17:13:04.236686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:fc7ed19b390aaf934ddec315ad7bec6c54c48bf40c8c6093719c8458b0f9ae89

Observation cb17fc77-4a29-4ce5-b5d5-a2ca876ba246 · outbound

This paper cites Grounding dino: Marrying dino with grounded pre-training for open-set object detection,.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models Grounding dino: Marrying dino with grounded pre-training for open-set object detection,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-06-27T17:13:04.236686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:4132281d7ef5652783b73d65b5f4066f779fa173292a43171b2d908cfbb07609

Observation 665f8116-a965-402a-8e22-270f8ac8d387 · outbound

This paper cites Simple online and realtime tracking,.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models Simple online and realtime tracking,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-06-27T17:13:04.236686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:10403bc6f489db8056b14e9a3fdc18623bb60db467ec8347afa2d3f080ddc554

Observation 39d58828-2788-459b-b4a4-2623688d79b3 · outbound

This paper cites Analyzing the behaviors of pedestrians and cyclists in interactions with autonomous systems using controlled experiments: A literature review,.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models Analyzing the behaviors of pedestrians and cyclists in interactions with autonomous systems using controlled experiments: A literature review,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-06-27T17:13:04.236686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:5cb4e002a0aa00fec868375c59bf5ebabea461e82f82d836a7b03e7c40dc3844

Observation 0fb200aa-dc60-4f49-a0d5-c456b8f9001d · outbound

This paper cites Challenges and trends in egocentric vision: A survey,.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models Challenges and trends in egocentric vision: A survey,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-06-27T17:13:04.236686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:e2349fc48bbbfb1915e070f8700fcc99a3b631e0e551eeb2f204defdfd011693

Observation bbb3c8c5-db49-4ea4-98f1-cfa1e63911b3 · outbound

This paper cites Eye Gaze-Informed and Context-Aware Pedestrian Trajectory Prediction in Shared Spaces with Automated Shuttles: A Virtual Reality Study.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models Eye Gaze-Informed and Context-Aware Pedestrian Trajectory Prediction in Shared Spaces with Automated Shuttles: A Virtual Reality Study

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-07-03T00:27:29.662992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:fd8992d7c816c4fb00c0284a4eca6a0e94da240fc7a9696ecc4fedf3145f093f

Observation 6a9dcb69-82a4-4c93-a649-fda35ff0141c · outbound

This paper cites Lora: Low-rank adaptation of large language models.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models Lora: Low-rank adaptation of large language models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-06-27T17:13:04.236686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:7da7ed93c405e64fa884c841d45b8ef91e8cb2dbffb906a8a81db19ab64cff7e

Observation fd29ecfd-42f1-46db-84c6-7f9c799a6921 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Su- pervision,.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models Learning Transferable Visual Models From Natural Language Su- pervision,

Reference 51

Resolution
unresolved
no resolver link, observed 2026-06-27T17:13:04.236686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:1fa3d3919806cdc89b49abdc337633c92753d206e575c6a9a7cd76d668927ff5

Observation a80980e7-0192-4524-a0ef-2b0480689392 · outbound

This paper cites Advancing Egocentric Video Question Answering with Multimodal Large Language Models.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models Advancing Egocentric Video Question Answering with Multimodal Large Language Models

Reference 52

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T00:27:29.683514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:c2a685bd84f405fa9bb3cff46d5c72f3b270e5af471749acd0b691628b5c6c38

Observation 28a30740-3f42-43dd-bc95-8acb2038985a · outbound

This paper cites LatentOmni: Rethinking Omni-Modal Understanding via Unified Audio-Visual Latent Reasoning.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models LatentOmni: Rethinking Omni-Modal Understanding via Unified Audio-Visual Latent Reasoning

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-07-03T00:27:29.673053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:2361f4a8768fda2c7244c11ee81a7b03728b656b826951521e81ed024cda2673

Observation 9494379a-d940-41a3-9c29-07a85af5f911 · outbound

This paper cites Chrono: A simple blueprint for representing time in mllms,.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models Chrono: A simple blueprint for representing time in mllms,

Reference 54

Resolution
unresolved
no resolver link, observed 2026-06-27T17:13:04.236686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:46f2ed398f16ae44ad70a875434d6d4a26990a6ded15738f9ad3079d95e20ce4

Pith citing papers

No inbound Pith citation observations are available.