Pith. sign in

Paper Citation Record · LEDGER

STAR: A Benchmark for Situated Reasoning in Real-World Videos

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 38 inbound Pith citation observations for arXiv:2405.09711.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.09711 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 38 of 38 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:33:32.602468Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

22
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 2c6edefc-924a-4814-96bb-d0ab6ccd99b4 · inbound

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models cites this paper.

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models STAR: A Benchmark for Situated Reasoning in Real-World Videos

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:01:53.971264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T06:01:53.730356Z digest=sha256:2847b2a649a25ee0253a20296b52a3faaf3219633af72fcb742b3f43b02f2bf6

Observation de30d0f5-b527-42d3-bfed-85727a24d0ad · inbound

LLaVA-Video: Video Instruction Tuning With Synthetic Data cites this paper.

LLaVA-Video: Video Instruction Tuning With Synthetic Data STAR: A Benchmark for Situated Reasoning in Real-World Videos

Reference 135

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:20:32.807920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:b4afb5cc9d1d4464e733700c7b08ed3d0ec539877995cdb8c25d4738909a4e97

Observation 64665e60-14af-4d31-81d3-bcfef72206ec · inbound

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling cites this paper.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling STAR: A Benchmark for Situated Reasoning in Real-World Videos

Reference 260

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:23:58.156138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:6cccb9151b8270aa9b9f304e892ba96dca472ea951c1738a68a29537f62fe5a0

Observation 062db950-583c-4fbb-871d-02d519762ad8 · inbound

Rethinking Causal Mask Attention for Vision-Language Inference cites this paper.

Rethinking Causal Mask Attention for Vision-Language Inference STAR: A Benchmark for Situated Reasoning in Real-World Videos

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T14:33:32.602468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:33:32.602468Z digest=sha256:2367e9eace2c32b85568a3c26a2ed3cd905b3cb135527b30ddc5a4581fffea99

Observation 7ee1a5e5-7b2c-498f-90ae-a196f9785ab4 · inbound

ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding cites this paper.

ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding STAR: A Benchmark for Situated Reasoning in Real-World Videos

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:40:46.242481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:40:46.242481Z digest=sha256:c41e7d4f675413fd82c7a2337b2060b27a3d50a784d8d360b3b05f26ae83f85e

Observation b23e32ea-f4c2-45bd-a21a-5b55d020807b · inbound

Reinforcing Video Reasoning with Focused Thinking cites this paper.

Reinforcing Video Reasoning with Focused Thinking STAR: A Benchmark for Situated Reasoning in Real-World Videos

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:13.702673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:13.702673Z digest=sha256:57a5d1d11518f44b161a1dc7a31218f56a3752e30689f51970606e58d52bc1b1

Observation 759cfd64-7b93-4d6e-9e2e-1b9fa7556250 · inbound

InterRVOS: Interaction-aware Referring Video Object Segmentation cites this paper.

InterRVOS: Interaction-aware Referring Video Object Segmentation STAR: A Benchmark for Situated Reasoning in Real-World Videos

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:29.288871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:29:29.288871Z digest=sha256:da15b5c82ae850729ad06be3a5f13ae5ebeacc22712c8a6807e350cb5bed3aff

Observation 28fd6f61-85cf-4000-9df3-5bdea4755f80 · inbound

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks cites this paper.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks STAR: A Benchmark for Situated Reasoning in Real-World Videos

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.357596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.357596Z digest=sha256:5524d75907c8f02e0cd24df9de60117c05d71c1e2e7bc34d3e746cddedf8558b

Observation b0beb357-0bc5-4803-9b93-2e37a27c3837 · inbound

TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision cites this paper.

TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision STAR: A Benchmark for Situated Reasoning in Real-World Videos

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T04:54:24.595020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:54:24.595020Z digest=sha256:d6f3a4a30f917c637f55bad73bf958ce6ed13f24df7e5f12b3f28a15c54ad281

Observation 9a8e578c-f3ed-4db6-b45f-5eba1d035eca · inbound

VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos cites this paper.

VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos STAR: A Benchmark for Situated Reasoning in Real-World Videos

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T04:22:56.051254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:22:56.051254Z digest=sha256:1089224d1e7839727df151610f47d42ced0978f0980e79366926fba78dc9348f

Observation 9fc06ce8-f3a8-4ffd-a395-9a0bd9cf8611 · inbound

IPFormer-VideoLLM: Enhancing Multi-modal Video Understanding for Multi-shot Scenes cites this paper.

IPFormer-VideoLLM: Enhancing Multi-modal Video Understanding for Multi-shot Scenes STAR: A Benchmark for Situated Reasoning in Real-World Videos

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T22:37:39.190020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:37:39.190020Z digest=sha256:e95c6eb63a902dc116c1e9224353946ed50ffa6d632ea266deae7acd695bafe1

Observation fd965cf1-4793-44c8-9b71-6bfdcfe742e1 · inbound

InterAct-Video: Reasoning-Rich Video QA for Urban Traffic cites this paper.

InterAct-Video: Reasoning-Rich Video QA for Urban Traffic STAR: A Benchmark for Situated Reasoning in Real-World Videos

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T15:53:53.840655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:53:53.840655Z digest=sha256:3505fbfe63deafc93e4e3d9c12a58ee0db2e6995157bdb731be0be7cb04ae0ef

Observation bccc7492-cf70-4dea-a402-134cf0aea9e6 · inbound

POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency cites this paper.

POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency STAR: A Benchmark for Situated Reasoning in Real-World Videos

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T13:22:35.455844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:22:35.455844Z digest=sha256:01f7fce953f4fcdf2b78825d26739d99d53f83d07c2b461fd81a19ebe1ff978e

Observation 905837be-2b9a-4bfc-bfc4-713a1d332ae9 · inbound

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs cites this paper.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs STAR: A Benchmark for Situated Reasoning in Real-World Videos

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:07.271934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:07.271934Z digest=sha256:8f762446fcc627d6639ad27de46a4ff78a5582e06d1bef6c0dfc3d4219466647

Observation 735132a7-13bc-4a9b-b53a-e64c0780247c · inbound

Agentic Physical AI toward a Domain-Specific Foundation Model for Nuclear Reactor Control cites this paper.

Agentic Physical AI toward a Domain-Specific Foundation Model for Nuclear Reactor Control STAR: A Benchmark for Situated Reasoning in Real-World Videos

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T17:00:24.028688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T16:57:19.490074Z digest=sha256:a3e88d95f678082ad528e881bb65be665cc2e873ffdcb69cecf04769b877bf26

Observation 120f12dd-3a86-4f5d-9f94-c11e9fbba5a2 · inbound

Mimic Human Cognition, Master Multi-Image Reasoning: A Meta-Action Framework for Enhanced Visual Understanding cites this paper.

Mimic Human Cognition, Master Multi-Image Reasoning: A Meta-Action Framework for Enhanced Visual Understanding STAR: A Benchmark for Situated Reasoning in Real-World Videos

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-03T11:13:16.935561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:13:16.935561Z digest=sha256:bc0af82278d170939157612122024da0ad32a63dca0370c8ee2a52cd2d7af8e5

Observation 66a596ee-822f-4326-9cc1-e40f7e93a353 · inbound

Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark cites this paper.

Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark STAR: A Benchmark for Situated Reasoning in Real-World Videos

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:08:04.316200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-14T22:05:07.326202Z digest=sha256:d79df9f7a9f90e9b424d7b5dce9935dc8e0134a182c78ab3a628017ffc27722e

Observation cc5e434d-a353-4b6b-8158-55286c0d98ed · inbound

VideoStir: Understanding Long Videos via Spatio-Temporally Structured and Intent-Aware RAG cites this paper.

VideoStir: Understanding Long Videos via Spatio-Temporally Structured and Intent-Aware RAG STAR: A Benchmark for Situated Reasoning in Real-World Videos

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:15:50.112085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T19:15:02.124035Z digest=sha256:95ec5727c1f924951b8d66cf792de8a80fb77e3de3b068554b275672159a537a

Observation 123b9b79-0ab7-486c-a028-10ec04398d9d · inbound

Chain-of-Glimpse: Search-Guided Progressive Object-Grounded Reasoning for Video Understanding cites this paper.

Chain-of-Glimpse: Search-Guided Progressive Object-Grounded Reasoning for Video Understanding STAR: A Benchmark for Situated Reasoning in Real-World Videos

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:05:22.204503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T12:03:09.408019Z digest=sha256:ac1e4c6606381a24b3aeece0c61cda753f8ccea806fd528ce91ae2f42cdd2e22

Observation 1ed24a58-b4f4-4690-898a-625f32d58e0e · inbound

Chain-of-Glimpse: Search-Guided Progressive Object-Grounded Reasoning for Video Understanding cites this paper.

Chain-of-Glimpse: Search-Guided Progressive Object-Grounded Reasoning for Video Understanding STAR: A Benchmark for Situated Reasoning in Real-World Videos

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-19T17:37:41.709145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T17:34:10.344111Z digest=sha256:7b6064c82c626ab541736c64ad49f975ae016c94953443df2c0d59425d61ab3e

Observation 5aa56276-6bf3-47d2-9082-8289ec639140 · inbound

NICE FACT: Diagnosing and Calibrating VLMs in Quantitative Reasoning for Kinematic Physics cites this paper.

NICE FACT: Diagnosing and Calibrating VLMs in Quantitative Reasoning for Kinematic Physics STAR: A Benchmark for Situated Reasoning in Real-World Videos

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:41:31.979277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T02:23:04.253007Z digest=sha256:9615c7244590626743b4ad52edbe69dfb758fcc3b338af9dc806e027cedb5c24

Observation fb947a5c-c059-4932-a558-f2947109e8d2 · inbound

EgoMemReason: A Memory-Driven Reasoning Benchmark for Long-Horizon Egocentric Video Understanding cites this paper.

EgoMemReason: A Memory-Driven Reasoning Benchmark for Long-Horizon Egocentric Video Understanding STAR: A Benchmark for Situated Reasoning in Real-World Videos

Reference 55

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:06:24.536554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-12T04:37:46.090777Z digest=sha256:768b2344536917ca7f3b7b4d1e8f74c8eec31c761e26b7726238bdaa08beed30

Observation f67abafe-e20d-4377-87db-50dc8928964e · inbound

TOC-Bench: A Temporal Object Consistency Benchmark for Video Large Language Models cites this paper.

TOC-Bench: A Temporal Object Consistency Benchmark for Video Large Language Models STAR: A Benchmark for Situated Reasoning in Real-World Videos

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:31:26.645656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T04:13:21.487431Z digest=sha256:c68c6c2bb71405b6ec87249623c53b2433b86051581c77afde348df58e342818

Observation 3da4fe23-a46b-445d-9362-f21ccd5487f8 · inbound

TOC-Bench: A Temporal Object Consistency Benchmark for Video Large Language Models cites this paper.

TOC-Bench: A Temporal Object Consistency Benchmark for Video Large Language Models STAR: A Benchmark for Situated Reasoning in Real-World Videos

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:57:28.352513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T06:53:42.726350Z digest=sha256:5602a84a28daf7e8ae091bed54a18ba2934f6feb41cc35c746d351384fe1d2ec

Observation 15ac86a5-8aaf-41b1-ad16-b6159506d969 · inbound

OProver: A Unified Framework for Agentic Formal Theorem Proving cites this paper.

OProver: A Unified Framework for Agentic Formal Theorem Proving STAR: A Benchmark for Situated Reasoning in Real-World Videos

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:48:23.101517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-20T14:43:46.517807Z digest=sha256:a641d57829979ecc6105e685f3e4ee356bfdc04053a5580632532f72c81b56b5

Observation 27867a6f-0e53-492e-93be-22038c983a16 · inbound

EvoVid: Temporal-Centric Self-Evolution for Video Large Language Models cites this paper.

EvoVid: Temporal-Centric Self-Evolution for Video Large Language Models STAR: A Benchmark for Situated Reasoning in Real-World Videos

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-22T07:21:12.820432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T07:19:30.508843Z digest=sha256:1c84d580c5e8939eacba610a8d9075a781286fac03bfa3cadf88f76d2fe88196

Observation 0e77163e-16be-4107-86bd-f14845461c73 · inbound

What-If World: A Causal Benchmark for General World Models in Embodied Scenarios cites this paper.

What-If World: A Causal Benchmark for General World Models in Embodied Scenarios STAR: A Benchmark for Situated Reasoning in Real-World Videos

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:23:50.384160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T18:23:22.987086Z digest=sha256:36df784344c192cf2c2fa74c7bca028e72499d54a82a49b0f753cda054e5a6dd

Observation f6646ba2-5194-4238-8731-f1f8dca3350d · inbound

StoryVideoQA: Scaling Deep Video Understanding with a Large-Scale, Multi-Genre and Auto-Generated Dataset cites this paper.

StoryVideoQA: Scaling Deep Video Understanding with a Large-Scale, Multi-Genre and Auto-Generated Dataset STAR: A Benchmark for Situated Reasoning in Real-World Videos

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T12:26:57.200530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T02:05:47.810096Z digest=sha256:07bd653e249e111eba00f5bd2345569b0484063972420f4b98763809c9fc7968

Observation a0677d9f-3162-40e3-93f3-a10432fe406a · inbound

ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? cites this paper.

ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? STAR: A Benchmark for Situated Reasoning in Real-World Videos

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-07-02T20:27:22.643440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T20:23:18.667677Z digest=sha256:a0f74587da4270ee3d7bdde691ff35e9574e48915bc293e0792280294e27d8b3

Observation c14895da-190b-4922-aee1-fe6cb16ea015 · inbound

FeVOS: Foresight Expression Video Object Segmentation cites this paper.

FeVOS: Foresight Expression Video Object Segmentation STAR: A Benchmark for Situated Reasoning in Real-World Videos

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T19:30:08.033476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-25T21:13:21.500674Z digest=sha256:710ecca6957089442080e53c4cf90d39d58980abbb7c4d67935d2e450fd2b237

Observation ae2fc7de-4929-4e7f-91f1-86a672fb19dc · inbound

Reflect-R1: Evidence-Driven Reflection for Self-Correction in Long Video Understanding cites this paper.

Reflect-R1: Evidence-Driven Reflection for Self-Correction in Long Video Understanding STAR: A Benchmark for Situated Reasoning in Real-World Videos

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T17:15:50.640611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T04:07:52.232959Z digest=sha256:316662775238e22877453246230161a12a4c75a04daa3e791b3a196e232aa609

Observation 735146d8-3a12-4ed6-b5a0-2d2961f19a2f · inbound

Reflect-R1: Evidence-Driven Reflection for Self-Correction in Long Video Understanding cites this paper.

Reflect-R1: Evidence-Driven Reflection for Self-Correction in Long Video Understanding STAR: A Benchmark for Situated Reasoning in Real-World Videos

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T09:54:34.485078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T09:53:49.542693Z digest=sha256:cbc151931e7f75c1cd4c53016c398a4b86ca612dd653644135b2c31e0e8d48e7

Observation 80288e54-4781-49cd-ad48-7730585a28a2 · inbound

Reflect-R1: Evidence-Driven Reflection for Self-Correction in Long Video Understanding cites this paper.

Reflect-R1: Evidence-Driven Reflection for Self-Correction in Long Video Understanding STAR: A Benchmark for Situated Reasoning in Real-World Videos

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T09:05:37.282597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-01T06:47:47.324684Z digest=sha256:85e100629f5833e3d23d8102b7c28a05df4cf2eaba26d855e60fdf9b9aa3bc10

Observation 13711394-58a0-4b2c-a88a-453d55d6f546 · inbound

HumanMoveVQA: Can Video MLLMs reason about human movement in videos? cites this paper.

HumanMoveVQA: Can Video MLLMs reason about human movement in videos? STAR: A Benchmark for Situated Reasoning in Real-World Videos

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-06-29T19:13:53.048614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T04:53:11.830488Z digest=sha256:07b3292ea44b0a7202fb223316acc7a2eb4009dc427802cecc3a549e6f939c8c

Observation 0978f925-f5ae-4cf8-97a3-fee8566e1573 · inbound

HumanMoveVQA: Can Video MLLMs reason about human movement in videos? cites this paper.

HumanMoveVQA: Can Video MLLMs reason about human movement in videos? STAR: A Benchmark for Situated Reasoning in Real-World Videos

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-06-30T09:54:35.318699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T09:46:56.063333Z digest=sha256:bceefd713acd99fd8d559f80caa25642fd3c39d2bbcc9306f4c114d721283130

Observation a4642844-0455-4f73-8e94-04cae7b4956a · inbound

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning cites this paper.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning STAR: A Benchmark for Situated Reasoning in Real-World Videos

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-06-30T06:24:19.569669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:08fe895e87325e9a19a0a308720ac96d4d8115544cc4c34f2e8c6ce6727e20e4

Observation 6bf62ec9-5422-40a6-9e7f-ff5da6e0f4dd · inbound

MuseBench: Benchmarking Intent-Level Audiovisual Arts Understanding in MLLMs cites this paper.

MuseBench: Benchmarking Intent-Level Audiovisual Arts Understanding in MLLMs STAR: A Benchmark for Situated Reasoning in Real-World Videos

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-06-30T06:34:19.497431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T06:25:38.593423Z digest=sha256:389efb9165e8db6183fce4a2141614804a0613ba2bd381da03edfb849c9a9bc8

Observation 88dcadad-d535-4981-8826-bc546bb0920b · inbound

IMBench: A Benchmark for Intuitive Robotic Manipulation cites this paper.

IMBench: A Benchmark for Intuitive Robotic Manipulation STAR: A Benchmark for Situated Reasoning in Real-World Videos

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T22:45:07.891633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:45:07.891633Z digest=sha256:af0e67383831c9e2c9f694d4164621c6f8a293c6900e2793a3a3d2c1a2aa5563