Pith. sign in

Paper Citation Record · LEDGER

STI-Bench: Are MLLMs Ready for Precise Spatial-Temporal World Understanding?

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 23 inbound Pith citation observations for arXiv:2503.23765.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.23765 v6

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 23 of 23 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:13:42.676327Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-10T12:15:01.137692Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e89c9018-c955-431b-b14c-e23c264773f3 · inbound

SpaceR: Reinforcing MLLMs in Video Spatial Reasoning cites this paper.

SpaceR: Reinforcing MLLMs in Video Spatial Reasoning STI-Bench: Are MLLMs Ready for Precise Spatial-Temporal World Understanding?

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-15T15:18:43.867380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T15:18:43.724432Z digest=sha256:5d2701af571076b98ca11c6a0d9c6638e7db0db0965f35862651ffc33b089fe1

Observation 2347f6aa-ac5a-4b8b-a3da-97312c56f293 · inbound

HCRMP: A LLM-Hinted Contextual Reinforcement Learning Framework for Autonomous Driving cites this paper.

HCRMP: A LLM-Hinted Contextual Reinforcement Learning Framework for Autonomous Driving STI-Bench: Are MLLMs Ready for Precise Spatial-Temporal World Understanding?

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T15:13:42.676327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:13:42.676327Z digest=sha256:7b12cc09369d39102fef40a9d1cc834a58ff9bb0c6abb2fb485d48df44b3f32f

Observation 41962da8-18cf-4446-99e1-a953da3e28de · inbound

VerIPO: Cultivating Long Reasoning in Video-LLMs via Verifier-Gudied Iterative Policy Optimization cites this paper.

VerIPO: Cultivating Long Reasoning in Video-LLMs via Verifier-Gudied Iterative Policy Optimization STI-Bench: Are MLLMs Ready for Precise Spatial-Temporal World Understanding?

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:36.284175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:36.284175Z digest=sha256:0e13df64a12699f6f0379f91552ec73e7411c9bbdb842a83f2c8b39e84af4b75

Observation 09605819-077f-4d67-9920-7b2a6992d53d · inbound

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence cites this paper.

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence STI-Bench: Are MLLMs Ready for Precise Spatial-Temporal World Understanding?

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:34:36.938612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T08:34:36.824053Z digest=sha256:45475cffada809546a50f6e3dda4ad7e74ea79d642d84a0ed612dbc7dc12ce68

Observation bb8d9f1c-d264-4eb8-8157-e750ada9dbc4 · inbound

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence cites this paper.

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence STI-Bench: Are MLLMs Ready for Precise Spatial-Temporal World Understanding?

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-22T01:00:51.280287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T00:59:13.826054Z digest=sha256:36325d36cb5e04afce7ad7b3d654e2cc3f5f5109deb8cb7afe0554b1939801c0

Observation 7c68a152-e7db-4071-b81e-cfc0a44eaa7f · inbound

EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World? cites this paper.

EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World? STI-Bench: Are MLLMs Ready for Precise Spatial-Temporal World Understanding?

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:42.647173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:42.647173Z digest=sha256:5a7ae6658c41a939f1d6cbb63d394dd14f38a5bc790b0a545670f0d2e03d2dcb

Observation 2d21fb99-1494-4ff9-9b69-2d621a7edce3 · inbound

The high-speed X-ray camera on AXIS: design and performance updates cites this paper.

The high-speed X-ray camera on AXIS: design and performance updates STI-Bench: Are MLLMs Ready for Precise Spatial-Temporal World Understanding?

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T18:48:03.400031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:48:03.400031Z digest=sha256:6c50c75b2bda185517b475131098704fc6883b081652ffb13c8509eb2cd54715

Observation 2bd7c452-e556-4171-8c3d-a06dcd6c19c2 · inbound

SpatialBench: Benchmarking Multimodal Large Language Models for Spatial Cognition cites this paper.

SpatialBench: Benchmarking Multimodal Large Language Models for Spatial Cognition STI-Bench: Are MLLMs Ready for Precise Spatial-Temporal World Understanding?

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-17T04:59:04.296499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T04:54:59.903644Z digest=sha256:b3c0addc20fcecc38e4a8cf869d0d52debb760458065ec31881e530a74b1954f

Observation 0db04dbc-4cc4-4ee5-8d21-1e2a28e23039 · inbound

SpaceDrive: Infusing Spatial Awareness into VLM-based Autonomous Driving cites this paper.

SpaceDrive: Infusing Spatial Awareness into VLM-based Autonomous Driving STI-Bench: Are MLLMs Ready for Precise Spatial-Temporal World Understanding?

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-22T12:21:31.382952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T12:17:54.325055Z digest=sha256:414dcee9ab26a3233cb5f22342eddf67d12261b9b9a43b3594fcfd6e089d025c

Observation c1482195-fb45-4349-9de8-55d321ef351d · inbound

SpaceDrive: Infusing Spatial Awareness into VLM-based Autonomous Driving cites this paper.

SpaceDrive: Infusing Spatial Awareness into VLM-based Autonomous Driving STI-Bench: Are MLLMs Ready for Precise Spatial-Temporal World Understanding?

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T17:07:43.319564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:07:43.319564Z digest=sha256:71bbe95db6adc2961106575f1c8c222d457ec198e36ce3ebf2dafd9e0cdbe7c0

Observation c4d1d119-d33f-4bd3-a308-897e33bc2a2c · inbound

$M^3-Verse$: A "Spot the Difference" Challenge for Large Multimodal Models cites this paper.

$M^3-Verse$: A "Spot the Difference" Challenge for Large Multimodal Models STI-Bench: Are MLLMs Ready for Precise Spatial-Temporal World Understanding?

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T14:57:29.129862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:57:29.129862Z digest=sha256:2ca061ef731f43874d4b65e997d06726825094a0c498d240fc7bda52cfee0c18

Observation 3bfd95b4-fbf8-4e1b-b685-8985553b8e4f · inbound

SCP: Spatial Causal Prediction in Video cites this paper.

SCP: Spatial Causal Prediction in Video STI-Bench: Are MLLMs Ready for Precise Spatial-Temporal World Understanding?

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-15T16:50:11.225647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T16:47:44.523606Z digest=sha256:098b9e64f3b77a32b3014abbac54c8140d6592944f1ae972248047a98f90985d

Observation d8f2f974-741a-475e-bb0e-f760f5eb2ab1 · inbound

SpatialStack: Layered Geometry-Language Fusion for 3D VLM Spatial Reasoning cites this paper.

SpatialStack: Layered Geometry-Language Fusion for 3D VLM Spatial Reasoning STI-Bench: Are MLLMs Ready for Precise Spatial-Temporal World Understanding?

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:08:04.700971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T22:04:18.591594Z digest=sha256:4db746e90a61c7a81ccad9a98e2d32ee1a6f28b08a9e0b470cc23fee86053b9f

Observation c84b84c9-beb7-4d1e-927c-fb0bc8bf13ba · inbound

LAST: Leveraging Tools as Hints to Enhance Spatial Reasoning for Multimodal Large Language Models cites this paper.

LAST: Leveraging Tools as Hints to Enhance Spatial Reasoning for Multimodal Large Language Models STI-Bench: Are MLLMs Ready for Precise Spatial-Temporal World Understanding?

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:00:47.753493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T19:26:23.660206Z digest=sha256:476171863cbae1c3c1d04bbca0918563a84c6a495d0f2a28309ea703265a4332

Observation 2ae6e612-d86f-48cd-af6d-69d084997e39 · inbound

ViSRA: A Video-based Spatial Reasoning Agent for Multi-modal Large Language Models cites this paper.

ViSRA: A Video-based Spatial Reasoning Agent for Multi-modal Large Language Models STI-Bench: Are MLLMs Ready for Precise Spatial-Temporal World Understanding?

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:46:35.632369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T04:00:23.681682Z digest=sha256:0a081f561b4591e09c3edb791298861209d5aef926d2db1596e433c6b5dc3d31

Observation 34bfc2b2-2a13-4f1c-b304-398a50681a7e · inbound

CaMo: Camera Motion Grounded Evaluation and Training for Vision-Language Models cites this paper.

CaMo: Camera Motion Grounded Evaluation and Training for Vision-Language Models STI-Bench: Are MLLMs Ready for Precise Spatial-Temporal World Understanding?

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T05:28:04.441525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-20T05:27:30.938311Z digest=sha256:f776b39d4461be59686f9dfc34abbc0f13257627fee715b66a264279ef891228

Observation d8b1592a-580c-4e30-ae30-f68d7d18df7f · inbound

Flat-Pack Bench: Evaluating Spatio-Temporal Understanding in Large Vision-Language Models through Furniture Assembly cites this paper.

Flat-Pack Bench: Evaluating Spatio-Temporal Understanding in Large Vision-Language Models through Furniture Assembly STI-Bench: Are MLLMs Ready for Precise Spatial-Temporal World Understanding?

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:21:20.665146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T09:20:32.920925Z digest=sha256:bae8ea1c4460b0bd3792c96b9b4fc9c534bbef8bc89a218988625bc069d40479

Observation dd3f9979-95c4-4dc4-a1f1-cd4eb25cc115 · inbound

Q-GeoMem: Question-Guided Geometric Memory for Video Spatial Reasoning cites this paper.

Q-GeoMem: Question-Guided Geometric Memory for Video Spatial Reasoning STI-Bench: Are MLLMs Ready for Precise Spatial-Temporal World Understanding?

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:23:50.980272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T18:15:28.261185Z digest=sha256:6705e12ebf1b191155f0d7885ab0263fe39b4d0b117e4c2faeb1c7b999173dd1

Observation 1920eedb-05f3-40f7-b2e6-3633131927b2 · inbound

Q-GeoMem: Question-Guided Geometric Memory for Video Spatial Reasoning cites this paper.

Q-GeoMem: Question-Guided Geometric Memory for Video Spatial Reasoning STI-Bench: Are MLLMs Ready for Precise Spatial-Temporal World Understanding?

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-12T15:54:12.909974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T15:54:12.909974Z digest=sha256:882eed63d40070019519531cca015629e0e748067e905ea2cd35b9f32e8cf0d7

Observation 41f59920-a41d-4fce-844a-118668491ab0 · inbound

Beyond 3D VQAs: Injecting 3D Spatial Priors into Vision-Language Models for Enhanced Geometric Reasoning cites this paper.

Beyond 3D VQAs: Injecting 3D Spatial Priors into Vision-Language Models for Enhanced Geometric Reasoning STI-Bench: Are MLLMs Ready for Precise Spatial-Temporal World Understanding?

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:53:13.667132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T07:47:52.739735Z digest=sha256:bb0e26f433d1cf191f9e9bc720aaa393b2b60e8cd293977746ffc3a20923f37e

Observation 29da8a06-3bb5-4983-9a59-bf92699eb05a · inbound

LongSpace: Exploring Long-Horizon Spatial Memory from Perception to Recall in Video cites this paper.

LongSpace: Exploring Long-Horizon Spatial Memory from Perception to Recall in Video STI-Bench: Are MLLMs Ready for Precise Spatial-Temporal World Understanding?

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T12:16:57.229619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T02:16:25.730555Z digest=sha256:829c0f3b191c88b3f98743df8ac135f30898ee8b56e78cc0247065c2eb86a1d7

Observation 3c809614-428b-44b5-80de-97df429d3699 · inbound

Reason, Then Re-reason: Cross-view Revisiting Improves Spatial Reasoning cites this paper.

Reason, Then Re-reason: Cross-view Revisiting Improves Spatial Reasoning STI-Bench: Are MLLMs Ready for Precise Spatial-Temporal World Understanding?

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-07-03T08:57:47.477127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T10:41:41.697216Z digest=sha256:7d2ce9067e52bbb5992f17a4c9b4b464ae0b86c3134cdee421bbd068b91f3ae7

Observation 124f1882-f6c3-404f-b165-96ff428db523 · inbound

AirGroundBench: Probing Spatial Intelligence in Multimodal Large Models under Heterogeneous Multi-View Embodied Collaboration cites this paper.

AirGroundBench: Probing Spatial Intelligence in Multimodal Large Models under Heterogeneous Multi-View Embodied Collaboration STI-Bench: Are MLLMs Ready for Precise Spatial-Temporal World Understanding?

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T04:43:06.920476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T04:34:40.202508Z digest=sha256:73e71ee4b382780061d5229c9d010bff8b089b3234682193bcf1a5ee2ba93eae