Pith. sign in

Paper Citation Record · LEDGER

ST-Think: How Multimodal Large Language Models Reason About 4D Worlds from Ego-Centric Videos

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2503.12542.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.12542 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:14:29.112728Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T17:27:15.721443Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation dde903a5-3660-47ac-b6f7-df37b6899a28 · inbound

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence cites this paper.

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence ST-Think: How Multimodal Large Language Models Reason About 4D Worlds from Ego-Centric Videos

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:34:36.943066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T08:34:36.824053Z digest=sha256:c8f90c75e00fdeab2055e6f6cfd6782aa083dfa503effaf942940807ed348a1a

Observation 6654a363-dc28-444d-82c0-a699d1dcfae2 · inbound

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence cites this paper.

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence ST-Think: How Multimodal Large Language Models Reason About 4D Worlds from Ego-Centric Videos

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-22T01:00:51.454558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T00:59:13.826054Z digest=sha256:b9fdb7978c3b5235c74231e86e45daf42ffbe674a46dadde5f8344d4db025cb8

Observation 7825ea91-ba7e-42fb-a506-0c2806d66248 · inbound

EgoVLM: Policy Optimization for Egocentric Video Understanding cites this paper.

EgoVLM: Policy Optimization for Egocentric Video Understanding ST-Think: How Multimodal Large Language Models Reason About 4D Worlds from Ego-Centric Videos

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:29.112728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:29.112728Z digest=sha256:bb06edd1c905bc307f8b60cdf96547f4fd94f1205a599a3afab69b44e5abd998

Observation 81f72875-eaac-4ac7-9633-7c3f7746bcc0 · inbound

SpatialBench: Benchmarking Multimodal Large Language Models for Spatial Cognition cites this paper.

SpatialBench: Benchmarking Multimodal Large Language Models for Spatial Cognition ST-Think: How Multimodal Large Language Models Reason About 4D Worlds from Ego-Centric Videos

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-17T04:59:04.287792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T04:54:59.903644Z digest=sha256:2d9bba8fb97e8472b0623056783e12d8dc48d96167a81ec422547b6cb7fe2646

Observation 53e5f12f-4bd8-436b-ac18-5a7cf1b0c975 · inbound

CT-1: Vision-Language-Camera Models Transfer Spatial Reasoning Knowledge to Camera-Controllable Video Generation cites this paper.

CT-1: Vision-Language-Camera Models Transfer Spatial Reasoning Knowledge to Camera-Controllable Video Generation ST-Think: How Multimodal Large Language Models Reason About 4D Worlds from Ego-Centric Videos

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:56:00.403306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T16:53:54.166457Z digest=sha256:82a7eda9a3b690991cf31f8b07bebd6731da39d5c2be2ac7ebee7f774fdcb2d2

Observation b2c6297c-2616-41dd-aa8e-d6ede03348d7 · inbound

ViSRA: A Video-based Spatial Reasoning Agent for Multi-modal Large Language Models cites this paper.

ViSRA: A Video-based Spatial Reasoning Agent for Multi-modal Large Language Models ST-Think: How Multimodal Large Language Models Reason About 4D Worlds from Ego-Centric Videos

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:46:35.883391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T04:00:23.681682Z digest=sha256:f10d29bad844b5b38c798bf9310286837d25f16c631f3bb3771cf63f073a7ccd

Observation 2cb0bb14-a96f-4199-a15b-2cfaab223378 · inbound

EgoProx: Evaluating MLLMs on Egocentric 3D Proximity Reasoning Across a Cognitive Hierarchy cites this paper.

EgoProx: Evaluating MLLMs on Egocentric 3D Proximity Reasoning Across a Cognitive Hierarchy ST-Think: How Multimodal Large Language Models Reason About 4D Worlds from Ego-Centric Videos

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:34:40.492377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T13:28:43.538541Z digest=sha256:cfbdfaec30adda6929ccf5ea736c873aa8f03c91c8c63741b66c12ee08d7fd8a

Observation c2728fcc-64a1-4f5c-b8ad-b65e11ea6ebc · inbound

Watch, Remember, Reason: Human-View Video Understanding with MLLMs cites this paper.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs ST-Think: How Multimodal Large Language Models Reason About 4D Worlds from Ego-Centric Videos

Reference 238

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:15.723158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:d0c0c0216503bd4729266f0e62654ac37da69e392a346c82e44d1473828fb09a