Pith. sign in

Paper Citation Record · LEDGER

ST-Think: How Multimodal Large Language Models Reason About 4D Worlds from Ego-Centric Videos

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2503.12542.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.12542 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:27:27.781024Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T17:27:15.721443Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 2cec31b1-09f4-451e-99bf-792fb76c1a22 · inbound

EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning cites this paper.

EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning ST-Think: How Multimodal Large Language Models Reason About 4D Worlds from Ego-Centric Videos

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T23:27:27.781024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:27:27.781024Z digest=sha256:e3638cffe8039e6e3096bed74ccf0581b23f2dba12f1d01d4322918ef4d026e3

Observation 96bb1d28-a914-49dc-8036-b391f3124010 · inbound

MIRAGE: A Multi-modal Benchmark for Spatial Perception, Reasoning, and Intelligence cites this paper.

MIRAGE: A Multi-modal Benchmark for Spatial Perception, Reasoning, and Intelligence ST-Think: How Multimodal Large Language Models Reason About 4D Worlds from Ego-Centric Videos

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T21:14:42.923386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:14:42.923386Z digest=sha256:6caee840a2b76ab727d1b15b88ce278fc224dde5aab61e6190284c5f430f75b7

Observation dde903a5-3660-47ac-b6f7-df37b6899a28 · inbound

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence cites this paper.

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence ST-Think: How Multimodal Large Language Models Reason About 4D Worlds from Ego-Centric Videos

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:34:36.943066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-16T08:34:36.824053Z digest=sha256:23bf3eeb64b1d8fb333ac711de64182b44acfa43ce7deb73aac1e9335f930d74

Observation 6654a363-dc28-444d-82c0-a699d1dcfae2 · inbound

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence cites this paper.

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence ST-Think: How Multimodal Large Language Models Reason About 4D Worlds from Ego-Centric Videos

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-22T01:00:51.454558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-22T00:59:13.826054Z digest=sha256:36e0b81d12a933deeb4bd6b9716cea605b456a64e8d726e82e44df622ea223c3

Observation 7825ea91-ba7e-42fb-a506-0c2806d66248 · inbound

EgoVLM: Policy Optimization for Egocentric Video Understanding cites this paper.

EgoVLM: Policy Optimization for Egocentric Video Understanding ST-Think: How Multimodal Large Language Models Reason About 4D Worlds from Ego-Centric Videos

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:29.112728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:29.112728Z digest=sha256:f454c24fa2ab616a3ba6ee80c88de8a84838b88f537b742646c743199daa2f2a

Observation 81f72875-eaac-4ac7-9633-7c3f7746bcc0 · inbound

SpatialBench: Benchmarking Multimodal Large Language Models for Spatial Cognition cites this paper.

SpatialBench: Benchmarking Multimodal Large Language Models for Spatial Cognition ST-Think: How Multimodal Large Language Models Reason About 4D Worlds from Ego-Centric Videos

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-17T04:59:04.287792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-17T04:54:59.903644Z digest=sha256:c20ebc28ceac3fce5314232b25c64e5db139ff50f01be8fb725f353de3decb23

Observation 53e5f12f-4bd8-436b-ac18-5a7cf1b0c975 · inbound

CT-1: Vision-Language-Camera Models Transfer Spatial Reasoning Knowledge to Camera-Controllable Video Generation cites this paper.

CT-1: Vision-Language-Camera Models Transfer Spatial Reasoning Knowledge to Camera-Controllable Video Generation ST-Think: How Multimodal Large Language Models Reason About 4D Worlds from Ego-Centric Videos

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:56:00.403306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T16:53:54.166457Z digest=sha256:280ae01799bd5c7420702a73962f362aab81d74c7222ec143f5f02aba1c26603

Observation b2c6297c-2616-41dd-aa8e-d6ede03348d7 · inbound

ViSRA: A Video-based Spatial Reasoning Agent for Multi-modal Large Language Models cites this paper.

ViSRA: A Video-based Spatial Reasoning Agent for Multi-modal Large Language Models ST-Think: How Multimodal Large Language Models Reason About 4D Worlds from Ego-Centric Videos

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:46:35.883391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-12T04:00:23.681682Z digest=sha256:7af6fe48e21b537a25a12c0853d4c4bb84846d3f85bff0a9613ce5fc076e1d53

Observation 2cb0bb14-a96f-4199-a15b-2cfaab223378 · inbound

EgoProx: Evaluating MLLMs on Egocentric 3D Proximity Reasoning Across a Cognitive Hierarchy cites this paper.

EgoProx: Evaluating MLLMs on Egocentric 3D Proximity Reasoning Across a Cognitive Hierarchy ST-Think: How Multimodal Large Language Models Reason About 4D Worlds from Ego-Centric Videos

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:34:40.492377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T13:28:43.538541Z digest=sha256:64b3306eb7c6ecd0519c9860d95449190c05dbe578dd4adfda2582cd8333c4b9

Observation c2728fcc-64a1-4f5c-b8ad-b65e11ea6ebc · inbound

Watch, Remember, Reason: Human-View Video Understanding with MLLMs cites this paper.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs ST-Think: How Multimodal Large Language Models Reason About 4D Worlds from Ego-Centric Videos

Reference 238

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:15.723158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:796ed75c8963351ee81e61454a38db1f10464d1f92d3fe8c50c9b5cc62d80f05