Pith. sign in

Paper Citation Record · LEDGER

EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 15 inbound Pith citation observations for arXiv:2412.04447.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.04447 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:29:11.117135Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T13:49:52.033876Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ec888612-2f53-4d8c-93fc-a4f9ef993653 · inbound

EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World? cites this paper.

EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World? EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:11.117135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:11.117135Z digest=sha256:c568ce8a2b44dc1aa5a49890729ea40f69eb678d3617e4c5b5064339b64a5d74

Observation 5900ee82-62fd-4757-bae9-cfd7a231a0fc · inbound

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding cites this paper.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.946680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.946680Z digest=sha256:ce01a2c14dc0de88f21af2c7959ab95732ad7733c2c97093e9244b2dc515b151

Observation fed62ce0-a034-485a-bf36-63085f2fbbd6 · inbound

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning cites this paper.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:33.993024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:33.993024Z digest=sha256:1457069abd86198780db1b2da4c0996f7a686d293d14f040e08966356b4f5d05

Observation 0cb8f075-81df-44af-a819-b23870384605 · inbound

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents cites this paper.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:43.779856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:43.779856Z digest=sha256:6a7ec700eb78f59fd00c90f2381f0e8eb38271ee8629bb47b6e9fd06400be158

Observation 0a77f47e-7049-499a-a220-3aa2d3e26933 · inbound

ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planning cites this paper.

ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planning EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-19T03:22:01.033928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T03:18:14.655384Z digest=sha256:3382471d82b461465a9e9b325c31c62d93959235c2b8f2b38793b48fd6583442

Observation 9b6c2843-4cbc-4135-a943-43880b395e63 · inbound

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts cites this paper.

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T13:12:40.150961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:12:40.150961Z digest=sha256:468df03aea790437955c75b72251d2e92fec2baba6b30b22fbf47c1f9a8f8f52

Observation f73b685f-a738-490e-a587-fa771a72ec12 · inbound

MiMo-Embodied: X-Embodied Foundation Model Technical Report cites this paper.

MiMo-Embodied: X-Embodied Foundation Model Technical Report EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:42:05.708154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T20:40:54.096289Z digest=sha256:9cf32eda98bace42f813d6f094195945a9e470dcd1e00c9e2fde48065b809e92

Observation aabe2527-e28c-4601-9756-9fb3e63a8cbe · inbound

Unmasking the Illusion of Embodied Reasoning in Vision-Language-Action Models cites this paper.

Unmasking the Illusion of Embodied Reasoning in Vision-Language-Action Models EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:56:28.347894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T04:27:18.284698Z digest=sha256:076806c5b87fd305047d80970abe387af0f1041b5d342763e1e594db07e57556

Observation 44818236-e821-4b69-a825-9897082ce3de · inbound

Long-Horizon Manipulation via Trace-Conditioned VLA Planning cites this paper.

Long-Horizon Manipulation via Trace-Conditioned VLA Planning EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios

Reference 52

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T14:41:39.077637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T21:10:03.554504Z digest=sha256:d591fb8620c7876a909a66150f112ffb44289348c251314d7a447b1287dad69c

Observation 61229f3d-2fe9-42ee-a71a-d73c8c0e3e7b · inbound

RECIPE: Procedural Planning via Grounding in Instructional Video cites this paper.

RECIPE: Procedural Planning via Grounding in Instructional Video EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:28:05.566482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T06:24:50.714859Z digest=sha256:d0b31d993d06f80c1a2465d7ea52fcd6e6af3fe70c9539771e6e1660026c9cac

Observation 2117b87c-cca6-43bf-8b2d-c1ace6e4cad7 · inbound

RoboStressBench: Benchmarking VLM Robustness to Physical Visual Stress in Embodied Scenes cites this paper.

RoboStressBench: Benchmarking VLM Robustness to Physical Visual Stress in Embodied Scenes EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:02:34.452888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T18:56:37.114948Z digest=sha256:631dbdde404ca5abd39190ddaaf843256abd1bb7a7a35ef5ad95cf0d35d05aa0

Observation 037bf477-997e-4777-a198-11fab6a70ab1 · inbound

When Video Misreads: Closed-Loop Distillation of Reading Heuristics for Exploratory Manipulation Trace QA cites this paper.

When Video Misreads: Closed-Loop Distillation of Reading Heuristics for Exploratory Manipulation Trace QA EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-02T23:07:27.135883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T18:26:47.393632Z digest=sha256:4d4ce16499e543723fb6d0645c44ba1e1f9a1cfd8c31014c39ead3826732a9b0

Observation 0f306271-80f7-4e66-b6de-0a52b681faef · inbound

Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models cites this paper.

Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-07-03T06:07:41.230001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T12:55:47.754632Z digest=sha256:d9b13e9e0fac2ef4169542acc6e4749f0a064bac55325c1aa655635a6a22d369

Observation adde64d9-d24a-4b7c-a1ed-84ed942ec750 · inbound

Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models cites this paper.

Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios

Reference 73

Resolution
unresolved
no resolver link, observed 2026-07-14T18:07:09.018997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T18:07:09.018997Z digest=sha256:b05df13bb0b9b109e0c54ed1bfff91aad436d75a3c716206cb44184429acc999

Observation 05c4a96d-0dd5-4fcd-b5f9-d2b44778827e · inbound

Advancing Omnimodal Embodied Agents from Isolated Skills to Everyday Physical Autonomy cites this paper.

Advancing Omnimodal Embodied Agents from Isolated Skills to Everyday Physical Autonomy EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:49:52.035634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T04:52:51.524022Z digest=sha256:620d3d730ec6fd94138a4aa18e3d6c6c8d9d0e6faaaf34a04ad82ca79e59d68f