Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T15:44:49.592008Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 1 inbound Pith citation observation for arXiv:2507.15130.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T15:44:49.592008Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-21T11:39:15.308355Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-21T11:40:03.350024Z
62 of 62 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation d1ef1572-a27c-4efd-b6ee-8a89deccadc8 · outbound
Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction When will you do what?-anticipating temporal occurrences of activities
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cb381446-e785-4243-9a8f-b90d5f812668 · outbound
Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction GPT-4 Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72238229-4323-4029-bb15-8fa4759981fb · outbound
Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Flamingo: a visual language model for few-shot learning
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fff40a00-3055-483b-99c0-73bf8ca694e5 · outbound
Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Hiervl: Learning hierarchical video- language embeddings
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f9f18253-afc6-42b3-9605-c383a5d72b26 · outbound
Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Procedure planning in instructional videos
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fbf579ff-4efe-43d4-889e-b9e8957a91b8 · outbound
Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Spatialvlm: Endow- ing vision-language models with spatial reasoning capabili- ties
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b1b39347-9368-4e03-9849-e3285bb02838 · outbound
Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction EgoPlan-Bench: Benchmarking Multimodal Large Language Models for Human-Level Planning
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92f198ec-ee75-48ac-8f2a-66d84f03eb99 · outbound
Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c23d4da-96e1-43a4-b120-7c4db34acdaa · outbound
Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Better & Faster Large Language Models via Multi-token Prediction
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b8f4eb0-eead-471e-a8b0-4397c62e683a · outbound
Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Ego4d: Around the world in 3,000 hours of egocentric video
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1e7b96ed-19db-489a-b249-63f02ce724b2 · outbound
Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Reasoning with Language Model is Planning with World Model
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66bf25b6-aa1e-4d23-82b3-1a50ffecafe2 · outbound
Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction LoRA: Low-Rank Adaptation of Large Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3231d02b-bfad-4a3e-9862-4454351e6988 · outbound
Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Vtimellm: Empower llm to grasp video moments
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b452a7b1-8d21-4b5b-a131-554a48270480 · outbound
Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Language models as zero-shot planners: Extract- ing actionable knowledge for embodied agents
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5ebe8f6a-baad-455f-a467-80ef575e3610 · outbound
Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Propose, Assess, Search: Harnessing LLMs for Goal-Oriented Planning in Instructional Videos
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0f03c10e-7a52-4b22-a0c9-7d99c64ba0c9 · outbound
Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Palm: Predicting actions through language models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation df76deea-f9a7-4c2b-9cea-6dc95f0b37bf · outbound
Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction LLaVA-OneVision: Easy Visual Task Transfer
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbd8abc5-45e6-4a6a-9c4d-1e2ad8ac49ec · outbound
Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 010fbb76-6cf4-4479-82d7-62955d2fc731 · outbound
Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction VideoChat: Chat-Centric Video Understanding
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e2c3f67-7f3d-463a-ae90-eab374619d4d · outbound
Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Mvbench: A comprehensive multi-modal video understand- ing benchmark
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 83e61c2f-2559-4abc-bed2-2b054d6e6cc0 · outbound
Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Llama-vid: An image is worth 2 tokens in large language models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e32811d9-557a-45dc-8a30-0dcf25808e09 · outbound
Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Skip-plan: Procedure plan- ning in instructional videos via condensed action space learn- ing
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5d031845-fb8f-4d37-83cd-620bed5edc2f · outbound
Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1826b2e1-b5d0-4166-95ad-9199fcbaf75d · outbound
Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction LLM+P: Empowering Large Language Models with Optimal Planning Proficiency
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d423aea-4a84-411c-8400-7a297d87cd58 · outbound
Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Visual instruction tuning
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 19b722b5-fe3a-422c-807c-cf83d362d10d · outbound
Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction A language-first approach for procedure planning
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ec54eccf-757e-4470-a178-8c20ac2ca022 · outbound
Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Intention-conditioned long-term human egocentric action anticipation
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e70b9c84-0c3a-4466-b36c-262583ebbfa4 · outbound
Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Can’t make an omelette without breaking some eggs: Plausible action anticipation using large video- language models
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7e57bf75-9560-4cb3-a881-8c99d592fa92 · outbound
Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Any- mal: An efficient and scalable any-modality augmented lan- guage model
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation acc44834-e809-4f34-9a37-ff9452287226 · outbound
Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Embodiedgpt: Vision-language pre-training via embodied chain of thought
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c0cd7f5f-bff3-49d8-b818-96adb1df2653 · outbound
Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Ego-topo: Environment affordances from egocentric video
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b8de1b13-6e6a-4cc0-b53e-59a503d393c0 · outbound
Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Why not use your text- book? knowledge-enhanced procedure planning of instruc- tional videos
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4d55b9dd-512e-4f59-9f33-0bc19cd72ffc · outbound
Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Re- thinking learning approaches for long-term action anticipa- tion
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 003f8245-b33f-43fc-8cb8-f28d80703fbb · outbound
Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Do Pre-trained Vision-Language Models Encode Object States?
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 23e267b8-7b9e-44fa-ba12-6ecdce73340d · outbound
Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction SCHEMA: State CHangEs MAtter for Procedure Planning in Instructional Videos
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb7d0d66-827a-41ca-975d-84c6678d8883 · outbound
Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Pretrained language models as visual planners for human assistance
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ad121c6a-9f47-4e65-820a-c8f927d7c9d2 · outbound
Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction EgoVideo: Exploring Egocentric Foundation Model and Downstream Adaptation
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 212abde4-2bc1-42ed-82a1-c02e251bec32 · outbound
Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d5e895f-1c46-43d7-9700-6300edf4e55b · outbound
Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Timechat: A time-sensitive multimodal large lan- guage model for long video understanding
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f2e9711-22f4-4fe6-98c5-e186b4b5e058 · outbound
Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Llm-planner: Few-shot grounded planning for embodied agents with large language models
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0332c79f-4bd2-4097-a37d-4d215e106510 · outbound
Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Moviechat: From dense token to sparse memory for long video understanding
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d54eda3-0659-45a6-9c88-e1e0f13aaf97 · outbound
Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Coin: A large-scale dataset for comprehensive instructional video analysis
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f219ca49-d55a-4d09-96f2-08470039b654 · outbound
Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Learning Multiple Object States from Actions via Large Language Models
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5f6f2c71-ab4a-44a8-a5de-3d2bf337fcc5 · outbound
Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ddc243e0-b54d-4218-ba0d-974bfd281f57 · outbound
Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction User-in-the-loop Evaluation of Multimodal LLMs for Activity Assistance
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3b647031-8ff2-4b39-9107-29dbc407fde3 · outbound
Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Event-guided procedure planning from in- structional videos with text supervision
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7bd79037-3e2b-4fd2-afcb-c52500d9e309 · outbound
Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Pdpp: Projected diffusion for procedure planning in instructional videos
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9cb4e834-73d3-49bc-aaee-44abd6ea5728 · outbound
Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Vamos: Versatile Action Models for Video Understanding
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a088458d-7634-4f1c-8ca8-51e26ea5be01 · outbound
Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Learn- ing object state changes in videos: An open-world perspec- tive
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 956f8e27-856e-4355-936d-0666a9808099 · outbound
Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Octopus: Embodied vision- language programmer from environmental feedback, 2023
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1d5423c0-02f8-49f9-9e76-ff78bfff4087 · outbound
Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction RAP: Retrieval-Augmented Planner for Adaptive Procedure Planning in Instructional Videos
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a03f4ecb-e404-4648-bf90-07472dcd2fc3 · outbound
Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Object-centric video representation for long-term action anticipation
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 23a278c6-bd3f-43e7-ab9b-259c65b1eb56 · outbound
Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b78c2393-32b0-4215-884f-cf024c8db353 · outbound
Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction P3iv: Prob- abilistic procedure planning from instructional videos with weak supervision
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 70571645-9ad2-4347-bea0-571aa42e7e8b · outbound
Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction AntGPT: Can Large Language Models Help Long-term Action Anticipation from Videos?
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 653059be-6758-4f8a-ad93-a510d079ded0 · outbound
Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Towards learning a generalist model for embod- ied navigation
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c8f553c1-b4d0-4bc1-9634-e562077195be · outbound
Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in Neural Information Processing Systems, 36:46595–46623, 2023
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed1535af-4470-4e8e-bb0b-5efafbe389f7 · outbound
Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58977a31-b981-4e08-ae42-3f13da79417e · outbound
Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Cross- task weakly supervised learning from instructional videos
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation aa874697-635c-4d25-8905-c864c4106a8e · outbound
Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction before” with “after
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation df875fdf-dfaf-490e-b2ae-bd12d17b3fe7 · outbound
Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Training We train our model for 1 epoch with a batch size of 1024
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 889eb96f-8bc6-4f6a-b0ec-b5d88ce41d6b · outbound
Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction dough”, “container
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 115e9f3b-1d83-4ccd-9d0b-7f6cbd293e61 · inbound
GeoWorld: Geometric World Models Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.