Pith. sign in

Paper Citation Record · LEDGER

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction

As of 20 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 1 inbound Pith citation observation for arXiv:2507.15130.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.15130 v1

Coverage vector

measured 62 of 62 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:44:49.592008Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-21T11:39:15.308355Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T11:40:03.350024Z

Reference resolution

62 of 62 outbound references displayed

  • verified exact6
  • verified fuzzy34
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d1ef1572-a27c-4efd-b6ee-8a89deccadc8 · outbound

This paper cites When will you do what?-anticipating temporal occurrences of activities.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction When will you do what?-anticipating temporal occurrences of activities

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:50.173046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T15:44:43.981738Z digest=sha256:02fce39374d7ade50cd9c436f21d30a640a59a0f28ba20319999f239b85aca8b

Observation cb381446-e785-4243-9a8f-b90d5f812668 · outbound

This paper cites GPT-4 Technical Report.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T15:44:44.065986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:44:44.065986Z digest=sha256:b99eee3ba12fc896a2d4edda66d4495e29859133832ae622127ff9aad194ba03

Observation 72238229-4323-4029-bb15-8fa4759981fb · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Flamingo: a visual language model for few-shot learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T15:44:44.155587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:44:44.155587Z digest=sha256:a15f8cb15dde4b948768b9a54f26c1d1be3046c43d682a9c8d259f022f9ac9bf

Observation fff40a00-3055-483b-99c0-73bf8ca694e5 · outbound

This paper cites Hiervl: Learning hierarchical video- language embeddings.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Hiervl: Learning hierarchical video- language embeddings

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:50.161138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T15:44:44.281924Z digest=sha256:d6b8a579d99568572d9eb0ea1d347a49296a1bc36d1bce21c24e6a8cd64a524b

Observation f9f18253-afc6-42b3-9605-c383a5d72b26 · outbound

This paper cites Procedure planning in instructional videos.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Procedure planning in instructional videos

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:50.153750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T15:44:44.356600Z digest=sha256:f04b84d286d622f45857ea0cc97ca4878d35fa7f84d4ff40efb7c41f60f8e542

Observation fbf579ff-4efe-43d4-889e-b9e8957a91b8 · outbound

This paper cites Spatialvlm: Endow- ing vision-language models with spatial reasoning capabili- ties.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Spatialvlm: Endow- ing vision-language models with spatial reasoning capabili- ties

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:50.146331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T15:44:44.497438Z digest=sha256:057ec1b36dd714febb83cc52db69b335f80f370ed610eccfce418a4c32931f8d

Observation b1b39347-9368-4e03-9849-e3285bb02838 · outbound

This paper cites EgoPlan-Bench: Benchmarking Multimodal Large Language Models for Human-Level Planning.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction EgoPlan-Bench: Benchmarking Multimodal Large Language Models for Human-Level Planning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T15:44:44.617672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:44:44.617672Z digest=sha256:3e34577ba923d153684f89fec2ab5514d3662d02b630d8250639fed469d14202

Observation 92f198ec-ee75-48ac-8f2a-66d84f03eb99 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T15:44:44.712467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:44:44.712467Z digest=sha256:92b52ad02f4231a1091d12a2b0ecb9633e6f7f40afb8e523c637f14196e51b25

Observation 6c23d4da-96e1-43a4-b120-7c4db34acdaa · outbound

This paper cites Better & Faster Large Language Models via Multi-token Prediction.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Better & Faster Large Language Models via Multi-token Prediction

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T15:44:44.806734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:44:44.806734Z digest=sha256:b44b5b5ce21f76d67ec8a76c7ff6f45388eabe61c381fa33cb0fb05bd7b0f9b5

Observation 9b8f4eb0-eead-471e-a8b0-4397c62e683a · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Ego4d: Around the world in 3,000 hours of egocentric video

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:50.138897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T15:44:44.914759Z digest=sha256:1d5915336a4fbd478e4eaad517cce0591f2cb983f3140c885f73c0205b189212

Observation 1e7b96ed-19db-489a-b249-63f02ce724b2 · outbound

This paper cites Reasoning with Language Model is Planning with World Model.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Reasoning with Language Model is Planning with World Model

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T15:44:45.036118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:44:45.036118Z digest=sha256:afde312757546f4cccbaa5b555282b3e78d20a22e042fe669596d786ba3718fb

Observation 66bf25b6-aa1e-4d23-82b3-1a50ffecafe2 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction LoRA: Low-Rank Adaptation of Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T15:44:45.143895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:44:45.143895Z digest=sha256:7fefb2f2c19906fd70d84d64413fddd76c5b6ad705aeb6b8fac80a3b91aaead9

Observation 3231d02b-bfad-4a3e-9862-4454351e6988 · outbound

This paper cites Vtimellm: Empower llm to grasp video moments.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Vtimellm: Empower llm to grasp video moments

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T15:44:45.259004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:44:45.259004Z digest=sha256:07df42139ec3e9569a7c70200aea9f9523dec97dcdf6d0abf5dd2f2b91e06c31

Observation b452a7b1-8d21-4b5b-a131-554a48270480 · outbound

This paper cites Language models as zero-shot planners: Extract- ing actionable knowledge for embodied agents.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Language models as zero-shot planners: Extract- ing actionable knowledge for embodied agents

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:50.066608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T15:44:45.361973Z digest=sha256:9dd93f086fd2cc18f2066d38337c40956746b147276b1dfb7e2dffc5fb78f06b

Observation 5ebe8f6a-baad-455f-a467-80ef575e3610 · outbound

This paper cites Propose, Assess, Search: Harnessing LLMs for Goal-Oriented Planning in Instructional Videos.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Propose, Assess, Search: Harnessing LLMs for Goal-Oriented Planning in Instructional Videos

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:44:49.792034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T15:44:45.482849Z digest=sha256:6c87bd5b4831daa94a4d502f4221781756a55e497927618ce90d209a4631f2aa

Observation 0f03c10e-7a52-4b22-a0c9-7d99c64ba0c9 · outbound

This paper cites Palm: Predicting actions through language models.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Palm: Predicting actions through language models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:50.059681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T15:44:45.617947Z digest=sha256:2d5b4ed368debf494264dcf400323aa5bdb8c99dd4e7a4cf148bfa58a1861d2d

Observation df76deea-f9a7-4c2b-9cea-6dc95f0b37bf · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction LLaVA-OneVision: Easy Visual Task Transfer

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T15:44:45.736431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:44:45.736431Z digest=sha256:c3a68c2994991bac39e80c11a7d4a319e8fefca638c1c965aebdcd1e22ae9214

Observation bbd8abc5-45e6-4a6a-9c4d-1e2ad8ac49ec · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:50.052622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T15:44:45.818880Z digest=sha256:7140554b699bd56a5ad0d1e8b267d05d1e26ce7feb8666cbbccbb3c923948fa0

Observation 010fbb76-6cf4-4479-82d7-62955d2fc731 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction VideoChat: Chat-Centric Video Understanding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T15:44:45.907158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:44:45.907158Z digest=sha256:61f7b7caceb44ad237d2625ed0f00f11b3243aa2a394667a7a2e90dca92ffcb2

Observation 8e2c3f67-7f3d-463a-ae90-eab374619d4d · outbound

This paper cites Mvbench: A comprehensive multi-modal video understand- ing benchmark.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Mvbench: A comprehensive multi-modal video understand- ing benchmark

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:50.045748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T15:44:45.981792Z digest=sha256:379acd596156f10656bfd4acebca31a348d807d5a520d9206f1520f05ab2d9d9

Observation 83e61c2f-2559-4abc-bed2-2b054d6e6cc0 · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Llama-vid: An image is worth 2 tokens in large language models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:50.038452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T15:44:46.075955Z digest=sha256:9a897a3b3ca53c00de4a358d32e7aff24a0c222db03be9cf44a9d30b48c8ad43

Observation e32811d9-557a-45dc-8a30-0dcf25808e09 · outbound

This paper cites Skip-plan: Procedure plan- ning in instructional videos via condensed action space learn- ing.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Skip-plan: Procedure plan- ning in instructional videos via condensed action space learn- ing

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:50.031304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T15:44:46.185631Z digest=sha256:7d219f37ddac23c201b470a8071d12e2922dcbedc948fa1bd172c1a678d75da2

Observation 5d031845-fb8f-4d37-83cd-620bed5edc2f · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T15:44:46.272623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:44:46.272623Z digest=sha256:e0ca66a07eff37de0d58d44fa8242182bca46198cc0c1243702ff2db5e70137f

Observation 1826b2e1-b5d0-4166-95ad-9199fcbaf75d · outbound

This paper cites LLM+P: Empowering Large Language Models with Optimal Planning Proficiency.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction LLM+P: Empowering Large Language Models with Optimal Planning Proficiency

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T15:44:46.355848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:44:46.355848Z digest=sha256:e68c2f87ca476dfffb25cc811ec91a717390153a1eac9252e7aac8cff670e766

Observation 0d423aea-4a84-411c-8400-7a297d87cd58 · outbound

This paper cites Visual instruction tuning.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Visual instruction tuning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:50.024103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T15:44:46.434456Z digest=sha256:bd415a088634f3cd2d7a1389c0bb6d39c59af30f60360f0dbba7a79c56bc0ff2

Observation 19b722b5-fe3a-422c-807c-cf83d362d10d · outbound

This paper cites A language-first approach for procedure planning.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction A language-first approach for procedure planning

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:50.017054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T15:44:46.528710Z digest=sha256:a09e2f6a633d40edb33e598d512f0cae9b47eed836df888488bf4bd6cf66a3a3

Observation ec54eccf-757e-4470-a178-8c20ac2ca022 · outbound

This paper cites Intention-conditioned long-term human egocentric action anticipation.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Intention-conditioned long-term human egocentric action anticipation

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:50.009663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T15:44:46.610609Z digest=sha256:d11d59979d27ebb8f162ce0c687a5afe1fcdc46532ec8e1cc34165ad964d561b

Observation e70b9c84-0c3a-4466-b36c-262583ebbfa4 · outbound

This paper cites Can’t make an omelette without breaking some eggs: Plausible action anticipation using large video- language models.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Can’t make an omelette without breaking some eggs: Plausible action anticipation using large video- language models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:50.002240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T15:44:46.680605Z digest=sha256:a91b2899b5cc7cdfea1bbd8b4aa4fe144d1111136e45db3285a844337d68b580

Observation 7e57bf75-9560-4cb3-a881-8c99d592fa92 · outbound

This paper cites Any- mal: An efficient and scalable any-modality augmented lan- guage model.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Any- mal: An efficient and scalable any-modality augmented lan- guage model

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:49.994645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T15:44:46.734541Z digest=sha256:3ba7db6afeb4c9bf3327648b326e8c0ecc4d0d11a4c4cd013b560f1a95be4d85

Observation acc44834-e809-4f34-9a37-ff9452287226 · outbound

This paper cites Embodiedgpt: Vision-language pre-training via embodied chain of thought.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Embodiedgpt: Vision-language pre-training via embodied chain of thought

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:49.986211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T15:44:46.817312Z digest=sha256:ca81b12cef70770b9ff85eedeaaedd047269545d5fd8fc4e099e86ec3f600569

Observation c0cd7f5f-bff3-49d8-b818-96adb1df2653 · outbound

This paper cites Ego-topo: Environment affordances from egocentric video.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Ego-topo: Environment affordances from egocentric video

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:49.979003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T15:44:46.889917Z digest=sha256:52d07928cea9ec77116bf3b661cae3b12425bf1138df26f45b1483c479286893

Observation b8de1b13-6e6a-4cc0-b53e-59a503d393c0 · outbound

This paper cites Why not use your text- book? knowledge-enhanced procedure planning of instruc- tional videos.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Why not use your text- book? knowledge-enhanced procedure planning of instruc- tional videos

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:49.971464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T15:44:46.976969Z digest=sha256:f473fc11f7cf042050483069f78c613a2491cd5b309a1ffad206df7711760b7a

Observation 4d55b9dd-512e-4f59-9f33-0bc19cd72ffc · outbound

This paper cites Re- thinking learning approaches for long-term action anticipa- tion.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Re- thinking learning approaches for long-term action anticipa- tion

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:49.964005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T15:44:47.046680Z digest=sha256:f4ea93e7130b0f6877ab2882f692a5d68fc32f575c573a2764a0968481e47419

Observation 003f8245-b33f-43fc-8cb8-f28d80703fbb · outbound

This paper cites Do Pre-trained Vision-Language Models Encode Object States?.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Do Pre-trained Vision-Language Models Encode Object States?

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:44:49.751802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T15:44:47.143277Z digest=sha256:f06b9f8a797cc8e37e3425b50306e126ee535808e833c7a2025e4bfd408b09c3

Observation 23e267b8-7b9e-44fa-ba12-6ecdce73340d · outbound

This paper cites SCHEMA: State CHangEs MAtter for Procedure Planning in Instructional Videos.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction SCHEMA: State CHangEs MAtter for Procedure Planning in Instructional Videos

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T15:44:47.232091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:44:47.232091Z digest=sha256:152972e63693b4d52d8f4b2e1f7bdf6090cfb4c53d79d2f880416787d52a7032

Observation eb7d0d66-827a-41ca-975d-84c6678d8883 · outbound

This paper cites Pretrained language models as visual planners for human assistance.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Pretrained language models as visual planners for human assistance

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:49.956596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T15:44:47.334176Z digest=sha256:3fa7e8805b1a843a0019e66f07f574fcd3f68d8d60c9ccc066350a458ff412c3

Observation ad121c6a-9f47-4e65-820a-c8f927d7c9d2 · outbound

This paper cites EgoVideo: Exploring Egocentric Foundation Model and Downstream Adaptation.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction EgoVideo: Exploring Egocentric Foundation Model and Downstream Adaptation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T15:44:47.422726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:44:47.422726Z digest=sha256:2a99a6637067cec48298ff082665f271991ba38b4cdd2fb5858bc88b53755ea0

Observation 212abde4-2bc1-42ed-82a1-c02e251bec32 · outbound

This paper cites Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T15:44:47.520419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:44:47.520419Z digest=sha256:d01039546ca0091dd629c3f022c0b8d0134108af90b4517425944f45e5dad466

Observation 8d5e895f-1c46-43d7-9700-6300edf4e55b · outbound

This paper cites Timechat: A time-sensitive multimodal large lan- guage model for long video understanding.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Timechat: A time-sensitive multimodal large lan- guage model for long video understanding

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T15:44:47.619011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:44:47.619011Z digest=sha256:8e2d3733a0371913acf1d1efab50de25ea9a6758fef7456bdc3a093d527f8fe6

Observation 9f2e9711-22f4-4fe6-98c5-e186b4b5e058 · outbound

This paper cites Llm-planner: Few-shot grounded planning for embodied agents with large language models.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Llm-planner: Few-shot grounded planning for embodied agents with large language models

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:49.944499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T15:44:47.715754Z digest=sha256:3eb375d5b80c81e0b77f351604164048c2741c7390e348e1f4d9b9e0cee3b3df

Observation 0332c79f-4bd2-4097-a37d-4d215e106510 · outbound

This paper cites Moviechat: From dense token to sparse memory for long video understanding.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Moviechat: From dense token to sparse memory for long video understanding

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T15:44:47.800874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:44:47.800874Z digest=sha256:3dc9e8af703869c8f0b054279e3c33950244e055526edc43a6160a6dd5e600a6

Observation 5d54eda3-0659-45a6-9c88-e1e0f13aaf97 · outbound

This paper cites Coin: A large-scale dataset for comprehensive instructional video analysis.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Coin: A large-scale dataset for comprehensive instructional video analysis

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:49.932602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T15:44:47.933683Z digest=sha256:29555d9c885effcefed9930089a7a0ca6d9257d98a016ef640cb4e1bbad1eac2

Observation f219ca49-d55a-4d09-96f2-08470039b654 · outbound

This paper cites Learning Multiple Object States from Actions via Large Language Models.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Learning Multiple Object States from Actions via Large Language Models

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:44:49.717370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T15:44:48.030903Z digest=sha256:78b99ab0b8f7d2b435afaa326476556df1400a686a29c4757276c971efac3f78

Observation 5f6f2c71-ab4a-44a8-a5de-3d2bf337fcc5 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T15:44:48.175250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:44:48.175250Z digest=sha256:f9bf0ac57daae92c103edf24a6685d9c093326d7d41a79543b1dc5d939bd92d0

Observation ddc243e0-b54d-4218-ba0d-974bfd281f57 · outbound

This paper cites User-in-the-loop Evaluation of Multimodal LLMs for Activity Assistance.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction User-in-the-loop Evaluation of Multimodal LLMs for Activity Assistance

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:44:49.697709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T15:44:48.352122Z digest=sha256:f98c210baa7ceeb40391f480761dda78802148e89f0ba06cd1f3f04be6aebe5a

Observation 3b647031-8ff2-4b39-9107-29dbc407fde3 · outbound

This paper cites Event-guided procedure planning from in- structional videos with text supervision.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Event-guided procedure planning from in- structional videos with text supervision

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:49.925468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T15:44:48.507610Z digest=sha256:e6df2ca959983fc65abc848ca3d120ff75b288efd54dce4f80de7b9b486d182d

Observation 7bd79037-3e2b-4fd2-afcb-c52500d9e309 · outbound

This paper cites Pdpp: Projected diffusion for procedure planning in instructional videos.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Pdpp: Projected diffusion for procedure planning in instructional videos

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:49.917737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T15:44:48.650999Z digest=sha256:cc9bbfb4349f0524d3cbd30a0c1e7c7066cb08b4b4afc9cd5d986ad96cbe9cfb

Observation 9cb4e834-73d3-49bc-aaee-44abd6ea5728 · outbound

This paper cites Vamos: Versatile Action Models for Video Understanding.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Vamos: Versatile Action Models for Video Understanding

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:44:49.687256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T15:44:48.688715Z digest=sha256:de59ffe89795c43a12eae212102bff656ebf05c39d9a43cbb38fdbfe451bd464

Observation a088458d-7634-4f1c-8ca8-51e26ea5be01 · outbound

This paper cites Learn- ing object state changes in videos: An open-world perspec- tive.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Learn- ing object state changes in videos: An open-world perspec- tive

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:49.909513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T15:44:48.748675Z digest=sha256:8ccf636acd6e39eab4de0dcd6d90836d505bc2ce4407a70449e450f876ebc936

Observation 956f8e27-856e-4355-936d-0666a9808099 · outbound

This paper cites Octopus: Embodied vision- language programmer from environmental feedback, 2023.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Octopus: Embodied vision- language programmer from environmental feedback, 2023

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:49.901873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T15:44:48.801693Z digest=sha256:e9a929af008200da1b5d737d1f7e1e2d8e323636e7a2411554425cee965eef7b

Observation 1d5423c0-02f8-49f9-9e76-ff78bfff4087 · outbound

This paper cites RAP: Retrieval-Augmented Planner for Adaptive Procedure Planning in Instructional Videos.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction RAP: Retrieval-Augmented Planner for Adaptive Procedure Planning in Instructional Videos

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:44:49.675996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T15:44:48.842813Z digest=sha256:283c63344ade1e293668e618d85e3c6f793a8b0071be747a556f4724bbca6483

Observation a03f4ecb-e404-4648-bf90-07472dcd2fc3 · outbound

This paper cites Object-centric video representation for long-term action anticipation.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Object-centric video representation for long-term action anticipation

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:49.894247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T15:44:48.910963Z digest=sha256:5cda454b330a467b308d228aad489f56fa38247c225035f2a03016ba9792a52c

Observation 23a278c6-bd3f-43e7-ab9b-259c65b1eb56 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T15:44:48.976855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:44:48.976855Z digest=sha256:e4016fff220a75ebae7593be8438f45caf9ecdf8015fb5f1cd3b1430c8c2a20e

Observation b78c2393-32b0-4215-884f-cf024c8db353 · outbound

This paper cites P3iv: Prob- abilistic procedure planning from instructional videos with weak supervision.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction P3iv: Prob- abilistic procedure planning from instructional videos with weak supervision

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:49.886095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T15:44:49.021366Z digest=sha256:4b78a184fde4c70cdcf00b3015e80589f867ee7e02f59df221f3a9974175ff46

Observation 70571645-9ad2-4347-bea0-571aa42e7e8b · outbound

This paper cites AntGPT: Can Large Language Models Help Long-term Action Anticipation from Videos?.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction AntGPT: Can Large Language Models Help Long-term Action Anticipation from Videos?

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T15:44:49.029086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:44:49.029086Z digest=sha256:3a8f7b9df0af3e8323ef0a3db7273dd193f67d54c3b5b13912d0b894110e5372

Observation 653059be-6758-4f8a-ad93-a510d079ded0 · outbound

This paper cites Towards learning a generalist model for embod- ied navigation.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Towards learning a generalist model for embod- ied navigation

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:49.878331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T15:44:49.034663Z digest=sha256:d0359cdd720bdfedd2f20e2c9c6d1d550bc410f23faf3777572f477e0498eea4

Observation c8f553c1-b4d0-4bc1-9634-e562077195be · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in Neural Information Processing Systems, 36:46595–46623, 2023.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in Neural Information Processing Systems, 36:46595–46623, 2023

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T15:44:49.120307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:44:49.120307Z digest=sha256:84d970533a76fcf35c9b32acc5426f020f8eb9898b4ec7debfe3b31f37706823

Observation ed1535af-4470-4e8e-bb0b-5efafbe389f7 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T15:44:49.190266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:44:49.190266Z digest=sha256:8ee88a4ee2fbe50e203b2dc8e30f4218332751c91e0e816dc9175c524926c024

Observation 58977a31-b981-4e08-ae42-3f13da79417e · outbound

This paper cites Cross- task weakly supervised learning from instructional videos.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Cross- task weakly supervised learning from instructional videos

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:49.866298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T15:44:49.288340Z digest=sha256:ed8aeaced5cc637bc998fc64185e9818164362dfd3dc3f5040e5ac24dc44c849

Observation aa874697-635c-4d25-8905-c864c4106a8e · outbound

This paper cites before” with “after.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction before” with “after

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:49.858942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T15:44:49.417662Z digest=sha256:6b39e1b57898add531f994367d4eac07445be428439f70dee8a3c4074d722662

Observation df875fdf-dfaf-490e-b2ae-bd12d17b3fe7 · outbound

This paper cites Training We train our model for 1 epoch with a batch size of 1024.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Training We train our model for 1 epoch with a batch size of 1024

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:49.851659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T15:44:49.495928Z digest=sha256:f8d348ea6ab2dc838c2c55614b94ea9277d92825dd06c7110bb10edb33f2d79f

Observation 889eb96f-8bc6-4f6a-b0ec-b5d88ce41d6b · outbound

This paper cites dough”, “container.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction dough”, “container

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:49.843986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T15:44:49.592008Z digest=sha256:a7ac3e2885aa226e5a91712dafb70aca0fc5d2a25bcf41bd92f1aa003baff58a

Pith citing papers

Observation 115e9f3b-1d83-4ccd-9d0b-7f6cbd293e61 · inbound

GeoWorld: Geometric World Models cites this paper.

GeoWorld: Geometric World Models Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction

Reference 85

Resolution
verified exact
arxiv_id, observed 2026-05-21T11:40:03.352082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T11:39:15.308355Z digest=sha256:09091e7a632b20678c938ff63d477b82d79c22e697612f40b72ba24a5df884b2