Pith. sign in

Paper Citation Record · LEDGER

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction

As of 7 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 1 inbound Pith citation observation for arXiv:2507.15130.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.15130 v1

Coverage vector

measured 62 of 62 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:44:49.592008Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-21T11:39:15.308355Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T11:40:03.350024Z

Reference resolution

62 of 62 outbound references displayed

  • verified exact6
  • verified fuzzy34
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d1ef1572-a27c-4efd-b6ee-8a89deccadc8 · outbound

This paper cites When will you do what?-anticipating temporal occurrences of activities.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction When will you do what?-anticipating temporal occurrences of activities

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:50.173046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:44:43.981738Z digest=sha256:356ed574f00ed6f0df0c127ed8d4f6d4d4b80601133c9d5ae89478d5ebfbdec1

Observation cb381446-e785-4243-9a8f-b90d5f812668 · outbound

This paper cites GPT-4 Technical Report.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T15:44:44.065986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:44:44.065986Z digest=sha256:ec755618ca5eef03392b3ac596616b82ded525758bfb3263e3f768626265591f

Observation 72238229-4323-4029-bb15-8fa4759981fb · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Flamingo: a visual language model for few-shot learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T15:44:44.155587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:44:44.155587Z digest=sha256:2dba9408d565c748d072d2868f5255c864bced774c6a9a760d3272059c5b9673

Observation fff40a00-3055-483b-99c0-73bf8ca694e5 · outbound

This paper cites Hiervl: Learning hierarchical video- language embeddings.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Hiervl: Learning hierarchical video- language embeddings

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:50.161138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:44:44.281924Z digest=sha256:19e7fa7fb1e0c0d13afc7d04dfc2331166f7b9724da6dd5f27ccb1676c383465

Observation f9f18253-afc6-42b3-9605-c383a5d72b26 · outbound

This paper cites Procedure planning in instructional videos.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Procedure planning in instructional videos

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:50.153750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:44:44.356600Z digest=sha256:af12af9f5f45ac638d4ea1669064811eb4f3fc1d757fbf099a1106f506b1f5f1

Observation fbf579ff-4efe-43d4-889e-b9e8957a91b8 · outbound

This paper cites Spatialvlm: Endow- ing vision-language models with spatial reasoning capabili- ties.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Spatialvlm: Endow- ing vision-language models with spatial reasoning capabili- ties

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:50.146331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:44:44.497438Z digest=sha256:44c08feeabc0c0436b8103353952ea9bb2f013450c2e2f6d7ea17f490098eeae

Observation b1b39347-9368-4e03-9849-e3285bb02838 · outbound

This paper cites EgoPlan-Bench: Benchmarking Multimodal Large Language Models for Human-Level Planning.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction EgoPlan-Bench: Benchmarking Multimodal Large Language Models for Human-Level Planning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T15:44:44.617672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:44:44.617672Z digest=sha256:8413351fd5345331b831b44872246632b0ea3117517e514b164499bbd4d99d03

Observation 92f198ec-ee75-48ac-8f2a-66d84f03eb99 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T15:44:44.712467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:44:44.712467Z digest=sha256:25335b3f061996b9bb7648eda91ccd9ef990296db5610ff8c4368e27381c78e2

Observation 6c23d4da-96e1-43a4-b120-7c4db34acdaa · outbound

This paper cites Better & Faster Large Language Models via Multi-token Prediction.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Better & Faster Large Language Models via Multi-token Prediction

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T15:44:44.806734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:44:44.806734Z digest=sha256:bb2e7d67a0261f12f77d5422e3d4b4e1467723c8a37b85267be0f635067817de

Observation 9b8f4eb0-eead-471e-a8b0-4397c62e683a · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Ego4d: Around the world in 3,000 hours of egocentric video

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:50.138897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:44:44.914759Z digest=sha256:2437bc417648b025ed74eddb85e8e3081cc353747fbb2f2fc7001c6704b1535f

Observation 1e7b96ed-19db-489a-b249-63f02ce724b2 · outbound

This paper cites Reasoning with Language Model is Planning with World Model.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Reasoning with Language Model is Planning with World Model

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T15:44:45.036118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:44:45.036118Z digest=sha256:155ec52c0e9f6063e1cbc719b7ba38e55634e35b292d4a74fb53adbb97a4f1c9

Observation 66bf25b6-aa1e-4d23-82b3-1a50ffecafe2 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction LoRA: Low-Rank Adaptation of Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T15:44:45.143895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:44:45.143895Z digest=sha256:118554cdf41fc7bcab6e2e4e52a349c3c258151360ec27785d6306b1528548a1

Observation 3231d02b-bfad-4a3e-9862-4454351e6988 · outbound

This paper cites Vtimellm: Empower llm to grasp video moments.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Vtimellm: Empower llm to grasp video moments

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T15:44:45.259004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:44:45.259004Z digest=sha256:9f8ac4472118468a870de30e6cd6380b2c83f7bf2b04797f0bcb3c0e03fb8d76

Observation b452a7b1-8d21-4b5b-a131-554a48270480 · outbound

This paper cites Language models as zero-shot planners: Extract- ing actionable knowledge for embodied agents.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Language models as zero-shot planners: Extract- ing actionable knowledge for embodied agents

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:50.066608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:44:45.361973Z digest=sha256:13db7e53fe4a41a68b9107149d97ce32e2212829534515a4b6444b13072d0d58

Observation 5ebe8f6a-baad-455f-a467-80ef575e3610 · outbound

This paper cites Propose, Assess, Search: Harnessing LLMs for Goal-Oriented Planning in Instructional Videos.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Propose, Assess, Search: Harnessing LLMs for Goal-Oriented Planning in Instructional Videos

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:44:49.792034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:44:45.482849Z digest=sha256:2876b95d23308c7c3a8d000fd3169afe4cfb5034315dd51e3d582fd88782f177

Observation 0f03c10e-7a52-4b22-a0c9-7d99c64ba0c9 · outbound

This paper cites Palm: Predicting actions through language models.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Palm: Predicting actions through language models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:50.059681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:44:45.617947Z digest=sha256:a8c6aecb12b7ed3ffb3058dd13fce6fe8e6fc2194dadd0b5b2209453b14a952f

Observation df76deea-f9a7-4c2b-9cea-6dc95f0b37bf · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction LLaVA-OneVision: Easy Visual Task Transfer

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T15:44:45.736431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:44:45.736431Z digest=sha256:a6815658d8fecb4e4a37f5664aa8f918dca73845f3dd62aa2d9e9e088c69d2a7

Observation bbd8abc5-45e6-4a6a-9c4d-1e2ad8ac49ec · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:50.052622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:44:45.818880Z digest=sha256:07d865e3694e8fc0d868d0454afaa9441a432a7e4d18db32793e5a84450784ca

Observation 010fbb76-6cf4-4479-82d7-62955d2fc731 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction VideoChat: Chat-Centric Video Understanding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T15:44:45.907158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:44:45.907158Z digest=sha256:fabd0d573d691f9d7eec1e5ae1c1f2ee6f03a468ce7f3b786847c60019f65928

Observation 8e2c3f67-7f3d-463a-ae90-eab374619d4d · outbound

This paper cites Mvbench: A comprehensive multi-modal video understand- ing benchmark.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Mvbench: A comprehensive multi-modal video understand- ing benchmark

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:50.045748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:44:45.981792Z digest=sha256:6105606430d612ada354541d7faceea7f97c0e98a77c339f59005d347bc558af

Observation 83e61c2f-2559-4abc-bed2-2b054d6e6cc0 · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Llama-vid: An image is worth 2 tokens in large language models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:50.038452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:44:46.075955Z digest=sha256:ed071ba5e7fda8caa8719c36deeefdb963f331438534c5514cd9c62528bc0c21

Observation e32811d9-557a-45dc-8a30-0dcf25808e09 · outbound

This paper cites Skip-plan: Procedure plan- ning in instructional videos via condensed action space learn- ing.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Skip-plan: Procedure plan- ning in instructional videos via condensed action space learn- ing

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:50.031304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:44:46.185631Z digest=sha256:52313399f70bd58bdb388778146f15669875155a802b95a85b0a1510fb0a5c65

Observation 5d031845-fb8f-4d37-83cd-620bed5edc2f · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T15:44:46.272623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:44:46.272623Z digest=sha256:3703a59dd77daaa6ddc249c5a8c671e736812b9f794bfc7918cb4bdf911100f2

Observation 1826b2e1-b5d0-4166-95ad-9199fcbaf75d · outbound

This paper cites LLM+P: Empowering Large Language Models with Optimal Planning Proficiency.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction LLM+P: Empowering Large Language Models with Optimal Planning Proficiency

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T15:44:46.355848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:44:46.355848Z digest=sha256:2e45c2e708a00866791f96fb3c2c48512c3dd7c6347d9cd9883e10666a43074e

Observation 0d423aea-4a84-411c-8400-7a297d87cd58 · outbound

This paper cites Visual instruction tuning.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Visual instruction tuning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:50.024103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:44:46.434456Z digest=sha256:ea82e6a6234a5869110e27077073faf1cdafdf11bd2a0bb6cba496573852a01f

Observation 19b722b5-fe3a-422c-807c-cf83d362d10d · outbound

This paper cites A language-first approach for procedure planning.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction A language-first approach for procedure planning

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:50.017054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:44:46.528710Z digest=sha256:3c3fd9df8a01e1581cac2ed269c6ed4c3cd6fac4b8b8b6a503298e778ad02281

Observation ec54eccf-757e-4470-a178-8c20ac2ca022 · outbound

This paper cites Intention-conditioned long-term human egocentric action anticipation.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Intention-conditioned long-term human egocentric action anticipation

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:50.009663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:44:46.610609Z digest=sha256:2f82b75aa69e3509d815b0f077e601680df489de6638c997b1e798552a67317d

Observation e70b9c84-0c3a-4466-b36c-262583ebbfa4 · outbound

This paper cites Can’t make an omelette without breaking some eggs: Plausible action anticipation using large video- language models.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Can’t make an omelette without breaking some eggs: Plausible action anticipation using large video- language models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:50.002240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:44:46.680605Z digest=sha256:5cf992606fbc362f82c62af103deaa3aa2c35dfd56e00041e5c5287d89b19dca

Observation 7e57bf75-9560-4cb3-a881-8c99d592fa92 · outbound

This paper cites Any- mal: An efficient and scalable any-modality augmented lan- guage model.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Any- mal: An efficient and scalable any-modality augmented lan- guage model

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:49.994645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:44:46.734541Z digest=sha256:722de099b2ce187b3950810721238dad0c5584479163325a1332e5f3edb67dbf

Observation acc44834-e809-4f34-9a37-ff9452287226 · outbound

This paper cites Embodiedgpt: Vision-language pre-training via embodied chain of thought.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Embodiedgpt: Vision-language pre-training via embodied chain of thought

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:49.986211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:44:46.817312Z digest=sha256:d78eb6459236f2ea83ed29363a5255cbdad515ce501a6a71c73c2b5bd0513024

Observation c0cd7f5f-bff3-49d8-b818-96adb1df2653 · outbound

This paper cites Ego-topo: Environment affordances from egocentric video.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Ego-topo: Environment affordances from egocentric video

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:49.979003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:44:46.889917Z digest=sha256:32fcaa7e05be610ec3312441c72f6d852a83747289ff77655faa1b32583ee088

Observation b8de1b13-6e6a-4cc0-b53e-59a503d393c0 · outbound

This paper cites Why not use your text- book? knowledge-enhanced procedure planning of instruc- tional videos.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Why not use your text- book? knowledge-enhanced procedure planning of instruc- tional videos

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:49.971464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:44:46.976969Z digest=sha256:52184d563783a9b22a0b56efe7075b48b06c0f7bb7d8f45aae11c864cd02cd05

Observation 4d55b9dd-512e-4f59-9f33-0bc19cd72ffc · outbound

This paper cites Re- thinking learning approaches for long-term action anticipa- tion.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Re- thinking learning approaches for long-term action anticipa- tion

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:49.964005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:44:47.046680Z digest=sha256:1c619b5020ec100e446d52c206d37f8aa52883f426002d88894b386ba386f64e

Observation 003f8245-b33f-43fc-8cb8-f28d80703fbb · outbound

This paper cites Do Pre-trained Vision-Language Models Encode Object States?.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Do Pre-trained Vision-Language Models Encode Object States?

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:44:49.751802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:44:47.143277Z digest=sha256:6cd047407ca2a9767082573d57f4ec80b76cfa9d34e93cff961d3afeef8b4164

Observation 23e267b8-7b9e-44fa-ba12-6ecdce73340d · outbound

This paper cites SCHEMA: State CHangEs MAtter for Procedure Planning in Instructional Videos.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction SCHEMA: State CHangEs MAtter for Procedure Planning in Instructional Videos

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T15:44:47.232091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:44:47.232091Z digest=sha256:c02d4a02e1a90f85c2844ca43668a6bbeaf00c8b2e1d79b71e93e1e295e6513f

Observation eb7d0d66-827a-41ca-975d-84c6678d8883 · outbound

This paper cites Pretrained language models as visual planners for human assistance.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Pretrained language models as visual planners for human assistance

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:49.956596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:44:47.334176Z digest=sha256:3290f91968a4a06df6a896b5fa91fe1cd6532d13a6260e535399f8e74fd537b1

Observation ad121c6a-9f47-4e65-820a-c8f927d7c9d2 · outbound

This paper cites EgoVideo: Exploring Egocentric Foundation Model and Downstream Adaptation.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction EgoVideo: Exploring Egocentric Foundation Model and Downstream Adaptation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T15:44:47.422726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:44:47.422726Z digest=sha256:2e0a1d865de839816ac3f7c32f5461f57f1eeabc0226dbee8c524c3c43c06b99

Observation 212abde4-2bc1-42ed-82a1-c02e251bec32 · outbound

This paper cites Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T15:44:47.520419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:44:47.520419Z digest=sha256:c73cb8216c215e27abb14a7fbf54186595a4f452123d0ed756732b4658b0ad81

Observation 8d5e895f-1c46-43d7-9700-6300edf4e55b · outbound

This paper cites Timechat: A time-sensitive multimodal large lan- guage model for long video understanding.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Timechat: A time-sensitive multimodal large lan- guage model for long video understanding

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T15:44:47.619011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:44:47.619011Z digest=sha256:189c3a728d2fb6c5da76ab6da07b9b62de247f1a9dd4bb52b7649fd86c9aa812

Observation 9f2e9711-22f4-4fe6-98c5-e186b4b5e058 · outbound

This paper cites Llm-planner: Few-shot grounded planning for embodied agents with large language models.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Llm-planner: Few-shot grounded planning for embodied agents with large language models

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:49.944499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:44:47.715754Z digest=sha256:5bdc16738477cd1ae55bcad6ed2fc160b7e2662581663dad91def01726256b88

Observation 0332c79f-4bd2-4097-a37d-4d215e106510 · outbound

This paper cites Moviechat: From dense token to sparse memory for long video understanding.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Moviechat: From dense token to sparse memory for long video understanding

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T15:44:47.800874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:44:47.800874Z digest=sha256:00e21d5ef9a9c12be9b782d99d8303f5cddd9c89e0c79721ea85049bcde86a17

Observation 5d54eda3-0659-45a6-9c88-e1e0f13aaf97 · outbound

This paper cites Coin: A large-scale dataset for comprehensive instructional video analysis.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Coin: A large-scale dataset for comprehensive instructional video analysis

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:49.932602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:44:47.933683Z digest=sha256:85c671f9f8b1fddc304f55294eaaf0b7c9fcacaf5d7ba12493d5548ce056662d

Observation f219ca49-d55a-4d09-96f2-08470039b654 · outbound

This paper cites Learning Multiple Object States from Actions via Large Language Models.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Learning Multiple Object States from Actions via Large Language Models

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:44:49.717370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:44:48.030903Z digest=sha256:9315eba287e9667e709ddf0fb49437528f6abccce7df3a76afdc46b9ae535ed4

Observation 5f6f2c71-ab4a-44a8-a5de-3d2bf337fcc5 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T15:44:48.175250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:44:48.175250Z digest=sha256:69b0cd83b9fc5341b509634fa1328c65499622c7d10e0d6fb62afcc02ea361b8

Observation ddc243e0-b54d-4218-ba0d-974bfd281f57 · outbound

This paper cites User-in-the-loop Evaluation of Multimodal LLMs for Activity Assistance.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction User-in-the-loop Evaluation of Multimodal LLMs for Activity Assistance

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:44:49.697709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:44:48.352122Z digest=sha256:f323af561e0a6ce60e880e21c15ecbc933b319d9899e55fe053db1ded8624286

Observation 3b647031-8ff2-4b39-9107-29dbc407fde3 · outbound

This paper cites Event-guided procedure planning from in- structional videos with text supervision.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Event-guided procedure planning from in- structional videos with text supervision

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:49.925468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:44:48.507610Z digest=sha256:dda42de8e7556d815789a3a9bd0806594cb7879697d353c7121e9cb195af5795

Observation 7bd79037-3e2b-4fd2-afcb-c52500d9e309 · outbound

This paper cites Pdpp: Projected diffusion for procedure planning in instructional videos.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Pdpp: Projected diffusion for procedure planning in instructional videos

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:49.917737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:44:48.650999Z digest=sha256:82e3583656cb076a0f4b2b5257b00eae8974deb15ab5df3ad60a7e6e6b16c0e2

Observation 9cb4e834-73d3-49bc-aaee-44abd6ea5728 · outbound

This paper cites Vamos: Versatile Action Models for Video Understanding.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Vamos: Versatile Action Models for Video Understanding

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:44:49.687256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:44:48.688715Z digest=sha256:1054145113276f167a67a4b079f352051273244e605f9cc22cb00f9a9163092d

Observation a088458d-7634-4f1c-8ca8-51e26ea5be01 · outbound

This paper cites Learn- ing object state changes in videos: An open-world perspec- tive.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Learn- ing object state changes in videos: An open-world perspec- tive

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:49.909513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:44:48.748675Z digest=sha256:42719dfd89e21ef23f477b5d7e6819d11bfb653f5f234492baadf06e3461cc34

Observation 956f8e27-856e-4355-936d-0666a9808099 · outbound

This paper cites Octopus: Embodied vision- language programmer from environmental feedback, 2023.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Octopus: Embodied vision- language programmer from environmental feedback, 2023

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:49.901873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:44:48.801693Z digest=sha256:31f009e987aefc23720b8820d4893d3d69ad85a35223784582b4f6327e6ad9fb

Observation 1d5423c0-02f8-49f9-9e76-ff78bfff4087 · outbound

This paper cites RAP: Retrieval-Augmented Planner for Adaptive Procedure Planning in Instructional Videos.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction RAP: Retrieval-Augmented Planner for Adaptive Procedure Planning in Instructional Videos

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:44:49.675996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:44:48.842813Z digest=sha256:4f87edb768966f57f27de5d0f1b3dbe3d7dee19bd8c6db31e9ee6283cda9266a

Observation a03f4ecb-e404-4648-bf90-07472dcd2fc3 · outbound

This paper cites Object-centric video representation for long-term action anticipation.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Object-centric video representation for long-term action anticipation

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:49.894247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:44:48.910963Z digest=sha256:0fcbf4ab8ecf63960a48e9ee982f868d74c9de06e53511f737b4707c14cadd66

Observation 23a278c6-bd3f-43e7-ab9b-259c65b1eb56 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T15:44:48.976855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:44:48.976855Z digest=sha256:9217795c567fddbba2d704f1671c501c624f0e41452bd1134a7f24addce22f9a

Observation b78c2393-32b0-4215-884f-cf024c8db353 · outbound

This paper cites P3iv: Prob- abilistic procedure planning from instructional videos with weak supervision.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction P3iv: Prob- abilistic procedure planning from instructional videos with weak supervision

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:49.886095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:44:49.021366Z digest=sha256:9f3d4dc264285f3242a574968b0534b99058fe97d3c5c9cfa41602ed2baeb7cf

Observation 70571645-9ad2-4347-bea0-571aa42e7e8b · outbound

This paper cites AntGPT: Can Large Language Models Help Long-term Action Anticipation from Videos?.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction AntGPT: Can Large Language Models Help Long-term Action Anticipation from Videos?

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T15:44:49.029086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:44:49.029086Z digest=sha256:636e7019b6034a8ab112d7fe2fb8b632edce5380c71d00ac12f22d44ce3be5ba

Observation 653059be-6758-4f8a-ad93-a510d079ded0 · outbound

This paper cites Towards learning a generalist model for embod- ied navigation.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Towards learning a generalist model for embod- ied navigation

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:49.878331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:44:49.034663Z digest=sha256:56c89f5009018aa9271e8db00115ef3dfc3ab3a3ca6dd7d3ac6b57f94b96c07b

Observation c8f553c1-b4d0-4bc1-9634-e562077195be · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in Neural Information Processing Systems, 36:46595–46623, 2023.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in Neural Information Processing Systems, 36:46595–46623, 2023

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T15:44:49.120307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:44:49.120307Z digest=sha256:ba15935ec5d2d4d3599cfe937b8cff2fd9528d493323edca451c2479274e13d0

Observation ed1535af-4470-4e8e-bb0b-5efafbe389f7 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T15:44:49.190266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:44:49.190266Z digest=sha256:e5bbd9120b262f37ceda47569a773659f679e4942df7a3c584d8e6408d48f308

Observation 58977a31-b981-4e08-ae42-3f13da79417e · outbound

This paper cites Cross- task weakly supervised learning from instructional videos.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Cross- task weakly supervised learning from instructional videos

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:49.866298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:44:49.288340Z digest=sha256:55c70f0c0422c15bd73332921014b0ba2287c2a4763df335090d58dae3abf3f8

Observation aa874697-635c-4d25-8905-c864c4106a8e · outbound

This paper cites before” with “after.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction before” with “after

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:49.858942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:44:49.417662Z digest=sha256:453f2c011dcb984fb74aed13c2736d18a33c09db5a3ca3d2e7ddca154d83d9f9

Observation df875fdf-dfaf-490e-b2ae-bd12d17b3fe7 · outbound

This paper cites Training We train our model for 1 epoch with a batch size of 1024.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Training We train our model for 1 epoch with a batch size of 1024

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:49.851659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:44:49.495928Z digest=sha256:cb7e182ca8b5247af627722488b93b5ac27225cf9e29b038368be9d5c0adc6e5

Observation 889eb96f-8bc6-4f6a-b0ec-b5d88ce41d6b · outbound

This paper cites dough”, “container.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction dough”, “container

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:49.843986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:44:49.592008Z digest=sha256:8e5e5cda2b643b8bfbc5a344a8a1ca49240e7d718dbcf9510ec48fc1701f8a39

Pith citing papers

Observation 115e9f3b-1d83-4ccd-9d0b-7f6cbd293e61 · inbound

GeoWorld: Geometric World Models cites this paper.

GeoWorld: Geometric World Models Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction

Reference 85

Resolution
verified exact
arxiv_id, observed 2026-05-21T11:40:03.352082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T11:39:15.308355Z digest=sha256:e5faa389fc8ebdaf4a8d63eb7164e6ca9f3a4fdc89258fb3d328482667459f4e