Pith. sign in

Paper Citation Record · LEDGER

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models

As of 19 August 2026, this Paper Citation Record lists 91 of 91 outbound references and 3 inbound Pith citation observations for arXiv:2501.12231.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.12231 v1

Coverage vector

measured 91 of 91 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T17:27:18.614377Z

measured 94 of 94 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T22:00:28.350003Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T17:27:15.161769Z

Reference resolution

91 of 91 outbound references displayed

  • verified exact1
  • verified fuzzy60
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0f054d61-0dba-499f-bdf7-2471f4a6d132 · outbound

This paper cites GPT-4 Technical Report.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.147094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.147094Z digest=sha256:fe40cd92ee557a554a4ae0f42e39729f3a9ff39d1e913748cc8418138c6dca2e

Observation bbc7e294-1f2e-42a4-baa8-c94c11f1c69b · outbound

This paper cites Ht-step: Aligning instructional articles with how-to videos.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Ht-step: Aligning instructional articles with how-to videos

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.153382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.153382Z digest=sha256:8c88b1b10bd91dfc489ac595c189173ac23a83dfc92d89029e63effe4b426be4

Observation 0fd5688d-92c4-439c-9b05-de697e4f6142 · outbound

This paper cites Unsuper- vised learning from narrated instruction videos.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Unsuper- vised learning from narrated instruction videos

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.159361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.159361Z digest=sha256:7d04ca9f6d809e531fc27264f6cd2e12483656aad0c44fdab0c36bbef7bc79a9

Observation 3a441990-3e0c-4b3b-9787-2254fbc2ab74 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Flamingo: a visual language model for few-shot learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.164584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.164584Z digest=sha256:3366bf1d9c945852c5b342fd946fa69ff5dc43297ae896fcccae8fb331c08842

Observation c9fb148e-45cb-41a0-b77a-feb5ee36bcc4 · outbound

This paper cites Localizing moments in video with natural language.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Localizing moments in video with natural language

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.169866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.169866Z digest=sha256:f188cfc22c0b6c9f9eac11d129934c1f5dc4dcce6f3fbbe9c9b1ba6b8f10691f

Observation f91d33f1-4595-4952-b0cb-f65cf552446a · outbound

This paper cites https://www.anthropic.com/news/developing- computer-use, 2024.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models https://www.anthropic.com/news/developing- computer-use, 2024

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.175183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.175183Z digest=sha256:0402675263960c3fc8d5808aea8983a4f8fe01b71e7261e83961823cf5361591

Observation f612e43d-53cd-4fbc-b5f4-4c457b0e5899 · outbound

This paper cites Video-mined task graphs for keystep recognition in instructional videos.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Video-mined task graphs for keystep recognition in instructional videos

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.180832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.180832Z digest=sha256:9200b44f71ba66075f8c2a5e764b7f63cad614b2f73ebe85b0de0623b47d30f2

Observation bd3ec09d-363e-4712-a532-65b93d9943f6 · outbound

This paper cites Detours for navigating instructional videos.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Detours for navigating instructional videos

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.186508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.186508Z digest=sha256:d4719fb091178ea020efb4c1a6c882e244f04f2f2347fa75c0c79229a630215a

Observation 6a74ed87-5bf5-4e06-b027-655d79e4f239 · outbound

This paper cites Is space-time attention all you need for video understanding? In ICML, page 4, 2021.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Is space-time attention all you need for video understanding? In ICML, page 4, 2021

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.191609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.191609Z digest=sha256:7d4ae9f8f0b4874bff9d3102bd873f3f736dd4cc3faaca45d895f9968b3b009b

Observation 5a9d14d8-935c-4e9b-a745-08498ce1091a · outbound

This paper cites Procedure planning in instructional videos via contextual modeling and model-based policy learning.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Procedure planning in instructional videos via contextual modeling and model-based policy learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.196739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.196739Z digest=sha256:2a921d4155b16f9bc289987c713fef2d1923277e13426dc32bdf4e2084816f62

Observation fe6c2e33-9f1f-4e96-bbd2-b2531c6b019d · outbound

This paper cites Activitynet: A large-scale video bench- mark for human activity understanding.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Activitynet: A large-scale video bench- mark for human activity understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.201677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.201677Z digest=sha256:d20c8173f73a2984887959edbd4b9de722ecf8c7e8db4d92fca0b2765cbb964e

Observation ef2fe9c7-d606-4e31-9992-1f3a3d34a5f7 · outbound

This paper cites Procedure planning in instructional videos.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Procedure planning in instructional videos

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.206432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.206432Z digest=sha256:ec8b9694bb06b987aae734b0cd7333cfc616435137574a453397edb548f0dd6c

Observation 9c129299-5f51-4509-88bb-069ccd546940 · outbound

This paper cites Temporally grounding natural sentence in video.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Temporally grounding natural sentence in video

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:20.077125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T17:27:18.211601Z digest=sha256:35da53696a7da4264f3f18f26710db486e0664ac64413a582479215dd6541d29

Observation 0564b7c4-b66d-427d-b20f-8925c9539acf · outbound

This paper cites Videollm-online: Online video large language model for streaming video.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Videollm-online: Online video large language model for streaming video

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:20.051969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T17:27:18.216331Z digest=sha256:ebf072fcd48ef57300d747410a00dbcb91321d9f1196deca1fcadb6f85c204bb

Observation 86bbecf3-2c85-4636-a1de-28804f6ceac4 · outbound

This paper cites Semantic proposal for activity localization in videos via sentence query.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Semantic proposal for activity localization in videos via sentence query

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:20.032877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T17:27:18.221166Z digest=sha256:de18db5815ae6e735bc3dcfbb353a0102375bac94f395d1d2671bfe11a1a096c

Observation f00acf7c-0e9b-4e87-bb24-a4de1732b50c · outbound

This paper cites KGPT: Knowledge-grounded pre-training for data-to-text generation.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models KGPT: Knowledge-grounded pre-training for data-to-text generation

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:20.014111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T17:27:18.226036Z digest=sha256:10e7db4257f707238b3d03003dac26b7025d83e48ff09f05ee3546d2db227e9e

Observation c78474aa-08c1-4bfe-9001-d22221ae367e · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models An image is worth 16x16 words: Transformers for image recognition at scale

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.995691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T17:27:18.230690Z digest=sha256:ffe73d382a47a04839de94bbd694443eb6b1056bc1c6849184c4906b0fc762b6

Observation 2ed87957-4a6a-47db-acb3-e2d79d793410 · outbound

This paper cites From Local to Global: A Graph RAG Approach to Query-Focused Summarization.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models From Local to Global: A Graph RAG Approach to Query-Focused Summarization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.235459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.235459Z digest=sha256:66d400cf4c760a375b72f3f150cfe951e21f05f105b7855fde5ba9c4a5ddd960

Observation d04cc3f4-08c4-41c6-bd1a-7323530553ce · outbound

This paper cites Masked Diffusion with Task-awareness for Procedure Planning in Instructional Videos.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Masked Diffusion with Task-awareness for Procedure Planning in Instructional Videos

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-10T17:27:18.797148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T17:27:18.240608Z digest=sha256:f6f7e89a45cd71c4d71f164676209f9e4752178e1eb1e549cb55707136871ac7

Observation b72e409f-5eea-483d-8074-0b85a7d1c6be · outbound

This paper cites Slowfast networks for video recognition.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Slowfast networks for video recognition

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.977834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T17:27:18.246444Z digest=sha256:dc7c569ed6b63d6ff33750c591133cae4f8627733e98df8259779eadef6097d7

Observation 248d6cc9-cd8e-4fe7-aeb6-f54c9509f8ed · outbound

This paper cites Tall: Temporal activity localization via language query.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Tall: Temporal activity localization via language query

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.251469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.251469Z digest=sha256:5831e4c810ed4aeaf8396ad64f2b88a53bc71de474269ac6f61183d8d4a3ed29

Observation 629d54ac-d2a0-4cf0-a721-22f9663b7c5e · outbound

This paper cites Retrieval-Augmented Generation for Large Language Models: A Survey.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Retrieval-Augmented Generation for Large Language Models: A Survey

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.256148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.256148Z digest=sha256:4e2c40862558475ed65d1b6873d7affd4b4f12c0f5cf09eb6a50625c1a95ca4f

Observation cc755edc-ea2f-4e1c-8cd0-93431f93e330 · outbound

This paper cites Mac: Mining activity concepts for language-based temporal local- ization.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Mac: Mining activity concepts for language-based temporal local- ization

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.261755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.261755Z digest=sha256:c2566b80b6217676c8a82bc62d004b99b2ee1c65f3c2ce3920cc4a7621de6b2a

Observation 537d9060-b12a-4da7-8cec-e84dec96dfc6 · outbound

This paper cites Imagebind: One embedding space to bind them all.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Imagebind: One embedding space to bind them all

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.266660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.266660Z digest=sha256:eec2e8a8a15844a1a3719297392c78e23e1d9d878f2e63ead49af19546ba1af9

Observation 81d1ae0f-f85d-46d6-b357-d1cc5c09400d · outbound

This paper cites Radar: automated task planning for proactive decision support.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Radar: automated task planning for proactive decision support

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.928432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T17:27:18.271492Z digest=sha256:5774b033042dc7949a3cf6f3451eabcc72c8e5647046c5a7de6f06fe478f7c0a

Observation 170d0ab5-d947-44b1-9229-e48d410d753d · outbound

This paper cites Retrieval augmented language model pre- training.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Retrieval augmented language model pre- training

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.911731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T17:27:18.276299Z digest=sha256:4a2037fa37d65217ab5818a9a0f7243285f72504eafc33735829914e7e4ab1eb

Observation 246d6b37-dad7-46cf-83e6-2d983773ddd9 · outbound

This paper cites Ma-lmm: Memory-augmented large multimodal model for long-term video understanding.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Ma-lmm: Memory-augmented large multimodal model for long-term video understanding

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.892002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T17:27:18.281075Z digest=sha256:45f3e205a15e4d722daae26ddd497471df4cf71905b3d00b37286e451c549b72

Observation fd43daab-b8d6-491e-b28b-8a999a1cfcae · outbound

This paper cites Denoising diffu- sion probabilistic models.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Denoising diffu- sion probabilistic models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.286179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.286179Z digest=sha256:f8add948e5a35abceb8d2058003b091fdac77066a4553b849f1cda716fdd7af6

Observation 2d50f50e-e48c-4099-b735-611dcdb025ad · outbound

This paper cites LoRA: Low-rank adaptation of large language models.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models LoRA: Low-rank adaptation of large language models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.864426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T17:27:18.291143Z digest=sha256:584c6a18241f02dd2ba8f89e6ae95d2853ea539b129d29d1af4736fdadcace76

Observation 048154c4-394f-48b6-a5e4-456912107e24 · outbound

This paper cites Audiogpt: Understanding and generating speech, music, sound, and talking head.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Audiogpt: Understanding and generating speech, music, sound, and talking head

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.846787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T17:27:18.296118Z digest=sha256:9eb7fabe58031b9a7fc4d9a1f07792ab98a77db5607f212e1c756077bf97a1ad

Observation b5ed1157-9431-4931-a58e-1f8b0c20658a · outbound

This paper cites Language is not all you need: Aligning perception with language models.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Language is not all you need: Aligning perception with language models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.829639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T17:27:18.301317Z digest=sha256:d561349df80918533d0fddbf619f8d3f7e14446468db547cd0338de9f2d403d1

Observation 300761ce-e88f-4c10-981d-959ac6b77067 · outbound

This paper cites Leveraging passage retrieval with generative models for open domain question answering.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Leveraging passage retrieval with generative models for open domain question answering

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.813004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T17:27:18.306835Z digest=sha256:bfc8219488d9afc2b2565f59bfef170f348156a2168295e0520ffe870d2618d1

Observation 3ef6a6de-5536-4722-b939-7b113c4ef4fa · outbound

This paper cites Mistral 7B.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Mistral 7B

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.311567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.311567Z digest=sha256:c71cdabb542de82727be3614d291bbec803699afc5b6b9cd2b4340de88a09805

Observation a0310a70-554e-4f9b-9f44-de75ffca7a9a · outbound

This paper cites Gen- erating images with multimodal language models.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Gen- erating images with multimodal language models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.317157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.317157Z digest=sha256:d997d8eb6da14bde1de470f895f5749ee5fedb624c9fbf30d6eac27ed31f0008

Observation 1aa40df6-34c4-4764-9a09-ecfe61dc25a9 · outbound

This paper cites WikiHow: A Large Scale Text Summarization Dataset.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models WikiHow: A Large Scale Text Summarization Dataset

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.322049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.322049Z digest=sha256:e0bb681aedf250a4907d893dc1cd625cdb9a832006b87ef0cd6ce3c00debf388

Observation 9bad06d2-7f42-44ae-96b8-95d227ef9d82 · outbound

This paper cites A survey on temporal sentence grounding in videos.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models A survey on temporal sentence grounding in videos

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.785086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T17:27:18.328096Z digest=sha256:79477115aedfadf8103cb5438f2ed9fae8c5a0a548b42ac0348111e761124827

Observation 4f859fb8-ec1e-47e5-bdac-e4c9aa136e71 · outbound

This paper cites Tvr: A large-scale dataset for video-subtitle moment retrieval.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Tvr: A large-scale dataset for video-subtitle moment retrieval

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.767356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T17:27:18.333044Z digest=sha256:7463b71ac021f2cc135036277627e2067ee80d5b32a3fe7318d65507afe38440

Observation b3dabebf-3f94-4b74-a5b4-6ec0448383ea · outbound

This paper cites Less is more: Clipbert for video-and-language learning via sparse sampling.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Less is more: Clipbert for video-and-language learning via sparse sampling

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.749524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T17:27:18.338050Z digest=sha256:3080174b93c16157191c951f3b068402a1ab367cd15e39f15dfce292770468ec

Observation 596d4a41-5f75-49d6-be14-de21ac0339ef · outbound

This paper cites Chatting makes perfect: Chat-based image retrieval.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Chatting makes perfect: Chat-based image retrieval

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.343420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.343420Z digest=sha256:e887fbee5b4751559059cb1c1d7582c32366efe9d90c011bf01cdaf6d54ccf53

Observation c7e9112a-b8ab-44e2-92bf-84a31838eade · outbound

This paper cites Retrieval- augmented generation for knowledge-intensive nlp tasks.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Retrieval- augmented generation for knowledge-intensive nlp tasks

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.721183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T17:27:18.348350Z digest=sha256:17bb49050923a250ef16be612ce1a06731d886411a5a04cf867bbf0f14d4f1be

Observation e8f42b6c-0d31-470d-8001-d8f90161d8e3 · outbound

This paper cites Blip- 2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Blip- 2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.353346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.353346Z digest=sha256:c9a11614cf9e699fd725885316b7c101ce8288a71e1bd47334dafba012906511

Observation 07b40416-83e2-4efc-bbd4-0a23843c2bbb · outbound

This paper cites Skip-plan: Procedure planning in instructional videos via condensed action space learning.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Skip-plan: Procedure planning in instructional videos via condensed action space learning

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.690149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T17:27:18.358932Z digest=sha256:7c91fe91389ca53d73201e1a3ce604f42aa2cda77a03ae69b89af3e952a419c1

Observation a7f9a262-5fb1-419d-b595-0fca782911c5 · outbound

This paper cites Cohn, and Janet B.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Cohn, and Janet B

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.665546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T17:27:18.363704Z digest=sha256:833052f5fd5515166aa3f7ae0c266b48ca04f20de2d89c2eb23c1ff231df4498

Observation 40b457fb-a7cd-44a8-a55b-8708313d02b8 · outbound

This paper cites Learning to recognize procedural activities with distant supervision.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Learning to recognize procedural activities with distant supervision

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.643121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T17:27:18.368912Z digest=sha256:427e2901b30f3dc05a8c443dfc37447e9acac7c84542f793b8dc67e1fe50ed1e

Observation 4713edc8-2e9f-4a33-acec-590ae14c0c9d · outbound

This paper cites Jointly cross-and self-modal graph attention network for query-based moment localization.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Jointly cross-and self-modal graph attention network for query-based moment localization

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.617330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T17:27:18.374456Z digest=sha256:17c62d47ed659cb384491432379f027285a80420b9e692c10156b9a8e1417349

Observation d8bfad8e-0392-4876-af25-c9613c35bd01 · outbound

This paper cites Improved baselines with visual instruction tuning.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Improved baselines with visual instruction tuning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.380143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.380143Z digest=sha256:c92cda699ed35d06a1a7180b1f04b935ac991ddf93faeb09255c208d4f737c9f

Observation b8aec459-d859-4fbd-9f72-631649fb012d · outbound

This paper cites Visual instruction tuning.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Visual instruction tuning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.385014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.385014Z digest=sha256:b689787dfca18914c5d94ff11e230abbeb6eb983a9f054c130804bcbf19add55

Observation 46da08f4-c521-47d8-8139-6a60145d8163 · outbound

This paper cites Attentive moment retrieval in videos.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Attentive moment retrieval in videos

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.569112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T17:27:18.390493Z digest=sha256:1cc605f7197103ace96af36ad02a49286379e53396e7df4ec8ba33153e2ccc26

Observation 26decaed-decf-42f3-85d4-e172b0f4d7a6 · outbound

This paper cites Language models of code are few-shot commonsense learners.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Language models of code are few-shot commonsense learners

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.545580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T17:27:18.395669Z digest=sha256:3d40b1b1d9c6692bbbd80c4e6c09ae7536d0c60689dbf28355d1d94f7419b290

Observation ebe4190f-1994-4be1-b33b-a277349828bf · outbound

This paper cites What’s cookin’? interpreting cooking videos using text, speech and vision.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models What’s cookin’? interpreting cooking videos using text, speech and vision

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.517289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T17:27:18.400639Z digest=sha256:2fb85cc6a57f8126c7588bf6c4e6222af25ec4e1954ebc5e499780b6faa916c3

Observation 8abe8ae2-99f4-4a9e-8f58-fab46928dc7c · outbound

This paper cites Howto100m: Learning a text-video embedding by watching hundred mil- lion narrated video clips.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Howto100m: Learning a text-video embedding by watching hundred mil- lion narrated video clips

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.496609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T17:27:18.405575Z digest=sha256:bf7886224a7b8dff8edfa30e2757e55c424cee63ea9ee780f9a8da8998048152

Observation 0f3350a2-a0a3-4501-ad92-c2c0d25d5f24 · outbound

This paper cites End-to-end learn- ing of visual representations from uncurated instructional videos.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models End-to-end learn- ing of visual representations from uncurated instructional videos

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.470230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T17:27:18.410996Z digest=sha256:5d383f68679581c99548007870b63d140ee702b5f8e09e670a4d2e834fa5be50

Observation ad815325-c0ba-41c6-a0dd-6595078478a7 · outbound

This paper cites Learning and Verification of Task Structure in Instructional Videos.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Learning and Verification of Task Structure in Instructional Videos

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.415992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.415992Z digest=sha256:8ff341ae5904a8330db614a017563d9758d409f3dc787726b099704da448112f

Observation ae29bc64-05ae-4610-b9f1-c961e70fb6e0 · outbound

This paper cites SCHEMA: State CHanges MAtter for procedure planning in instructional videos.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models SCHEMA: State CHanges MAtter for procedure planning in instructional videos

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.451365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T17:27:18.421653Z digest=sha256:1573d26b7e5252feb2c123c8b05c6c94c03b5aa81fa7c16abfdbbdf6f3403173

Observation 46042b3c-b577-4cd8-a223-dec461577ac3 · outbound

This paper cites Virtualhome: Sim- ulating household activities via programs.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Virtualhome: Sim- ulating household activities via programs

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.430224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T17:27:18.427256Z digest=sha256:3d0e8185509c778e309ec6143b408d6221ac0506de34c6e331293ab9a0cd0086

Observation 1d462551-7fc9-4673-824b-d3ef509b4255 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Learning transferable visual models from natural language supervi- sion

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.432722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.432722Z digest=sha256:9b9ba1c48db502020e44672267e0700520af1fe946a14cc1ae3f211103b91802

Observation f28cd14f-2aa5-49f4-b2f8-1edd37df038a · outbound

This paper cites Grounding action descriptions in videos.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Grounding action descriptions in videos

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.398146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T17:27:18.437992Z digest=sha256:3185f73c288204dc407a04e623f6fa7e5dfb0c1f664ed92dedb03705d6af8de5

Observation 8f268d0b-267a-4bcd-b7ac-7362cd2a945d · outbound

This paper cites FLAP: Flow-adhering planning with constrained decoding in LLMs.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models FLAP: Flow-adhering planning with constrained decoding in LLMs

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.380232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T17:27:18.443152Z digest=sha256:48d52ead5402681e124282c83fe1ca324b851f8bbb31553c67ebe49cb3508594

Observation 69608492-668a-46ee-a34e-8077fe24d1a5 · outbound

This paper cites proScript: Par- tially ordered scripts generation.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models proScript: Par- tially ordered scripts generation

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.363004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T17:27:18.448309Z digest=sha256:32cf1dbd9b4fa35c8bde9a0aee4d77c8377aaf3a375fd0b3437d00a53dda36b2

Observation 88d099c9-b89b-4ee9-b1a2-b817958e3567 · outbound

This paper cites As- sembly101: A large-scale multi-view video dataset for un- derstanding procedural activities.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models As- sembly101: A large-scale multi-view video dataset for un- derstanding procedural activities

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.345705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T17:27:18.453462Z digest=sha256:7a24e57e6b28c05df5139a1b4026d1ad386475eb610d5f5151b581807ccf4300

Observation 374b58e8-0949-41aa-ae82-225144ebb97e · outbound

This paper cites Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.327594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T17:27:18.458789Z digest=sha256:2a8ed0a5be748fe586067f64aa9fed9c8608704470aa375ea3496241f43779da

Observation 5c00df61-fe98-42ae-aa1e-54de161d9b76 · outbound

This paper cites Moviechat: From dense token to sparse memory for long video understanding.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Moviechat: From dense token to sparse memory for long video understanding

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.310703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T17:27:18.464824Z digest=sha256:6fc4b93dc281d1efbaadf9230fd734ac4ac8d94e6097f239266c04838bb7f8ba

Observation ebf6b249-9a9b-4921-b8c6-45eb5f078a1a · outbound

This paper cites Mpnet: Masked and permuted pre-training for language understanding.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Mpnet: Masked and permuted pre-training for language understanding

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.293438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T17:27:18.469567Z digest=sha256:767fd55ecac5c24d6d40c4d0b014e0e3a983142d2308a55ff54f3f87f914478e

Observation c3a0b2e7-5820-41df-a9ef-2d276b799ea7 · outbound

This paper cites Language Models Can See: Plugging Visual Controls in Text Generation.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Language Models Can See: Plugging Visual Controls in Text Generation

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.474723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.474723Z digest=sha256:c6d8e91f06b268dc33a19de1700eca8de40e23658be25c39f65a489066e183b4

Observation 6019f744-9b5d-4a78-9201-ffe1265ca81f · outbound

This paper cites PandaGPT: One model to instruction-follow them all.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models PandaGPT: One model to instruction-follow them all

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.276653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T17:27:18.479891Z digest=sha256:28b9541d59eec91c1261a2a8148ac6ba39058b82840f31867a3d9b649ce25589

Observation f9f1a374-9afe-4810-83be-9ba1448966c0 · outbound

This paper cites Plate: Visually-grounded planning with transformers in procedural tasks.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Plate: Visually-grounded planning with transformers in procedural tasks

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.258622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T17:27:18.484978Z digest=sha256:450ce2ef45f56a412f50973c9520ba39a6b580e4ac98ae53dec70028e5225a62

Observation 76415b0a-33bb-4fbd-98bb-489517242b8a · outbound

This paper cites Coin: A large-scale dataset for comprehensive instructional video analysis.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Coin: A large-scale dataset for comprehensive instructional video analysis

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.239508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T17:27:18.489713Z digest=sha256:cb48807754efcd984f4faa65a0eda61688fcbf4bac29a6ba70dcaa45a13b8169

Observation a2a0d350-3c6b-4c88-ad2a-6c639902c601 · outbound

This paper cites On the planning abilities of large language models-a critical investigation.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models On the planning abilities of large language models-a critical investigation

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.222261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T17:27:18.495158Z digest=sha256:bd20be951f6fa596fd821f38fb1964f02c102500305630e8af0260c891674dd8

Observation ecf651f8-0f62-4594-85cc-a4041e7331b5 · outbound

This paper cites Gomez, Lukasz Kaiser, and Illia Polosukhin.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Gomez, Lukasz Kaiser, and Illia Polosukhin

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.203062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T17:27:18.500019Z digest=sha256:c4e9e03f08be8b097e6b032ed455bdf7aab427a97c83a8a34e83c439676d14fd

Observation d7d46fa4-61ed-4a8f-8c20-036ce1a5e244 · outbound

This paper cites Event-guided procedure planning from in- structional videos with text supervision.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Event-guided procedure planning from in- structional videos with text supervision

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.185574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T17:27:18.504726Z digest=sha256:cb49b164c16b2a51a7c7605466d355d1efe62941979c40aedd9d8a2512e06656

Observation aa1704bf-798f-4e21-9317-a7a01628461f · outbound

This paper cites Pdpp: Projected diffusion for procedure planning in instructional videos.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Pdpp: Projected diffusion for procedure planning in instructional videos

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.168578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T17:27:18.509417Z digest=sha256:d2f7e911456f64d04ee5295e93d0004227989d2dfb700efeeed9b8e4e5e055c5

Observation 641a44df-1b47-42d3-b018-956b5133109d · outbound

This paper cites Temporal segment networks: Towards good practices for deep action recognition.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Temporal segment networks: Towards good practices for deep action recognition

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.150365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T17:27:18.514360Z digest=sha256:4026b367ff9cad6d9e88985521c34f31a59ae7a92c52109948def9ea615c0661

Observation 5155ce14-7170-4533-a1bc-32587b290c6d · outbound

This paper cites Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.519058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.519058Z digest=sha256:6823fa83c1e1b82e3ab8f908c07e8b3de94040587fc5532c31adaef622637e65

Observation 9f82821f-3eaa-4d58-9597-5dec1e9624c7 · outbound

This paper cites NExt-GPT: Any-to-any multimodal LLM.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models NExt-GPT: Any-to-any multimodal LLM

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.133657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T17:27:18.524104Z digest=sha256:26f0e29d7261871d368ae2af5d122d3e3c0ef4efc56ff51dd32c90e2c5a7b95e

Observation bfe34a2f-97a6-47ab-8583-2618443845c0 · outbound

This paper cites Rethinking spatiotemporal feature learning: Speed-accuracy trade-offs in video classification.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Rethinking spatiotemporal feature learning: Speed-accuracy trade-offs in video classification

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.116500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T17:27:18.528664Z digest=sha256:73f53e76416a5be8a6360f1eec250f9ae308cb2ce23a40147a25c721cc04ad55

Observation b4329a38-a98a-4794-bac9-b97886dcf609 · outbound

This paper cites Translating Natural Language to Planning Goals with Large-Language Models.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Translating Natural Language to Planning Goals with Large-Language Models

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.533727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.533727Z digest=sha256:ddf76df60519ad097039deed195954ceb7c79df673dec3db60f1851da83c7ba3

Observation d198cda2-5116-439d-ac92-f0bafaac0ec3 · outbound

This paper cites Multilevel language and vision integration for text-to-clip retrieval.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Multilevel language and vision integration for text-to-clip retrieval

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.097622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T17:27:18.539015Z digest=sha256:732046553f267375b4c44512ff85f7835b6e5e57c0a97976bb1dcffdcb27c776

Observation bb148e8c-adcf-4f31-b9d8-0bd1fa4bb170 · outbound

This paper cites VideoCLIP: Contrastive pre- training for zero-shot video-text understanding.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models VideoCLIP: Contrastive pre- training for zero-shot video-text understanding

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.079625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T17:27:18.544244Z digest=sha256:aec8b3f505a87c16505162d048aad7ac80331afb01187e7b8b8aee4fb4cc26f4

Observation 6cc9b393-3050-4d91-a6f8-f455fe2de311 · outbound

This paper cites Retrieval- augmented generation with knowledge graphs for customer service question answering.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Retrieval- augmented generation with knowledge graphs for customer service question answering

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.060275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T17:27:18.549134Z digest=sha256:9ecbab1b76d9090e210917b058772bc2adc81fdd0665f6d0012f7899d0517cda

Observation f320f8fd-dd08-47a9-9add-799ba20f929d · outbound

This paper cites Activitynet-qa: A dataset for understanding complex web videos via question answering.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Activitynet-qa: A dataset for understanding complex web videos via question answering

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.043613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T17:27:18.554112Z digest=sha256:7b3ec2a14d2a9e1ee46ff0bb0f66bc7f3a8648876b5aaad7f1a104d3640865e7

Observation 4a9e3b04-e34b-4f3f-bd07-db480e1e6fe9 · outbound

This paper cites Distilling script knowledge from large language models for constrained language planning.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Distilling script knowledge from large language models for constrained language planning

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.025535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T17:27:18.559096Z digest=sha256:1bb8294422d13595591baab804aa94e2351a045f0a73fd15d18600b7a7668e53

Observation 1e5d354d-f8e9-4e08-88f6-567a49c7bad2 · outbound

This paper cites Man: Moment alignment network for natural language moment retrieval via iterative graph adjustment.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Man: Moment alignment network for natural language moment retrieval via iterative graph adjustment

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.007476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T17:27:18.563941Z digest=sha256:4ce95b805c22b3561ca09ef637c8b2c07d21e4167ed9555855226d22e7cc9c0f

Observation e47a14f2-1409-470a-b597-6372db2921c9 · outbound

This paper cites SpeechGPT: Empowering large language models with intrinsic cross-modal conversa- tional abilities.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models SpeechGPT: Empowering large language models with intrinsic cross-modal conversa- tional abilities

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:18.990834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T17:27:18.568729Z digest=sha256:5fc3f78e5875acca3449d7ef7340bfa7d91cc09bf16c9c4a4255bf647e28cc00

Observation 65d682a9-ad24-4ef0-9aff-985af49ebc05 · outbound

This paper cites Video-LLaMA: An instruction-tuned audio-visual language model for video un- derstanding.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Video-LLaMA: An instruction-tuned audio-visual language model for video un- derstanding

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:18.974322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T17:27:18.574278Z digest=sha256:cf284ae88065aa6e774058184d45cbd3783924ea8a6794b545a592683994ffe0

Observation d628f934-164d-4b8b-a513-33d7a7b43b42 · outbound

This paper cites Temporal sentence grounding in videos: A survey and fu- ture directions.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Temporal sentence grounding in videos: A survey and fu- ture directions

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:18.957496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T17:27:18.579525Z digest=sha256:29d8d352d9395a702ecb0623a351c440bc33f03edd7b8f85f703154890238f24

Observation 703bc398-48a0-48c7-9795-3cb775a7a0ae · outbound

This paper cites P3iv: Probabilistic procedure planning from instructional videos with weak su- pervision.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models P3iv: Probabilistic procedure planning from instructional videos with weak su- pervision

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:18.939426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T17:27:18.584443Z digest=sha256:b05b41bbdf6137cf0dad6d24fbdcb341a8c8278863c34b3f7a7d86b15a076807

Observation 1bfea070-78ba-47be-a292-a0398e67453e · outbound

This paper cites Learning procedure-aware video repre- sentation from instructional videos and their narrations.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Learning procedure-aware video repre- sentation from instructional videos and their narrations

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:18.921158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T17:27:18.590544Z digest=sha256:6dd66b5667247775bd6d658a05daba2c3d9411023c4986fd06cfcddd26d84f8c

Observation cd6107ab-a413-41f7-ad15-d3bcb5ac5f94 · outbound

This paper cites Procedure-aware pretraining for instructional video understanding.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Procedure-aware pretraining for instructional video understanding

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:18.904145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T17:27:18.596552Z digest=sha256:b8f0e5a0dc6114a39107250a60d510ed8eef56b3d1b1bd80e8a72c278e58965d

Observation 7901844d-b6bf-424c-b788-b3bd49dd1263 · outbound

This paper cites Towards auto- matic learning of procedures from web instructional videos.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Towards auto- matic learning of procedures from web instructional videos

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:18.887341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T17:27:18.602763Z digest=sha256:507dd95c8cb97a86ec384c34366783b1c9e12b56b5339af1529339de91a5bc3d

Observation d130331b-ff3b-4241-903e-a583413649ee · outbound

This paper cites MiniGPT-4: Enhancing vision-language understanding with advanced large language models.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models MiniGPT-4: Enhancing vision-language understanding with advanced large language models

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:18.869455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T17:27:18.608913Z digest=sha256:6aef557e95b98ae8f526893a97e38e908863ae651c4c08ab0a8a046284b18221

Observation f8de0d6b-998d-4e0c-bf16-2b794b0714b4 · outbound

This paper cites Cross- task weakly supervised learning from instructional videos.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Cross- task weakly supervised learning from instructional videos

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:18.851520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T17:27:18.614377Z digest=sha256:b427ee7e84c916f8b727874f37c800e374325078fb342dd466535a2be587979e

Pith citing papers

Observation 1a817673-2a56-4d5d-916d-4a9b06e41a2b · inbound

SFHand: Learning Embodied Manipulation by Streaming Egocentric 3D Hand Forecasting cites this paper.

SFHand: Learning Embodied Manipulation by Streaming Egocentric 3D Hand Forecasting InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-21T17:54:18.278334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-21T17:53:19.002603Z digest=sha256:306521eb94258f08e399993cc439fcf60ab7fc7faec92f69176869a15d2a9cfe

Observation 94abaaac-ea99-486a-89b4-b240171202ed · inbound

VisionClaw: Always-On AI Agents through Smart Glasses cites this paper.

VisionClaw: Always-On AI Agents through Smart Glasses InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-13T18:13:05.498964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-13T18:12:34.183092Z digest=sha256:e7e55725d98a77cac02590e3124290ac664aa4b03b2161369832b99078a50c91

Observation 3b18f4b1-b9d9-41b9-9fb7-55667d3f9008 · inbound

Watch, Remember, Reason: Human-View Video Understanding with MLLMs cites this paper.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models

Reference 256

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:15.163371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:7c054bdb31ebce32b4d8f488b45d5f1bfb30af6853bbe81344caa7c062c9399f