Pith. sign in

Paper Citation Record · LEDGER

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models

As of 19 August 2026, this Paper Citation Record lists 91 of 91 outbound references and 3 inbound Pith citation observations for arXiv:2501.12231.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.12231 v1

Coverage vector

measured 91 of 91 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T17:27:18.614377Z

measured 94 of 94 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T22:00:28.350003Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T17:27:15.161769Z

Reference resolution

91 of 91 outbound references displayed

  • verified exact1
  • verified fuzzy60
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0f054d61-0dba-499f-bdf7-2471f4a6d132 · outbound

This paper cites GPT-4 Technical Report.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.147094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.147094Z digest=sha256:fe40cd92ee557a554a4ae0f42e39729f3a9ff39d1e913748cc8418138c6dca2e

Observation bbc7e294-1f2e-42a4-baa8-c94c11f1c69b · outbound

This paper cites Ht-step: Aligning instructional articles with how-to videos.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Ht-step: Aligning instructional articles with how-to videos

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.153382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.153382Z digest=sha256:8c88b1b10bd91dfc489ac595c189173ac23a83dfc92d89029e63effe4b426be4

Observation 0fd5688d-92c4-439c-9b05-de697e4f6142 · outbound

This paper cites Unsuper- vised learning from narrated instruction videos.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Unsuper- vised learning from narrated instruction videos

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.159361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.159361Z digest=sha256:7d04ca9f6d809e531fc27264f6cd2e12483656aad0c44fdab0c36bbef7bc79a9

Observation 3a441990-3e0c-4b3b-9787-2254fbc2ab74 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Flamingo: a visual language model for few-shot learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.164584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.164584Z digest=sha256:3366bf1d9c945852c5b342fd946fa69ff5dc43297ae896fcccae8fb331c08842

Observation c9fb148e-45cb-41a0-b77a-feb5ee36bcc4 · outbound

This paper cites Localizing moments in video with natural language.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Localizing moments in video with natural language

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.169866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.169866Z digest=sha256:f188cfc22c0b6c9f9eac11d129934c1f5dc4dcce6f3fbbe9c9b1ba6b8f10691f

Observation f91d33f1-4595-4952-b0cb-f65cf552446a · outbound

This paper cites https://www.anthropic.com/news/developing- computer-use, 2024.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models https://www.anthropic.com/news/developing- computer-use, 2024

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.175183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.175183Z digest=sha256:0402675263960c3fc8d5808aea8983a4f8fe01b71e7261e83961823cf5361591

Observation f612e43d-53cd-4fbc-b5f4-4c457b0e5899 · outbound

This paper cites Video-mined task graphs for keystep recognition in instructional videos.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Video-mined task graphs for keystep recognition in instructional videos

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.180832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.180832Z digest=sha256:9200b44f71ba66075f8c2a5e764b7f63cad614b2f73ebe85b0de0623b47d30f2

Observation bd3ec09d-363e-4712-a532-65b93d9943f6 · outbound

This paper cites Detours for navigating instructional videos.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Detours for navigating instructional videos

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.186508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.186508Z digest=sha256:d4719fb091178ea020efb4c1a6c882e244f04f2f2347fa75c0c79229a630215a

Observation 6a74ed87-5bf5-4e06-b027-655d79e4f239 · outbound

This paper cites Is space-time attention all you need for video understanding? In ICML, page 4, 2021.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Is space-time attention all you need for video understanding? In ICML, page 4, 2021

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.191609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.191609Z digest=sha256:7d4ae9f8f0b4874bff9d3102bd873f3f736dd4cc3faaca45d895f9968b3b009b

Observation 5a9d14d8-935c-4e9b-a745-08498ce1091a · outbound

This paper cites Procedure planning in instructional videos via contextual modeling and model-based policy learning.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Procedure planning in instructional videos via contextual modeling and model-based policy learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.196739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.196739Z digest=sha256:2a921d4155b16f9bc289987c713fef2d1923277e13426dc32bdf4e2084816f62

Observation fe6c2e33-9f1f-4e96-bbd2-b2531c6b019d · outbound

This paper cites Activitynet: A large-scale video bench- mark for human activity understanding.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Activitynet: A large-scale video bench- mark for human activity understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.201677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.201677Z digest=sha256:d20c8173f73a2984887959edbd4b9de722ecf8c7e8db4d92fca0b2765cbb964e

Observation ef2fe9c7-d606-4e31-9992-1f3a3d34a5f7 · outbound

This paper cites Procedure planning in instructional videos.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Procedure planning in instructional videos

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.206432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.206432Z digest=sha256:ec8b9694bb06b987aae734b0cd7333cfc616435137574a453397edb548f0dd6c

Observation 9c129299-5f51-4509-88bb-069ccd546940 · outbound

This paper cites Temporally grounding natural sentence in video.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Temporally grounding natural sentence in video

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:20.077125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:27:18.211601Z digest=sha256:87b89a98fb142ab3c7554b4d74d038dc7ff0e418e2a321c6049dc9f827195c6c

Observation 0564b7c4-b66d-427d-b20f-8925c9539acf · outbound

This paper cites Videollm-online: Online video large language model for streaming video.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Videollm-online: Online video large language model for streaming video

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:20.051969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:27:18.216331Z digest=sha256:097f3ddb90a243804633d5844813cd50f71bcdd4761d8da045c1395788500732

Observation 86bbecf3-2c85-4636-a1de-28804f6ceac4 · outbound

This paper cites Semantic proposal for activity localization in videos via sentence query.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Semantic proposal for activity localization in videos via sentence query

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:20.032877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:27:18.221166Z digest=sha256:91f6e394bad75752c13dc88ac1c5ae94cb534175953648538a926389fd0d0953

Observation f00acf7c-0e9b-4e87-bb24-a4de1732b50c · outbound

This paper cites KGPT: Knowledge-grounded pre-training for data-to-text generation.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models KGPT: Knowledge-grounded pre-training for data-to-text generation

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:20.014111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:27:18.226036Z digest=sha256:2dabd5ce4940a01387e0de23601c48a5eddf03e8fc84900271154f582620a0b5

Observation c78474aa-08c1-4bfe-9001-d22221ae367e · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models An image is worth 16x16 words: Transformers for image recognition at scale

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.995691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:27:18.230690Z digest=sha256:a074920cfde6b39040b34750dbe520bd878c9594e3958f86331393f60ba61e0e

Observation 2ed87957-4a6a-47db-acb3-e2d79d793410 · outbound

This paper cites From Local to Global: A Graph RAG Approach to Query-Focused Summarization.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models From Local to Global: A Graph RAG Approach to Query-Focused Summarization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.235459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.235459Z digest=sha256:66d400cf4c760a375b72f3f150cfe951e21f05f105b7855fde5ba9c4a5ddd960

Observation d04cc3f4-08c4-41c6-bd1a-7323530553ce · outbound

This paper cites Masked Diffusion with Task-awareness for Procedure Planning in Instructional Videos.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Masked Diffusion with Task-awareness for Procedure Planning in Instructional Videos

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-10T17:27:18.797148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:27:18.240608Z digest=sha256:43b9ccaae284c2f615fd61caf9260e8a02c946501be233996d09d36939005070

Observation b72e409f-5eea-483d-8074-0b85a7d1c6be · outbound

This paper cites Slowfast networks for video recognition.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Slowfast networks for video recognition

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.977834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:27:18.246444Z digest=sha256:5e7b2e0d354ce3b50c2765d2378a884cfed183486dbaebde75932010d687c0bd

Observation 248d6cc9-cd8e-4fe7-aeb6-f54c9509f8ed · outbound

This paper cites Tall: Temporal activity localization via language query.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Tall: Temporal activity localization via language query

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.251469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.251469Z digest=sha256:5831e4c810ed4aeaf8396ad64f2b88a53bc71de474269ac6f61183d8d4a3ed29

Observation 629d54ac-d2a0-4cf0-a721-22f9663b7c5e · outbound

This paper cites Retrieval-Augmented Generation for Large Language Models: A Survey.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Retrieval-Augmented Generation for Large Language Models: A Survey

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.256148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.256148Z digest=sha256:4e2c40862558475ed65d1b6873d7affd4b4f12c0f5cf09eb6a50625c1a95ca4f

Observation cc755edc-ea2f-4e1c-8cd0-93431f93e330 · outbound

This paper cites Mac: Mining activity concepts for language-based temporal local- ization.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Mac: Mining activity concepts for language-based temporal local- ization

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.261755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.261755Z digest=sha256:c2566b80b6217676c8a82bc62d004b99b2ee1c65f3c2ce3920cc4a7621de6b2a

Observation 537d9060-b12a-4da7-8cec-e84dec96dfc6 · outbound

This paper cites Imagebind: One embedding space to bind them all.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Imagebind: One embedding space to bind them all

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.266660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.266660Z digest=sha256:eec2e8a8a15844a1a3719297392c78e23e1d9d878f2e63ead49af19546ba1af9

Observation 81d1ae0f-f85d-46d6-b357-d1cc5c09400d · outbound

This paper cites Radar: automated task planning for proactive decision support.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Radar: automated task planning for proactive decision support

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.928432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:27:18.271492Z digest=sha256:98719f7b85e9c043c85b0852fdee1cd49e15a07f2e82bbdaf39f1f1db8c1b90b

Observation 170d0ab5-d947-44b1-9229-e48d410d753d · outbound

This paper cites Retrieval augmented language model pre- training.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Retrieval augmented language model pre- training

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.911731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:27:18.276299Z digest=sha256:10f314c9e9728b9c41f0187751c43d3d7d042e7c9b83a69b6ed3f093b784e0a9

Observation 246d6b37-dad7-46cf-83e6-2d983773ddd9 · outbound

This paper cites Ma-lmm: Memory-augmented large multimodal model for long-term video understanding.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Ma-lmm: Memory-augmented large multimodal model for long-term video understanding

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.892002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:27:18.281075Z digest=sha256:9945b39f519a43319920c32145ef91214b2dc78283b5df0d77c8ffc2adabb8d3

Observation fd43daab-b8d6-491e-b28b-8a999a1cfcae · outbound

This paper cites Denoising diffu- sion probabilistic models.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Denoising diffu- sion probabilistic models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.286179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.286179Z digest=sha256:f8add948e5a35abceb8d2058003b091fdac77066a4553b849f1cda716fdd7af6

Observation 2d50f50e-e48c-4099-b735-611dcdb025ad · outbound

This paper cites LoRA: Low-rank adaptation of large language models.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models LoRA: Low-rank adaptation of large language models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.864426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:27:18.291143Z digest=sha256:64dd6d371cabb0133b5ed1022a6d17dc8cb7b99c3724285c353606dbd6fe4f95

Observation 048154c4-394f-48b6-a5e4-456912107e24 · outbound

This paper cites Audiogpt: Understanding and generating speech, music, sound, and talking head.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Audiogpt: Understanding and generating speech, music, sound, and talking head

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.846787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:27:18.296118Z digest=sha256:8b32119781f0a23d31379f9a4b05cb9009415187e58375f6482c92de9a39ce0b

Observation b5ed1157-9431-4931-a58e-1f8b0c20658a · outbound

This paper cites Language is not all you need: Aligning perception with language models.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Language is not all you need: Aligning perception with language models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.829639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:27:18.301317Z digest=sha256:9e476fd7ff36db41e474372e828ec14fced09d45d42cdb440333b0b9d095c085

Observation 300761ce-e88f-4c10-981d-959ac6b77067 · outbound

This paper cites Leveraging passage retrieval with generative models for open domain question answering.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Leveraging passage retrieval with generative models for open domain question answering

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.813004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:27:18.306835Z digest=sha256:3ef5eb7dbcdc6da2a6c3b6e0c0beb561a33dee49fdc70e912c0c0c2b15c57ece

Observation 3ef6a6de-5536-4722-b939-7b113c4ef4fa · outbound

This paper cites Mistral 7B.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Mistral 7B

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.311567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.311567Z digest=sha256:c71cdabb542de82727be3614d291bbec803699afc5b6b9cd2b4340de88a09805

Observation a0310a70-554e-4f9b-9f44-de75ffca7a9a · outbound

This paper cites Gen- erating images with multimodal language models.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Gen- erating images with multimodal language models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.317157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.317157Z digest=sha256:d997d8eb6da14bde1de470f895f5749ee5fedb624c9fbf30d6eac27ed31f0008

Observation 1aa40df6-34c4-4764-9a09-ecfe61dc25a9 · outbound

This paper cites WikiHow: A Large Scale Text Summarization Dataset.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models WikiHow: A Large Scale Text Summarization Dataset

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.322049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.322049Z digest=sha256:e0bb681aedf250a4907d893dc1cd625cdb9a832006b87ef0cd6ce3c00debf388

Observation 9bad06d2-7f42-44ae-96b8-95d227ef9d82 · outbound

This paper cites A survey on temporal sentence grounding in videos.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models A survey on temporal sentence grounding in videos

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.785086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:27:18.328096Z digest=sha256:c3aa49655767e3a17c06030780af52d35689d00330033c9411b810a9a9d296f0

Observation 4f859fb8-ec1e-47e5-bdac-e4c9aa136e71 · outbound

This paper cites Tvr: A large-scale dataset for video-subtitle moment retrieval.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Tvr: A large-scale dataset for video-subtitle moment retrieval

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.767356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:27:18.333044Z digest=sha256:2f2fc06f23feac40cdcdfca20ed430ac505903447c1207f8b8a17fa92f5a6c4e

Observation b3dabebf-3f94-4b74-a5b4-6ec0448383ea · outbound

This paper cites Less is more: Clipbert for video-and-language learning via sparse sampling.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Less is more: Clipbert for video-and-language learning via sparse sampling

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.749524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:27:18.338050Z digest=sha256:25b4c39a8979ed499658526a9d3afa6b71855495bf5c5638df16b40205e57e36

Observation 596d4a41-5f75-49d6-be14-de21ac0339ef · outbound

This paper cites Chatting makes perfect: Chat-based image retrieval.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Chatting makes perfect: Chat-based image retrieval

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.343420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.343420Z digest=sha256:e887fbee5b4751559059cb1c1d7582c32366efe9d90c011bf01cdaf6d54ccf53

Observation c7e9112a-b8ab-44e2-92bf-84a31838eade · outbound

This paper cites Retrieval- augmented generation for knowledge-intensive nlp tasks.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Retrieval- augmented generation for knowledge-intensive nlp tasks

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.721183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:27:18.348350Z digest=sha256:b8eaa0a481035babfe6c0bc3f98eab09fb2657947d9afc283b2f1845bbd165f8

Observation e8f42b6c-0d31-470d-8001-d8f90161d8e3 · outbound

This paper cites Blip- 2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Blip- 2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.353346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.353346Z digest=sha256:c9a11614cf9e699fd725885316b7c101ce8288a71e1bd47334dafba012906511

Observation 07b40416-83e2-4efc-bbd4-0a23843c2bbb · outbound

This paper cites Skip-plan: Procedure planning in instructional videos via condensed action space learning.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Skip-plan: Procedure planning in instructional videos via condensed action space learning

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.690149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:27:18.358932Z digest=sha256:4b483f2b33fa314eedaa92073fc52b8329e627165e4b541eb4314b20308f4940

Observation a7f9a262-5fb1-419d-b595-0fca782911c5 · outbound

This paper cites Cohn, and Janet B.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Cohn, and Janet B

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.665546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:27:18.363704Z digest=sha256:b4f172221824f3ec8abc7d402217f65ced2066d63657c9ab1a2b0c564e4b35f6

Observation 40b457fb-a7cd-44a8-a55b-8708313d02b8 · outbound

This paper cites Learning to recognize procedural activities with distant supervision.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Learning to recognize procedural activities with distant supervision

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.643121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:27:18.368912Z digest=sha256:e800f24f500b2284b3f085c43f4695e41e2a780425d04a18a6a79b80534a35c4

Observation 4713edc8-2e9f-4a33-acec-590ae14c0c9d · outbound

This paper cites Jointly cross-and self-modal graph attention network for query-based moment localization.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Jointly cross-and self-modal graph attention network for query-based moment localization

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.617330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:27:18.374456Z digest=sha256:eac0967eae125ed4ec0117893ccd35bcf0e0e82eff61867f453800d764981dbe

Observation d8bfad8e-0392-4876-af25-c9613c35bd01 · outbound

This paper cites Improved baselines with visual instruction tuning.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Improved baselines with visual instruction tuning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.380143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.380143Z digest=sha256:c92cda699ed35d06a1a7180b1f04b935ac991ddf93faeb09255c208d4f737c9f

Observation b8aec459-d859-4fbd-9f72-631649fb012d · outbound

This paper cites Visual instruction tuning.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Visual instruction tuning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.385014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.385014Z digest=sha256:b689787dfca18914c5d94ff11e230abbeb6eb983a9f054c130804bcbf19add55

Observation 46da08f4-c521-47d8-8139-6a60145d8163 · outbound

This paper cites Attentive moment retrieval in videos.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Attentive moment retrieval in videos

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.569112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:27:18.390493Z digest=sha256:1d3672671d73757ed1bc93577d9ad2e2f18d18ab585e37916fba42ae4f2cd1ce

Observation 26decaed-decf-42f3-85d4-e172b0f4d7a6 · outbound

This paper cites Language models of code are few-shot commonsense learners.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Language models of code are few-shot commonsense learners

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.545580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:27:18.395669Z digest=sha256:cc76b7952af36fb56312c488ebbd4b2f2f22a7d1f5f00978a2bafe83235ce4af

Observation ebe4190f-1994-4be1-b33b-a277349828bf · outbound

This paper cites What’s cookin’? interpreting cooking videos using text, speech and vision.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models What’s cookin’? interpreting cooking videos using text, speech and vision

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.517289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:27:18.400639Z digest=sha256:0a921c7fdc385bd5078943a456bb55c732fcd38029768e8b505a9b3cee41ed00

Observation 8abe8ae2-99f4-4a9e-8f58-fab46928dc7c · outbound

This paper cites Howto100m: Learning a text-video embedding by watching hundred mil- lion narrated video clips.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Howto100m: Learning a text-video embedding by watching hundred mil- lion narrated video clips

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.496609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:27:18.405575Z digest=sha256:6cb5e1867000355c2d2a4c8365440355be565625b888fa1c09316ca4722197e7

Observation 0f3350a2-a0a3-4501-ad92-c2c0d25d5f24 · outbound

This paper cites End-to-end learn- ing of visual representations from uncurated instructional videos.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models End-to-end learn- ing of visual representations from uncurated instructional videos

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.470230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:27:18.410996Z digest=sha256:f63bc4c671c9378080667a230b138c8590209983675fa112acb2b1e49b49595e

Observation ad815325-c0ba-41c6-a0dd-6595078478a7 · outbound

This paper cites Learning and Verification of Task Structure in Instructional Videos.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Learning and Verification of Task Structure in Instructional Videos

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.415992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.415992Z digest=sha256:8ff341ae5904a8330db614a017563d9758d409f3dc787726b099704da448112f

Observation ae29bc64-05ae-4610-b9f1-c961e70fb6e0 · outbound

This paper cites SCHEMA: State CHanges MAtter for procedure planning in instructional videos.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models SCHEMA: State CHanges MAtter for procedure planning in instructional videos

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.451365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:27:18.421653Z digest=sha256:89a5c7b2cac3c8d7785e30faa20b15acf32b9727f333e6d61e059207b541d649

Observation 46042b3c-b577-4cd8-a223-dec461577ac3 · outbound

This paper cites Virtualhome: Sim- ulating household activities via programs.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Virtualhome: Sim- ulating household activities via programs

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.430224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:27:18.427256Z digest=sha256:cfd49e59b42246cf73cfd2ffb4297a85b2587ba8be37ff087eafc0fb9f5f0a34

Observation 1d462551-7fc9-4673-824b-d3ef509b4255 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Learning transferable visual models from natural language supervi- sion

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.432722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.432722Z digest=sha256:9b9ba1c48db502020e44672267e0700520af1fe946a14cc1ae3f211103b91802

Observation f28cd14f-2aa5-49f4-b2f8-1edd37df038a · outbound

This paper cites Grounding action descriptions in videos.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Grounding action descriptions in videos

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.398146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:27:18.437992Z digest=sha256:7606b087bcb74201149fdd42232208924cf741d2661df752e5aafefa19f5337f

Observation 8f268d0b-267a-4bcd-b7ac-7362cd2a945d · outbound

This paper cites FLAP: Flow-adhering planning with constrained decoding in LLMs.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models FLAP: Flow-adhering planning with constrained decoding in LLMs

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.380232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:27:18.443152Z digest=sha256:b66f861a8a182166e3f5b9819f3e516e0b12c2a06d2a6ba465b0578e53812ce5

Observation 69608492-668a-46ee-a34e-8077fe24d1a5 · outbound

This paper cites proScript: Par- tially ordered scripts generation.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models proScript: Par- tially ordered scripts generation

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.363004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:27:18.448309Z digest=sha256:3c45260f817e84f8300a6cc54e41e8fc086addb7aa68c2ac2c3c2dc1997fa26f

Observation 88d099c9-b89b-4ee9-b1a2-b817958e3567 · outbound

This paper cites As- sembly101: A large-scale multi-view video dataset for un- derstanding procedural activities.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models As- sembly101: A large-scale multi-view video dataset for un- derstanding procedural activities

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.345705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:27:18.453462Z digest=sha256:889860f4f53a77f1725d0e64f00c3ffe23ac018d5180641da0c2736ad15bb5fd

Observation 374b58e8-0949-41aa-ae82-225144ebb97e · outbound

This paper cites Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.327594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:27:18.458789Z digest=sha256:2ff3f2b6dd5848b78f1aa54af874792e2062a76d6eee4576a58aa2b2a5fb5190

Observation 5c00df61-fe98-42ae-aa1e-54de161d9b76 · outbound

This paper cites Moviechat: From dense token to sparse memory for long video understanding.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Moviechat: From dense token to sparse memory for long video understanding

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.310703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:27:18.464824Z digest=sha256:b763afc5c3a70715f76dbeaa4dc0d3ab6cd82d01812a3101308f8f7fc8fbb6b2

Observation ebf6b249-9a9b-4921-b8c6-45eb5f078a1a · outbound

This paper cites Mpnet: Masked and permuted pre-training for language understanding.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Mpnet: Masked and permuted pre-training for language understanding

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.293438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:27:18.469567Z digest=sha256:91822d3524c1fda4c9992fe5254e0c7ed71231c9c473c3500e78c7762a6e0bc8

Observation c3a0b2e7-5820-41df-a9ef-2d276b799ea7 · outbound

This paper cites Language Models Can See: Plugging Visual Controls in Text Generation.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Language Models Can See: Plugging Visual Controls in Text Generation

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.474723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.474723Z digest=sha256:c6d8e91f06b268dc33a19de1700eca8de40e23658be25c39f65a489066e183b4

Observation 6019f744-9b5d-4a78-9201-ffe1265ca81f · outbound

This paper cites PandaGPT: One model to instruction-follow them all.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models PandaGPT: One model to instruction-follow them all

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.276653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:27:18.479891Z digest=sha256:c4c35349bc3c206f6c1a5f717e4507c1e57eb5cba3c12782fda2d4c3c384933c

Observation f9f1a374-9afe-4810-83be-9ba1448966c0 · outbound

This paper cites Plate: Visually-grounded planning with transformers in procedural tasks.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Plate: Visually-grounded planning with transformers in procedural tasks

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.258622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:27:18.484978Z digest=sha256:f36f68696c2a71c29b81b7842a9a03c282a5ed0aad335e6ec5cc70cb3b262402

Observation 76415b0a-33bb-4fbd-98bb-489517242b8a · outbound

This paper cites Coin: A large-scale dataset for comprehensive instructional video analysis.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Coin: A large-scale dataset for comprehensive instructional video analysis

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.239508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:27:18.489713Z digest=sha256:59e38ea10cfbcadaf1789e7ba146863f679ea017aa5593ca37b8c5a284457a50

Observation a2a0d350-3c6b-4c88-ad2a-6c639902c601 · outbound

This paper cites On the planning abilities of large language models-a critical investigation.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models On the planning abilities of large language models-a critical investigation

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.222261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:27:18.495158Z digest=sha256:ae23748204dbeb6b6a1c116e679a3b2a3b29e87f1c45375dd6e498f815c3b318

Observation ecf651f8-0f62-4594-85cc-a4041e7331b5 · outbound

This paper cites Gomez, Lukasz Kaiser, and Illia Polosukhin.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Gomez, Lukasz Kaiser, and Illia Polosukhin

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.203062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:27:18.500019Z digest=sha256:779096e11e5b9519c10053761cd64daac9295cf52d44f0ee977df4eef2ddc25c

Observation d7d46fa4-61ed-4a8f-8c20-036ce1a5e244 · outbound

This paper cites Event-guided procedure planning from in- structional videos with text supervision.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Event-guided procedure planning from in- structional videos with text supervision

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.185574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:27:18.504726Z digest=sha256:61108a2daf74f24efca52107047205eead5a2b54c321138c54f6261fab83f5d3

Observation aa1704bf-798f-4e21-9317-a7a01628461f · outbound

This paper cites Pdpp: Projected diffusion for procedure planning in instructional videos.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Pdpp: Projected diffusion for procedure planning in instructional videos

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.168578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:27:18.509417Z digest=sha256:58b2659a0e44d61660a945805d0eb67d427597c86656325fbf1f43886f5edeb7

Observation 641a44df-1b47-42d3-b018-956b5133109d · outbound

This paper cites Temporal segment networks: Towards good practices for deep action recognition.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Temporal segment networks: Towards good practices for deep action recognition

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.150365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:27:18.514360Z digest=sha256:d15f83077b16b5ae49284c4768a5ef26e4ebc4ea2ecdf85d1a4abdbe7ecccb3d

Observation 5155ce14-7170-4533-a1bc-32587b290c6d · outbound

This paper cites Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.519058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.519058Z digest=sha256:6823fa83c1e1b82e3ab8f908c07e8b3de94040587fc5532c31adaef622637e65

Observation 9f82821f-3eaa-4d58-9597-5dec1e9624c7 · outbound

This paper cites NExt-GPT: Any-to-any multimodal LLM.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models NExt-GPT: Any-to-any multimodal LLM

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.133657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:27:18.524104Z digest=sha256:26f80a14e6f095d5e74892fd6292cec81ae4d3babca861ce6107f1418eaab84b

Observation bfe34a2f-97a6-47ab-8583-2618443845c0 · outbound

This paper cites Rethinking spatiotemporal feature learning: Speed-accuracy trade-offs in video classification.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Rethinking spatiotemporal feature learning: Speed-accuracy trade-offs in video classification

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.116500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:27:18.528664Z digest=sha256:64516b86cd173a2284ada16702fce724f0500a98d84760baa803e2186752d613

Observation b4329a38-a98a-4794-bac9-b97886dcf609 · outbound

This paper cites Translating Natural Language to Planning Goals with Large-Language Models.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Translating Natural Language to Planning Goals with Large-Language Models

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.533727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.533727Z digest=sha256:ddf76df60519ad097039deed195954ceb7c79df673dec3db60f1851da83c7ba3

Observation d198cda2-5116-439d-ac92-f0bafaac0ec3 · outbound

This paper cites Multilevel language and vision integration for text-to-clip retrieval.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Multilevel language and vision integration for text-to-clip retrieval

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.097622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:27:18.539015Z digest=sha256:11a2120881f716da59d06145854c4567614014f991f776be93416275ad6a160f

Observation bb148e8c-adcf-4f31-b9d8-0bd1fa4bb170 · outbound

This paper cites VideoCLIP: Contrastive pre- training for zero-shot video-text understanding.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models VideoCLIP: Contrastive pre- training for zero-shot video-text understanding

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.079625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:27:18.544244Z digest=sha256:048bd22919c1aeabfc763e9c4fa64f7b21be4c5dde837e630643dd4d4d1ce91f

Observation 6cc9b393-3050-4d91-a6f8-f455fe2de311 · outbound

This paper cites Retrieval- augmented generation with knowledge graphs for customer service question answering.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Retrieval- augmented generation with knowledge graphs for customer service question answering

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.060275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:27:18.549134Z digest=sha256:3d67461d62eb165f28a16dc0ff85f68693154ce309b2ad4072375ad7a57312dc

Observation f320f8fd-dd08-47a9-9add-799ba20f929d · outbound

This paper cites Activitynet-qa: A dataset for understanding complex web videos via question answering.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Activitynet-qa: A dataset for understanding complex web videos via question answering

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.043613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:27:18.554112Z digest=sha256:edec1172bcd29347d613af6520c1bfe68b96a68649569ed6b5246ec11f3edc52

Observation 4a9e3b04-e34b-4f3f-bd07-db480e1e6fe9 · outbound

This paper cites Distilling script knowledge from large language models for constrained language planning.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Distilling script knowledge from large language models for constrained language planning

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.025535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:27:18.559096Z digest=sha256:eec0a7bdd2750d9541b638ca49a564ac3d9d9dd837e238156468254b5315d58d

Observation 1e5d354d-f8e9-4e08-88f6-567a49c7bad2 · outbound

This paper cites Man: Moment alignment network for natural language moment retrieval via iterative graph adjustment.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Man: Moment alignment network for natural language moment retrieval via iterative graph adjustment

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.007476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:27:18.563941Z digest=sha256:4f0a7af43c59c19ae8efbdaa8b82692cce8eda9111e1f6898a4060bbd0049ab6

Observation e47a14f2-1409-470a-b597-6372db2921c9 · outbound

This paper cites SpeechGPT: Empowering large language models with intrinsic cross-modal conversa- tional abilities.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models SpeechGPT: Empowering large language models with intrinsic cross-modal conversa- tional abilities

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:18.990834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:27:18.568729Z digest=sha256:dc0cdec6bc99d955c1e2e80eb2bf6c879f055232f0bf1192a408588f32e977b4

Observation 65d682a9-ad24-4ef0-9aff-985af49ebc05 · outbound

This paper cites Video-LLaMA: An instruction-tuned audio-visual language model for video un- derstanding.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Video-LLaMA: An instruction-tuned audio-visual language model for video un- derstanding

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:18.974322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:27:18.574278Z digest=sha256:12a9f3b407daa59c9b3cb1ad4b1fb7a3629eaabb6ece32f7e5a3fd8aa7e75be3

Observation d628f934-164d-4b8b-a513-33d7a7b43b42 · outbound

This paper cites Temporal sentence grounding in videos: A survey and fu- ture directions.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Temporal sentence grounding in videos: A survey and fu- ture directions

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:18.957496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:27:18.579525Z digest=sha256:deae0759b53c84d2c748e7921e6a5f6c6823304e54c833b5e833de368fa0e3d3

Observation 703bc398-48a0-48c7-9795-3cb775a7a0ae · outbound

This paper cites P3iv: Probabilistic procedure planning from instructional videos with weak su- pervision.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models P3iv: Probabilistic procedure planning from instructional videos with weak su- pervision

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:18.939426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:27:18.584443Z digest=sha256:ca70bc4ee73e5495d37477cd761274fafc08052d54bd93d500d80dd7efa002b5

Observation 1bfea070-78ba-47be-a292-a0398e67453e · outbound

This paper cites Learning procedure-aware video repre- sentation from instructional videos and their narrations.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Learning procedure-aware video repre- sentation from instructional videos and their narrations

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:18.921158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:27:18.590544Z digest=sha256:b9768e5c32b111b33996a87de35b25f5aa7490b7aa2c0490f6667712a3a52e11

Observation cd6107ab-a413-41f7-ad15-d3bcb5ac5f94 · outbound

This paper cites Procedure-aware pretraining for instructional video understanding.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Procedure-aware pretraining for instructional video understanding

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:18.904145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:27:18.596552Z digest=sha256:87a9192eeb299d8f23bb217212fcac62f6b298a4b54846f0e1bd15fa5f22c512

Observation 7901844d-b6bf-424c-b788-b3bd49dd1263 · outbound

This paper cites Towards auto- matic learning of procedures from web instructional videos.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Towards auto- matic learning of procedures from web instructional videos

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:18.887341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:27:18.602763Z digest=sha256:7c6951855e27263be6d4de4011c7803f32dfe87ad2b10b52f573b69591559e91

Observation d130331b-ff3b-4241-903e-a583413649ee · outbound

This paper cites MiniGPT-4: Enhancing vision-language understanding with advanced large language models.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models MiniGPT-4: Enhancing vision-language understanding with advanced large language models

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:18.869455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:27:18.608913Z digest=sha256:66feb9c407ed88a2ca2d1f85f84600f7ea690fee50ff4c7d97322bd42acf16ff

Observation f8de0d6b-998d-4e0c-bf16-2b794b0714b4 · outbound

This paper cites Cross- task weakly supervised learning from instructional videos.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Cross- task weakly supervised learning from instructional videos

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:18.851520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:27:18.614377Z digest=sha256:e449a57a34a29a04210722a1e3836b5ca796f9cad3af859a2c68f5fe332c8996

Pith citing papers

Observation 1a817673-2a56-4d5d-916d-4a9b06e41a2b · inbound

SFHand: Learning Embodied Manipulation by Streaming Egocentric 3D Hand Forecasting cites this paper.

SFHand: Learning Embodied Manipulation by Streaming Egocentric 3D Hand Forecasting InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-21T17:54:18.278334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-21T17:53:19.002603Z digest=sha256:45b459e1080d0cd901ad30a9bc596798741049ab3f671aab1ae24260bf51d784

Observation 94abaaac-ea99-486a-89b4-b240171202ed · inbound

VisionClaw: Always-On AI Agents through Smart Glasses cites this paper.

VisionClaw: Always-On AI Agents through Smart Glasses InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-13T18:13:05.498964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-13T18:12:34.183092Z digest=sha256:d75075bdf05f6542a0ef01730fa14356f05522fbc8e53997441fa8b90c7040c1

Observation 3b18f4b1-b9d9-41b9-9fb7-55667d3f9008 · inbound

Watch, Remember, Reason: Human-View Video Understanding with MLLMs cites this paper.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models

Reference 256

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:15.163371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:766d02f0116ce6dc69f7ed813d35a82afbcc4061ae8cc567aefcace0f1748fd6