Pith. sign in

Paper Citation Record · LEDGER

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos

As of 15 August 2026, this Paper Citation Record lists 71 of 71 outbound references and 1 inbound Pith citation observation for arXiv:2411.14901.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.14901 v1

Coverage vector

measured 71 of 71 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T14:51:17.742891Z

measured 72 of 72 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T23:32:17.962997Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T23:32:18.789296Z

Reference resolution

71 of 71 outbound references displayed

  • verified exact1
  • verified fuzzy37
  • unresolved33
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f6de1641-e8d3-49f1-9180-66134e2807f3 · outbound

This paper cites needle in a haystack.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos needle in a haystack

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.633581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:51:17.480428Z digest=sha256:96adf53d858951d8b105d24e71c011d764cbb723647b886a7be8d817b5fd2f08

Observation 14f8428d-d7b1-4c95-9845-40a18e4758a1 · outbound

This paper cites needle in a haystack.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos needle in a haystack

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.485527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.485527Z digest=sha256:d945f07abd427fdd817bc97362809a40b7bddf355d15f10f3498a0972932e2b9

Observation 932cbd73-d287-489b-86e0-771313651af9 · outbound

This paper cites https://sharegpt.com/.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos https://sharegpt.com/

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.621707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:51:17.489695Z digest=sha256:cf98b23b9a47b8cb1a8a64436b2bed6a19133b161b10553849126b1026f2d845

Observation 2494fcf0-7316-44fd-8728-03845a06f407 · outbound

This paper cites https://github.com/gkamradt/LLMTest_ NeedleInAHaystack.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos https://github.com/gkamradt/LLMTest_ NeedleInAHaystack

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.610355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:51:17.493940Z digest=sha256:41f2d3e71c28a80991aa8164add1995d9a9d0304e50e0c2527e121f0d646ad70

Observation 514194ae-6741-4e06-b166-858bb214e022 · outbound

This paper cites Lo- calizing moments in long video via multimodal guidance.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Lo- calizing moments in long video via multimodal guidance

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.598205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:51:17.497775Z digest=sha256:d960d83137872c171c3395ca2204b12f241a4851b0c56ae9dfe3f7ada9b5d791

Observation 70c6f739-fbc3-4dfb-9a5b-18bf90ef7e20 · outbound

This paper cites Functional brain organization of preparatory attentional control in visual search.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Functional brain organization of preparatory attentional control in visual search

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.586723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:51:17.501746Z digest=sha256:bf7a2681a4069c202e6cef6c5e6440d508ec22752ca3c4407ebe788952f2341c

Observation c3a58dae-6d8d-4df8-9152-0553858fa38f · outbound

This paper cites End-to- end object detection with transformers.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos End-to- end object detection with transformers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.505724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.505724Z digest=sha256:a3af36a7db12c13a00dd30f98f5c9ad919261d56346235e88a382fc67996439d

Observation f6a37a7e-6852-4567-ba07-10d991286069 · outbound

This paper cites VideoLLM: Modeling Video Sequence with Large Language Models.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos VideoLLM: Modeling Video Sequence with Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.509494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.509494Z digest=sha256:eba85e218e829c274e6c6c6033f64d0ee73929887eac1b695f5b87206334d113

Observation b8ee4b52-d852-41af-b32a-ee00329f0aad · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Gonzalez, Ion Stoica, and Eric P

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.567451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:51:17.513431Z digest=sha256:8aa2717bc342436b10c94e3a167f16dc1eba0a8761f5d49e26452443483e0f49

Observation 8425f23a-9f25-42db-b812-904a74ca0259 · outbound

This paper cites Bert: Pre-training of deep bidirectional trans- formers for language understanding, 2019.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Bert: Pre-training of deep bidirectional trans- formers for language understanding, 2019

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.556051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:51:17.516640Z digest=sha256:9fb2b740792d3e260bc2b316d8533125cae113be88881b201462adad12ee1f47

Observation 7bfc6ce1-d9db-4510-b89a-3b9e40b894c4 · outbound

This paper cites Uatvr: Uncertainty-adaptive text-video retrieval,.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Uatvr: Uncertainty-adaptive text-video retrieval,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.543891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:51:17.519848Z digest=sha256:66b46d71b61bdb8a32321f07f0cf405fba438defca69156f59ec77c15d67cc97

Observation b54cee9a-901a-4e69-b862-fad754abe3e1 · outbound

This paper cites Multi-modal transformer for video retrieval.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Multi-modal transformer for video retrieval

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.531878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:51:17.523083Z digest=sha256:a3d83d947b7bf124a9f53374637939a8143ffb594f497bb6ac407c21ce18c14d

Observation d59df371-f9dc-4faf-9877-95bd9c219341 · outbound

This paper cites AssistGPT: A General Multi-modal Assistant that can Plan, Execute, Inspect, and Learn.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos AssistGPT: A General Multi-modal Assistant that can Plan, Execute, Inspect, and Learn

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.526246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.526246Z digest=sha256:899efae49eff9a20965695366f80fdfbfa27952f16d4c1327a23d9d47b1e5a09

Observation c23fb26c-f5c4-464b-8e35-9534ba353132 · outbound

This paper cites X-pool: Cross-modal language-video attention for text- video retrieval, 2022.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos X-pool: Cross-modal language-video attention for text- video retrieval, 2022

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.520335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:51:17.529565Z digest=sha256:67c7af30b7de52176afe73bfdb017c76ee5434310b0c629a49fb85eda0efe726

Observation bc23c703-d543-4919-8f23-5065e7fcfed0 · outbound

This paper cites RGNet: A Unified Clip Retrieval and Grounding Network for Long Videos.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos RGNet: A Unified Clip Retrieval and Grounding Network for Long Videos

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-12T14:51:18.046182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:51:17.532531Z digest=sha256:849dd37a9bc3a5b2578c8ee372182d35915f53459603424e674d7c8564950bc7

Observation 255d90f0-9d27-4813-ad21-2f1b032353c9 · outbound

This paper cites CONE: An Efficient COarse-to-fiNE Alignment Framework for Long Video Temporal Grounding.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos CONE: An Efficient COarse-to-fiNE Alignment Framework for Long Video Temporal Grounding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.536220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.536220Z digest=sha256:557dd2398bc114d8ae0f7de46c5b2859f18489c2a0fefbed434711d94e0ad58f

Observation 4395324c-3dd4-4766-adfa-b8ea1ccf4037 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos LoRA: Low-Rank Adaptation of Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.539848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.539848Z digest=sha256:2f6066efed4578740668b2caa9d16e4a3a131242924d595310776f58dc9b7988

Observation 857ffaaa-1626-4923-8342-37e56075e9d4 · outbound

This paper cites Vtimellm: Empower llm to grasp video moments.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Vtimellm: Empower llm to grasp video moments

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.509395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:51:17.543649Z digest=sha256:34004cb0aa334596a042e910d8140ff0f70173d9167bc2c892b7a6cf041d195a

Observation 5760acec-746b-45f6-bfbc-d5a67ef8412f · outbound

This paper cites LITA: Language Instructed Temporal-Localization Assistant.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos LITA: Language Instructed Temporal-Localization Assistant

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.547199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.547199Z digest=sha256:45bddd52f967a37fb5b0f71471b4f37728e121be0d20fd7637d339562580b00c

Observation 6e092eaf-b30b-4139-ac67-d61fcb73725d · outbound

This paper cites Audio- enhanced text-to-video retrieval using text-conditioned fea- ture alignment, 2023.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Audio- enhanced text-to-video retrieval using text-conditioned fea- ture alignment, 2023

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.498330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:51:17.551108Z digest=sha256:651fb48e0496d9c3f52d2f74bdec6522f64c26aeaeefeaecdaa9b08ec523925b

Observation 5fdfe32c-e0ae-4d09-97d9-3009ec7483f9 · outbound

This paper cites Efficient long- text understanding with short-text models.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Efficient long- text understanding with short-text models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.487239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:51:17.554425Z digest=sha256:35870c806c44dddc6fab180a461362cf43f5951fdeb00f4f7b6ff465450d4fa6

Observation 859ddc87-c4f6-4c55-a492-ae3650642b4b · outbound

This paper cites Diffusionret: Generative text-video retrieval with diffusion model, 2023.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Diffusionret: Generative text-video retrieval with diffusion model, 2023

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.476521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:51:17.557860Z digest=sha256:af81fd617a5ab6db21f54b1bf4e109fb6ac18bb96f7d11d078bbb328d2bdc042

Observation 7e84c5c3-f984-47f6-8364-31ec71538ba3 · outbound

This paper cites Chat-univi: Unified visual representation em- powers large language models with image and video un- derstanding.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Chat-univi: Unified visual representation em- powers large language models with image and video un- derstanding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.561282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.561282Z digest=sha256:0116015ecf70af756b01b70c59f46f6b22a2cc6746665532d2dd2a1aad6fe0bb

Observation f31af705-9116-4cde-8097-d2707e6b1939 · outbound

This paper cites Large Language Models Must Be Taught to Know What They Don't Know.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Large Language Models Must Be Taught to Know What They Don't Know

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.565722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.565722Z digest=sha256:2985d57d64df1b249e3b1dd0ab6b724a581492fa836d372e941aba2f03cac8c1

Observation 4f91482b-c689-4763-a06a-5d7a9054805b · outbound

This paper cites Uncertainty-Aware Evaluation for Vision-Language Models.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Uncertainty-Aware Evaluation for Vision-Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.569740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.569740Z digest=sha256:11b63bca13ec0afa1eee8131845cf8b78b42b229e1332e1a26e392814628ae6f

Observation 1793bfdc-bdeb-44ae-980b-102e7a7189f6 · outbound

This paper cites Detecting mo- ments and highlights in videos via natural language queries.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Detecting mo- ments and highlights in videos via natural language queries

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.459089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:51:17.574528Z digest=sha256:20a207cd09c48a93ac4ef750b6ae6f2d63284ce0eee8b3430355dc6e87442f59

Observation 46ec3526-134e-46ce-bf8d-cdbd49dce8be · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.578279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.578279Z digest=sha256:91958b38f9e641b5d498cbf43b0c1a620544a41f3e2ff33e308ffa2e88ca2ada

Observation 76d82182-d59a-49bd-8510-a33aeb440e24 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos VideoChat: Chat-Centric Video Understanding

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.582466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.582466Z digest=sha256:511c84ac224464ecefea8248f9fba7c4e1f39d0e6e921ccfd7e80819f2021175

Observation 2eb60991-0dce-4779-9f93-9019603535f1 · outbound

This paper cites Ground- inggpt: Language enhanced multi-modal grounding model.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Ground- inggpt: Language enhanced multi-modal grounding model

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.441298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:51:17.586499Z digest=sha256:bd0ff6c41df8f32aff7db4ac5dd47b50b8319b036a7a679947d4ecba992c56d2

Observation cd4f3baa-467a-4ef6-bce8-b42ab4540ad7 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.590376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.590376Z digest=sha256:69c4bd69503d373a65f7890a40cff2c4e73e10bd596636a3b2aa087414412e9d

Observation 05cbee80-bb5a-4306-b1f2-fb7b8f0f0578 · outbound

This paper cites Univtg: Towards unified video- language temporal grounding.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Univtg: Towards unified video- language temporal grounding

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.430557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:51:17.594249Z digest=sha256:a4ffe6186d312065e8a5d70385ad2566d6834fe1b3b4fa1e8c917a639618a448

Observation 39bb5fd4-231b-40aa-a575-3eb309801b08 · outbound

This paper cites Visual Instruction Tuning.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Visual Instruction Tuning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.597917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.597917Z digest=sha256:74b9c40f9697bbd28a956b70c5b6a168c9e4c1beaccfa45646aae9a0ce0a8f5c

Observation 34c89be5-11c2-4918-a3c3-3aba6a6ddf88 · outbound

This paper cites ReLER@ZJU-Alibaba Submission to the Ego4D Natural Language Queries Challenge 2022.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos ReLER@ZJU-Alibaba Submission to the Ego4D Natural Language Queries Challenge 2022

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.602209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.602209Z digest=sha256:979fb431e8deb4141fc640da80aabf7c0c5fb22fc545f263d2f3a55877fb1871

Observation 96bdaf0f-68b8-4ce2-b281-74fd7486bc94 · outbound

This paper cites Lost in the Middle: How Language Models Use Long Contexts.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Lost in the Middle: How Language Models Use Long Contexts

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.605648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.605648Z digest=sha256:cc45c7eaf10c2b0bc7763f6d5f47b55576dfd68c8ea9a3ce0fa127ef99b58043

Observation 353f11b3-9fc8-4847-b641-021352ca4cbd · outbound

This paper cites Umt: Unified multi-modal transformers for joint video moment retrieval and highlight detection.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Umt: Unified multi-modal transformers for joint video moment retrieval and highlight detection

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.420410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:51:17.609270Z digest=sha256:572d75b5a1b769a80462c082804ba82d81521fd78f30f3c9e7d311defdad320c

Observation ab253223-2440-4d2e-9dcf-5f41221f69d6 · outbound

This paper cites Decoupled weight decay regularization, 2019.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Decoupled weight decay regularization, 2019

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.409303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:51:17.612626Z digest=sha256:f188d69b1a3b91731e47339309b3d88677118cae179aed93c2ee37e4839ca709

Observation 37251822-9afd-464b-9f2f-b3981631e3be · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.615743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.615743Z digest=sha256:29488089f01fc1af44a2d59ea973b48e36f72f791ceca6a59658be7aba24d2e9

Observation 19ff04d7-30bc-4f41-b283-d45d25fb7c41 · outbound

This paper cites Query-dependent video representa- tion for moment retrieval and highlight detection.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Query-dependent video representa- tion for moment retrieval and highlight detection

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.397946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:51:17.618988Z digest=sha256:c381398c1590aeba4dfcecf26ec1960bc432df9771f97278160208983362c03b

Observation ebb83402-3994-44ec-9a7b-18b8f6713bf9 · outbound

This paper cites Snag: Scalable and accurate video grounding.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Snag: Scalable and accurate video grounding

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.386923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:51:17.622614Z digest=sha256:cd940123aef22cd59da161e8618199ac0dc690f603d1cd7178b9a0fbf117c1d3

Observation 68f00f37-6ea9-412c-aede-02cffb35ab00 · outbound

This paper cites Towards Calibrated Robust Fine-Tuning of Vision-Language Models.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Towards Calibrated Robust Fine-Tuning of Vision-Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.626259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.626259Z digest=sha256:c2e77f983309b9bc7e200582ef824d0f5b6009fba47dc8775a1c4ccd4a36b30a

Observation 2d85814d-1246-4674-b31f-fe6f5d6d28f6 · outbound

This paper cites Obtaining well calibrated probabilities using bayesian binning.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Obtaining well calibrated probabilities using bayesian binning

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.375813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:51:17.630348Z digest=sha256:eda0b23c5cd43a2c0351a6cb381ce50c85664f88ca52de374e8f7e0fd4272575

Observation 8c9c6ddd-4e66-4dc0-bd29-dfb737e37993 · outbound

This paper cites Scanning Only Once: An End-to-end Framework for Fast Temporal Grounding in Long Videos.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Scanning Only Once: An End-to-end Framework for Fast Temporal Grounding in Long Videos

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.634152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.634152Z digest=sha256:cef9680e7b4d41632d9fdd8f0024d270fab3526c9d8b709cfc4aaa2b3501a66f

Observation f894d7f6-1da3-4fff-8285-c0f3687dff1e · outbound

This paper cites Momen- tor: Advancing video large language model with fine-grained temporal reasoning, 2024.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Momen- tor: Advancing video large language model with fine-grained temporal reasoning, 2024

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.363249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:51:17.638044Z digest=sha256:236180cddc0bca178c378f5ba0a52a086aac5acb892f463369c44e73ce0e13ef

Observation 385ffb20-aa04-4efe-9481-fc8712c9a1ea · outbound

This paper cites Chatvtg: Video temporal grounding via chat with video dialogue large language models.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Chatvtg: Video temporal grounding via chat with video dialogue large language models

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.351289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:51:17.641966Z digest=sha256:b46788e67c0c05ddc6391c7d1d1b355577e4763fc27c5b5d9d61b1f999296093

Observation 369d7e0f-0aba-4b96-8864-7d0504029881 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Learn- ing transferable visual models from natural language super- vision

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.339416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:51:17.645735Z digest=sha256:24f660b72b978f5789c5471d42b14c55d8ca9b03f0a76953ffdf8951cbbef281

Observation bd3e9ee6-103f-49aa-92b8-11e2a1e79dad · outbound

This paper cites Timechat: A time-sensitive multimodal large lan- guage model for long video understanding.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Timechat: A time-sensitive multimodal large lan- guage model for long video understanding

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.327828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:51:17.649213Z digest=sha256:3bf30188cd1fa4ef98ada5b3f1439823067accc4a8ddaa784ac0fa3cabe11344

Observation fa1786fc-b54f-462c-9a64-d3e8d146a890 · outbound

This paper cites Vlg-net: Video-language graph matching network for video grounding.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Vlg-net: Video-language graph matching network for video grounding

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.315273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:51:17.652731Z digest=sha256:40648b966b111af2e9afff39531d1727900e677ea33fe0d63d4d95573da4ca32

Observation 3cb809b6-898d-4661-918b-d83729b0dc55 · outbound

This paper cites Mad: A scalable dataset for language grounding in videos from movie audio descriptions.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Mad: A scalable dataset for language grounding in videos from movie audio descriptions

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.656723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.656723Z digest=sha256:6ae92cc6579f1a1447821f3c1377f369a69b91384503e9e50c9bc9d0781675c9

Observation 961da660-c1f6-4d3b-9626-7d6cb1786d99 · outbound

This paper cites Moviechat: From dense token to sparse memory for long video understanding.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Moviechat: From dense token to sparse memory for long video understanding

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.660574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.660574Z digest=sha256:5d5ab7299384687e1316b50c7ebeb4fef74167a13f24676645c28d63275a20f2

Observation e7322f6a-c423-4ba0-b032-eb1cf4246819 · outbound

This paper cites Avicuna: Audio-visual llm with interleaver and context- boundary alignment for temporal referential dialogue.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Avicuna: Audio-visual llm with interleaver and context- boundary alignment for temporal referential dialogue

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.664377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.664377Z digest=sha256:0e46e82675253e41ce5a1d964ec3deb162a00f8f82f0fa890dafeb5e447e68e3

Observation ec48ea8d-2526-42e0-8254-dd4948a38cd6 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos LLaMA: Open and Efficient Foundation Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.667958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.667958Z digest=sha256:b8abd6c3a447ec79630a35dd5298eb380de989e28cce076b00287cd05cfd0a34

Observation 65b3fd40-def7-460e-bfd8-1db356291c92 · outbound

This paper cites ChatVideo: A Tracklet-centric Multimodal and Versatile Video Understanding System.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos ChatVideo: A Tracklet-centric Multimodal and Versatile Video Understanding System

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.672191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.672191Z digest=sha256:2b5d704b941c381b776e1ed29a1d531204a535db2634e73d037eca1028a8f73d

Observation a4943ab4-0d48-4b0a-9490-36a41f56a9a5 · outbound

This paper cites Omnivid: A generative framework for universal video understanding.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Omnivid: A generative framework for universal video understanding

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.675817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.675817Z digest=sha256:000c0cb364d9a490e803a8be617a26da2a984047c6602a5ac1a376c5d6885925

Observation f34862bb-cd24-414a-b9a5-ddde6d1215ee · outbound

This paper cites Text is mass: Modeling as stochastic embedding for text-video retrieval, 2024.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Text is mass: Modeling as stochastic embedding for text-video retrieval, 2024

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.283505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:51:17.679428Z digest=sha256:7820105e3b61cf3ac938db9821fc62fd5327b1f55320fcb8cbe6a569200a2d8d

Observation 09b2e3e7-88b7-464a-aabb-3bfc16f0a067 · outbound

This paper cites HawkEye: Training Video-Text LLMs for Grounding Text in Videos.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.683167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.683167Z digest=sha256:23336be109d9a2b04d199ee67cfe1bf656b4bb382625a6485c312623167a8bb3

Observation b96e8328-abb2-4c62-89d9-c9ad92851365 · outbound

This paper cites Five factors that guide attention in visual search.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Five factors that guide attention in visual search

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.270664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:51:17.687158Z digest=sha256:4d6541b0a55b390726535a074071ea93d6523efa49b031e8374607ca948c5bdc

Observation 33b0a0e2-eeea-45ab-ba58-bc38164d8fb1 · outbound

This paper cites Msr-vtt: A large video description dataset for bridging video and language.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Msr-vtt: A large video description dataset for bridging video and language

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.258309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:51:17.690819Z digest=sha256:e2324a9a598b9fd527056a2eb3044e2d5925eefab5994d3837daca4c12feab12

Observation 1a570ca5-d499-4696-9f06-b73cd1f2e861 · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.695172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.695172Z digest=sha256:1e8848ccfec68afa1467f47511f8515a0043c26589b53f884ca72789f950ba01

Observation 3d11ef86-3c09-4298-adee-355fb136229e · outbound

This paper cites SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.699205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.699205Z digest=sha256:24b96c802b2f90124e13b3c809c242c466a65ac9709f040cdae682b20c402399

Observation 9e40798d-db98-4886-944f-82aedba8fb24 · outbound

This paper cites Clip-vip: Adapting pre- trained image-text model to video-language representation alignment.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Clip-vip: Adapting pre- trained image-text model to video-language representation alignment

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.247005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:51:17.703027Z digest=sha256:8798a7171297255f8e883f2da252e9bb6b1fd87febf00423c6733e28fa54c0bd

Observation 6290a9e5-877f-4046-b121-d32a102f6e91 · outbound

This paper cites Clip-vip: Adapting pre- trained image-text model to video-language representation alignment, 2023.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Clip-vip: Adapting pre- trained image-text model to video-language representation alignment, 2023

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.235064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:51:17.706082Z digest=sha256:af2079bcd7b838f8b0e850e3b8c9e2fd38dbed21bf14f92e53b0714b7468e46c

Observation 39f0e8a5-21c9-470e-b485-db4f92b1617e · outbound

This paper cites Vidchapters-7m: Video chapters at scale,.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Vidchapters-7m: Video chapters at scale,

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.223591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:51:17.709255Z digest=sha256:02850efdc0326e116b7b8341732d402c3bcdb382b8058a26e0315b33b28eb555

Observation c9001ed8-d3d1-4cf2-89d8-19583cb0605e · outbound

This paper cites Vid2seq: Large-scale pretraining of a vi- sual language model for dense video captioning.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Vid2seq: Large-scale pretraining of a vi- sual language model for dense video captioning

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.212240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:51:17.712563Z digest=sha256:087afdc1a849b289e0b4964d47a36f30fa12acaebba27dadf690f09f03c2d901

Observation fe54dfd1-c547-4831-b9f9-96fa5c24a3a1 · outbound

This paper cites Self-chained image-language model for video localization and question answering.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Self-chained image-language model for video localization and question answering

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.715867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.715867Z digest=sha256:f0dc80f88c6c7dc387a562b200f90056e8beadcb9a1d1a3374e8623ca4b726f5

Observation 6dc5565c-50d5-4e59-9bab-48ea7cf2f19e · outbound

This paper cites A joint se- quence fusion model for video question answering and re- trieval.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos A joint se- quence fusion model for video question answering and re- trieval

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.718994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.718994Z digest=sha256:a84ce1d9f9bf611d7597e01a3ea7cfa06ba83fa250f4fd90a7d8af72048c7e1a

Observation 4db7b359-bc0c-425b-8507-1cbba8219f49 · outbound

This paper cites A sim- ple llm framework for long-range video question-answering,.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos A sim- ple llm framework for long-range video question-answering,

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.722170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.722170Z digest=sha256:317e35f681fd14993e44933bf946e5ac14a4d8a1b94878c08295f02188143077

Observation 11d10cb2-8c7c-4929-a6b0-fef46608327a · outbound

This paper cites Span-based Localizing Network for Natural Language Video Localization.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Span-based Localizing Network for Natural Language Video Localization

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.726089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.726089Z digest=sha256:decd796bdb47c784dd9e86ccdef4c7b24e52d68e81c0560f28ee5a42c5853b09

Observation b79c7ce4-cf4f-45e0-be7b-fba95f2d2af5 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.730863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.730863Z digest=sha256:87901ac2a483fffb80f8d2b4c9908769b1761e00ec951ba4fab8a74387a5b466

Observation ed8e90c1-123a-42a6-8552-8f642f137f6f · outbound

This paper cites Learning 2d temporal adjacent networks for moment local- ization with natural language.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Learning 2d temporal adjacent networks for moment local- ization with natural language

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.179532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:51:17.734969Z digest=sha256:da1c6fa176b9e83aedcc3d5430c47aab6acd1b1ebd0ef4ae1e6671940ab6cc37

Observation 4458d16d-b41f-4419-8d00-ddaebcc8bdbc · outbound

This paper cites Learning video representations from large lan- guage models.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Learning video representations from large lan- guage models

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.167474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:51:17.739263Z digest=sha256:6e5a02d2ad80f57a95eb0ddc09860659006149f22f47cd958a71f01b0fed0d66

Observation 6de24dff-be19-4a52-86a4-6078f8679cf5 · outbound

This paper cites <video> Does the <event> happen in the video? Answer yes or no.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos <video> Does the <event> happen in the video? Answer yes or no

Reference 4096

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.154762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:51:17.742891Z digest=sha256:7e01b1305924ec709631c95b50dd2b23fb0fc60277889eb027d9bcb9796f1efc

Pith citing papers

Observation fbe52eca-4221-4173-a058-749e81744064 · inbound

A Survey on Video Temporal Grounding with Multimodal Large Language Model cites this paper.

A Survey on Video Temporal Grounding with Multimodal Large Language Model ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos

Reference 111

Resolution
verified exact
local_arxiv, observed 2026-08-05T23:32:18.794543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-05T23:32:17.962997Z digest=sha256:0bf5ba80a1b771f6d6d49d3bdefd8c3b2e76aad2ea144ca3faa9b7f0e91bbc9d