Pith. sign in

Paper Citation Record · LEDGER

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval

As of 15 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 7 inbound Pith citation observations for arXiv:2411.14505.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.14505 v1

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T15:46:34.180369Z

measured 62 of 62 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:32:52.859271Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T17:27:14.988080Z

Reference resolution

55 of 55 outbound references displayed

  • verified exact1
  • verified fuzzy40
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 29aea9b4-9f57-4e3d-9a78-1ba0dd04daee · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T15:46:33.987486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:46:33.987486Z digest=sha256:94dafabfb8d29b37edd4b71351d2f864a17599eb9d8ebe4043ef00f30e782900

Observation 5e137d13-e512-4518-8e9e-ad2bc1e43c5e · outbound

This paper cites Localizing mo- ments in video with natural language.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Localizing mo- ments in video with natural language

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:46:34.826376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:46:33.992151Z digest=sha256:06086dfa0b469f2a46e5c2a013727754a14591f0cf726581f5763db28e804a9c

Observation 74001564-cdce-4479-9e12-72b00375fe3b · outbound

This paper cites Openflamingo: An open-source frame- work for training large autoregressive vision-language mod- els, 2023.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Openflamingo: An open-source frame- work for training large autoregressive vision-language mod- els, 2023

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:46:34.815024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:46:33.995449Z digest=sha256:2324f1dcfbe35779bd4720c719f9fb729684df36fad3e5bfd37b17400fc33f67

Observation df9a1c82-2d1e-4c25-bfc5-f7fdc5c9745f · outbound

This paper cites CTRN: Class-Temporal Relational Network for Action Detection.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval CTRN: Class-Temporal Relational Network for Action Detection

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-12T15:46:34.294023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:46:33.999034Z digest=sha256:86ae2dacc8bc2acfcc49e944fde414bf90f741246cb7eb3deb24fa57b2c9b8ac

Observation 49dd12b1-cddf-4b10-8a3f-e4b8a4f74c17 · outbound

This paper cites Ms-tct: Multi-scale temporal con- vtransformer for action detection.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Ms-tct: Multi-scale temporal con- vtransformer for action detection

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:46:34.803597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:46:34.003000Z digest=sha256:a8692f55d312c4e58a3c20f2b97c2b8faba7cd18fb4e6402683c72a51a7c16df

Observation 52de721a-5de5-496d-8f39-41286100187e · outbound

This paper cites Instructblip: Towards general- purpose vision-language models with instruction tuning,.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Instructblip: Towards general- purpose vision-language models with instruction tuning,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T15:46:34.006592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:46:34.006592Z digest=sha256:1299f764c628f2b4ba2a3341bd37beeda4b43c9b570f284c16ad388f3a8f222d

Observation 30268258-4e5a-45fa-8a3a-6bb5ef239cde · outbound

This paper cites Tall: Temporal activity localization via language query.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Tall: Temporal activity localization via language query

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:46:34.784282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:46:34.010324Z digest=sha256:9f3133d194935e0427d4dcd43c7b59ef97985a6f8eb306e134d20a1afff44b37

Observation 0b4df684-294d-4393-afa5-0363afeb3d15 · outbound

This paper cites Video action transformer network.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Video action transformer network

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:46:34.772785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:46:34.014040Z digest=sha256:bbc9ba9d3fdda5770237ede837c1eba1d1c1ae07d1fdd055f543f5136fd0765e

Observation 9aedc72a-7358-42f7-bd3f-c848bd9d4586 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval LoRA: Low-Rank Adaptation of Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T15:46:34.017653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:46:34.017653Z digest=sha256:cf0f1a59ed48a0c9923501fde5b39075e4b47ee190cc2e0dc89730154d90d77a

Observation 557fffbd-95b8-4070-9a83-357119e67536 · outbound

This paper cites Knowing where to focus: Event-aware transformer for video grounding, 2023.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Knowing where to focus: Event-aware transformer for video grounding, 2023

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:46:34.761744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:46:34.021559Z digest=sha256:1215b7fb5bd052e25fefbf6ad4e75d21a8d0099ccf697bfc78ea5444e9426fcd

Observation 3488fcf5-b860-4968-b767-a0b37bab20db · outbound

This paper cites Efficient multimodal large language models: A survey.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Efficient multimodal large language models: A survey

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T15:46:34.025371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:46:34.025371Z digest=sha256:7a5950542b202f434fb0a4f9047940f63c99eab47487fda8a75c238a4016c251

Observation 2adcebb4-0b78-45eb-a74f-b2fb462a1e8c · outbound

This paper cites Dense-captioning events in videos,.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Dense-captioning events in videos,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T15:46:34.029064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:46:34.029064Z digest=sha256:6a1fc2d4bb67f4485aadb100d4466cf62d30862b6b214345e31057aa66866512

Observation 40028e2e-2684-48b2-aea1-10f5e0b6ea3d · outbound

This paper cites Temporal convolutional networks for ac- tion segmentation and detection.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Temporal convolutional networks for ac- tion segmentation and detection

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:46:34.743835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:46:34.032763Z digest=sha256:87a5050f2daa92ec5cee4c3b09a7939b72e6c5626cb0bc53de8606ae49fc74d9

Observation 1351cd3b-6637-480e-94f8-918d8740b673 · outbound

This paper cites Berg, and Mohit Bansal.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Berg, and Mohit Bansal

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:46:34.732984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:46:34.036780Z digest=sha256:a6d6e55e1d06c29678d47d527e300fc8f58babad1837321e96f162fe4f7cf866

Observation caaaa026-fd5e-4e66-b865-ec6984da59e1 · outbound

This paper cites Detecting mo- ments and highlights in videos via natural language queries.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Detecting mo- ments and highlights in videos via natural language queries

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:46:34.721157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:46:34.040363Z digest=sha256:bb26c9e481b3a33757243c8630422370c8959e545f12880a889a2af7d6800059

Observation 7d3f1942-2f5b-44ea-a7a4-3c544590db18 · outbound

This paper cites Mimic-it: Multi-modal in-context instruction tuning, 2023.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Mimic-it: Multi-modal in-context instruction tuning, 2023

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T15:46:34.043808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:46:34.043808Z digest=sha256:a3d4f6b74bff6ede7b65e89c712ea333ce84e5bd5432421137cf8ff687f27a31

Observation 4fb2af8c-fcc8-4136-9ccf-17de71e758c7 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models, 2023.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models, 2023

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T15:46:34.047579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:46:34.047579Z digest=sha256:72723e06b2e7e68f67f5798574dc7297740db65416dbb5b4a2d8e56174a471a6

Observation 48d2f88f-bc4d-4379-bd57-59a0eb3a210b · outbound

This paper cites A survey on benchmarks of multimodal large language models,.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval A survey on benchmarks of multimodal large language models,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:46:34.695108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:46:34.051139Z digest=sha256:e690c11d55821c12aee83bf5503df003753be610896e644117015e059c14065e

Observation 7bc2fa66-4b8e-400c-b54a-8c4a5a35b30a · outbound

This paper cites Videochat: Chat-centric video understanding, 2024.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Videochat: Chat-centric video understanding, 2024

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T15:46:34.054794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:46:34.054794Z digest=sha256:feb0e7e56954a5c8647b35a07acb005a2b0e43e70dfcea563840eed91106454e

Observation 5d95690e-881a-43eb-8591-e22017a0498d · outbound

This paper cites Fast learning of temporal action proposal via dense boundary generator.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Fast learning of temporal action proposal via dense boundary generator

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:46:34.676050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:46:34.058301Z digest=sha256:1a662e457d44c172bf2d9b4d2c388f209c74b81717db898ee372e88e92d7fb09

Observation 9c8b8761-daf7-43ea-880d-fa85cf37a993 · outbound

This paper cites Univtg: Towards unified video- language temporal grounding, 2023.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Univtg: Towards unified video- language temporal grounding, 2023

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:46:34.664660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:46:34.061972Z digest=sha256:b2a354685f5e90dec0c13707e7f83dadf771c73f24e7f84916bbf0d22c2790c2

Observation 71aee2b2-1ce9-4d9b-901d-e25aa33cfbf3 · outbound

This paper cites Bsn: Boundary sensitive network for temporal action proposal generation.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Bsn: Boundary sensitive network for temporal action proposal generation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:46:34.654031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:46:34.066083Z digest=sha256:071ba403584ac45d752ec8cae2f8a591f2755d0197323e6ba466d1fa8b19ec7e

Observation 8c4996de-1346-4b70-a7d6-06e9f565ddd7 · outbound

This paper cites Visual instruction tuning, 2023.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Visual instruction tuning, 2023

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T15:46:34.069591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:46:34.069591Z digest=sha256:805be8bea845ac011339d0faf182129243efca6c059611f1ba712c78931f89b1

Observation 1174f455-31c5-4ed0-948c-9c7b642a567f · outbound

This paper cites Umt: Unified multi-modal transformers for joint video moment retrieval and highlight detection.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Umt: Unified multi-modal transformers for joint video moment retrieval and highlight detection

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:46:34.635444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:46:34.072992Z digest=sha256:536aea648a7dbe2e84df5ad37cc12ad362f12af92da85f234ed806fba7764e76

Observation 71cbe1a5-0956-4492-b4c4-7d6b114981f4 · outbound

This paper cites Decoupled weight decay regularization, 2019.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Decoupled weight decay regularization, 2019

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T15:46:34.076530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:46:34.076530Z digest=sha256:0f7facc095d705a02fdde47f6dafa08c4f834527d3aa8772a2418d5783a173ab

Observation 246e3dfe-457e-4f15-ae4a-7c4edb6b7597 · outbound

This paper cites Valley: Video assistant with large language model enhanced ability, 2023.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Valley: Video assistant with large language model enhanced ability, 2023

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T15:46:34.079946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:46:34.079946Z digest=sha256:9a5783ee9cbb31a425509bb3539c61fa9e55d44ec20f9c04cb299c555d843244

Observation 19f0d48a-eb54-4200-943a-d996aed2b9e5 · outbound

This paper cites Video-chatgpt: Towards detailed video understanding via large vision and language models, 2024.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Video-chatgpt: Towards detailed video understanding via large vision and language models, 2024

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:46:34.609308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:46:34.083458Z digest=sha256:29f7697212cde571c07fae003635fae48317ee58c4fa98a317ff8524d8ed71ae

Observation 53832ca4-6755-4aa6-bf0a-05ee22e796cf · outbound

This paper cites The surprising effectiveness of multimodal large language models for video moment retrieval, 2024.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval The surprising effectiveness of multimodal large language models for video moment retrieval, 2024

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:46:34.598219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:46:34.086941Z digest=sha256:e86d77b6e919ac269dd259703bd0b9e00f0ba547fbabc80b769bd43c279a67d2

Observation b2bb8faa-3cb8-4b76-9c4a-eb24a7479be3 · outbound

This paper cites Query-dependent video representa- tion for moment retrieval and highlight detection, 2023.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Query-dependent video representa- tion for moment retrieval and highlight detection, 2023

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:46:34.586889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:46:34.090381Z digest=sha256:b4d349496a10fa2110f45df0fdbdd234f50538f31af02824aea2290307d5d698

Observation fd5efcf6-9441-4802-8a47-4cacf6a24796 · outbound

This paper cites Correlation-guided query-dependency calibration for video temporal grounding, 2024.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Correlation-guided query-dependency calibration for video temporal grounding, 2024

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:46:34.575592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:46:34.094038Z digest=sha256:26820e1950f9076e2736fb446ba520ce01efe03ba8e653917299a6ff631df395

Observation 8770941d-1da9-4a76-9b6a-9742bc443b62 · outbound

This paper cites Local- global video-text interactions for temporal grounding.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Local- global video-text interactions for temporal grounding

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:46:34.564426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:46:34.097533Z digest=sha256:d65fe6bba709a7767b5fa765cebc46546d5afa4d3e62926d9e9d103fcca8c6bf

Observation 04bc2f47-7523-40fb-8322-ba84ceab1104 · outbound

This paper cites Pat: Position-aware transformer for dense multi-label action detection.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Pat: Position-aware transformer for dense multi-label action detection

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:46:34.553015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:46:34.101057Z digest=sha256:79b0980522d99b3a5bf69f02dd1ab3ca7df7094805bbde4ed10817336cfa058e

Observation 0fd47542-75b4-408d-8759-a2762961bf45 · outbound

This paper cites Temporal action localization in untrimmed videos via multi-stage cnns.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Temporal action localization in untrimmed videos via multi-stage cnns

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:46:34.541275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:46:34.104973Z digest=sha256:a6877b4175b9e231fa31292c5db447d98fef47ee31178572a708c2245dedbe32

Observation 06cf80aa-567e-4be8-b43b-a2f20b646461 · outbound

This paper cites Vlg-net: Video-language graph matching network for video grounding, 2021.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Vlg-net: Video-language graph matching network for video grounding, 2021

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:46:34.530151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:46:34.108323Z digest=sha256:8b41ef1361336ebea9fbad94db749aa614b2563b876b0ab19861aecb92c2455e

Observation 2c06cbe0-62ce-4cf3-9857-2154ddbc99f1 · outbound

This paper cites Learning grounded vision-language representation for versatile understanding in untrimmed videos, 2023.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Learning grounded vision-language representation for versatile understanding in untrimmed videos, 2023

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:46:34.519074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:46:34.111758Z digest=sha256:59878540c27b6a6388bca93d5c6f7713b5def6ad7cac898082f8d91cb6ea7562

Observation a557d774-50bd-472c-a27d-9ebbf57e4a13 · outbound

This paper cites Internvideo2: Scaling foundation models for multimodal video understanding, 2024.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Internvideo2: Scaling foundation models for multimodal video understanding, 2024

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:46:34.508781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:46:34.115164Z digest=sha256:be7b2f985de668480469d49b6ef9b560dbe0b9fe3ab010698474d4a14a7e7eb4

Observation 70727b54-4d97-42c4-8ec7-9bdd134efb16 · outbound

This paper cites Unloc: A unified framework for video localization tasks, 2023.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Unloc: A unified framework for video localization tasks, 2023

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:46:34.498670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:46:34.118300Z digest=sha256:2749a02ce04eae495dd690a606b012a46b65e5c8e401aa3f9d738921dc5ff2c4

Observation b175bed0-7f32-4ac1-836d-2ab0537235bd · outbound

This paper cites Unloc: A unified framework for video localization tasks, 2023.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Unloc: A unified framework for video localization tasks, 2023

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:46:34.489049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:46:34.121669Z digest=sha256:2f6490872addb18f23a86fb2e2dc3b692e5a1ec5bbf47dc28537ced75124eeb5

Observation 99495fa3-ae96-4f4d-b2a3-a1b32e1025b3 · outbound

This paper cites Self-chained image-language model for video localization and question answering, 2023.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Self-chained image-language model for video localization and question answering, 2023

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:46:34.478605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:46:34.124891Z digest=sha256:d072d03fbaa310ddb3d9dba9744aef9a3b7b1a8f71fdeb4a83f7d95fbafcae63

Observation 53683c38-2ee5-42a1-8334-a5ccb8a990e7 · outbound

This paper cites Semantic conditioned dynamic modulation for tempo- ral sentence grounding in videos.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Semantic conditioned dynamic modulation for tempo- ral sentence grounding in videos

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:46:34.466924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:46:34.128070Z digest=sha256:06728d395161d61fdd5938f6209631af7eb2b43829e164c59ad32d1c6f1ed113

Observation 3c4a4bbc-5325-4133-a3cd-864e8e3d07e4 · outbound

This paper cites Graph con- volutional networks for temporal action localization.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Graph con- volutional networks for temporal action localization

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:46:34.454748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:46:34.131043Z digest=sha256:a1c3ccfe77ba3d16a59eee7ea9cf8dffedd2ae0aae895ef417df06c2cb1da498

Observation abe0109e-633f-4829-b1b3-f1ac251df33b · outbound

This paper cites Dense regression network for video grounding, 2020.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Dense regression network for video grounding, 2020

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:46:34.443580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:46:34.134371Z digest=sha256:8892d361f7d32e0fd1ca1ee9a385371bf13a2c33c502be8767352162d98fd0e4

Observation 16ead2ca-91f5-4d79-b4ce-19ee17466dc8 · outbound

This paper cites Unimd: Towards unifying moment retrieval and temporal ac- tion detection, 2024.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Unimd: Towards unifying moment retrieval and temporal ac- tion detection, 2024

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:46:34.432557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:46:34.137745Z digest=sha256:54e025b530e07efb4211687f6f187537b89a9942a53f8170a266f5f3cee4a78a

Observation 6ffe070f-1deb-4aa9-bb1e-af6a2e1b46bf · outbound

This paper cites Actionformer: Localizing moments of actions with transformers.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Actionformer: Localizing moments of actions with transformers

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:46:34.420874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:46:34.141472Z digest=sha256:da3fe89aee14dba0618f5f4e8cc2a3aca77e3780176e1442ea4d75c8fb6e68eb

Observation 88ccc7a2-1faf-4697-a385-08e9d72ead34 · outbound

This paper cites Man: Moment alignment network for natural language moment retrieval via iterative graph adjustment.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Man: Moment alignment network for natural language moment retrieval via iterative graph adjustment

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:46:34.409042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:46:34.145028Z digest=sha256:8b1d4f7bec64c1210eadd401d44b88c5623350afdc3e75926b9c180801463052

Observation fb139d1d-1b80-4db8-8346-44fa833f7abb · outbound

This paper cites Video-llama: An instruction-tuned audio-visual language model for video un- derstanding, 2023.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Video-llama: An instruction-tuned audio-visual language model for video un- derstanding, 2023

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T15:46:34.148600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:46:34.148600Z digest=sha256:7d776f870064f2c009c8cf20649bd146483ea78aba5068a446bab2d421b6b868

Observation d70a4ef0-1d32-41b6-810f-abe85ea0bf88 · outbound

This paper cites Llama-adapter: Efficient fine-tuning of language models with zero-init attention, 2024.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Llama-adapter: Efficient fine-tuning of language models with zero-init attention, 2024

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:46:34.390918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:46:34.151866Z digest=sha256:f7609c9c168b676ef67424b220357b1fc177c60ed15ffcdb29f1dbd548807701

Observation 30d3c546-e572-4a63-9f0f-244e97c6613a · outbound

This paper cites Learning 2d temporal adjacent networks for moment local- ization with natural language.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Learning 2d temporal adjacent networks for moment local- ization with natural language

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T15:46:34.155307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:46:34.155307Z digest=sha256:b2f0b9c179876efcb725ad7d45c088c14c42c94c69029ed7d8257a39ec11c526

Observation 0ce4def2-55cd-43a5-945e-ae64f5ff87d1 · outbound

This paper cites Temporal action detection with structured segment networks.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Temporal action detection with structured segment networks

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:46:34.372984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:46:34.158791Z digest=sha256:ef2d327191cf7e0abffd53a3e30e2e7777e688eaa36b7e942b35d8c94c9b8e35

Observation 8b0f7f74-0fa9-4704-9eb2-de4b3c8e6e6d · outbound

This paper cites Minigpt-4: Enhancing vision-language understanding with advanced large language models, 2023.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Minigpt-4: Enhancing vision-language understanding with advanced large language models, 2023

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:46:34.361880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:46:34.162165Z digest=sha256:08fd5953ff842c75220da9cafacd18393b916de9b43badbe1703a233091ff42e

Observation f7d3ba55-d9ba-42ae-ad0c-cea89989a227 · outbound

This paper cites Enriching local and global contexts for temporal action localization.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Enriching local and global contexts for temporal action localization

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:46:34.350436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:46:34.165559Z digest=sha256:427b68ee02d46c04b512891061b76482a4bcd8c692151012f41858bafd41a006

Observation a8e99c9d-6d27-43a4-8ffe-5dd551089ade · outbound

This paper cites an unresolved cited work.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:46:34.339358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:46:34.168963Z digest=sha256:9ebb50153cdbe5fba68ed1b7af0ab358dc5c7757f88f45192bfca25b42a12d2b

Observation 5eab3430-918f-4c82-ae81-aed16842eb78 · outbound

This paper cites [[-1, -1]].

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval [[-1, -1]]

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:46:34.328503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:46:34.172906Z digest=sha256:ff2fa8fe70df22d19b85dacdabbea30346042de2dda271906f9e2bbc663ab76a

Observation 7c2b5e37-9af4-4a0b-9e1d-3aae182d0e2f · outbound

This paper cites automated devices operating in a modern factory.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval automated devices operating in a modern factory

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:46:34.317235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:46:34.176858Z digest=sha256:89e16cdfe5d1b8ebda3f6d5ea268bbbc285fcf3f372987e1c3012cc61b779bcc

Observation 8500c787-c420-47b0-81e9-08b5dc4e9d88 · outbound

This paper cites This section explores po- tential future directions for enhancing the performance of MLLMs in moment retrieval tasks.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval This section explores po- tential future directions for enhancing the performance of MLLMs in moment retrieval tasks

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:46:34.305958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:46:34.180369Z digest=sha256:110f07687789647afa06c3d6b570dfe050fafc5139525d4d062d6f14ca69696b

Pith citing papers

Observation a8c5eb12-65ec-487d-8628-b0f2bed9b148 · inbound

DisTime: Distribution-based Time Representation for Video Large Language Models cites this paper.

DisTime: Distribution-based Time Representation for Video Large Language Models LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:52.859271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:52.859271Z digest=sha256:16d7a6b7ae2cf409aff0921f6862e60778793e655f6c27533f71a7d8e5583828

Observation 34b3b1a3-7981-425a-a102-7663f633fb81 · inbound

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding cites this paper.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T18:40:02.771931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:40:02.771931Z digest=sha256:121a18032b847b888222e5baa5818be4375cd77141ef7e17776c638924699cf0

Observation e4436d2b-52f0-4d91-9374-14e51158659b · inbound

A Survey on Video Temporal Grounding with Multimodal Large Language Model cites this paper.

A Survey on Video Temporal Grounding with Multimodal Large Language Model LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval

Reference 106

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.938393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.938393Z digest=sha256:884954aa3b54181cecad0ce774f46463d86e74acf4e4ac8c75d4afe101835720

Observation 66daadb5-120d-4ba9-985d-3d3e1ae18f45 · inbound

SMART: Shot-Aware Multimodal Video Moment Retrieval with Audio-Enhanced MLLM cites this paper.

SMART: Shot-Aware Multimodal Video Moment Retrieval with Audio-Enhanced MLLM LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T21:42:45.907432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:42:45.907432Z digest=sha256:5fd9227ac9a14a965fe9e22ac7e714d98b58ae7faf88678b1022f8029f77bc84

Observation 4b3d4251-7538-4694-85e7-b6fa0d1c2d9c · inbound

SiMing-Bench: Evaluating Procedural Correctness from Continuous Interactions in Clinical Skill Videos cites this paper.

SiMing-Bench: Evaluating Procedural Correctness from Continuous Interactions in Clinical Skill Videos LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:30:58.574381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-10T17:37:40.373211Z digest=sha256:55f0648fe79d7a94c52c47a5551fd1b8fb2985e25a63ec32d176d42239581b8b

Observation 905d3179-9bae-429f-9fb5-bafc0ad4fb4e · inbound

Towards One-to-Many Temporal Grounding cites this paper.

Towards One-to-Many Temporal Grounding LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T12:16:57.671960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-28T02:11:48.455492Z digest=sha256:c54fc6d25c46bda85ec1d265f5d4e21b87f29dd83f5c3fb1c5fd7fb78a4014e9

Observation 0b2b1842-8324-47dd-9e39-7e9d9215ef89 · inbound

Watch, Remember, Reason: Human-View Video Understanding with MLLMs cites this paper.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:14.990641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:1e4b0fec965e81e8bc535b81dac81cfea09f50fdb116e215fad379c95eb74a03