Pith. sign in

Paper Citation Record · LEDGER

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding

As of 23 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 2 inbound Pith citation observations for arXiv:2506.07600.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.07600 v1

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:35:51.010785Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T19:15:02.124035Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T20:06:13.270265Z

Reference resolution

58 of 58 outbound references displayed

  • verified exact0
  • verified fuzzy23
  • unresolved35
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0f78863e-e058-499f-8271-8b3df8218a37 · outbound

This paper cites Ask in Any Modality: A Comprehensive Survey on Multimodal Retrieval-Augmented Generation.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Ask in Any Modality: A Comprehensive Survey on Multimodal Retrieval-Augmented Generation

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:50.716342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:50.716342Z digest=sha256:c40933ce7081025fb18557e4ecbd324a97e8c1650fb9f9ab1c243871a2537b00

Observation b63a5283-71c9-45e2-a079-fc9ef20f04ec · outbound

This paper cites Flamingo: a visual language model for few-shot learning.Advances in neural information processing systems, 35: 23716–23736, 2022.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Flamingo: a visual language model for few-shot learning.Advances in neural information processing systems, 35: 23716–23736, 2022

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:50.722098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:50.722098Z digest=sha256:8683c4d585342e2ec78486d8f39b1dcde5d6a2ec55afff7c51332c4d84fda3e0

Observation 2b4c3222-d80a-40e0-9ff5-6713bb3c1325 · outbound

This paper cites Vivit: A video vision transformer.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Vivit: A video vision transformer

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:50.726908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:50.726908Z digest=sha256:391430f13147495f149d5d4d07d153c84c2f3a5354db29db0240354e8d44d1f5

Observation 4049a5bd-46e6-4991-a9e3-2b1c5d84afce · outbound

This paper cites Unified graph structured models for video understanding.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Unified graph structured models for video understanding

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:52.098046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:35:50.732034Z digest=sha256:6c8c57c053a4bea4a80167b36fc3d507d06f3f2311ae20bdd85c021158124d06

Observation b3c25b8a-437b-4dc0-9001-eca88c2a9f25 · outbound

This paper cites Self-rag: Learning to retrieve, generate, and critique through self-reflection.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Self-rag: Learning to retrieve, generate, and critique through self-reflection

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:50.736730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:50.736730Z digest=sha256:852abdfb5964c5f5b3cf2df9a9d91c5d8db5a00ba5190c12dd475ee8a09fce28

Observation 4aa681f7-bfa0-43cc-9892-b054964886a7 · outbound

This paper cites Sharegpt4video: Improving video understanding and generation with better captions.Advances in Neural Information Processing Systems, 37:19472–19495, 2024.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Sharegpt4video: Improving video understanding and generation with better captions.Advances in Neural Information Processing Systems, 37:19472–19495, 2024

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:52.070182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:35:50.741138Z digest=sha256:70d3a16eb0b4454193cde982db5c7c5ab72e250091fea2b9e1c204203fe25dcf

Observation d0bd1e4f-7665-4a89-852b-f07dce6a7fa9 · outbound

This paper cites WavRAG: Audio-Integrated Retrieval Augmented Generation for Spoken Dialogue Models.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding WavRAG: Audio-Integrated Retrieval Augmented Generation for Spoken Dialogue Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:50.746422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:50.746422Z digest=sha256:a5760780d0cabcbc4b1f2cddf8e156170543df97242e06fc9116e60ad8e6648e

Observation 787bbd23-4312-46c8-a59a-1bdd76534d1b · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:50.751331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:50.751331Z digest=sha256:4c8a425f848d315d260a40751b4b71585cf8696209fe1c2e053e9e002fbfcd21

Observation 1d00498a-0c60-4167-9e19-2234800adfc9 · outbound

This paper cites an unresolved cited work.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:35:52.052581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:35:50.756117Z digest=sha256:dacc1b63b293d84f6dc2db03a3a25ab7411a5e5723557737f527fd2e759c1f17

Observation a5ce122e-a137-4c49-a7ca-877690d956bb · outbound

This paper cites Sscformer: Push the limit of chunk-wise conformer for streaming asr using sequentially sampled chunks and chunked causal convolution.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Sscformer: Push the limit of chunk-wise conformer for streaming asr using sequentially sampled chunks and chunked causal convolution

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:52.034449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:35:50.762345Z digest=sha256:e6666ad4fe4fb4c9ba34dd5e64ff286f334d41a3a14e8d4676110c4cccc46bba

Observation 353333cb-08da-4433-80ec-c3f2e8d06b12 · outbound

This paper cites Event segmentation and seven types of narrative discontinuity in popular movies.Acta psychologica, 149:69–77, 2014.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Event segmentation and seven types of narrative discontinuity in popular movies.Acta psychologica, 149:69–77, 2014

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:52.017494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:35:50.766627Z digest=sha256:77a3cd03e7d7ac1221728abf636bb911d394a23773104266578825b87d0611de

Observation 53b133fe-8f81-4f76-ac9a-d7e0e3ebc3b4 · outbound

This paper cites Large-scale narrative events in popular cinema.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Large-scale narrative events in popular cinema

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:52.000990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:35:50.770754Z digest=sha256:3fd68437e4f2b28e7c64c35961fda4af3a43a9a2574ae7e658d200d422759ad9

Observation 42d938a9-4bf6-435d-ad1f-ca8e98c3aa5d · outbound

This paper cites From Local to Global: A Graph RAG Approach to Query-Focused Summarization.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding From Local to Global: A Graph RAG Approach to Query-Focused Summarization

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:50.775173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:50.775173Z digest=sha256:bbb8f392195f3c5d285b2ce03c18aee1f865b73a1f52f32d696f98df5a1dceaf

Observation 16cffcf3-1e1b-45a6-83ef-a3abc8e033f4 · outbound

This paper cites Videoagent: A memory-augmented multimodal agent for video understanding.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Videoagent: A memory-augmented multimodal agent for video understanding

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:51.985406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:35:50.779943Z digest=sha256:54b83f89a010bfb7a406890c7fdfd1247d11a09f08327f95946873480d596936

Observation 2616490e-330d-4164-ad63-ef68cacedb31 · outbound

This paper cites Distil-Whisper: Robust Knowledge Distillation via Large-Scale Pseudo Labelling.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Distil-Whisper: Robust Knowledge Distillation via Large-Scale Pseudo Labelling

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:50.784040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:50.784040Z digest=sha256:267641b5b2758d90a4a40213a9c031937b86c79de3ced81c5372b61b3d79c678

Observation c1cab8fa-f191-40aa-9038-3dd0efdbfaf9 · outbound

This paper cites Retrieval-Augmented Generation for Large Language Models: A Survey.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Retrieval-Augmented Generation for Large Language Models: A Survey

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:50.788520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:50.788520Z digest=sha256:86f917ed51176b02c3caa3b3b1271fdea55d9b0ba523741148b045b12e306513

Observation de8acc01-3b82-452f-92dd-3ba6f04c347f · outbound

This paper cites Imagebind: One embedding space to bind them all.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Imagebind: One embedding space to bind them all

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:50.793225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:50.793225Z digest=sha256:8abb77a915cea10b9ee1c7a600a01800557e7fa0514bf1656049ecccae3e68cc

Observation 1c4d502e-b1bd-405a-8c0c-27d89342ab2e · outbound

This paper cites Lightrag: Simple and fast retrieval-augmented generation.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Lightrag: Simple and fast retrieval-augmented generation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:50.797523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:50.797523Z digest=sha256:14f811ad8a5a9847409e7264f294ecdbace026e75c654122e4e47d59e3c845e4

Observation 8590ccea-a560-4f54-a566-bd67cde45dde · outbound

This paper cites VideoRAG: Retrieval-Augmented Generation over Video Corpus.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding VideoRAG: Retrieval-Augmented Generation over Video Corpus

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:50.801689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:50.801689Z digest=sha256:2ef97a36dd9e57ad4b5d6e701975de7a6206837266da469cd1a99f335da30d30

Observation 5fa4fc56-0f84-4871-994c-fa8357254b9c · outbound

This paper cites Learning temporal video procedure segmentation from an automatically collected large dataset.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Learning temporal video procedure segmentation from an automatically collected large dataset

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:51.948243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:35:50.806583Z digest=sha256:7f9366de7df2a6f6fdf6a8dca85b61290b0c4608ccf93ce763c5dc81ea32063d

Observation 3c713945-46a5-408d-acb2-a0ecf0d7a903 · outbound

This paper cites Diffusionret: Generative text-video retrieval with diffusion model.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Diffusionret: Generative text-video retrieval with diffusion model

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:51.932126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:35:50.810986Z digest=sha256:b9a4317eed6007ddfaaa6d2e24b3e48a1c1b45b214cc637d32643a6066d99c8f

Observation ed6b7bfa-2166-4a9b-93b6-5f8e8ff72f6e · outbound

This paper cites Scene Graph Generation Strategy with Co-occurrence Knowledge and Learnable Term Frequency.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Scene Graph Generation Strategy with Co-occurrence Knowledge and Learnable Term Frequency

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:50.815034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:50.815034Z digest=sha256:b0458481b6872a4e9bbfbdeda1b793b84278e4554c7dd0cc5e57d31d06a5e7a8

Observation 6c089fd8-d280-4021-97cd-dbdd99f5ce81 · outbound

This paper cites Multimodal Reasoning with Multimodal Knowledge Graph.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Multimodal Reasoning with Multimodal Knowledge Graph

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:50.819345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:50.819345Z digest=sha256:85a9a3db1022f39764bbbace76802847fc2b98777459f01f336068ea2843a865

Observation 351fcb6e-6148-4978-b96e-70ca201633ab · outbound

This paper cites Retrieval-augmented generation for knowledge-intensive nlp tasks.Advances in neural information processing systems, 33:9459–9474, 2020.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Retrieval-augmented generation for knowledge-intensive nlp tasks.Advances in neural information processing systems, 33:9459–9474, 2020

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:50.823707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:50.823707Z digest=sha256:d57090340ba7c3cb0606cfa209854c66944974c21b71c7346f3a14edaacdac78

Observation b98855b0-59d2-4594-862a-32d7436c411f · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:50.828025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:50.828025Z digest=sha256:21e41e453342a22cfa7f45a13c9bc3a0130ff6791a0a7c935867a8e407eadff1

Observation c19432be-2e3c-44ab-82f5-9ad99d1453ae · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding VideoChat: Chat-Centric Video Understanding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:50.832358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:50.832358Z digest=sha256:a1028568d86838bf942e95339171c224ef38a4b04057ac422e1edefaff36bdb7

Observation 5542f591-d7eb-4a35-9ee4-6c374c36ef95 · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Llama-vid: An image is worth 2 tokens in large language models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:50.837354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:50.837354Z digest=sha256:c238e29aa2a419562bc23a05b0290eea2318c73f67fb9b19070116ebff85991c

Observation 8bde4f57-de10-42fb-a65a-61b852cbd1d7 · outbound

This paper cites UniVL: A Unified Video and Language Pre-Training Model for Multimodal Understanding and Generation.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding UniVL: A Unified Video and Language Pre-Training Model for Multimodal Understanding and Generation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:50.842483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:50.842483Z digest=sha256:be43644bc2447daca3037a12e64c2436e4e68ae1dfc6f577a56dc5e4cf590506

Observation 13c512e3-d06d-4697-80a8-633f478325e7 · outbound

This paper cites The impact of continuity editing in narrative film on event segmentation.Cognitive science, 35(8):1489–1517, 2011.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding The impact of continuity editing in narrative film on event segmentation.Cognitive science, 35(8):1489–1517, 2011

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:51.883271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:35:50.847978Z digest=sha256:16fe04b27a9e8d62603f719a4daf9d4d3fb769d0113d60d63bfa7c46b5a36447

Observation 06cd2825-1733-4f1c-b7ce-5c6ed7b1e87f · outbound

This paper cites Learning joint embedding with multimodal cues for cross-modal video-text retrieval.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Learning joint embedding with multimodal cues for cross-modal video-text retrieval

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:51.866744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:35:50.852386Z digest=sha256:29930c44cf0410e50e6c2c2730a36a24875a657badb265501b15599362d9017c

Observation bc233881-faed-4732-8fd9-8fe0c8ade76d · outbound

This paper cites Boundary-aware Self-supervised Learning for Video Scene Segmentation.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Boundary-aware Self-supervised Learning for Video Scene Segmentation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:50.858475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:50.858475Z digest=sha256:73ed103e41b9d44779a8ff50a95e6e3c2a8f493846ef2e1285e420e90f8d1447

Observation 4ffcb655-c661-401a-9aee-a54baf56b1ce · outbound

This paper cites Video shot boundary detection: a review.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Video shot boundary detection: a review

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:51.849862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:35:50.863274Z digest=sha256:3da859b89d25d0dd036c748ec6b6a27661ed7c9d742daecd263aac751e942895

Observation 62a16ffe-9666-4ddb-8c19-34f9ce839849 · outbound

This paper cites What is a multi-modal knowledge graph: A survey.Big Data Research, 32:100380, 2023.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding What is a multi-modal knowledge graph: A survey.Big Data Research, 32:100380, 2023

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:51.832449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:35:50.867733Z digest=sha256:5166ca82740fd8a4e084f6b973d5f83ed5e13cd22045189a3246adbc496ecd7e

Observation 0e5e5683-a5e7-4451-a79b-9e0c6bf79174 · outbound

This paper cites Learning transferable visual models from natural language supervision.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Learning transferable visual models from natural language supervision

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:50.872103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:50.872103Z digest=sha256:4999105ac353e5d41a9274f15630b6b7de89ff5cbec2af2e7308a0e44372b01e

Observation 0c2e3e7c-3ea5-4db7-ba22-9e6a489e87bc · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Robust speech recognition via large-scale weak supervision

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:50.876870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:50.876870Z digest=sha256:0a4f942ed618c3fb4f58f720217235fb9509c8a7238560f4ca62dc746a8261b9

Observation 569ffada-7c69-455a-ba47-b05e7ad7e021 · outbound

This paper cites VideoRAG: Retrieval-Augmented Generation with Extreme Long-Context Videos.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding VideoRAG: Retrieval-Augmented Generation with Extreme Long-Context Videos

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:50.881284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:50.881284Z digest=sha256:8e1df94c27666e83c1c5e328e6fd2e0804904f47f4654ab3213b83ed4ea69d89

Observation 3bf78b05-a0ad-4f6c-bd63-585137abeb79 · outbound

This paper cites Temporal video segmentation to scenes using high-level audiovisual features.IEEE Transactions on Circuits and Systems for Video Technology, 21(8):1163–1177, 2011.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Temporal video segmentation to scenes using high-level audiovisual features.IEEE Transactions on Circuits and Systems for Video Technology, 21(8):1163–1177, 2011

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:51.792701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:35:50.885872Z digest=sha256:28e494033570ef627fbce7c8b9425cfd0a781a02985417143758d483ea4aeb5b

Observation dfd66a0f-e71e-4f30-b259-61b9958a83e3 · outbound

This paper cites Videobert: A joint model for video and language representation learning.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Videobert: A joint model for video and language representation learning

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:51.775488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:35:50.890403Z digest=sha256:f2e7826a848475d65d8cb4f9869079a4ed7e95182578fae457cea569cad3f537

Observation 9a62d48e-553a-420c-b665-e452fc124e47 · outbound

This paper cites Temporal scene montage for self- supervised video scene boundary detection.ACM Transactions on Multimedia Computing, Communi- cations and Applications, 20(7):1–19, 2024.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Temporal scene montage for self- supervised video scene boundary detection.ACM Transactions on Multimedia Computing, Communi- cations and Applications, 20(7):1–19, 2024

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:51.758787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:35:50.894622Z digest=sha256:960cd6b19ac09d03b73c040bcb021964b5c8b028760350c274860c5b05eac428

Observation 7d489f82-4221-48ca-93fe-4945e8906294 · outbound

This paper cites Videoagent: Long-form video understanding with large language model as agent.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Videoagent: Long-form video understanding with large language model as agent

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:50.898518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:50.898518Z digest=sha256:5f08dd385560dcb366eabdba376875c71a602a12ef0b2fddb886360f451f7671

Observation 539bc9a7-3595-46ab-8a8e-871564008167 · outbound

This paper cites Scene consistency representation learning for video scene segmentation.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Scene consistency representation learning for video scene segmentation

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:51.731995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:35:50.902567Z digest=sha256:ae5e4082fb8445728f3340b2a577a52358071b24c062f73116ce6719381fe328

Observation bbecc0ac-db34-4231-bf89-b3a95d52059a · outbound

This paper cites VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:50.906801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:50.906801Z digest=sha256:6e03a204a7d8447ee5c85ef25a0c7def54ec15cb1e623cf14ac98b6219b8f43a

Observation 417c172f-13e5-4862-81ef-0eff5edd6e6a · outbound

This paper cites Retrieval- augmented egocentric video captioning.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Retrieval- augmented egocentric video captioning

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:51.715997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:35:50.911215Z digest=sha256:3da2e76292e8ed62147fdaf6a67d0b00d28a86d825c36bd2285ecf90d1532a06

Observation b844bde7-bfd0-42a9-b90a-90514f00b513 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:50.915326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:50.915326Z digest=sha256:ac9a072c032c989005b7d5124bf8f3813689b2e887751cc5205bdd172ae5c1ed

Observation e7a23ea1-dcb5-4a6d-a212-4b6c0bf8ac72 · outbound

This paper cites Momentseeker: A comprehensive benchmark and a strong baseline for moment retrieval within long videos.arXiv preprint arXiv:2502.12558, 2025.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Momentseeker: A comprehensive benchmark and a strong baseline for moment retrieval within long videos.arXiv preprint arXiv:2502.12558, 2025

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:50.920388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:50.920388Z digest=sha256:8c1a79e61a9569b581b036b206a6bb39035cc9b99c328d315889502d0e7a1912

Observation 3d972c0f-14f9-4ff4-87b8-55fc80973776 · outbound

This paper cites A formal study of shot boundary detection.IEEE transactions on circuits and systems for video technology, 17(2):168–186, 2007.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding A formal study of shot boundary detection.IEEE transactions on circuits and systems for video technology, 17(2):168–186, 2007

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:51.699263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:35:50.924862Z digest=sha256:865c165ce46f015751d1dbfef840cd4178d7c81e52f79baab7339e35930ecfcf

Observation 8cdd8dc1-7a4b-429b-b457-13c683723b0f · outbound

This paper cites The brain’s cutting- room floor: Segmentation of narrative cinema.Frontiers in human neuroscience, 4:168, 2010.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding The brain’s cutting- room floor: Segmentation of narrative cinema.Frontiers in human neuroscience, 4:168, 2010

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:51.683472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:35:50.929441Z digest=sha256:a73d05dfebedd1b4518f96d4fbce395f072be28e1df017b95704af25c7bcc822

Observation 94805a66-2b73-4012-b5ab-e94a4f8e294d · outbound

This paper cites Merlot reserve: Neural script knowledge through vision and language and sound.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Merlot reserve: Neural script knowledge through vision and language and sound

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:51.667190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:35:50.935121Z digest=sha256:1adbaf3bc2458a74e9f4461b963a7b9fcf9a4d25f4ad0b75069721c0a827fdd7

Observation 74bf6da5-efe4-415c-8d58-fe897a7d3590 · outbound

This paper cites RAKG:Document-level Retrieval Augmented Knowledge Graph Construction.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding RAKG:Document-level Retrieval Augmented Knowledge Graph Construction

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:50.939358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:50.939358Z digest=sha256:d399eac333237893e62f8b1125bbc242875ae5e3558e2845d29414304d14d83f

Observation 6722b5b0-1c8d-4a75-a47b-60ab28d58aa8 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:50.943960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:50.943960Z digest=sha256:c82c69d715dcf9d5c0a1f34365662a208dd8756e243e434a4c0cca58e265567c

Observation 8ff9f0a4-ca16-4e0f-a52a-b64249c31166 · outbound

This paper cites Please maintain the required format in your response.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Please maintain the required format in your response

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:51.595091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:35:50.967095Z digest=sha256:db070c3f22ceeb1a9c6439698b7232ff7d0eb3984d89ce1f953c27a710bd8447

Observation 66e0bc2a-77d6-4210-af0d-cb3bf0ebfd8c · outbound

This paper cites an unresolved cited work.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:35:51.650621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:35:50.971971Z digest=sha256:76734eca0a76ac22e949c4ab0b2d9069f646bdfff7c570d8480ed807d620c0c1

Observation 5b9ec369-cc81-4154-bc82-1272e2439b41 · outbound

This paper cites Please maintain the required format in your response.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Please maintain the required format in your response

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:51.560866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:35:50.989912Z digest=sha256:86157259150bb9519778dd764a308ba3de953e52b360a0129246040adc219741

Observation 7353939f-dbf5-4101-88c7-32311ade0fd3 · outbound

This paper cites an unresolved cited work.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:35:51.536324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:35:50.994061Z digest=sha256:373d74d252cdec93eecaa02f9b3f3ec9acfb9baf02a3c8c75c34edb329cc7e83

Observation 39d875e6-f1c9-4291-9ae0-0c50abad866e · outbound

This paper cites an unresolved cited work.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:35:51.629732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:35:50.998519Z digest=sha256:cfae9cfbac00653a7e275aec52662815dc6b55901456d4e8f8c245c1eeb309b4

Observation bf189050-0e90-4dd1-835d-073c1343bfd4 · outbound

This paper cites an unresolved cited work.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:35:51.612920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:35:51.002426Z digest=sha256:c2d966fc191ecc55d7e13acb56b1306ddae9ff2d93fe358e7e7edd389bf89707

Observation dcc53b3c-47fd-46e1-8b8d-b7856737744b · outbound

This paper cites an unresolved cited work.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:35:51.578428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:35:51.006523Z digest=sha256:fc3f0e6036842a5dc0b7c4dad3c9a028a7ad3841ecc6ed4cfbf90eefa1694298

Observation 56d795ad-6a2c-412d-b5f2-5fdaba24c93f · outbound

This paper cites Is This the End of RAG? Anthropic’s NEW Prompt Caching.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Is This the End of RAG? Anthropic’s NEW Prompt Caching

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:51.519683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:35:51.010785Z digest=sha256:3aedb94d7b62331505dad8a6da43a9c215f18ddf8efb345699c187fcca73d21d

Pith citing papers

Observation 421e6316-0ac4-469d-b56b-9ab6d9809818 · inbound

VideoStir: Understanding Long Videos via Spatio-Temporally Structured and Intent-Aware RAG cites this paper.

VideoStir: Understanding Long Videos via Spatio-Temporally Structured and Intent-Aware RAG SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:15:50.244866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T19:15:02.124035Z digest=sha256:caf0e2de2ffaedc10799db1a86967b3eda9e42d712734368019581fc9bed9f1b

Observation 47a983ba-dc27-4128-a3a7-eec95b5f3594 · inbound

Event-Causal RAG: A Retrieval-Augmented Generation Framework for Long Video Reasoning in Complex Scenarios cites this paper.

Event-Causal RAG: A Retrieval-Augmented Generation Framework for Long Video Reasoning in Complex Scenarios SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:06:13.272901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-08T10:15:15.129358Z digest=sha256:49d0c4ceb98b043442d38d92ba2c8896b3cae6011a94ad7b9db410b48cdd9d4a