Pith. sign in

Paper Citation Record · LEDGER

SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models

As of 20 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 7 inbound Pith citation observations for arXiv:2505.12589.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.12589 v1

Coverage vector

measured 62 of 62 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:35:54.610338Z

measured 69 of 69 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:15:47.263691Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T22:44:01.321727Z

Reference resolution

62 of 62 outbound references displayed

  • verified exact0
  • verified fuzzy36
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fffc2ec5-bd33-452b-84d3-763e3051b91e · outbound

This paper cites Robust real-time unusual event detection using multiple fixed-location monitors.

SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models Robust real-time unusual event detection using multiple fixed-location monitors

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:35:55.603806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:35:54.330944Z digest=sha256:3cf47531a4aa900b5d6dcdf1014483815d0e2c31d20b081e26af359427808688

Observation 7cdaba83-6491-4828-b94e-3b4772471824 · outbound

This paper cites Col- laborative learning of anomalies with privacy (clap) for unsupervised video anomaly detection: A new baseline.

SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models Col- laborative learning of anomalies with privacy (clap) for unsupervised video anomaly detection: A new baseline

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:35:55.588574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:35:54.336245Z digest=sha256:3e8aa7f584538cfcbcd2958f1f0686b74c41c3b6eb5dd49d13f5b89bfbe5f5a2

Observation bf5b81ce-ba1d-4dc2-9024-f2c810c0daf0 · outbound

This paper cites Expert video-surveillance system for real-time detection of suspicious behaviors in shopping malls.

SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models Expert video-surveillance system for real-time detection of suspicious behaviors in shopping malls

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:35:55.572276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:35:54.342088Z digest=sha256:365fb09e8fb6f0df7a03e30e0966c3ee208a2a1d9e72857883198f99799483e3

Observation 2d21df78-cea7-458f-bc59-e64033008ec5 · outbound

This paper cites Qwen2.5-VL Technical Report.

SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models Qwen2.5-VL Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T20:35:54.346783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:35:54.346783Z digest=sha256:b8e4c5757b3d3fd92801ae1c4f76f60ff4b50aeb133a8225b786774970cc9b97

Observation 872cfcbe-b287-4059-bdfd-480ccac673a6 · outbound

This paper cites A new benchmark dataset for semi-supervised video anomaly detection and anticipation in complex scenes.

SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models A new benchmark dataset for semi-supervised video anomaly detection and anticipation in complex scenes

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:35:55.556566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:35:54.351594Z digest=sha256:d53314fc993562c66d0de4125c7a75d9e58583327fe1635afa4cef9cda4e55f9

Observation c8f70ae5-f7ad-4e0f-8d2d-fbd3975a9b5c · outbound

This paper cites A Benchmark for Crime Surveillance Video Analysis with Large Models.

SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models A Benchmark for Crime Surveillance Video Analysis with Large Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T20:35:54.356178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:35:54.356178Z digest=sha256:2b1d081c5c0d0adc78295f2a70416bc2da1ae8f7af02557c4d2b3d5c9dbfbd91

Observation 3637e468-1fd4-4c26-a45d-a4c56628d4b6 · outbound

This paper cites PCA-Bench: Evaluating Multimodal Large Language Models in Perception-Cognition-Action Chain.

SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models PCA-Bench: Evaluating Multimodal Large Language Models in Perception-Cognition-Action Chain

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T20:35:54.361381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:35:54.361381Z digest=sha256:87e2c7a1a20663c0e547690dc59b03e6f154479e7cbf88154c4940f08eda1fb7

Observation 3a20c8e5-9658-4394-8b3d-e13cbc5bd47c · outbound

This paper cites ShareGPT4Video: Improving Video Understanding and Generation with Better Captions.

SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T20:35:54.366150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:35:54.366150Z digest=sha256:e9a86b5376ebfca4134c5836fec7ccdf5eee4a5adaebef3e00e2294ec5df7757

Observation 32c19a5f-1b8a-4617-8919-70bbadeae1ae · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T20:35:54.371294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:35:54.371294Z digest=sha256:05f8b1f0bb9850a5a2e4d964ea3213b40cd80a8640d3f8ceb89dd6161f04eb49

Observation c79afe55-5a95-4e55-9a4e-eef1245e05b3 · outbound

This paper cites Meva: A large-scale multiview, multimodal video dataset for activity detection.

SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models Meva: A large-scale multiview, multimodal video dataset for activity detection

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:35:55.541965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:35:54.375977Z digest=sha256:ff7e05ab990ee54c14a6fe9b584e44d9e867c1bce4e4f25c9ebb783ada0f198f

Observation 7684a559-1945-4c3e-8832-c2b8d34ee97f · outbound

This paper cites A survey on multimodal large language models for autonomous driving.

SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models A survey on multimodal large language models for autonomous driving

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:35:55.527010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:35:54.380864Z digest=sha256:73bb41208506c3cd00a35cd6fee1d1ac873b80c9413385b17ffba59c5759476d

Observation 50243912-abb3-4ea7-a11b-c8d22706a5d7 · outbound

This paper cites Any-shot sequential anomaly detection in surveillance videos.

SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models Any-shot sequential anomaly detection in surveillance videos

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:35:55.511792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:35:54.385314Z digest=sha256:774cc60d41b5acf768f29065cbd341251c03e4b2d3db89cb916ea692f6028e43

Observation b05d6f78-5e5f-4c83-9fe6-b3e44bc82297 · outbound

This paper cites MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding.

SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T20:35:54.389906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:35:54.389906Z digest=sha256:f450496c18fd9f6d10e5cfa43d8d2f106cda6784b7d9cba0c36fa0ea7beef330

Observation b7f6d8a5-763d-4ad0-8b3a-9d7f1d421039 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T20:35:54.394435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:35:54.394435Z digest=sha256:0c7458ece7827db7cc64ae7d82b799c5d6f3558c90f304450ae583074e5abcee

Observation 441499e3-c931-4ace-b567-22ca5bfcc405 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T20:35:54.399176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:35:54.399176Z digest=sha256:203bfeff7e782e4463f8ad4b4aff05194f93dc03805e3f19bbafcb483e71eeb0

Observation 751cdc9f-2f7b-4700-9239-60058aaf2876 · outbound

This paper cites Chatglm: A family of large language models from glm-130b to glm-4 all tools, 2024.

SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models Chatglm: A family of large language models from glm-130b to glm-4 all tools, 2024

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T20:35:54.403876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:35:54.403876Z digest=sha256:8a9e83a2742d212d7553418abaa9cfc7754aba080568c64e00a4f18a41784b08

Observation df00b970-8667-45ef-b1e6-a5e22a608ddf · outbound

This paper cites Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark.

SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T20:35:54.408983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:35:54.408983Z digest=sha256:7480e28769fa9d89facb05037c8544a82f04896911acf656f022e956801ec1d7

Observation 4dcd1813-98bb-4a7c-9d07-376288f0226a · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T20:35:54.413693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:35:54.413693Z digest=sha256:121c520b4c024d6bd34bbfaae045ac20c03625f7880bd0033eea45c6fd536889

Observation 72ced7d9-32e8-4287-881c-afd8a511e650 · outbound

This paper cites Smart city as a smart service system: Human-computer interaction and smart city surveillance systems.

SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models Smart city as a smart service system: Human-computer interaction and smart city surveillance systems

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:35:55.487316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:35:54.418772Z digest=sha256:4ccb8596465426b43a2f70050d6872cc3b4cfcc2419b0723222d4b51908633a4

Observation 3c71c165-dea2-407a-acf1-83244904fbeb · outbound

This paper cites Visual question answering: A survey of methods, datasets, evaluation, and challenges.

SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models Visual question answering: A survey of methods, datasets, evaluation, and challenges

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:35:55.472017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:35:54.423159Z digest=sha256:0ecdec351b74b2327adabccf3d57800f7ce824d54ce90a5274074f0bb922375b

Observation 93c585a4-5ece-44f2-938b-37d03c7f1fe6 · outbound

This paper cites an unresolved cited work.

SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:35:55.457585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:35:54.427971Z digest=sha256:f4b3470fe0f70e6c913a71903337b3e0af6dfa7105ff93a769c9d2edeb40e766

Observation 458fdf29-7cf1-49fe-bee5-6d48478212ea · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models LLaVA-OneVision: Easy Visual Task Transfer

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T20:35:54.432337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:35:54.432337Z digest=sha256:f25574c1691734be4a90aa29711e8ece449b0e41b9847361528e1264b3b0234f

Observation d089f308-6008-4597-9f3f-2b20730d426f · outbound

This paper cites Seed-bench: Benchmarking multimodal large language models.

SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models Seed-bench: Benchmarking multimodal large language models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T20:35:54.437086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:35:54.437086Z digest=sha256:42124de87a7bd766d3eda637d022740d6e6057fb3587ec89038fd5d8267da857

Observation 6bff518f-a5ba-4fb6-8921-1ec9cde4f3ad · outbound

This paper cites A Survey on Benchmarks of Multimodal Large Language Models.

SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models A Survey on Benchmarks of Multimodal Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T20:35:54.441663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:35:54.441663Z digest=sha256:f237ec5d741c75e7c6a71d5a88d5efa5928609246b01c142bb9b3f993f6c4ca2

Observation 19e2cc2b-dcdc-4f50-ac31-f8742995830e · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark.

SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models Mvbench: A comprehensive multi-modal video understanding benchmark

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T20:35:54.446249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:35:54.446249Z digest=sha256:c4fe276001a5b98887cfd36307a4a7d762b03af49b58bde0cd2c98d262c0ee92

Observation de62ac9d-ae02-49fc-9b21-f813fa3989da · outbound

This paper cites Anomaly detection and localization in crowded scenes.

SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models Anomaly detection and localization in crowded scenes

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:35:55.425019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:35:54.450825Z digest=sha256:cd0b26dd5602de9190b0544a1e6373952e3426b84b89df0219e907f3e5a86177

Observation 966a1e2a-11cb-49eb-ba76-deda696ce2ee · outbound

This paper cites A survey of state of the art large vision language models: Alignment, benchmark, evaluations and challenges.

SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models A survey of state of the art large vision language models: Alignment, benchmark, evaluations and challenges

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:35:55.410065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:35:54.455316Z digest=sha256:93e537ca1817cfff8b14e8305a962d419866e652086860e497e300f1bcc4b102

Observation cbd70aa2-e512-4f34-9072-c2a7612c2f7a · outbound

This paper cites Revealing spatio-temporal evolution of urban visual environments with street view imagery.

SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models Revealing spatio-temporal evolution of urban visual environments with street view imagery

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:35:55.394176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:35:54.459643Z digest=sha256:4730771a63ccf1d7016037cd38ae20dd8b102b6cc5266ab0d1a744262564e1fb

Observation b69cc2d9-d977-4da0-bb2d-6b4248a6a762 · outbound

This paper cites A case for distributed multilevel storage infrastructure for visual surveillance in intelligent transportation networks.

SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models A case for distributed multilevel storage infrastructure for visual surveillance in intelligent transportation networks

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:35:55.379501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:35:54.464329Z digest=sha256:e392b7aab91d0538c295c09c525706ee879dce34b08048c3e87c8513764d08da

Observation 0b2df4de-ff4b-4be6-86bd-f90b0a3e9ece · outbound

This paper cites Mm-safetybench: A benchmark for safety evaluation of multimodal large language models.

SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models Mm-safetybench: A benchmark for safety evaluation of multimodal large language models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:35:55.364280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:35:54.469218Z digest=sha256:a6fc9cacec19aa8bd2aad2b19a8b70a999b136a95d0bcf180bd622372fb5fc50

Observation 2055165c-aa15-4c86-9c9d-dbb29ba1c03a · outbound

This paper cites Liu, Dingkang Yang, Yan Wang, Jing Liu, and Liang Song.

SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models Liu, Dingkang Yang, Yan Wang, Jing Liu, and Liang Song

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:35:55.349726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:35:54.473640Z digest=sha256:d7c1f4b763f92e65902b77e01aac86b24b020d3cf843928703c81de869f3fdd5

Observation 1afb5bc6-715b-4bfd-8e8b-8bbe36f872c7 · outbound

This paper cites Generalized video anomaly event detection: Systematic taxonomy and comparison of deep models.

SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models Generalized video anomaly event detection: Systematic taxonomy and comparison of deep models

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:35:55.335099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:35:54.477962Z digest=sha256:7b0023e2a3abf4555845cec8e500a5cd9407bad01365c594875dbc711c5b553b

Observation 72abb966-3b1a-4657-b69d-c7dbc5246c68 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pages 216–233.

SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pages 216–233

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T20:35:54.482413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:35:54.482413Z digest=sha256:a8b0212886153f5a07fc1225e8a8028096a84bd2de28a433093f50c16e837500

Observation 39187d1d-2899-4278-94ab-f53c03e634df · outbound

This paper cites Abnormal event detection at 150 fps in matlab.

SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models Abnormal event detection at 150 fps in matlab

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:35:55.311435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:35:54.486690Z digest=sha256:1f12b174968ffdf8ca3654b03ed8565965e23f2eb3ea3e3ae8125286a77dd7c3

Observation 35045b39-f3d0-4d7e-b990-49b91e6818e2 · outbound

This paper cites A revisit of sparse coding based anomaly detection in stacked rnn framework.

SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models A revisit of sparse coding based anomaly detection in stacked rnn framework

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:35:55.297474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:35:54.491721Z digest=sha256:9ed26f682c7bc9d117990e059fd99df127f3f36ff6c99fda1ecb29241e49d92d

Observation 93891901-1c52-4eba-ae0c-293f280fc4d4 · outbound

This paper cites Data sets, modeling, and decision making in smart cities: A survey.

SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models Data sets, modeling, and decision making in smart cities: A survey

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:35:55.283576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:35:54.496141Z digest=sha256:b2b9c4416bad1bec4d67d7048f1ef6a8dfa2952f7e33954ceae062571326c3a0

Observation a2f5efdf-269e-461b-a8b9-db50c14e8818 · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T20:35:54.500387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:35:54.500387Z digest=sha256:6a4ea3d5cda0b00d20b278cf39abbbbccabcb31e2564857f26293cebab72a0f0

Observation dc87103b-ee54-46de-8669-02c42f0da851 · outbound

This paper cites VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding.

SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T20:35:54.504822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:35:54.504822Z digest=sha256:fe90491d62e5d63f739ed380d19dd6b5d8d8cc17d4b2c3ca0e10d0f72ea0cc38

Observation 058c3887-9ac9-41f1-98a1-5eb098542237 · outbound

This paper cites Spatiotemporal anomaly detection using deep learning for real-time video surveillance.

SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models Spatiotemporal anomaly detection using deep learning for real-time video surveillance

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:35:55.270103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:35:54.509551Z digest=sha256:6c48eb969bbe579e3dd99bd0369bbf7b00ed03678a7dddf0ebe0f6ba7e991f14

Observation 93471766-3d1a-4cff-a46d-f969a82a7f7c · outbound

This paper cites A comprehensive analysis of real-time video anomaly detection methods for human and vehicular movement.

SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models A comprehensive analysis of real-time video anomaly detection methods for human and vehicular movement

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:35:55.255406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:35:54.513784Z digest=sha256:d162f9fe2a5cf1470715ac6223cd539e8a1e216338a0e8250098b34a9cf9e52d

Observation 9044befd-ecbe-4b6f-9cd3-39cd3421c0b7 · outbound

This paper cites Deep learning approaches for video-based anomalous activity detection.

SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models Deep learning approaches for video-based anomalous activity detection

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:35:55.241210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:35:54.518178Z digest=sha256:6797d155f41ba6953ba521984bf7e5d9aa773be72075f499845e4228028d4bac

Observation 984c4a15-2edf-45e2-8fcd-6ae83989f3a3 · outbound

This paper cites Sreenu and M.

SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models Sreenu and M

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:35:55.226627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:35:54.522402Z digest=sha256:b695ff7f78d6d40f93173ceeac3a504d3c8f82b646f59741201d3a840686387f

Observation 0301d947-9ba1-43b4-aee3-f3e614732e5f · outbound

This paper cites Real-world anomaly detection in surveillance videos.

SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models Real-world anomaly detection in surveillance videos

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:35:55.211797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:35:54.526812Z digest=sha256:34a5cd89cec5bc47d0a6d3dbf9ba97f6f7c792a1990359f6ea12588cf958e09d

Observation 22c5a9c3-0b60-4b6d-b098-63f62eb8d450 · outbound

This paper cites Video surveillance systems-current status and future trends.

SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models Video surveillance systems-current status and future trends

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:35:55.197674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:35:54.531125Z digest=sha256:0d58c2e7563106bfae6bd2a26104b544a4907f458207a21ab0048cf9c6c12a9f

Observation c6e80720-ec20-467b-bc97-974b2c9d474e · outbound

This paper cites Multimodal Needle in a Haystack: Benchmarking Long-Context Capability of Multimodal Large Language Models.

SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models Multimodal Needle in a Haystack: Benchmarking Long-Context Capability of Multimodal Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T20:35:54.535439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:35:54.535439Z digest=sha256:7147cd6877057551895f153f57c0a56d4a323801b0e19c7182949136f873f67b

Observation c0a856c2-9cd8-4e45-84f9-e27d9c14ff99 · outbound

This paper cites Exploring the Reasoning Abilities of Multimodal Large Language Models (MLLMs): A Comprehensive Survey on Emerging Trends in Multimodal Reasoning.

SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models Exploring the Reasoning Abilities of Multimodal Large Language Models (MLLMs): A Comprehensive Survey on Emerging Trends in Multimodal Reasoning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T20:35:54.540185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:35:54.540185Z digest=sha256:31cc21ce6a06c4647ce34e84aaca0b59ff887cc894e26c914d298acd691a367e

Observation 530bff89-0940-4346-be89-1238686f1c15 · outbound

This paper cites Longvlm: Efficient long video understanding via large language models.

SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models Longvlm: Efficient long video understanding via large language models

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:35:55.183761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:35:54.544652Z digest=sha256:6c1151128ee9923d0f46e0401da900528d4c0d3e9885a7e0c1f424b25e4a185a

Observation cf851d3b-a9d3-4cdd-b709-32609fc23f28 · outbound

This paper cites Open-vocabulary video anomaly detection.

SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models Open-vocabulary video anomaly detection

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:35:55.169057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:35:54.548822Z digest=sha256:37b536078efe4bdf9b903aef596f0ee51581d032b28d1ab5c269914113da013e

Observation bb2a0f43-1aa0-480a-b3db-3a4f28a86858 · outbound

This paper cites Next-qa: Next phase of question- answering to explaining temporal actions.

SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models Next-qa: Next phase of question- answering to explaining temporal actions

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:35:55.154028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:35:54.553297Z digest=sha256:08d4769e377a5b0cefaaf62eda869589a7333552d17a99e17b7537164b7b274d

Observation aaa5798e-b23e-451a-9175-d945872969d5 · outbound

This paper cites Funqa: Towards surprising video comprehension.

SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models Funqa: Towards surprising video comprehension

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:35:55.139360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:35:54.557456Z digest=sha256:764c3b740187a3c007fa8d8af3955f3f4d58e70741feeffd19f59402e6017099

Observation b5f1bb64-ec22-434c-990f-6dad8a9b7069 · outbound

This paper cites Video structured description technology based intelligence analysis of surveillance videos for public security applications.

SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models Video structured description technology based intelligence analysis of surveillance videos for public security applications

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:35:55.125668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:35:54.562164Z digest=sha256:de0d4aa30cd6f5716913d6e9d3d41390c1374a05583318d5677fbd3133bfd809

Observation bc85f8e4-ca3d-4284-be05-3d32b96b39f1 · outbound

This paper cites Qwen2.5 Technical Report.

SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models Qwen2.5 Technical Report

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T20:35:54.566643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:35:54.566643Z digest=sha256:0aecbbc534efd6aa09eb99ee8986b14d78c64e9523e259ff6f64ae2e0bf32d63

Observation de00a680-c873-4025-903c-fda999fe2241 · outbound

This paper cites Svbench: A benchmark with temporal multi-turn dialogues for streaming video understanding.

SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models Svbench: A benchmark with temporal multi-turn dialogues for streaming video understanding

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T20:35:54.571037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:35:54.571037Z digest=sha256:8136f0289b7abd78a2ffa76463d49400dd740479a197bc311cc4079ebdc96bcd

Observation a8c32b1e-3de4-436f-ae63-f2edb5426ee5 · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T20:35:54.575145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:35:54.575145Z digest=sha256:30576cf79ceab83fc572f36f8d35d7be84d1f804f0663d6bfa453375775a172c

Observation d622a28d-c4c1-44e6-8dda-9fcb9ef1e7c2 · outbound

This paper cites Surveillance video-and-language understanding: from small to large multimodal models.

SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models Surveillance video-and-language understanding: from small to large multimodal models

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:35:55.110627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:35:54.579704Z digest=sha256:160ad50a64d9b07201f5ce6be1084e3dfd369b74273ece06b451ca0d986d1e1e

Observation 70340733-8589-4245-8597-4c503cbd8d26 · outbound

This paper cites Towards surveillance video-and-language understanding: New dataset baselines and challenges.

SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models Towards surveillance video-and-language understanding: New dataset baselines and challenges

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:35:55.094492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:35:54.583900Z digest=sha256:83f748dec6f62d9b9e80d2f0573fe3f29790186773bdf595352299960733a814

Observation b083faaf-25b0-4dcd-94d6-7b2fbd5ce6c4 · outbound

This paper cites Har- nessing large language models for training-free video anomaly detection.

SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models Har- nessing large language models for training-free video anomaly detection

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:35:55.079581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:35:54.588215Z digest=sha256:97fa493b91d87b4f59721ac0bcad0d76f61b4783f582929ca080de392dea45e0

Observation c421cddf-bb31-478e-a1b2-e2bf9fc59f63 · outbound

This paper cites VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding.

SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-15T20:35:54.592357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:35:54.592357Z digest=sha256:1ea45cae76de5c14478a21ad1bd16362026142246894bd01e4531017e78fd739

Observation 0d21578a-3594-46d0-a7ff-78af9574f860 · outbound

This paper cites Llava-next: A strong zero-shot video understanding model, April 2024.

SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models Llava-next: A strong zero-shot video understanding model, April 2024

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:35:55.063966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:35:54.596967Z digest=sha256:8dc4fb25ea725c7803ca279e73e36b662ae10fbe739532963fe4855c38f4f194

Observation c50dad0f-113a-4b9e-be24-b35dbb991139 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-15T20:35:54.601567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:35:54.601567Z digest=sha256:bbeed233baf9d68e8b47f1efdf125a550dbd53742e71f33254a9e768be822d01

Observation 4eba873f-5cad-4182-9561-273a4df7b66b · outbound

This paper cites Anomalynet: An anomaly detection network for video surveillance.

SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models Anomalynet: An anomaly detection network for video surveillance

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:35:55.048972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:35:54.606198Z digest=sha256:57eb459ec5784ca44a85eaa8c6bed0a4960869b79987739271f8e00251941f3d

Observation 7632f66e-b6dd-4a48-954d-ff2c0ecde234 · outbound

This paper cites 2 0 1 8 -0 3 -1 5 _ 1 0 _ 1 3 1.

SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models 2 0 1 8 -0 3 -1 5 _ 1 0 _ 1 3 1

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:35:55.032878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:35:54.610338Z digest=sha256:e15d64a67106c2b71519c4b9705fda6a11d5a7f05cc50b2c304a91701ac979a7

Pith citing papers

Observation 4b115999-15e5-4575-8c98-d5ceb4a12853 · inbound

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models cites this paper.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T19:06:23.701939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:06:23.701939Z digest=sha256:0c013a69e6861476c0b25097bbc3be8e3ede8451010822e1c119ed32923e00ca

Observation 239e5431-98ec-4e19-9a00-164b5ca3ca63 · inbound

ProcObject-10K: Benchmarking Object-Centric Procedural Understanding in Instructional Videos cites this paper.

ProcObject-10K: Benchmarking Object-Centric Procedural Understanding in Instructional Videos SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:08:56.600848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-17T03:05:42.773193Z digest=sha256:391d47dd85bc350bec91787010ead19f2f242af17409c3dc5b295719949209e1

Observation b1c66789-5806-49e3-8774-61e562d6ebc8 · inbound

CrashSight: A Phase-Aware, Infrastructure-Centric Video Benchmark for Traffic Crash Scene Understanding and Reasoning cites this paper.

CrashSight: A Phase-Aware, Infrastructure-Centric Video Benchmark for Traffic Crash Scene Understanding and Reasoning SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:30:57.908140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T17:09:57.040135Z digest=sha256:8e057b06dc3b2affe7e5ccc64061d8c63e8e5d771dddbe117f89e1e5d45cbc06

Observation 3eb130d2-a3de-47cf-ae0d-9389acf53a4b · inbound

MAVEN: A Multi-stage Agentic Annotation Pipeline for Video Reasoning Tasks cites this paper.

MAVEN: A Multi-stage Agentic Annotation Pipeline for Video Reasoning Tasks SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-22T07:36:13.989290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T07:36:09.562362Z digest=sha256:e1bf3af7fe49bf5cd0a0f45d3e058817f76ed507f313337af30c87bc7178b27e

Observation 29ab977d-45d4-4862-a0a2-88ef6011c0f5 · inbound

MAVEN: A Multi-stage Agentic Annotation Pipeline for Video Reasoning Tasks cites this paper.

MAVEN: A Multi-stage Agentic Annotation Pipeline for Video Reasoning Tasks SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-12T16:17:21.074502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T16:17:21.074502Z digest=sha256:c78ebeeb3f6eb5b722daab754ee8202b3cbbb3a04195dbe605bd14ea4aea9e87

Observation d59e0931-11dc-4425-8c61-dece5328e698 · inbound

MetaphorVU: Towards Metaphorical Video Understanding cites this paper.

MetaphorVU: Towards Metaphorical Video Understanding SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:44:01.323382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-29T22:43:20.101830Z digest=sha256:bf473692af5bc423e93afe3533f4e71b91f130ce8902287d0de53c4b5df5e063

Observation 149f5536-42ec-4d19-9965-3ea8ab966e7b · inbound

From Detection to Understanding: TAR and TAR-Bench for Multi-Task Traffic Anomaly Reasoning cites this paper.

From Detection to Understanding: TAR and TAR-Bench for Multi-Task Traffic Anomaly Reasoning SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-14T04:15:47.263691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:15:47.263691Z digest=sha256:8e9a3145bdf82b23b3a3bc395c2ebd92dd33cd4a6194062b651e284d2d7a558e