Pith. sign in

Paper Citation Record · LEDGER

A Benchmark for Crime Surveillance Video Analysis with Large Models

As of 9 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 1 inbound Pith citation observation for arXiv:2502.09325.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.09325 v1

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T21:57:23.605752Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T12:34:00.291412Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T12:34:01.473793Z

Reference resolution

23 of 23 outbound references displayed

  • verified exact0
  • verified fuzzy12
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 260d3569-4a70-4f89-a4e4-1849d1b19b18 · outbound

This paper cites Real-world anomaly detection in surveillance videos,.

A Benchmark for Crime Surveillance Video Analysis with Large Models Real-world anomaly detection in surveillance videos,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T21:57:23.967451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T21:57:23.499488Z digest=sha256:8457c7e5255bf206de524850da13419deb215c97bf598e1e1c6714e80a768bee

Observation 9fb02311-00a7-4e94-b0b5-eddf21228def · outbound

This paper cites End-to-end dense video captioning with parallel decoding,.

A Benchmark for Crime Surveillance Video Analysis with Large Models End-to-end dense video captioning with parallel decoding,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T21:57:23.952942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T21:57:23.505187Z digest=sha256:cb030b4784f063807ec41459171e5924caf25056f76ab7190152a5b795d85191

Observation 0208e327-6542-4ef8-a3f8-639f7a0b3570 · outbound

This paper cites Batchnorm-based weakly supervised video anomaly detection,.

A Benchmark for Crime Surveillance Video Analysis with Large Models Batchnorm-based weakly supervised video anomaly detection,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T21:57:23.937711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T21:57:23.510266Z digest=sha256:a2c7277068031038bd900a05b7080d182ad4ba13fc369c73017669f221eb1d15

Observation e27c7e7d-dadf-42b1-a397-cd55746f579b · outbound

This paper cites Negative sample matters: A renaissance of metric learning for temporal grounding,.

A Benchmark for Crime Surveillance Video Analysis with Large Models Negative sample matters: A renaissance of metric learning for temporal grounding,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T21:57:23.922758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T21:57:23.515602Z digest=sha256:cda2d664e0d751a4904d30bd20cf4b169810d60785ac99ce928d637459ba9a93

Observation 679274b7-5041-47db-bbad-dbe7f88f2517 · outbound

This paper cites Qwen Technical Report.

A Benchmark for Crime Surveillance Video Analysis with Large Models Qwen Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T21:57:23.520432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T21:57:23.520432Z digest=sha256:eae7c750a011e23f5613a0f0853aa7f0243ecc3302cbe2b9eaad45f1c073285b

Observation 58cde45a-7803-4029-8cba-4a3418c7745b · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

A Benchmark for Crime Surveillance Video Analysis with Large Models LLaVA-OneVision: Easy Visual Task Transfer

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T21:57:23.526197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T21:57:23.526197Z digest=sha256:0590abe2f2583899fe9c7c960edc5f23db0ed271973c85ab426566ae4783251d

Observation 841be4f3-bb80-498c-9357-4296f1561561 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks,.

A Benchmark for Crime Surveillance Video Analysis with Large Models Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T21:57:23.907277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T21:57:23.531972Z digest=sha256:adbc0e6ae3b1994794624b9c82400bf893af632798013bf3072aaae53ff8e2b7

Observation b7b05c2c-a212-4279-ad41-ef48d3472378 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

A Benchmark for Crime Surveillance Video Analysis with Large Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T21:57:23.536528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T21:57:23.536528Z digest=sha256:fa3b9ca22ee48a97a5323871d9e01170c521201d43346851a01f1fb7ffbec38e

Observation 23d7d4a6-a240-4ad5-96c4-b6711ee42ddc · outbound

This paper cites Learning transferable visual models from natural language supervision,.

A Benchmark for Crime Surveillance Video Analysis with Large Models Learning transferable visual models from natural language supervision,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T21:57:23.889826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T21:57:23.541353Z digest=sha256:ba2dc57d28a9cdfd40e9fb7409e25787cde68091a1b2f764b63702b962b7af11

Observation 2a63086e-d396-48cb-8ed7-f0352f29fa7c · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

A Benchmark for Crime Surveillance Video Analysis with Large Models MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T21:57:23.546071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T21:57:23.546071Z digest=sha256:2e9e8de7af62846343fe7faa2c8e2273f127f2f0c0d2a5d94c2842f64b1e8205

Observation 150d137c-bdca-4a4b-b301-00436d5484cd · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

A Benchmark for Crime Surveillance Video Analysis with Large Models MMBench: Is Your Multi-modal Model an All-around Player?

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T21:57:23.551182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T21:57:23.551182Z digest=sha256:491587ab6fb2c3d82fbce8020a864947e3a5ee940722f2afed309c5927a7388d

Observation ccd52aed-5402-4615-9018-e7e59f80c64b · outbound

This paper cites MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding.

A Benchmark for Crime Surveillance Video Analysis with Large Models MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T21:57:23.556854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T21:57:23.556854Z digest=sha256:3b016f5bbf64eaf093bd9dfd10910343c6f4d1a07f1a792f1b7317d125f8c7dc

Observation 2700f3e8-dc5e-4b0b-a31a-6eb540e83db8 · outbound

This paper cites Mvbench: A comprehen- sive multi-modal video understanding benchmark,.

A Benchmark for Crime Surveillance Video Analysis with Large Models Mvbench: A comprehen- sive multi-modal video understanding benchmark,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T21:57:23.873404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T21:57:23.561397Z digest=sha256:b11e031b9bd22754903a71109f4c8f7dabb820a58b06dab708d8c3d7a9901738

Observation a56f90a3-807d-4a24-b421-f909a039c61a · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

A Benchmark for Crime Surveillance Video Analysis with Large Models Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T21:57:23.566060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T21:57:23.566060Z digest=sha256:4da21ffc2f71446b3c4308edfebf22e038beb853843c4cf65a7572244b7073c1

Observation eef0a864-8074-42b3-ad52-ace1fe4c0f51 · outbound

This paper cites Taisu: A 166m large-scale high-quality dataset for chinese vision-language pre- training,.

A Benchmark for Crime Surveillance Video Analysis with Large Models Taisu: A 166m large-scale high-quality dataset for chinese vision-language pre- training,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T21:57:23.857864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T21:57:23.570422Z digest=sha256:a94924a4a4af0a4b103c6af86308c5389569cecb61cd39221f504ba721044ff8

Observation 61a203ac-3c54-4d1c-ba9f-c341e539c17d · outbound

This paper cites Not only look, but also listen: Learning multimodal violence detection under weak supervision,.

A Benchmark for Crime Surveillance Video Analysis with Large Models Not only look, but also listen: Learning multimodal violence detection under weak supervision,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T21:57:23.842682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T21:57:23.574833Z digest=sha256:97851edbf5f9774d1369c1e38bd70f832867edd46ca2d772c363063e7fef085c

Observation bfddacc4-5271-4d89-b7a9-626a0f7db05c · outbound

This paper cites Towards surveillance video-and-language understanding: New dataset baselines and challenges,.

A Benchmark for Crime Surveillance Video Analysis with Large Models Towards surveillance video-and-language understanding: New dataset baselines and challenges,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T21:57:23.827753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T21:57:23.578853Z digest=sha256:510285e103770a418e21da9f8fa745fbb14f6cfabfb17496c2e7bbdc76df5a09

Observation 6a22cef7-fb11-42e3-ac22-2787ee1981d0 · outbound

This paper cites GPT-4 Technical Report.

A Benchmark for Crime Surveillance Video Analysis with Large Models GPT-4 Technical Report

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T21:57:23.583300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T21:57:23.583300Z digest=sha256:bf8351c8316a30fb006b2ceddd1106da8847d7f41b9ca76c3c2eb76abddc83c2

Observation baad31a9-9e0c-4e4d-89c1-04f6e3dfe3ca · outbound

This paper cites Quo vadis, action recognition? a new model and the kinetics dataset,.

A Benchmark for Crime Surveillance Video Analysis with Large Models Quo vadis, action recognition? a new model and the kinetics dataset,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T21:57:23.587747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T21:57:23.587747Z digest=sha256:ae3b836139d6cafe7fd8a4a08195a9f9bd4d543d1bb48c14b08624efd41eb7f6

Observation 863f404e-4d47-460d-99e7-f1cb722d6664 · outbound

This paper cites Learning spatiotemporal features with 3d convolu- tional networks,.

A Benchmark for Crime Surveillance Video Analysis with Large Models Learning spatiotemporal features with 3d convolu- tional networks,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T21:57:23.802040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T21:57:23.592363Z digest=sha256:eabdd3d4494e7368aa4b75b28125e1f270951be13d825d65d6198d044eb390d9

Observation 824f3a7a-d6e4-4750-85cd-ff501340ede9 · outbound

This paper cites DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter.

A Benchmark for Crime Surveillance Video Analysis with Large Models DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T21:57:23.596697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T21:57:23.596697Z digest=sha256:d3fa7af4e97f4388c5ee05f40b1b6b58f9da285ec11fc1eeb36851c7f36f2b42

Observation f5c2967f-63a1-4b81-b0b2-66236210753d · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation,.

A Benchmark for Crime Surveillance Video Analysis with Large Models Bleu: a method for automatic evaluation of machine translation,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T21:57:23.601536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T21:57:23.601536Z digest=sha256:19b131e7aabad55c31a450afc82b02dbd1d79b9fefc8e767e1d9bbbe64d185b7

Observation 7d9ef54e-ff4f-4fda-b658-04a72035301b · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering,.

A Benchmark for Crime Surveillance Video Analysis with Large Models Making the v in vqa matter: Elevating the role of image understanding in visual question answering,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T21:57:23.776354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T21:57:23.605752Z digest=sha256:369fa623bf08a17da9b29dc2ad0a4b3765e48a7006b90b31364b848aa711271e

Pith citing papers

Observation 5c9d601a-255d-4600-98bf-4cd34de9293f · inbound

The Evolution of Video Anomaly Detection: A Unified Framework from DNN to MLLM cites this paper.

The Evolution of Video Anomaly Detection: A Unified Framework from DNN to MLLM A Benchmark for Crime Surveillance Video Analysis with Large Models

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-08-06T12:34:01.477068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:34:00.291412Z digest=sha256:2e375b8b44b0982b38ebd96c9cfca9a511f6787b6798da84e93295636c207cd3