Pith. sign in

Paper Citation Record · LEDGER

VEU-Bench: Towards Comprehensive Understanding of Video Editing

As of 18 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 2 inbound Pith citation observations for arXiv:2504.17828.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.17828 v1

Coverage vector

measured 53 of 53 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:50:29.935626Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:08:41.965687Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-13T21:18:17.187055Z

Reference resolution

53 of 53 outbound references displayed

  • verified exact2
  • verified fuzzy21
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation df1542c8-2c67-4624-a15c-24caeb6f6067 · outbound

This paper cites The anatomy of video editing: A dataset and benchmark suite for ai-assisted video editing.

VEU-Bench: Towards Comprehensive Understanding of Video Editing The anatomy of video editing: A dataset and benchmark suite for ai-assisted video editing

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:50:30.681590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:50:29.719063Z digest=sha256:2c3431f0e74c3b2ed1955fe79028df3ba526a7d719985ca40faf17807afaabe8

Observation beb144f7-db10-4bfe-8dd2-a481b1edba75 · outbound

This paper cites Theory of film practice.

VEU-Bench: Towards Comprehensive Understanding of Video Editing Theory of film practice

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:50:30.663217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:50:29.724028Z digest=sha256:99e1c15d3c39358cb487149242762694d9ec7289276961adf56802f91313b543

Observation 5595545d-4e5a-44c4-ab37-17a59303d935 · outbound

This paper cites Reframe Anything: LLM Agent for Open World Video Reframing.

VEU-Bench: Towards Comprehensive Understanding of Video Editing Reframe Anything: LLM Agent for Open World Video Reframing

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T10:50:29.728425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:50:29.728425Z digest=sha256:422c4a931520ea5ef7d9815bcda419f129551caf982a38f21970525eb9bc32cb

Observation 2a79a1f4-8a33-416d-a960-f09d74109829 · outbound

This paper cites Match cutting: Finding cuts with smooth visual transitions.

VEU-Bench: Towards Comprehensive Understanding of Video Editing Match cutting: Finding cuts with smooth visual transitions

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:50:30.649683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:50:29.733054Z digest=sha256:9b94ea4039cff8bb846ee65decc7fa9930ea34ab2efac0007344e08350475018

Observation c7dfa941-13be-4f45-b591-a253d06c7c01 · outbound

This paper cites ShareGPT4Video: Improving Video Understanding and Generation with Better Captions.

VEU-Bench: Towards Comprehensive Understanding of Video Editing ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T10:50:29.737303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:50:29.737303Z digest=sha256:4cfc18328b6eaecd8e11fb0ce3fa9c289f3da9b5d3b4ced57df748c29f4894de

Observation b3db770a-b655-4bb7-aa36-8e01199c7ae3 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

VEU-Bench: Towards Comprehensive Understanding of Video Editing How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T10:50:29.741646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:50:29.741646Z digest=sha256:e001894f5d4f8c36ce4696dafc60a8594bb24aece207f55a4cf93e38b74b5793

Observation 3351857e-fa43-432f-bef0-26ec0cc6b496 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

VEU-Bench: Towards Comprehensive Understanding of Video Editing VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T10:50:29.746377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:50:29.746377Z digest=sha256:2c5412dd028be5b331e0b2fdf824399c928046a8b4c9391c23e25f50a9584eef

Observation a754fa32-211a-45d2-8038-8f7b487133f4 · outbound

This paper cites The technique of film and video editing: history, theory, and practice.

VEU-Bench: Towards Comprehensive Understanding of Video Editing The technique of film and video editing: history, theory, and practice

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:50:30.636715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:50:29.750248Z digest=sha256:a24bb580be9002e15d98720ce5de871f48c1d82b0b2cf221538f41fb692b0285

Observation 3a9f687c-a050-4582-a0ea-c9b0cf0357e5 · outbound

This paper cites Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos.

VEU-Bench: Towards Comprehensive Understanding of Video Editing Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T10:50:29.754241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:50:29.754241Z digest=sha256:e6418cfdfd7af6069b69a533f43e42978dede3d4a45cf33dd9f49ac1380ea0eb

Observation 0c919074-4ccd-4001-8559-1770e74ba05e · outbound

This paper cites Automatic Non-Linear Video Editing Transfer.

VEU-Bench: Towards Comprehensive Understanding of Video Editing Automatic Non-Linear Video Editing Transfer

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-16T10:50:30.323439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:50:29.758395Z digest=sha256:de8184935f562f7684068987733ab9f59076e68926762f19f60a7bb0e93bb957

Observation edd27ce0-1296-4e94-a87b-ef66ac6f0892 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

VEU-Bench: Towards Comprehensive Understanding of Video Editing Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T10:50:29.763000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:50:29.763000Z digest=sha256:e66d69f6e4f2278aa0218c39303f7a0d5eb085ece2fff49f00ba0e304953e27b

Observation 8ed59f42-0b90-4f8f-9ef0-30986f331394 · outbound

This paper cites Edit3K: Universal Representation Learning for Video Editing Components.

VEU-Bench: Towards Comprehensive Understanding of Video Editing Edit3K: Universal Representation Learning for Video Editing Components

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T10:50:29.767070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:50:29.767070Z digest=sha256:ce959b547317a2b145313e92bb75692387a0a5387dc8458c2cdd83c58e0484af

Observation d4005da3-0840-4488-aed2-b1e65b38636b · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

VEU-Bench: Towards Comprehensive Understanding of Video Editing LoRA: Low-Rank Adaptation of Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T10:50:29.771160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:50:29.771160Z digest=sha256:1fc49f72efa2a5affe974432e2d8d77c69cc7eb2c70b3619db1d45a0796b825d

Observation e261ab51-bb24-4474-9371-e2408318bfbf · outbound

This paper cites Vtimellm: Empower llm to grasp video moments.

VEU-Bench: Towards Comprehensive Understanding of Video Editing Vtimellm: Empower llm to grasp video moments

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:50:30.623631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:50:29.775357Z digest=sha256:1840c88bec7543d72ee23ed4432a9a1daca7d31bc6bf3275ccf0fdb79e90db8c

Observation 73678a1a-6d14-422b-8d89-2a769e4d20a2 · outbound

This paper cites Movienet: A holistic dataset for movie understanding.

VEU-Bench: Towards Comprehensive Understanding of Video Editing Movienet: A holistic dataset for movie understanding

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:50:30.610505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:50:29.779201Z digest=sha256:ba437940fbc1408b2c8f01a593bee32860bfe941f56d5f604ef9119eb129da74

Observation 68d33af1-7e79-4102-a671-fd5fade83331 · outbound

This paper cites Study of vari- ous video annotation techniques.

VEU-Bench: Towards Comprehensive Understanding of Video Editing Study of vari- ous video annotation techniques

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:50:30.597848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:50:29.783094Z digest=sha256:b647cab8cd6fb18cfb0bfb253e53a68a35650434f1ddd8f950168164d0ae4ddb

Observation 4a8e3af3-20fc-4ea2-a60b-8b3d86b630f6 · outbound

This paper cites Automatic color scheme extraction from movies.

VEU-Bench: Towards Comprehensive Understanding of Video Editing Automatic color scheme extraction from movies

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:50:30.585033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:50:29.786897Z digest=sha256:a8dec7f28382fe57c63f28afb9c693cbee36e38d09e6e764674d0ae807672a21

Observation e89586d6-630c-443f-b1bc-0021e5cc66f7 · outbound

This paper cites Frame Order Matters: A Temporal Sequence-Aware Model for Few-Shot Action Recognition.

VEU-Bench: Towards Comprehensive Understanding of Video Editing Frame Order Matters: A Temporal Sequence-Aware Model for Few-Shot Action Recognition

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T10:50:29.790543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:50:29.790543Z digest=sha256:67e9e5ee99de1075fa3813116235ef917dd761cd7e54420fe17f2a5b7b8b0c78

Observation 4b0d1287-fdb6-4ba6-9f4e-21bb5ea600d7 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

VEU-Bench: Towards Comprehensive Understanding of Video Editing LLaVA-OneVision: Easy Visual Task Transfer

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T10:50:29.794977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:50:29.794977Z digest=sha256:f5d0e0e809fe000f4bfc0456363e599f1865347043918d60676fe6ce9786fb53

Observation 2d400170-9e9e-4a89-9a06-69b9794aa5bf · outbound

This paper cites Mvbench: A comprehensive multi-modal video understand- ing benchmark.

VEU-Bench: Towards Comprehensive Understanding of Video Editing Mvbench: A comprehensive multi-modal video understand- ing benchmark

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:50:30.572030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:50:29.799069Z digest=sha256:bb0b030acbaefe68abb5ca194a3f428799b1d47c2ad7c5f44cc60a5d89d5b5b4

Observation 8f0db0df-d6d7-4151-967e-e4e583c20c2c · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

VEU-Bench: Towards Comprehensive Understanding of Video Editing Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T10:50:29.803058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:50:29.803058Z digest=sha256:b661cdf1eb56fa955e74d2ab88126fa122275166f222473b4dba4e1805c071b4

Observation 7aea31df-fcd5-40dc-a327-bff6b98039b0 · outbound

This paper cites Vila: On pre-training for visual language models.

VEU-Bench: Towards Comprehensive Understanding of Video Editing Vila: On pre-training for visual language models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:50:30.559138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:50:29.807109Z digest=sha256:43aa1f30347c3c914ca258e2f977762ff42775ce9fd4dada0759fc4818d4b04f

Observation 73917e7a-15a5-4439-9230-725c43bc659a · outbound

This paper cites OmniCLIP: Adapting CLIP for Video Recognition with Spatial-Temporal Omni-Scale Feature Learning.

VEU-Bench: Towards Comprehensive Understanding of Video Editing OmniCLIP: Adapting CLIP for Video Recognition with Spatial-Temporal Omni-Scale Feature Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T10:50:29.811027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:50:29.811027Z digest=sha256:4de0b59835b83b3f286bb14478109c4e66bf5bc9503caead4d9e74357d1f71eb

Observation 468a9f86-0350-4cb1-827d-ae8a6ee09610 · outbound

This paper cites TempCompass: Do Video LLMs Really Understand Videos?.

VEU-Bench: Towards Comprehensive Understanding of Video Editing TempCompass: Do Video LLMs Really Understand Videos?

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T10:50:29.815623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:50:29.815623Z digest=sha256:9878d081a24e9ab40e515685282a00a1aad07a0c67925a5a553a40907e3c5557

Observation c9a6248b-b904-44db-aff0-78109d8466af · outbound

This paper cites Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution.

VEU-Bench: Towards Comprehensive Understanding of Video Editing Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T10:50:29.819644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:50:29.819644Z digest=sha256:2e6735f1b313d3d083da2c797f081324bc2a57dd9c3db2904e014d6035473a7c

Observation 4d84cdba-c51d-4582-84fe-e8517857df54 · outbound

This paper cites Decoupled Weight Decay Regularization.

VEU-Bench: Towards Comprehensive Understanding of Video Editing Decoupled Weight Decay Regularization

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T10:50:29.824192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:50:29.824192Z digest=sha256:f0e47965b05c4b14b4777f61ed5c6ad8d7b0ea60665b286f30e774cabd76d145

Observation dc7bf747-f927-41c1-bf94-44379ec167ec · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

VEU-Bench: Towards Comprehensive Understanding of Video Editing Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T10:50:29.828630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:50:29.828630Z digest=sha256:8f6566e8a19b17201cdf9098fed6a7170bdf7ac3514c36f99d391a0ad9c5221b

Observation 6cd29212-5daf-4d76-96ad-008046d093cd · outbound

This paper cites Video-chatgpt: Towards detailed video understanding via large vision and language models.

VEU-Bench: Towards Comprehensive Understanding of Video Editing Video-chatgpt: Towards detailed video understanding via large vision and language models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:50:30.546166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:50:29.832573Z digest=sha256:64b59dff5458a23734ae5d9a6b1644feebb8fe65c377b6617ec8cd22b190d55c

Observation f7beb7ce-554f-4a46-aff1-e5cb0db0f250 · outbound

This paper cites Film language: A semiotics of the cinema.

VEU-Bench: Towards Comprehensive Understanding of Video Editing Film language: A semiotics of the cinema

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:50:30.533118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:50:29.836635Z digest=sha256:e3b7342f811f8be8084f05db567bab694850ca6f87c31826316216c67d56de47

Observation 6fc9c3f4-7bf4-4c71-a0c4-9899432b3a79 · outbound

This paper cites Hello gpt-4o.

VEU-Bench: Towards Comprehensive Understanding of Video Editing Hello gpt-4o

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:50:30.520155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:50:29.840466Z digest=sha256:8de76144c574adee50506fe23ff3367c04014caaa11e6a1c8ef60f74e8f83008

Observation 428a71d4-79e1-4026-8e88-f4986202a2c6 · outbound

This paper cites Learning to cut by watch- ing movies.

VEU-Bench: Towards Comprehensive Understanding of Video Editing Learning to cut by watch- ing movies

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:50:30.506467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:50:29.844481Z digest=sha256:eb617cc8f7b0a07f3b72a9c1b4236089a835c24609bd16a14c481eedaaf87046

Observation 65b45f01-ef13-46c5-bce8-816fb3f72caf · outbound

This paper cites Moviecuts: A new dataset and benchmark for cut type recognition.

VEU-Bench: Towards Comprehensive Understanding of Video Editing Moviecuts: A new dataset and benchmark for cut type recognition

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:50:30.493170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:50:29.848450Z digest=sha256:999e93df926268381de6d5415e67ee54c66c3255e7751aad60ddee2b9b7046ac

Observation 541fefd0-a7db-4db1-859d-e0fc4c8a4870 · outbound

This paper cites Perception test: A diagnostic benchmark for multimodal video models.

VEU-Bench: Towards Comprehensive Understanding of Video Editing Perception test: A diagnostic benchmark for multimodal video models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:50:30.478920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:50:29.852455Z digest=sha256:9aac331a49af654e35236e8696d0a100d75ec7fb73bc0b510385fec13e874c0d

Observation 53ddda76-f0fc-4a95-a994-a3dc9da77f97 · outbound

This paper cites Autotran- sition: Learning to recommend video transition effects.

VEU-Bench: Towards Comprehensive Understanding of Video Editing Autotran- sition: Learning to recommend video transition effects

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:50:30.463313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:50:29.856455Z digest=sha256:c4eec50523569e3d000b014eb48f305558552e091966b925f0f04f2a2b4debe8

Observation f1e1463d-c381-4a44-ac2b-d25a5958e5ac · outbound

This paper cites Harnessing ai for augmenting creativity: Appli- cation to movie trailer creation.

VEU-Bench: Towards Comprehensive Understanding of Video Editing Harnessing ai for augmenting creativity: Appli- cation to movie trailer creation

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:50:30.448755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:50:29.860575Z digest=sha256:ca9b344aeeef3908c4704d2c6504ae6f8e76f43fae13e228706cc8b27197fe7b

Observation 7ad4a3b3-9abe-4f14-ac99-62a99f5824c9 · outbound

This paper cites Diffusion Model-Based Video Editing: A Survey.

VEU-Bench: Towards Comprehensive Understanding of Video Editing Diffusion Model-Based Video Editing: A Survey

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T10:50:29.864489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:50:29.864489Z digest=sha256:33e3cf067873b79139b07be8f19b6951b7516e3e56080b7005bf8c324f2a0539

Observation e8576dfe-7d15-40e7-adcc-918118619203 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

VEU-Bench: Towards Comprehensive Understanding of Video Editing Gemini: A Family of Highly Capable Multimodal Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-16T10:50:29.868612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:50:29.868612Z digest=sha256:5e3a65951deb8aedb9861f8ccd1fce96ffb96756cf3a6b7e36fd0f2c519c803b

Observation c53bf548-3e40-45fe-b2c8-ff5bc67b71da · outbound

This paper cites Converting video formats with ffmpeg.

VEU-Bench: Towards Comprehensive Understanding of Video Editing Converting video formats with ffmpeg

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:50:30.434949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:50:29.872819Z digest=sha256:3e60cc712071159c7223fe1ff7218be762e8d4fd933e92a1bdab3f04fa5be9f7

Observation e317ea7e-28aa-4d83-87a0-e22badb7d63b · outbound

This paper cites Movie lens: Discovering and characterizing editing patterns in the anal- ysis of short movie sequences.

VEU-Bench: Towards Comprehensive Understanding of Video Editing Movie lens: Discovering and characterizing editing patterns in the anal- ysis of short movie sequences

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:50:30.421455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:50:29.876820Z digest=sha256:6a307216ebfb6e9f42d98b9a22ea55239da4badcf83bd4baf858a13a88c952eb

Observation 83adc78b-320e-415b-9731-e83f95bff8ac · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

VEU-Bench: Towards Comprehensive Understanding of Video Editing Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T10:50:29.880662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:50:29.880662Z digest=sha256:d448b90f06b88c8fb727f610cbd71bd9518d4cc30bce5199e021dc5a2a18329d

Observation cb1685b3-a5b2-45f8-b21b-0ab8809b8b54 · outbound

This paper cites LVBench: An Extreme Long Video Understanding Benchmark.

VEU-Bench: Towards Comprehensive Understanding of Video Editing LVBench: An Extreme Long Video Understanding Benchmark

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T10:50:29.884649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:50:29.884649Z digest=sha256:8235182393c219126b627aee6078f16fe355f70073f613ed83ab2e48e7d558c2

Observation 45929c96-d7fa-48d0-ab14-aa397f109fc6 · outbound

This paper cites Editing techniques with final cut pro.

VEU-Bench: Towards Comprehensive Understanding of Video Editing Editing techniques with final cut pro

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:50:30.407011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:50:29.888880Z digest=sha256:3bb63004100582a1e61cceff32a5617a4cd8e81b48aa9de4bb472471aac80662

Observation 8e86f818-dcd2-4741-a4ce-57d825c415fe · outbound

This paper cites LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding.

VEU-Bench: Towards Comprehensive Understanding of Video Editing LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T10:50:29.893249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:50:29.893249Z digest=sha256:ec6c733c64f97ee298695222e624aa89a35a60ce6a61012acd41901d48d05d97

Observation 16a12aa8-eff9-408a-b1d3-77ba01e12531 · outbound

This paper cites Zero-Shot Long-Form Video Understanding through Screenplay.

VEU-Bench: Towards Comprehensive Understanding of Video Editing Zero-Shot Long-Form Video Understanding through Screenplay

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-08-16T10:50:30.095171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:50:29.897427Z digest=sha256:e65b01c8175a468c58886f449747c92fa7119ae8c32217ea275bcffcf2fc368a

Observation c31df461-7bc2-43e2-a2a2-480719ea4c3b · outbound

This paper cites Video Repurposing from User Generated Content: A Large-scale Dataset and Benchmark.

VEU-Bench: Towards Comprehensive Understanding of Video Editing Video Repurposing from User Generated Content: A Large-scale Dataset and Benchmark

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-16T10:50:29.901553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:50:29.901553Z digest=sha256:af9d5e8bc41831c9211e30637a725b10d8460566e1a205f94ca0dd23c1ab4d5b

Observation 227d5305-fddf-43ff-aa02-e6f2529b05f7 · outbound

This paper cites Beyond Raw Videos: Understanding Edited Videos with Large Multimodal Model.

VEU-Bench: Towards Comprehensive Understanding of Video Editing Beyond Raw Videos: Understanding Edited Videos with Large Multimodal Model

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-16T10:50:29.905591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:50:29.905591Z digest=sha256:6a0e1620366e947533f1feea98efc1b13f96c205e71b46aae0f27fedbc13322c

Observation 18bda182-7892-4052-bb7a-c37ea413cccc · outbound

This paper cites LongVILA: Scaling Long-Context Visual Language Models for Long Videos.

VEU-Bench: Towards Comprehensive Understanding of Video Editing LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-16T10:50:29.910764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:50:29.910764Z digest=sha256:c1c9f464b8d88c4dd266b108e4d9103d8d3a85f92b78333051e32e383547440b

Observation 2d406e3a-6e76-425a-aefd-ba39fb6a3d8c · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

VEU-Bench: Towards Comprehensive Understanding of Video Editing MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-16T10:50:29.915272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:50:29.915272Z digest=sha256:7cea2dfaec4e008cc9e507c47056d38f3029128ceaafc15772a4475dcc94b440

Observation c4431211-cab5-45ba-86fa-5796e0e7adb9 · outbound

This paper cites mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models.

VEU-Bench: Towards Comprehensive Understanding of Video Editing mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-16T10:50:29.918962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:50:29.918962Z digest=sha256:2e933eddce45fa2733122adfd4477b029ee4e6488b1395604e7994a86e4a015b

Observation 436a70dd-7435-4c9d-af02-12d9b814f80d · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

VEU-Bench: Towards Comprehensive Understanding of Video Editing Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-16T10:50:29.923137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:50:29.923137Z digest=sha256:0e7c5187544fde0b2cd9efbfaea34d901a9ae34ba32b1efa89843980bc40c4b8

Observation d761f01e-d6ef-4c4b-9afb-59c0bba72a75 · outbound

This paper cites InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output.

VEU-Bench: Towards Comprehensive Understanding of Video Editing InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-16T10:50:29.927424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:50:29.927424Z digest=sha256:0276fe539daa2a09cf8915a07dbc6d662cb2c19d4488c67aebe7974908447dbd

Observation 7dee8977-a33c-49e8-9fc7-dd7533c02385 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

VEU-Bench: Towards Comprehensive Understanding of Video Editing LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-16T10:50:29.931575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:50:29.931575Z digest=sha256:84c044a1aad2f2a5975ddcbf320ae4d826d7df238ee2e822517191b1a1870888

Observation 8f12a558-3aa6-4041-9e8e-bf707c9c3e40 · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

VEU-Bench: Towards Comprehensive Understanding of Video Editing MLVU: Benchmarking Multi-task Long Video Understanding

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-16T10:50:29.935626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:50:29.935626Z digest=sha256:4109a95e13b55977e087d1a4c0d5c8ed906b6abd388ede5932f86b7b9e1bd2fb

Pith citing papers

Observation 86bf2cb7-94da-403f-8ed4-05db4472013b · inbound

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought cites this paper.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought VEU-Bench: Towards Comprehensive Understanding of Video Editing

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:41.965687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:41.965687Z digest=sha256:0ae36c3c5376a2027140a4b630af0d8f1967c5ac1341ba18be903e1422e349bf

Observation d47f2d1d-dfd3-4cfa-b3df-33df2b972b46 · inbound

VERTIGO: Visual Preference Optimization for Cinematic Camera Trajectory Generation cites this paper.

VERTIGO: Visual Preference Optimization for Cinematic Camera Trajectory Generation VEU-Bench: Towards Comprehensive Understanding of Video Editing

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-13T21:18:17.189424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T21:14:42.021240Z digest=sha256:531ac70cc2d9d264a6e4912faa6d2d2361fd15661e36167e591b61790b646ef0