Pith. sign in

Paper Citation Record · LEDGER

MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing

As of 11 August 2026, this Paper Citation Record lists 66 of 66 outbound references and 0 inbound Pith citation observations for arXiv:2607.25300.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.25300 v1

Coverage vector

measured 66 of 66 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T02:57:35.417802Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

66 of 66 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved65
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ecced3ec-8956-43b8-8873-8915804c5eca · outbound

This paper cites Adopting self- supervised learning into unsupervised video summarization through restorative score.

MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing Adopting self- supervised learning into unsupervised video summarization through restorative score

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T02:57:33.526010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:57:33.526010Z digest=sha256:21efced3d385fd1cfa73fcff48cb00d30f1ad22297bda653e0cd054c155bcb64

Observation acda662c-6162-4503-9b3c-fd3458d436e4 · outbound

This paper cites Combining global and local attention with positional encoding for video summarization.

MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing Combining global and local attention with positional encoding for video summarization

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T02:57:33.531817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:57:33.531817Z digest=sha256:142d008fd5f69a470b02ac5ac49ef1eaf1c365af058da26602a6000f5b41e8ec

Observation 27d5d21a-0523-4123-a40c-ae1ba4b9053d · outbound

This paper cites Summarizing videos using con- centrated attention and considering the uniqueness and diver- sity of the video frames.

MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing Summarizing videos using con- centrated attention and considering the uniqueness and diver- sity of the video frames

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T02:57:33.537215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:57:33.537215Z digest=sha256:4dd6888599405af71e0b769ad59030bbaa045b822b93b254f4670042f2246e50

Observation 59baaf17-8b2b-4b27-b110-c428ee0f6d7c · outbound

This paper cites Scaling up video summarization pretraining with large language models.

MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing Scaling up video summarization pretraining with large language models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T02:57:33.542503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:57:33.542503Z digest=sha256:e4150b62a6d38f461db5d30a7d8240b56f0a30cc1d12012953d8e0e658055bb5

Observation 1de320e2-cd3a-4139-94e1-2109e9b99123 · outbound

This paper cites Qwen3-VL Technical Report.

MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing Qwen3-VL Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T02:57:33.547372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:57:33.547372Z digest=sha256:beb5beffcc52167f7ef0324f7d10804b9b9cc2927c63fdbb13c68808b0562c50

Observation ff241b79-9bbc-432b-ad24-16efc41ab4d6 · outbound

This paper cites Blender studio films.https://studio.

MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing Blender studio films.https://studio

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T02:57:33.553920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:57:33.553920Z digest=sha256:4d89b7011a20894b23c352c038955c54d0b3464fd1fe24334b69976e921b7fb1

Observation 093d8ffd-64fb-4612-9cf7-98c1eedbb5ec · outbound

This paper cites VSUMM: A mechanism designed to produce static video summaries and a novel evaluation method.Pattern Recog- nit.

MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing VSUMM: A mechanism designed to produce static video summaries and a novel evaluation method.Pattern Recog- nit

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T02:57:33.559812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:57:33.559812Z digest=sha256:fe0a78e974e8f78f8df5af2f1fee1df8e07b40951b62a006212391d33d81656f

Observation 87394ed8-fd48-4735-977e-5a16af907239 · outbound

This paper cites Summarizing videos with attention.

MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing Summarizing videos with attention

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T02:57:33.566410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:57:33.566410Z digest=sha256:dc672a7b4482efc5e54961f2006b8da503039cb8a4c33aac849461c60207ad72

Observation f517d32b-6f65-4f09-81b4-b8854a870cae · outbound

This paper cites Video-R1: Reinforcing video reasoning in MLLMs.NeurIPS, 38:99114–99137,.

MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing Video-R1: Reinforcing video reasoning in MLLMs.NeurIPS, 38:99114–99137,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T02:57:33.573564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:57:33.573564Z digest=sha256:477850837532f7bab7be2fe3362b41871a46d11fc828b85fa04becead69d828b

Observation 1526b89c-f3bf-481d-a0f6-4b3f32a69e5b · outbound

This paper cites Au- tomatic non-linear video editing transfer.

MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing Au- tomatic non-linear video editing transfer

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T02:57:33.577949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:57:33.577949Z digest=sha256:6f9a0ef4440a0752482c84da6e1235b0be68230eb569523c832f3800d9c2aa93

Observation 4612b575-b863-4be5-8268-5df9bfe7066b · outbound

This paper cites Video-MME: The first-ever comprehensive evaluation benchmark of multi- modal LLMs in video analysis.

MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing Video-MME: The first-ever comprehensive evaluation benchmark of multi- modal LLMs in video analysis

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T02:57:33.581992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:57:33.581992Z digest=sha256:d04b261fb37adf82c2af2bf5791d2b550ccb8efdafcca649bdef1d020c2dbaae

Observation 1049e734-6252-40a7-b747-b7c66e8b30f5 · outbound

This paper cites Training-free language-guided video summarization via multi-grained saliency scoring.

MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing Training-free language-guided video summarization via multi-grained saliency scoring

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T02:57:33.586663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:57:33.586663Z digest=sha256:bc47cbee2ec16fd5e84d578386da4833fc04ba6d4b6ba7f6afae0054706d91d1

Observation 9aa6d92c-f5ab-46ce-84bf-d1d51241a196 · outbound

This paper cites Supervised video summarization via multiple feature sets with parallel attention.

MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing Supervised video summarization via multiple feature sets with parallel attention

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T02:57:33.591243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:57:33.591243Z digest=sha256:c1f04b992da3769e3192351f85e7b5547bde602393b3d17c6ab7925164477e85

Observation 5a58f19a-82d1-447d-a574-17ed5d27a40b · outbound

This paper cites A Survey on LLM-as-a-Judge.

MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing A Survey on LLM-as-a-Judge

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T02:57:33.595868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:57:33.595868Z digest=sha256:ca3c0c16f93ac906b1fe8091b654679a3670a3513ec22577bd68a7794211e7b6

Observation c728aa61-d9b1-4197-b6a2-5da281883f28 · outbound

This paper cites VTG-LLM: Integrating timestamp knowledge into video LLMs for enhanced video temporal grounding.

MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing VTG-LLM: Integrating timestamp knowledge into video LLMs for enhanced video temporal grounding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T02:57:33.600907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:57:33.600907Z digest=sha256:da5e596edbfdaf86160d5d3c5dca605d4c7ad7135318daa0e3a6b5ea5147f02c

Observation 3ca9901d-096e-4061-ab4c-97a148843892 · outbound

This paper cites Creating summaries from user videos.

MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing Creating summaries from user videos

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T02:57:33.606418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:57:33.606418Z digest=sha256:b658114d499e91ed72fc6e62d9c321b6e1bdf2956063fe958563091ffb1d7c86

Observation a7d90a2c-4405-41d9-aae1-0bf39a09d96b · outbound

This paper cites Align and attend: Multimodal summarization with dual contrastive losses.

MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing Align and attend: Multimodal summarization with dual contrastive losses

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T02:57:33.611332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:57:33.611332Z digest=sha256:06720f1084097ec845faa0fec263cb09a08c5a1d73b622429f1485a102258129

Observation acfa2fc3-39d0-46a4-9ed5-1b8035c2a4de · outbound

This paper cites MovieNet: A holistic dataset for movie under- standing.

MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing MovieNet: A holistic dataset for movie under- standing

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T02:57:33.617191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:57:33.617191Z digest=sha256:063a42533c81dc905fc6fda50f6b60043aeb25dfa491ad7030b2701824445002

Observation 6fd57df8-6057-4080-ac17-65e9d06b8542 · outbound

This paper cites Joint video summarization and moment localization by cross-task sample transfer.

MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing Joint video summarization and moment localization by cross-task sample transfer

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T02:57:33.623434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:57:33.623434Z digest=sha256:9132409ec03110ed233924dcb504edb71721d143a492738e92fcab83b39e7c69

Observation a192e377-fd68-4f53-8236-3dcd1e8de382 · outbound

This paper cites Discriminative feature learning for unsu- pervised video summarization.

MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing Discriminative feature learning for unsu- pervised video summarization

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T02:57:33.633320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:57:33.633320Z digest=sha256:8c247bd6470e9446ed2d1feddc0c7e2a6698834616f4d5b0aefaa2400f1fcaff

Observation e914b6eb-0492-4220-8877-28d176408329 · outbound

This paper cites Dense-captioning events in videos.

MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing Dense-captioning events in videos

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T02:57:33.646692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:57:33.646692Z digest=sha256:b8346d4d76e94f568b830d9c91fda30bcae725c8dec14940114e940db70e4c1f

Observation 2da01de5-bbbf-4d77-8beb-e4ccfb1c4176 · outbound

This paper cites Computational video editing for dialogue-driven scenes.ACM Trans.

MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing Computational video editing for dialogue-driven scenes.ACM Trans

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T02:57:33.659453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:57:33.659453Z digest=sha256:edb6601c4e43032c0f5d93ecf68166da0d087a773c35fd68baa44177fff83661

Observation f3f5c708-8ebf-4238-8d97-484837367b00 · outbound

This paper cites Video sum- marization with large language models.

MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing Video sum- marization with large language models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T02:57:33.680677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:57:33.680677Z digest=sha256:9824aedf1769fdcb94b826e5a9b66adfbe5d270b621bfa7f460cc71df75b7a01

Observation 86900be1-3d86-456f-a173-501ca56e32a4 · outbound

This paper cites Detecting mo- ments and highlights in videos via natural language queries.

MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing Detecting mo- ments and highlights in videos via natural language queries

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T02:57:33.700419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:57:33.700419Z digest=sha256:0a3832c2ba50a9a53edfd2c62dfa3ce042660d79bdadda34eb22936c1c1aea5a

Observation b63513ae-f924-4abf-ad7b-743101a434e0 · outbound

This paper cites Progressive video summarization via multimodal self- supervised learning.

MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing Progressive video summarization via multimodal self- supervised learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T02:57:33.721473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:57:33.721473Z digest=sha256:a1b06c154d154faee42b150c1b405cc3ba870f0e7599b545ca0f8319947ef6fa

Observation 062e7a7a-18d2-4467-adf9-01a0a304ad21 · outbound

This paper cites Progressive video summarization via multimodal self- supervised learning.

MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing Progressive video summarization via multimodal self- supervised learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T02:57:33.747254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:57:33.747254Z digest=sha256:6bdb86ca1e315a5c701b0935f750b6cf6a240a9811bc026e5be9fadeee006fe0

Observation da640c86-dd3c-4611-b5c5-b1c815b40c3a · outbound

This paper cites VideoChat-R1: Enhancing spatio-temporal percep- tion via reinforcement fine-tuning.NeurIPS, 2025.

MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing VideoChat-R1: Enhancing spatio-temporal percep- tion via reinforcement fine-tuning.NeurIPS, 2025

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T02:57:33.771469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:57:33.771469Z digest=sha256:2fd6d22d187ee744d7110fe7fcb94edf7781691fc7b0d8eaf51ecd190bb0ca5b

Observation 41c3e89d-2252-4411-b6d1-10fcd170bf27 · outbound

This paper cites VideoXum: cross- modal visual and textural summarization of videos.IEEE Trans.

MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing VideoXum: cross- modal visual and textural summarization of videos.IEEE Trans

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T02:57:33.793614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:57:33.793614Z digest=sha256:0644fc630648fe1ce52d451d443d29630b671fe1c2913ec9678cec9c9bf46bfc

Observation b70bf6a0-9822-46b6-8405-91157efb81a6 · outbound

This paper cites Visual instruction tuning.NeurIPS, 36:34892–34916, 2023.

MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing Visual instruction tuning.NeurIPS, 36:34892–34916, 2023

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T02:57:33.819796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:57:33.819796Z digest=sha256:7063815ca1d627b236b007b9cefc8dd5cb226db94477f8cf3f6b45b14c4673e9

Observation 4fcf045e-d7f8-4251-8ddc-1a187e0b7952 · outbound

This paper cites G-Eval: NLG evaluation using gpt-4 with better human alignment.

MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing G-Eval: NLG evaluation using gpt-4 with better human alignment

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T02:57:33.841722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:57:33.841722Z digest=sha256:d9d014ef0c7739e38633c22dcbe9443b6c84b999702a7fda160445dddf640bdd

Observation 49c0c03c-75b9-41ff-81c8-1c126a6ae26b · outbound

This paper cites an unresolved cited work.

MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T02:57:33.910460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:57:33.910460Z digest=sha256:c90cd9a2c1c689e84cf8959a18158bd397d2541e83a7edfa6d9d4ceb1531d5f4

Observation 1ebf3e82-6a7c-431d-8d72-8db30a0f133b · outbound

This paper cites Chrono: A simple blueprint for representing time in MLLMs.arXiv preprint arXiv:2406.18113, 2024.

MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing Chrono: A simple blueprint for representing time in MLLMs.arXiv preprint arXiv:2406.18113, 2024

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T02:57:34.008992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:57:34.008992Z digest=sha256:1d85cebcb706a6a1a322c7b2a6a44b79691b98ec1aa1a031e7e2669f1876e4f8

Observation ec5547e9-99dc-4c2c-b7f1-3d9f24c42fdf · outbound

This paper cites CLIP-It! language-guided video summarization.

MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing CLIP-It! language-guided video summarization

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T02:57:34.127674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:57:34.127674Z digest=sha256:4916923e5ad89f49264f765e2f43c4e96dd7d4712ddaaddb916027586739b7d3

Observation b48724bc-2013-4ce5-afd3-1dd42a55a6ee · outbound

This paper cites Rethinking the evaluation of video summaries.

MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing Rethinking the evaluation of video summaries

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T02:57:34.285272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:57:34.285272Z digest=sha256:1eb776ac145e15cb8bde8cce55e7c66943fc83327639c6358a2eebe43474b0f3

Observation f334b4cd-eb25-4f6f-bb75-206a797ed8ca · outbound

This paper cites Contrastive losses are natural criteria for unsu- pervised video summarization.

MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing Contrastive losses are natural criteria for unsu- pervised video summarization

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T02:57:34.407233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:57:34.407233Z digest=sha256:0ad442ce8b7b359ea63d6bba8d28a81fcd54fa89a7f6bac5e2fbfaeab1368e98

Observation 6b51240d-382e-4790-b7b0-646b837f855a · outbound

This paper cites Mea- sure Twice, Cut Once: A semantic-oriented approach to video temporal localization with video LLMs.

MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing Mea- sure Twice, Cut Once: A semantic-oriented approach to video temporal localization with video LLMs

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T02:57:34.538157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:57:34.538157Z digest=sha256:f5050a6891b3605bc182434c0d51a3607fd788258e11a6a8618e170531824f71

Observation d5debe9c-38ca-407e-bc5a-c918a90f6f4f · outbound

This paper cites MovieCuts: A new dataset and benchmark for cut type recognition.

MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing MovieCuts: A new dataset and benchmark for cut type recognition

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T02:57:34.652076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:57:34.652076Z digest=sha256:c000d474c6664e90bf35e93d6bd4b7ec8ef868687ae83171c89b790b33b1b3b0

Observation 5922fd12-7f72-49b0-ae3d-b7d02d6b5cc7 · outbound

This paper cites Generative timelines for instructed visual assembly.

MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing Generative timelines for instructed visual assembly

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T02:57:34.724308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:57:34.724308Z digest=sha256:0e72b238d67a99ff685665ff74b00793d228b5b8332f49d00dba0eef48887f74

Observation 340a0024-cb23-46fc-ad60-2e35af0d907e · outbound

This paper cites MMSum: A dataset for multimodal summarization and thumbnail gen- eration of videos.

MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing MMSum: A dataset for multimodal summarization and thumbnail gen- eration of videos

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T02:57:34.884059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:57:34.884059Z digest=sha256:57fdf7248cbb227189e538c0856a85d5adb6358014aac9a4b29645d9144ad801

Observation 0449790b-45b0-41e3-9677-b1c772a3243a · outbound

This paper cites TimeChat: A time-sensitive multimodal large language model for long video understanding.

MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing TimeChat: A time-sensitive multimodal large language model for long video understanding

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T02:57:34.950203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:57:34.950203Z digest=sha256:51689e6dcfb681d3f6a0224292d048e106e7583d215265e30d82bea9247d5aac

Observation 41add252-88da-4500-82b0-0b2e80965a3d · outbound

This paper cites an unresolved cited work.

MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T02:57:35.052912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:57:35.052912Z digest=sha256:d64e9fb4ed43c0ec91a08ae87388e8c234a4c46898e6bb2576b5104bba8006b4

Observation 71a2a1fc-0ad2-49c5-b9d6-3f60a8bfdb18 · outbound

This paper cites Judging the judges: A system- atic study of position bias in LLM-as-a-Judge.

MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing Judging the judges: A system- atic study of position bias in LLM-as-a-Judge

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T02:57:35.119099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:57:35.119099Z digest=sha256:4838313024ae14b10e6e40ebd23059500f3a00627492f586c8093df379cdda08

Observation 05e84f1d-c50e-42e2-abff-77a382ff51e9 · outbound

This paper cites Generic event boundary de- tection: A benchmark for event segmentation.

MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing Generic event boundary de- tection: A benchmark for event segmentation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T02:57:35.224975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:57:35.224975Z digest=sha256:6860df764320fad7184a51ac2c3597fd3e11750f28c043d2696f3573c51956c6

Observation 129aa3b8-ae6f-44d0-b473-8fe1787e0a99 · outbound

This paper cites CSTA: Cnn- based spatiotemporal attention for video summarization.

MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing CSTA: Cnn- based spatiotemporal attention for video summarization

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-01T02:57:35.325730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:57:35.325730Z digest=sha256:6b96c543962d08a169a6235aa866ce7cbcf5517f276b638cdd3b59e6b72e8154

Observation 79796f32-ba54-46e4-9a8b-1e47c22205b8 · outbound

This paper cites Csta: Cnn- based spatiotemporal attention for video summarization.

MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing Csta: Cnn- based spatiotemporal attention for video summarization

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T02:57:35.330070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:57:35.330070Z digest=sha256:8f9d015215933a80b67cbfbe2dbb9180521c732d582156d548d51a69433c4635

Observation e7cc1cd1-8c6a-4491-aa08-5d0f4bdeeb89 · outbound

This paper cites TVSum: Summarizing web videos using titles.

MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing TVSum: Summarizing web videos using titles

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-01T02:57:35.334102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:57:35.334102Z digest=sha256:a9a77fa20e41f5bca9eb3547b2a505ef34ee4db23578c1544b72911398587329

Observation 4094f2cd-4abe-4731-b9d8-8afde550e466 · outbound

This paper cites Language-guided self-supervised video summarization using text semantic matching considering the diversity of the video.

MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing Language-guided self-supervised video summarization using text semantic matching considering the diversity of the video

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-01T02:57:35.338185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:57:35.338185Z digest=sha256:a1c6138c3a1ea168aaf8d3f42e6ad00a3e4ae5aeb052830d691e17c00f4d5db8

Observation 04c50e48-d688-4b89-b878-b6073eca71b3 · outbound

This paper cites Gemini: A family of highly capable multimodal models, 2025.

MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing Gemini: A family of highly capable multimodal models, 2025

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-01T02:57:35.342337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:57:35.342337Z digest=sha256:71bb79d50dede82bc527892e2be9f37816b8dea6d24d9e568ceebedd0be0d2f8

Observation 9aeab3b7-f291-467e-9636-b198e9faf1b7 · outbound

This paper cites QuickCut: An interactive tool for editing narrated video.

MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing QuickCut: An interactive tool for editing narrated video

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-01T02:57:35.347096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:57:35.347096Z digest=sha256:3f94b257719c9c60c382a42409298303694b7b32a369d08dd9ae03e1b09ab8a5

Observation 91f8cce7-7d4a-472f-a3f5-f575e5e5fff6 · outbound

This paper cites Query Twice: Dual mixture attention meta learning for video summarization.

MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing Query Twice: Dual mixture attention meta learning for video summarization

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-01T02:57:35.351267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:57:35.351267Z digest=sha256:486b4be90d9393422cb11ac344d9c67135a4bfd0ca9dd3b6945d331cc1abee0b

Observation 3fb8ca8c-cf18-4fcf-839d-204a33781b98 · outbound

This paper cites Write-A-Video: Computational video montage from themed text.ACM Trans.

MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing Write-A-Video: Computational video montage from themed text.ACM Trans

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-01T02:57:35.356284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:57:35.356284Z digest=sha256:1a1d21a21d6e43c6bee80e63f12011d9e6e75fb43c990ff5a66a09e511fcacee

Observation c2aa9254-6cd7-4399-80e8-efc9746a1a6f · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-01T02:57:35.360849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:57:35.360849Z digest=sha256:e5765b18031bba4938fc50ebe85e4aead5a08d2bff721fb739db2ab6e7dc5f75

Observation a1397759-3ae4-4575-a7bc-fe20cf4c8182 · outbound

This paper cites Time-R1: Post-training large vision language model for temporal video grounding.NeurIPS, 38:83330– 83364, 2026.

MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing Time-R1: Post-training large vision language model for temporal video grounding.NeurIPS, 38:83330– 83364, 2026

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-01T02:57:35.364941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:57:35.364941Z digest=sha256:8e9095bf0a2394689fa1683db186cd65e56aa0382c1cdb13a568fd07680a4fae

Observation a07e9033-6de5-465c-8aff-cee5b8e24bde · outbound

This paper cites LongVideoBench: a benchmark for long-context interleaved video-language understanding.

MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing LongVideoBench: a benchmark for long-context interleaved video-language understanding

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-01T02:57:35.368984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:57:35.368984Z digest=sha256:a4484740f09fa2c6e2284533241dea81767e393d3266b37489beb7f0d7df4499

Observation 2a774dfc-600c-434f-899c-2654861e571b · outbound

This paper cites Transcript to Video: Efficient clip sequencing from texts.

MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing Transcript to Video: Efficient clip sequencing from texts

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-01T02:57:35.372928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:57:35.372928Z digest=sha256:15d03324fac12165b3d7f0ff32903e147d5279c3bedb80bf1c598eceff58a512

Observation 000544e0-d069-4b55-a5de-b07b9c020bcb · outbound

This paper cites Vid2seq: Large-scale pretraining of a vi- sual language model for dense video captioning.

MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing Vid2seq: Large-scale pretraining of a vi- sual language model for dense video captioning

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-01T02:57:35.376928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:57:35.376928Z digest=sha256:4d6d5df3c5018d11aa51538c0391e839e588f2e7bc8e7be9ce35da0e4f066bcb

Observation d2dd359e-a90e-45f9-bfb7-3443b4f3b7f7 · outbound

This paper cites TimeLens: Rethinking video temporal grounding with multimodal LLMs.

MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing TimeLens: Rethinking video temporal grounding with multimodal LLMs

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-01T02:57:35.380934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:57:35.380934Z digest=sha256:fd062a446b940412497dbe6e7f248bfeeb1755ac212bd56b262b8ed5743b3e9c

Observation 355f80e7-92a6-489f-bca3-d1ee81bf4c9d · outbound

This paper cites GPT-4V(ision) as a Generalist Evaluator for Vision-Language Tasks.

MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing GPT-4V(ision) as a Generalist Evaluator for Vision-Language Tasks

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-01T02:57:35.385182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:57:35.385182Z digest=sha256:5f0deaeaf18c8ca49312f9e232464f9a7e05b1117f2416da6c9aed06fe90b6d1

Observation 6d41d762-6f0e-471f-bb80-5be68a0be5f8 · outbound

This paper cites Re- constructive sequence-graph network for video summariza- tion.IEEE Trans.

MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing Re- constructive sequence-graph network for video summariza- tion.IEEE Trans

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-01T02:57:35.389309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:57:35.389309Z digest=sha256:3eaebd0a5357c48a974fe0f46a88dcf00f317e338eda69e680ca2c397a2ede87

Observation 1c464df8-65b5-4080-a5f1-2c97c1e014db · outbound

This paper cites Xing, Hao Zhang, Joseph E.

MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing Xing, Hao Zhang, Joseph E

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-01T02:57:35.393918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:57:35.393918Z digest=sha256:aaf15f21273d3515e1ba7b995b9eef1969217aab8add44018deaf13a0274bf58

Observation 38eaef64-af30-48cd-ab3a-81913c92698b · outbound

This paper cites Deep semantic and attentive network for unsupervised video summarization.ACM Trans.

MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing Deep semantic and attentive network for unsupervised video summarization.ACM Trans

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-01T02:57:35.397619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:57:35.397619Z digest=sha256:8552e99534229a11cc26410f4e9ffdf6a9d3e2fec6c598524f39c89cec887d77

Observation 4d47acc8-b7f2-4d90-8e10-8ea365fe22cc · outbound

This paper cites MLVU: Benchmarking multi-task long video understanding.

MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing MLVU: Benchmarking multi-task long video understanding

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-01T02:57:35.401484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:57:35.401484Z digest=sha256:8eb3b8d83ba21a7e384c8be17c2120c8bf80309fa1f19d6b44fb6b2ec1da790f

Observation 975ca5e2-7685-46d7-9015-5f25ee07038b · outbound

This paper cites Deep reinforce- ment learning for unsupervised video summarization with diversity-representativeness reward.

MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing Deep reinforce- ment learning for unsupervised video summarization with diversity-representativeness reward

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-01T02:57:35.405324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:57:35.405324Z digest=sha256:36e4ea2636014447d93e934fcc283745d8bdfd104d60c28abd721894fae8727a

Observation a73ca798-edcd-4bac-baef-7dc7cf444de0 · outbound

This paper cites Edits” counts edits with at least one temporal reversal; “Cuts.

MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing Edits” counts edits with at least one temporal reversal; “Cuts

Reference 64

Resolution
malformed identifier
no resolver link, observed 2026-08-01T02:57:35.408991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:57:35.408991Z digest=sha256:e59c40fc992681db138c4e468c5d1be4a4949265a587a8f6a66e33a824830610

Observation ec6c8238-0f87-4c20-b4c3-11eabcdd9401 · outbound

This paper cites an unresolved cited work.

MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing Unresolved cited work

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-01T02:57:35.413753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:57:35.413753Z digest=sha256:5203396d7e4248551b9e9226f799e4505174cfb6603cdc23495f8470ef552e04

Observation 5f239f22-9ad7-4f3d-83af-9b65558bdbf3 · outbound

This paper cites 01:30:50.

MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing 01:30:50

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-01T02:57:35.417802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:57:35.417802Z digest=sha256:0ccecc348f2eaab934dc28370d2361efb05f79b90900bf39307ba82c2a83164d

Pith citing papers

No inbound Pith citation observations are available.