Pith. sign in

Paper Citation Record · LEDGER

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs

As of 19 August 2026, this Paper Citation Record lists 71 of 71 outbound references and 0 inbound Pith citation observations for arXiv:2505.19535.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.19535 v1

Coverage vector

measured 71 of 71 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:15:48.995779Z

measured 71 of 71 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

71 of 71 outbound references displayed

  • verified exact0
  • verified fuzzy55
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 49259069-4357-497b-af56-bb497360a7ed · outbound

This paper cites Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:58.979116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:15:43.651761Z digest=sha256:59601e6daf57996386b80d00ec8218741a99d3b1f2eb23b3b320501165a2a40d

Observation d726bc0c-741b-47db-9faa-edd19846277f · outbound

This paper cites Tokenflow: Unified image tokenizer for multimodal understanding and generation,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Tokenflow: Unified image tokenizer for multimodal understanding and generation,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:58.676070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:15:43.742055Z digest=sha256:e26b5e08b9d776c1362654be79e7e109d7c31952d9c45729eece6638c224f654

Observation 2f1b4a46-ee54-44d6-9c4c-13ed27702cc6 · outbound

This paper cites Text2video-zero: Text-to-image diffusion models are zero-shot video generators,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Text2video-zero: Text-to-image diffusion models are zero-shot video generators,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:58.429716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:15:43.841821Z digest=sha256:225181124b56d8b81f679f980d1a32bf005bf5a146bcff6c6f1aacb29eeb6f50

Observation 60d1a271-6100-4bc2-8999-e9635db4082b · outbound

This paper cites Ccedit: Creative and controllable video editing via diffusion models,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Ccedit: Creative and controllable video editing via diffusion models,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:58.111035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:15:43.901242Z digest=sha256:db1288de9be7cf414196c1e8d262c8a0c053bf985fb6c6e0553f8999390732d0

Observation fb830f27-c3ed-4215-ad0d-1187587f0dab · outbound

This paper cites Controlvideo: Training-free controllable text-to-video generation,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Controlvideo: Training-free controllable text-to-video generation,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:57.919016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:15:44.024176Z digest=sha256:0ed2420f09a51121effe414c4c6a0905e777b78d1b22eb746d230e54f12708be

Observation b9a50780-125f-42b2-988f-e6e5d402579e · outbound

This paper cites Fatezero: Fusing attentions for zero-shot text-based video editing,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Fatezero: Fusing attentions for zero-shot text-based video editing,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:57.713443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:15:44.195828Z digest=sha256:6e3447f49afc238571df77965f166fc784d127f37d2be8c0104b8543fdbae97b

Observation 05c9366b-5047-46ac-8824-75e3d23a5552 · outbound

This paper cites Flatten: optical flow-guided attention for consistent text-to-video editing,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Flatten: optical flow-guided attention for consistent text-to-video editing,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:57.555158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:15:44.339750Z digest=sha256:a15a82a580bf9271b47caa8aadc5a51b26ed2b37463165ca0beb00d79fd2864c

Observation 2941ad73-01ea-4f6a-b494-4538eea50770 · outbound

This paper cites Fresco: Spatial-temporal correspondence for zero-shot video translation,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Fresco: Spatial-temporal correspondence for zero-shot video translation,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:57.301203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:15:44.443157Z digest=sha256:946147c0552a9e44c624dd083273d4ca0e54508be10a8b2e719caac71fe19c6e

Observation 47e9861c-8a2d-4354-b017-20dbd203fab1 · outbound

This paper cites Pix2video: Video editing using image diffusion,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Pix2video: Video editing using image diffusion,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:57.154855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:15:44.574580Z digest=sha256:a31d6d7062bf909d1add48a4ec9f702c8651361654afcc872c87c7d37548197c

Observation 77254bda-3784-410c-9557-33efcdda1cfb · outbound

This paper cites Rave: Randomized noise shuffling for fast and consistent video editing with diffusion models,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Rave: Randomized noise shuffling for fast and consistent video editing with diffusion models,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:56.972919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:15:44.666645Z digest=sha256:7a6001f17a1ef0a2a8f537cb747737861699686c6759549b95a86feb8b65bc76

Observation 6d9acb26-1789-4300-a8ff-33f813aa1edb · outbound

This paper cites Slicedit: Zero-shot video editing with text-to-image diffusion models using spatio-temporal slices,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Slicedit: Zero-shot video editing with text-to-image diffusion models using spatio-temporal slices,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:56.824036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:15:44.788333Z digest=sha256:27017b393ddfebf783c7e3bd616c0a3e36dab9c51de2a5d42611bfde3b6e4feb

Observation ad892360-db24-4767-b16a-1ac2c692ca94 · outbound

This paper cites Zero-shot video editing using off-the-shelf image diffusion models,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Zero-shot video editing using off-the-shelf image diffusion models,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:56.720375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:15:44.869181Z digest=sha256:3c5fd6b5d74295cca475c3af6376c2367d9ce98f2d922f055af8bb29a098e4b1

Observation 76e4353c-1fb6-48f1-ab9d-c5f18c8268da · outbound

This paper cites Video quality assessment: A comprehensive survey,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Video quality assessment: A comprehensive survey,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:56.573558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:15:44.964072Z digest=sha256:122f00e5d90114d4ce56865110426f0788d686b3f7d9b94def0817ae41683501

Observation e4796152-584e-4f9a-81d4-47c9342fe8d8 · outbound

This paper cites Exploring video quality assessment on user generated contents from aesthetic and technical perspectives,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Exploring video quality assessment on user generated contents from aesthetic and technical perspectives,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:56.401381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:15:45.062144Z digest=sha256:f2c9ec1aed523399630495b9a118bfed152cca70796ed17eed986158712ae9de

Observation a5d76651-a57b-4067-ab29-cab838afc2b4 · outbound

This paper cites Fast-vqa: Efficient end-to-end video quality assessment with fragment sampling,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Fast-vqa: Efficient end-to-end video quality assessment with fragment sampling,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:56.275262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:15:45.140757Z digest=sha256:5bd341e9b11ae8920d75141e7e29c1d0cec367821ab4867f701ca85822dfc78a

Observation 141736e6-6844-4070-9a96-8377261db06e · outbound

This paper cites A deep learning based no-reference quality assessment model for ugc videos,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs A deep learning based no-reference quality assessment model for ugc videos,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:56.115372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:15:45.223546Z digest=sha256:bd6926d9a56a2317edbe3ddf9416007198ff4fcb9a371ced7d576eaa191961a9

Observation 075bcaea-ff09-4a68-b72e-52c2829fc70e · outbound

This paper cites Quality assessment of in-the-wild videos,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Quality assessment of in-the-wild videos,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:55.923996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:15:45.303727Z digest=sha256:44d6abc578261cb94371d30dc28ff80a71f4796c06c270cae1b540995bfca50e

Observation db5585b8-a3c6-4ad3-8baa-48519105547e · outbound

This paper cites Ugc-vqa: Benchmarking blind video quality assessment for user generated content,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Ugc-vqa: Benchmarking blind video quality assessment for user generated content,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:55.755136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:15:45.396067Z digest=sha256:f23621006a3d23fc51f7fbd91acfe65ded88575545ac1212cf13671d8c5d6805

Observation 015e49cc-725b-4faa-83a4-9c08c135c34f · outbound

This paper cites Ve-bench: Subjective-aligned benchmark suite for text-driven video editing quality assessment,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Ve-bench: Subjective-aligned benchmark suite for text-driven video editing quality assessment,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:55.658082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:15:45.493803Z digest=sha256:e2b0a49860d21949f512931466203ebe18c588e106ec9500204c04a67e4e4a3a

Observation 67dc21db-bbe0-4bd6-a807-dea04770e5b2 · outbound

This paper cites Subjective-aligned dataset and metric for text-to-video quality assessment,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Subjective-aligned dataset and metric for text-to-video quality assessment,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:55.566921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:15:45.569424Z digest=sha256:311f22e546c1ca823be280116e90d215cedc200b0982143446aa3fb34aef9b07

Observation 382ed968-9186-4e37-a265-2431fb21154b · outbound

This paper cites Aigv-assessor: Benchmarking and evaluating the perceptual quality of text-to-video generation with lmm,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Aigv-assessor: Benchmarking and evaluating the perceptual quality of text-to-video generation with lmm,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:55.397791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:15:45.660369Z digest=sha256:84a781ecbabe47ebcd579358e96ae66358fa1fe9cc4ff5fadde5e5322929fd75

Observation 37e63f66-13fb-4635-8466-3d8fe8f257ed · outbound

This paper cites Cvpr 2023 text guided video editing competition,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Cvpr 2023 text guided video editing competition,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:55.263392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:15:45.715156Z digest=sha256:c83e0652baf04bcafe25800c7b5697c5fc4ede42e9d4a980fec1fa27e2f67559

Observation 8567a7a2-5497-4f38-bc97-57584fa67d93 · outbound

This paper cites Harnessing the power of llms in practice: A survey on chatgpt and beyond,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Harnessing the power of llms in practice: A survey on chatgpt and beyond,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:55.127769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:15:45.772617Z digest=sha256:13dcd09e862d73917322e77d4b746501d440ad5ae0e293a756ccb8298987a67f

Observation 909a4cb3-8658-48fb-9a50-7f876484b7a9 · outbound

This paper cites Señorita- 2m: A high-quality instruction-based dataset for general video editing by video specialists,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Señorita- 2m: A high-quality instruction-based dataset for general video editing by video specialists,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:54.957500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:15:45.864925Z digest=sha256:5935a21a2ea377c3779d3ee971a9fd6884d992ef36820d23547e68f9007b5661

Observation fc602e93-00f8-4104-9789-58f37f95a2e1 · outbound

This paper cites The 2017 davis challenge on video object segmentation,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs The 2017 davis challenge on video object segmentation,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:54.811565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:15:45.952544Z digest=sha256:7f1baf6d2636a14904785f05082a9080c4d5bbc36ed63e089e69450531e180a8

Observation b3a43984-edbb-4e33-8d52-60aa4f3a2ae7 · outbound

This paper cites The kinetics human action video dataset,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs The kinetics human action video dataset,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:54.691934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:15:46.050651Z digest=sha256:639d02bddbbefbe7fb379d493f2478c0fea45fb6335083a40525c1df8ee09f33

Observation a159a87b-60dd-4c58-8742-d0082cd328ed · outbound

This paper cites JimengAI.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs JimengAI

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:54.573138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:15:46.117830Z digest=sha256:4c880d1aee3301f055e0ce08083c2fdf4bf0d9dc104281a42a243a46424087db

Observation e6e4e421-f9e6-476f-9d65-a954de5c28f3 · outbound

This paper cites Methodology for the subjective assessment of the quality of television pictures,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Methodology for the subjective assessment of the quality of television pictures,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:54.428466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:15:46.186786Z digest=sha256:740622dacda92d2d77279a7c09bd5f93bc189228a23bd6bdd276be06d237587e

Observation c9cdea49-6a70-47ab-9afd-3a6c8b0baa73 · outbound

This paper cites Blind image quality assessment based on high order statistics aggregation,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Blind image quality assessment based on high order statistics aggregation,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:54.235971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:15:46.255534Z digest=sha256:5a7377e6de5efb0da7ce71f521c0034bf254780113f27f3933cb3a742064ee33

Observation 0fb5fcef-ca42-4cee-ad5b-f95dd1083574 · outbound

This paper cites Learning without human scores for blind image quality assessment,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Learning without human scores for blind image quality assessment,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:54.125353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:15:46.336487Z digest=sha256:87757457bf62e930e649c748442e2f9bfd7139e7e3102cffb6ff7bac07ba4622

Observation 58668ad9-9f78-4c1b-b000-54e70d6d5511 · outbound

This paper cites Making a “completely blind.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Making a “completely blind

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:54.055030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:15:46.415983Z digest=sha256:7ba87f561a2a904e2c7962fde2f55093c1d3c7d5bb03fbedd05fb97af12ae2d5

Observation 147b5037-7061-4135-a851-e193d80a0556 · outbound

This paper cites Clipscore: A reference-free evaluation metric for image captioning,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Clipscore: A reference-free evaluation metric for image captioning,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:53.999272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:15:46.508413Z digest=sha256:b609fa7402dc49e9ac616ab7330d3d98005b448613b541b119b8a5cd52eafed1

Observation 0082d370-eec2-4389-a745-f24588fd8b08 · outbound

This paper cites Evaluating Text-to-Visual Generation with Image-to-Text Generation.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Evaluating Text-to-Visual Generation with Image-to-Text Generation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:46.589608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:46.589608Z digest=sha256:580eedf11698143f8f37dfe57ce27c61d7fc3ba824c234013c60d619a3ea0e93

Observation fa6aed1c-e178-4085-a73d-c7cc251bcf1c · outbound

This paper cites LMM4LMM: Benchmarking and Evaluating Large-multimodal Image Generation with LMMs.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs LMM4LMM: Benchmarking and Evaluating Large-multimodal Image Generation with LMMs

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:46.647524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:46.647524Z digest=sha256:4b2b82ec39927f252b9f76d55b7a37628a9ee91f575fd7e013f063eca0edada2

Observation 9720f9e0-4f87-4be1-ad52-edb8df975867 · outbound

This paper cites Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:46.707838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:46.707838Z digest=sha256:69bab29da3ae51a581ddccc0f1e220fdce93bbc400299ab56eea774f12280df4

Observation b6ab865e-ded1-44f6-b9ce-1167a8e66290 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs An image is worth 16x16 words: Transformers for image recognition at scale,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:53.957123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:15:46.796154Z digest=sha256:3c9d87ba33f73397df8a97f6509a0bbac4c30e1af0055f0b8e057c28202deae3

Observation 514cfcbf-2390-4329-9e73-c59ea492ffa2 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs An image is worth 16x16 words: Transformers for image recognition at scale,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:53.808845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:15:46.889880Z digest=sha256:359c5d8b9d75e92894d9882d48fde1f010a3a12364c462ce44003befb14df6a1

Observation 2a2ac725-4e2f-4ca5-9db3-93e2c7a3cef6 · outbound

This paper cites MLP-Mixer: An all-MLP Architecture for Vision.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs MLP-Mixer: An all-MLP Architecture for Vision

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:46.972275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:46.972275Z digest=sha256:6d3edbbcb2af849cb86e81967599aab92c8d14d8b2432d7330061a62f642e14e

Observation 9fc44e78-5762-44f8-b03f-022e60c60f95 · outbound

This paper cites How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:47.047930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:47.047930Z digest=sha256:6f767ee878181e156b065488ec0963609d8dd21c9e3ac8cd53c1c69b43625670

Observation 8febaff3-2ab2-4aa5-b1d0-617cdeb1442c · outbound

This paper cites When Vision Transformers Outperform ResNets without Pre-training or Strong Data Augmentations.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs When Vision Transformers Outperform ResNets without Pre-training or Strong Data Augmentations

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:47.098749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:47.098749Z digest=sha256:a451a244cee771f06cf7754d13ea19ab3ab96c19656d1f33a61c370305ec11ed

Observation 0db7e7a8-44cb-495a-88e1-b7ee9eb8c641 · outbound

This paper cites Surrogate gap minimization improves sharpness-aware training,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Surrogate gap minimization improves sharpness-aware training,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:53.597132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:15:47.146349Z digest=sha256:c5faf616cc437d417d255c7fb9b5ca45f17be78c1c3979997a2ef92f846be60e

Observation 761438d2-410b-48c9-a6b6-59cb3f6a55f7 · outbound

This paper cites Lit: Zero-shot transfer with locked-image text tuning,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Lit: Zero-shot transfer with locked-image text tuning,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:53.438080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:15:47.196903Z digest=sha256:2aec6cddf56f32e99524dd4dc45da15922ef86e96d13eb08fa513f181ee5c606

Observation 1664f839-5ebe-4a7c-bb8f-b579c62ce79f · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:47.249382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:47.249382Z digest=sha256:00db9c1b2716dc7cb2b639391954d84c5ffbfa3bfab4d334429b04bd5ec2a421

Observation 8adb38d0-4237-4e12-b5bd-e0eb6b3fe445 · outbound

This paper cites Mlp-net: Multilayer perceptron fusion network for infrared small target detection,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Mlp-net: Multilayer perceptron fusion network for infrared small target detection,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:53.213820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:15:47.303914Z digest=sha256:c8531d3f1fd61a50513961093526ebfd79b556076fefa07333e0b17df90c7148

Observation 65a60f14-dd82-401e-a048-46d7db18602a · outbound

This paper cites Lora: Low-rank adaptation of large language models,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Lora: Low-rank adaptation of large language models,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:52.987257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:15:47.364459Z digest=sha256:dfddcddc33d99b84d1e7b7e2bffec70c38de6aa08d282907b627818398fe802c

Observation d53a406b-4189-488a-99ed-9272ada4c374 · outbound

This paper cites Imagereward: learning and evaluating human preferences for text-to-image generation,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Imagereward: learning and evaluating human preferences for text-to-image generation,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:52.781780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:15:47.419599Z digest=sha256:acf5b9854e5adc6993b43f2f4fe806e1310191d1e26a631ea0ef7ca3c01de95f

Observation c95245d3-7d6c-402a-aa39-aec719e0a493 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision- language understanding and generation,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Blip: Bootstrapping language-image pre-training for unified vision- language understanding and generation,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:52.561473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:15:47.497508Z digest=sha256:933c1ac0057b34c1a83a28fe438941d597c009c408d0799bc4cdb23a8ef52843

Observation 4f49277c-9c6d-4030-a140-a3ee6b1f4dde · outbound

This paper cites Pick-a-pic: An open dataset of user preferences for text-to-image generation,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Pick-a-pic: An open dataset of user preferences for text-to-image generation,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:52.397957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:15:47.556046Z digest=sha256:8579d43245fe31e2c45b75bfc421ce444e5291e37a4becc53876676e2ca9d7ae

Observation 11517637-b9b8-4fb5-a164-d3be77ecd053 · outbound

This paper cites Building cnn-based models for image aesthetic score prediction using an ensemble,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Building cnn-based models for image aesthetic score prediction using an ensemble,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:52.174546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:15:47.618742Z digest=sha256:b352e5b652de48dbd7d19c4a477e2726a9831cb6c47702af17262253d944cb11

Observation f8b2128c-3ff8-4a5c-b537-99a95b3bab61 · outbound

This paper cites Llava-next: A strong zero-shot video understanding model,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Llava-next: A strong zero-shot video understanding model,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:51.951951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:15:47.671411Z digest=sha256:87b162e0642b300a81f4313acd342cda2ec5e2b4d0cca143fbd6df1477a68b85

Observation 92317989-0820-4ead-b432-384fdfa782e0 · outbound

This paper cites Internvideo: General video foundation models via generative and discriminative learning,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Internvideo: General video foundation models via generative and discriminative learning,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:51.692478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:15:47.723918Z digest=sha256:1bb08c0d008da32e052d056ac8c78e64a40dc25db5920bb27dcd2c013f4cacff

Observation b97954b8-5d3e-4053-a390-0a9036c01158 · outbound

This paper cites VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:47.793673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:47.793673Z digest=sha256:c6a16b61e147be3d799274a955b9a9969eea632f5d60bc3a129209dd193a67d8

Observation 16b6cae0-883d-4759-b36e-d73e8eec2a68 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:51.511413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:15:47.842241Z digest=sha256:bc1007c1367259ebcef0d56f88d3fa06215f77ee0adb4fb03f4c2f6ad2abfff2

Observation dbc1d83f-1c7a-43b1-995f-c66bac60f979 · outbound

This paper cites mplug-owl3: Towards long image-sequence understanding in multi-modal large language models,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs mplug-owl3: Towards long image-sequence understanding in multi-modal large language models,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:51.344286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:15:47.905970Z digest=sha256:cf0aae3359786ba270d70ddd48604e3d51e421066b11040b1a99477b827cfdc0

Observation caaec614-1d36-46d3-b353-25615a657f3e · outbound

This paper cites No-reference image quality assessment in the spatial domain,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs No-reference image quality assessment in the spatial domain,

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:47.959442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:47.959442Z digest=sha256:55a4c8daf19f8e2f6493d84baed0ed638363407e90e984b389e727876e5a9961

Observation e78e9d04-ab79-454a-9fff-8fdf3e641069 · outbound

This paper cites Blind image quality estimation via distortion aggravation,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Blind image quality estimation via distortion aggravation,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:51.198918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:15:48.029727Z digest=sha256:0ca73d224bd219aa368a45195adef9c7e3d9280c909556ae9d71c16998581c32

Observation 79cbeca4-a14e-441e-be03-e04a34a93f5b · outbound

This paper cites Blind quality assessment based on pseudo- reference image,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Blind quality assessment based on pseudo- reference image,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:51.004024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:15:48.090491Z digest=sha256:662d2ffdf6eb8ec60534fc80198b859cbb0a4efd373a3c5859a8a28ffa2fb66c

Observation 403b8b0a-6ae7-426a-a442-ceabac5f51bc · outbound

This paper cites Visual instruction tuning,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Visual instruction tuning,

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:48.138760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:48.138760Z digest=sha256:dbdec09ff488c806b70032172ab0b246e40052a5893e34f0b8cb507cbbae4636

Observation 1f8005a3-abc0-497d-a663-8d9b958afd0b · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:48.190049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:48.190049Z digest=sha256:5cc43dc02d16ba79ef6d5fad9c346b2f363980133dc9a443a79a3e8bec61141a

Observation f883cf17-f2ed-408c-9b71-becbad8dd081 · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Llava-next: Improved reasoning, ocr, and world knowledge,

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:48.245414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:48.245414Z digest=sha256:e77cb6cae68e14f837e4400e65aa380d3ba0bab3056709ac7d39d36e1773888a

Observation a9df28ab-111c-4a68-821e-a7aaf4172c19 · outbound

This paper cites Improved baselines with visual instruction tuning,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Improved baselines with visual instruction tuning,

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:48.349429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:48.349429Z digest=sha256:ea2d575b165b1f40f5c216bc4514d0146e46ca88335f77c0a7f35027a1f3493b

Observation 865709a4-8de1-4b48-be07-011908404bcd · outbound

This paper cites Llava-next: Stronger llms supercharge multimodal capabilities in the wild,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Llava-next: Stronger llms supercharge multimodal capabilities in the wild,

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:50.758104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:15:48.412809Z digest=sha256:83100efb4b90fd77af5433630bae48553d9d60f52ec21f5479a07aa19643485a

Observation 8b7fa63b-bfa7-40b9-993b-d750b86428ca · outbound

This paper cites Llava-next: What else influences visual instruction tuning beyond data?,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Llava-next: What else influences visual instruction tuning beyond data?,

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:50.534207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:15:48.473759Z digest=sha256:bb7328b46c1133c02a78f0bcce86ea8c11d8bfd6ec16bb9d202b12cb7e38c89d

Observation 18370840-77bd-42b7-9442-253997febc9a · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:48.537796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:48.537796Z digest=sha256:70cc4411303b27827020b19a8584246a6228782aac8809d9a5f5fa8a118ff9b7

Observation 2a484005-937b-47b3-912e-9e815524a744 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:48.595108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:48.595108Z digest=sha256:7cf74d1ad2a1e2632551365e6f4cfe9402287921a0e80c6db8269930e038f404

Observation 3e219ea3-4a27-42f1-9288-8bbe7e7f46c3 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Learning transferable visual models from natural language supervision,

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:50.332408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:15:48.673221Z digest=sha256:e02530bd0e6235b1a43e55f7b956fd60bcdb3cb4b45bd3ab04e200cb3a2442d4

Observation dba5ab64-a687-4600-bbb6-142a5c5a19b2 · outbound

This paper cites Unmasked teacher: Towards training- efficient video foundation models,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Unmasked teacher: Towards training- efficient video foundation models,

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:50.106302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:15:48.732449Z digest=sha256:5f1384fa2030dba35e77c759498055b0349cd94b6bc8d54f37349f7f2c035da2

Observation 7e7a514f-68bf-41ce-aea6-7b24568ecc26 · outbound

This paper cites Stablevqa: A deep no-reference quality assessment model for video stability,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Stablevqa: A deep no-reference quality assessment model for video stability,

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:49.870172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:15:48.810810Z digest=sha256:654f44edaf9c75620f72a73204730ba3dc48fd539b2715b2b794abfab545940e

Observation f8d41227-00e3-4ca1-8fdf-781525ee55a7 · outbound

This paper cites Q-Align: Teaching LMMs for Visual Scoring via Discrete Text-Defined Levels.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Q-Align: Teaching LMMs for Visual Scoring via Discrete Text-Defined Levels

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:48.881635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:48.881635Z digest=sha256:54a9089a0a754602d5c37206418e3c5b979861144ca16927a497cd447bd944dc

Observation b589e717-2e61-441e-bf33-a1400fd02fa0 · outbound

This paper cites Excellent.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Excellent

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:49.655959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:15:48.932160Z digest=sha256:663486a51addd13459b7d05de34cfae367fb09f7d802cb6f603f73a9dc53912e

Observation 4d17d73d-ae3a-467f-bceb-a0c55182ad0f · outbound

This paper cites Excellent.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Excellent

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:49.409764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:15:48.995779Z digest=sha256:aa2f2a7b7f8ebbc38fb9b2856e736f68179d4688c52bb0834221ec02dacd947c

Pith citing papers

No inbound Pith citation observations are available.