Pith. sign in

Paper Citation Record · LEDGER

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs

As of 9 August 2026, this Paper Citation Record lists 71 of 71 outbound references and 0 inbound Pith citation observations for arXiv:2505.19535.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.19535 v1

Coverage vector

measured 71 of 71 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:15:48.995779Z

measured 71 of 71 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

71 of 71 outbound references displayed

  • verified exact0
  • verified fuzzy55
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 49259069-4357-497b-af56-bb497360a7ed · outbound

This paper cites Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:58.979116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:15:43.651761Z digest=sha256:86b3397596ebc2f4a1bcc0b2312fb3aad755ba82187731752057fff45b5e87c6

Observation d726bc0c-741b-47db-9faa-edd19846277f · outbound

This paper cites Tokenflow: Unified image tokenizer for multimodal understanding and generation,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Tokenflow: Unified image tokenizer for multimodal understanding and generation,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:58.676070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:15:43.742055Z digest=sha256:4dd1db2f6f7901f399999ba8d81d49dbdde2cd87bd8d91e475babc5709c3466f

Observation 2f1b4a46-ee54-44d6-9c4c-13ed27702cc6 · outbound

This paper cites Text2video-zero: Text-to-image diffusion models are zero-shot video generators,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Text2video-zero: Text-to-image diffusion models are zero-shot video generators,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:58.429716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:15:43.841821Z digest=sha256:b6eada42bbb2e4abc3fdb96c0dcd2012c9ff136d298be640a986c289a8707d13

Observation 60d1a271-6100-4bc2-8999-e9635db4082b · outbound

This paper cites Ccedit: Creative and controllable video editing via diffusion models,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Ccedit: Creative and controllable video editing via diffusion models,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:58.111035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:15:43.901242Z digest=sha256:0294c8561597247bce76cf14e3d51b648e5c3c9102615344b0451bb8959905d7

Observation fb830f27-c3ed-4215-ad0d-1187587f0dab · outbound

This paper cites Controlvideo: Training-free controllable text-to-video generation,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Controlvideo: Training-free controllable text-to-video generation,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:57.919016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:15:44.024176Z digest=sha256:25b2614abfbbfc33b9381b71e396bfeeb864f08582eaa9ee61a8b1b6065248f5

Observation b9a50780-125f-42b2-988f-e6e5d402579e · outbound

This paper cites Fatezero: Fusing attentions for zero-shot text-based video editing,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Fatezero: Fusing attentions for zero-shot text-based video editing,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:57.713443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:15:44.195828Z digest=sha256:d176d2cd2e6e10e072b4a17a054e1259e7a61f98c69a1312d0f6e57681278eb1

Observation 05c9366b-5047-46ac-8824-75e3d23a5552 · outbound

This paper cites Flatten: optical flow-guided attention for consistent text-to-video editing,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Flatten: optical flow-guided attention for consistent text-to-video editing,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:57.555158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:15:44.339750Z digest=sha256:78674f23db027748524a47fab9c3e8c902dc7ff8065808b825112926c06ec4fe

Observation 2941ad73-01ea-4f6a-b494-4538eea50770 · outbound

This paper cites Fresco: Spatial-temporal correspondence for zero-shot video translation,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Fresco: Spatial-temporal correspondence for zero-shot video translation,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:57.301203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:15:44.443157Z digest=sha256:003241f1f89b7882e7b1f7e3621208d14682e6a01f4ea07306bcbfc325afd30a

Observation 47e9861c-8a2d-4354-b017-20dbd203fab1 · outbound

This paper cites Pix2video: Video editing using image diffusion,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Pix2video: Video editing using image diffusion,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:57.154855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:15:44.574580Z digest=sha256:f722f571db53adb9b85b0ef42de668c450c9984466f6a2dc92ec800b6246a47f

Observation 77254bda-3784-410c-9557-33efcdda1cfb · outbound

This paper cites Rave: Randomized noise shuffling for fast and consistent video editing with diffusion models,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Rave: Randomized noise shuffling for fast and consistent video editing with diffusion models,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:56.972919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:15:44.666645Z digest=sha256:b0bcba4476cb10108b5b637e2ff7ed5e57e5ab33064b3c501826f0624bc0705f

Observation 6d9acb26-1789-4300-a8ff-33f813aa1edb · outbound

This paper cites Slicedit: Zero-shot video editing with text-to-image diffusion models using spatio-temporal slices,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Slicedit: Zero-shot video editing with text-to-image diffusion models using spatio-temporal slices,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:56.824036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:15:44.788333Z digest=sha256:7b241ac1198b75e1fd0fd5eec9eddae8b716087b7f29dc940c26bb484a37e941

Observation ad892360-db24-4767-b16a-1ac2c692ca94 · outbound

This paper cites Zero-shot video editing using off-the-shelf image diffusion models,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Zero-shot video editing using off-the-shelf image diffusion models,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:56.720375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:15:44.869181Z digest=sha256:3decae3f51d4a71e13923d6566c171b8dfd894f9d52149a6a1b855d78e675e09

Observation 76e4353c-1fb6-48f1-ab9d-c5f18c8268da · outbound

This paper cites Video quality assessment: A comprehensive survey,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Video quality assessment: A comprehensive survey,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:56.573558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:15:44.964072Z digest=sha256:dffa4914ade43f2597c90b36a9ac2e6d7ae13dbbe4ad707c3ae35d18b284fbd9

Observation e4796152-584e-4f9a-81d4-47c9342fe8d8 · outbound

This paper cites Exploring video quality assessment on user generated contents from aesthetic and technical perspectives,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Exploring video quality assessment on user generated contents from aesthetic and technical perspectives,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:56.401381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:15:45.062144Z digest=sha256:dc4fc678dd0f0de8147657f01c6ed88942461ebd298ec8ae2f2a3cd2e064f7b3

Observation a5d76651-a57b-4067-ab29-cab838afc2b4 · outbound

This paper cites Fast-vqa: Efficient end-to-end video quality assessment with fragment sampling,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Fast-vqa: Efficient end-to-end video quality assessment with fragment sampling,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:56.275262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:15:45.140757Z digest=sha256:db6284523e3a4ee9a38087f952b941052a847ee59d2f824c5c9a78d7a61d7bd2

Observation 141736e6-6844-4070-9a96-8377261db06e · outbound

This paper cites A deep learning based no-reference quality assessment model for ugc videos,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs A deep learning based no-reference quality assessment model for ugc videos,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:56.115372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:15:45.223546Z digest=sha256:b93686e25e316927b7e31164a537f462ac25591603d1e2256ed03c5e3f20777e

Observation 075bcaea-ff09-4a68-b72e-52c2829fc70e · outbound

This paper cites Quality assessment of in-the-wild videos,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Quality assessment of in-the-wild videos,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:55.923996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:15:45.303727Z digest=sha256:93753a3a39d5917666251432cffa9d221788dbc34691b24fa5fd8f8716481ef9

Observation db5585b8-a3c6-4ad3-8baa-48519105547e · outbound

This paper cites Ugc-vqa: Benchmarking blind video quality assessment for user generated content,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Ugc-vqa: Benchmarking blind video quality assessment for user generated content,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:55.755136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:15:45.396067Z digest=sha256:5113dfda584e3fb0d8ef07862bd2e0cc8193b6bc6467a13b86251f063197fe9e

Observation 015e49cc-725b-4faa-83a4-9c08c135c34f · outbound

This paper cites Ve-bench: Subjective-aligned benchmark suite for text-driven video editing quality assessment,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Ve-bench: Subjective-aligned benchmark suite for text-driven video editing quality assessment,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:55.658082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:15:45.493803Z digest=sha256:8623e385ad6e2337dfd13da4923792021f6d93f3a01d8696406caed99f017ff6

Observation 67dc21db-bbe0-4bd6-a807-dea04770e5b2 · outbound

This paper cites Subjective-aligned dataset and metric for text-to-video quality assessment,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Subjective-aligned dataset and metric for text-to-video quality assessment,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:55.566921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:15:45.569424Z digest=sha256:6f2b25c30f315d9d4e18ec65ee2a9b95fb505ae5daf3b5464a44192ae3e0fc57

Observation 382ed968-9186-4e37-a265-2431fb21154b · outbound

This paper cites Aigv-assessor: Benchmarking and evaluating the perceptual quality of text-to-video generation with lmm,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Aigv-assessor: Benchmarking and evaluating the perceptual quality of text-to-video generation with lmm,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:55.397791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:15:45.660369Z digest=sha256:89f3f44bee29b15f23106b310477499736f45271d54df2d51050206f0112ad16

Observation 37e63f66-13fb-4635-8466-3d8fe8f257ed · outbound

This paper cites Cvpr 2023 text guided video editing competition,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Cvpr 2023 text guided video editing competition,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:55.263392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:15:45.715156Z digest=sha256:eff94d2473b9ef18918e39b42bcf6312ae72df076622736908833443072cfe87

Observation 8567a7a2-5497-4f38-bc97-57584fa67d93 · outbound

This paper cites Harnessing the power of llms in practice: A survey on chatgpt and beyond,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Harnessing the power of llms in practice: A survey on chatgpt and beyond,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:55.127769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:15:45.772617Z digest=sha256:94d704059cb1033ad3ca89f60b7507bbd0d12c5fdc8856d43166fa953b78cebd

Observation 909a4cb3-8658-48fb-9a50-7f876484b7a9 · outbound

This paper cites Señorita- 2m: A high-quality instruction-based dataset for general video editing by video specialists,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Señorita- 2m: A high-quality instruction-based dataset for general video editing by video specialists,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:54.957500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:15:45.864925Z digest=sha256:1fbba0df3eef68891fa64f87f396d665aff3df0449d825d8f364fb4b86854c14

Observation fc602e93-00f8-4104-9789-58f37f95a2e1 · outbound

This paper cites The 2017 davis challenge on video object segmentation,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs The 2017 davis challenge on video object segmentation,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:54.811565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:15:45.952544Z digest=sha256:38d912419e1f3d087ef4bf5527a99d1607c39cd5203e2b9cd1353e9ab5ba75aa

Observation b3a43984-edbb-4e33-8d52-60aa4f3a2ae7 · outbound

This paper cites The kinetics human action video dataset,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs The kinetics human action video dataset,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:54.691934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:15:46.050651Z digest=sha256:2a4739f24b148df3e2b68b8a9d6b1f778ed44962835873c68d957e31a8013ec3

Observation a159a87b-60dd-4c58-8742-d0082cd328ed · outbound

This paper cites JimengAI.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs JimengAI

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:54.573138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:15:46.117830Z digest=sha256:430667f23be9b29c957ffe5bd8e327593466aac5b61fc81d393716e2834c85a3

Observation e6e4e421-f9e6-476f-9d65-a954de5c28f3 · outbound

This paper cites Methodology for the subjective assessment of the quality of television pictures,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Methodology for the subjective assessment of the quality of television pictures,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:54.428466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:15:46.186786Z digest=sha256:55fb97d88aba706d970fe23fa07c8078b466b0d30ac616df189f0b0335332e0f

Observation c9cdea49-6a70-47ab-9afd-3a6c8b0baa73 · outbound

This paper cites Blind image quality assessment based on high order statistics aggregation,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Blind image quality assessment based on high order statistics aggregation,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:54.235971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:15:46.255534Z digest=sha256:24e5d0ee7a9e1cb1dfb9459c5e542dd73f8ff5ba87169ef005f657d4e97ea292

Observation 0fb5fcef-ca42-4cee-ad5b-f95dd1083574 · outbound

This paper cites Learning without human scores for blind image quality assessment,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Learning without human scores for blind image quality assessment,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:54.125353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:15:46.336487Z digest=sha256:879ce8d70950c8dc0998fa7900efacaf5ad0e9cb88ff59252392b0ecf4cd93f1

Observation 58668ad9-9f78-4c1b-b000-54e70d6d5511 · outbound

This paper cites Making a “completely blind.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Making a “completely blind

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:54.055030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:15:46.415983Z digest=sha256:15d9da7a3041ecfc9f7413e7871c338768c70dc24a665290fe484e64b85b490d

Observation 147b5037-7061-4135-a851-e193d80a0556 · outbound

This paper cites Clipscore: A reference-free evaluation metric for image captioning,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Clipscore: A reference-free evaluation metric for image captioning,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:53.999272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:15:46.508413Z digest=sha256:df0c13dfabc493e6bfe563317237959a20c7a596f2064d2ae756d10b84ecbb69

Observation 0082d370-eec2-4389-a745-f24588fd8b08 · outbound

This paper cites Evaluating Text-to-Visual Generation with Image-to-Text Generation.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Evaluating Text-to-Visual Generation with Image-to-Text Generation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:46.589608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:46.589608Z digest=sha256:9316688af8c68797152efb8a36f1e0271c51474ac12b4227115363d8be5e9366

Observation fa6aed1c-e178-4085-a73d-c7cc251bcf1c · outbound

This paper cites LMM4LMM: Benchmarking and Evaluating Large-multimodal Image Generation with LMMs.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs LMM4LMM: Benchmarking and Evaluating Large-multimodal Image Generation with LMMs

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:46.647524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:46.647524Z digest=sha256:118565e82816b9a3ec9a6b7cef568c46445d436bc086c4dc4e221a6900419fb4

Observation 9720f9e0-4f87-4be1-ad52-edb8df975867 · outbound

This paper cites Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:46.707838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:46.707838Z digest=sha256:87f3537eecc4cebe39b00947086b3f2592822f14ed071371a3a8924cb86932f2

Observation b6ab865e-ded1-44f6-b9ce-1167a8e66290 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs An image is worth 16x16 words: Transformers for image recognition at scale,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:53.957123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:15:46.796154Z digest=sha256:df24d764d748805b8da1108d410a9582959c16a047abd31a53ba8d25808d06b9

Observation 514cfcbf-2390-4329-9e73-c59ea492ffa2 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs An image is worth 16x16 words: Transformers for image recognition at scale,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:53.808845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:15:46.889880Z digest=sha256:492a91ef41fb754396c4fb855d8113d15ca57cb14af68921a69ddc10d6bb88a8

Observation 2a2ac725-4e2f-4ca5-9db3-93e2c7a3cef6 · outbound

This paper cites MLP-Mixer: An all-MLP Architecture for Vision.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs MLP-Mixer: An all-MLP Architecture for Vision

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:46.972275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:46.972275Z digest=sha256:511b2043f4e5cca5636392aee8f0b09dad346819091d16fde68c316c8631d91d

Observation 9fc44e78-5762-44f8-b03f-022e60c60f95 · outbound

This paper cites How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:47.047930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:47.047930Z digest=sha256:16a713823ff18ca57fdae33011ff27afacf8ab1cc87ff8131fa1c2eaa230801d

Observation 8febaff3-2ab2-4aa5-b1d0-617cdeb1442c · outbound

This paper cites When Vision Transformers Outperform ResNets without Pre-training or Strong Data Augmentations.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs When Vision Transformers Outperform ResNets without Pre-training or Strong Data Augmentations

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:47.098749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:47.098749Z digest=sha256:6e309605fd924d5f78a4ce15eb7b342f10bd6bbc66e5bdeb4b54d65feb91be11

Observation 0db7e7a8-44cb-495a-88e1-b7ee9eb8c641 · outbound

This paper cites Surrogate gap minimization improves sharpness-aware training,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Surrogate gap minimization improves sharpness-aware training,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:53.597132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:15:47.146349Z digest=sha256:aa2804c3ea55106238f4ffcc425322049504eb434c331f6550e21aa133a4b13e

Observation 761438d2-410b-48c9-a6b6-59cb3f6a55f7 · outbound

This paper cites Lit: Zero-shot transfer with locked-image text tuning,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Lit: Zero-shot transfer with locked-image text tuning,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:53.438080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:15:47.196903Z digest=sha256:8a8698dbcfdc634e97da8ae4de40a10e2781544f9808aaab4955e0443be79558

Observation 1664f839-5ebe-4a7c-bb8f-b579c62ce79f · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:47.249382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:47.249382Z digest=sha256:c4e3ae451350f323babf1e343b71f798a1e841814972fe4939c60a072e6cad14

Observation 8adb38d0-4237-4e12-b5bd-e0eb6b3fe445 · outbound

This paper cites Mlp-net: Multilayer perceptron fusion network for infrared small target detection,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Mlp-net: Multilayer perceptron fusion network for infrared small target detection,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:53.213820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:15:47.303914Z digest=sha256:cff4fe8761bfc3b462b74c228c6f93352e88e34207eae1f08503d97fd548a868

Observation 65a60f14-dd82-401e-a048-46d7db18602a · outbound

This paper cites Lora: Low-rank adaptation of large language models,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Lora: Low-rank adaptation of large language models,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:52.987257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:15:47.364459Z digest=sha256:2ad3f796d601e9340c9cb087fb417186daf94cc17dff279dc4bf0d13b256081e

Observation d53a406b-4189-488a-99ed-9272ada4c374 · outbound

This paper cites Imagereward: learning and evaluating human preferences for text-to-image generation,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Imagereward: learning and evaluating human preferences for text-to-image generation,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:52.781780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:15:47.419599Z digest=sha256:8956f2539bc127398e566ca370b3d45a2b3641a317ce26d50b4c548760857687

Observation c95245d3-7d6c-402a-aa39-aec719e0a493 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision- language understanding and generation,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Blip: Bootstrapping language-image pre-training for unified vision- language understanding and generation,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:52.561473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:15:47.497508Z digest=sha256:82281d1643ffb1ad06993f6676d001bbaad5d9117de93b41139e227c8ceb104a

Observation 4f49277c-9c6d-4030-a140-a3ee6b1f4dde · outbound

This paper cites Pick-a-pic: An open dataset of user preferences for text-to-image generation,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Pick-a-pic: An open dataset of user preferences for text-to-image generation,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:52.397957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:15:47.556046Z digest=sha256:7ef153947a1e37d3803c184633548ac4b9423e34616ed16b9b64a94ba1873df8

Observation 11517637-b9b8-4fb5-a164-d3be77ecd053 · outbound

This paper cites Building cnn-based models for image aesthetic score prediction using an ensemble,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Building cnn-based models for image aesthetic score prediction using an ensemble,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:52.174546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:15:47.618742Z digest=sha256:534c428aa01ba25899ce8515b9788a12017920cb18b2655242409c792fe2ceac

Observation f8b2128c-3ff8-4a5c-b537-99a95b3bab61 · outbound

This paper cites Llava-next: A strong zero-shot video understanding model,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Llava-next: A strong zero-shot video understanding model,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:51.951951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:15:47.671411Z digest=sha256:2f9968d82ebe2bfb8efd5a245fa98a87a2b949d3abbed5aa93fe69368993d989

Observation 92317989-0820-4ead-b432-384fdfa782e0 · outbound

This paper cites Internvideo: General video foundation models via generative and discriminative learning,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Internvideo: General video foundation models via generative and discriminative learning,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:51.692478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:15:47.723918Z digest=sha256:49a231ba3251d58166e6d640d71a3afe216784c4fc12f78944913f4cadb3d46f

Observation b97954b8-5d3e-4053-a390-0a9036c01158 · outbound

This paper cites VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:47.793673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:47.793673Z digest=sha256:1f47bb354f461b0a9163ad707d96b8492b45461077e4404b97b778d89772cad4

Observation 16b6cae0-883d-4759-b36e-d73e8eec2a68 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:51.511413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:15:47.842241Z digest=sha256:58971f240ef2a8c2014520d97c6dc0e4c3d6e707a6dd3428814776fbd8a675c8

Observation dbc1d83f-1c7a-43b1-995f-c66bac60f979 · outbound

This paper cites mplug-owl3: Towards long image-sequence understanding in multi-modal large language models,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs mplug-owl3: Towards long image-sequence understanding in multi-modal large language models,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:51.344286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:15:47.905970Z digest=sha256:bdf35a049ea6ad1d2f867f868bf38173e13abbfade12d428763d5035575d2df9

Observation caaec614-1d36-46d3-b353-25615a657f3e · outbound

This paper cites No-reference image quality assessment in the spatial domain,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs No-reference image quality assessment in the spatial domain,

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:47.959442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:47.959442Z digest=sha256:b94dcdaf15034e0d5580178bc19c16bfa2cc717cdb20cf3f80044639139fbb8b

Observation e78e9d04-ab79-454a-9fff-8fdf3e641069 · outbound

This paper cites Blind image quality estimation via distortion aggravation,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Blind image quality estimation via distortion aggravation,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:51.198918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:15:48.029727Z digest=sha256:e3166d66fd6432880e107e4867df152b887d19cf6d284b3a89226b753539e87e

Observation 79cbeca4-a14e-441e-be03-e04a34a93f5b · outbound

This paper cites Blind quality assessment based on pseudo- reference image,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Blind quality assessment based on pseudo- reference image,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:51.004024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:15:48.090491Z digest=sha256:85ebe7a7379eccc7f705d85367e8ac98cd357b39742293d0ba4a7b7c6b386145

Observation 403b8b0a-6ae7-426a-a442-ceabac5f51bc · outbound

This paper cites Visual instruction tuning,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Visual instruction tuning,

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:48.138760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:48.138760Z digest=sha256:09e1c1256d236196d6377b689e8891e3925784808b95ccff1f9ded7a6b4ff3e7

Observation 1f8005a3-abc0-497d-a663-8d9b958afd0b · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:48.190049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:48.190049Z digest=sha256:a7499575a7e11a26a1a0ce1e6e613321bb9535c6ceae1707576f88d01ab7806e

Observation f883cf17-f2ed-408c-9b71-becbad8dd081 · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Llava-next: Improved reasoning, ocr, and world knowledge,

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:48.245414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:48.245414Z digest=sha256:ae6068f6c2d5af7099b64efce3d388364290772590fb8be74d2d3270ad8d3c7e

Observation a9df28ab-111c-4a68-821e-a7aaf4172c19 · outbound

This paper cites Improved baselines with visual instruction tuning,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Improved baselines with visual instruction tuning,

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:48.349429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:48.349429Z digest=sha256:58902ebe5cdfc583340065bf609b44e4865ee3b55d066cb2856d842a4d978239

Observation 865709a4-8de1-4b48-be07-011908404bcd · outbound

This paper cites Llava-next: Stronger llms supercharge multimodal capabilities in the wild,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Llava-next: Stronger llms supercharge multimodal capabilities in the wild,

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:50.758104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:15:48.412809Z digest=sha256:78b0423316d37d36ad434e19c20f498a2e1052502b26b44b1020c00806079c8e

Observation 8b7fa63b-bfa7-40b9-993b-d750b86428ca · outbound

This paper cites Llava-next: What else influences visual instruction tuning beyond data?,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Llava-next: What else influences visual instruction tuning beyond data?,

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:50.534207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:15:48.473759Z digest=sha256:d27263c99f4c941f3089a60c116b1f2e1b58f0417e1b2ea53d8b4d40aeeacd07

Observation 18370840-77bd-42b7-9442-253997febc9a · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:48.537796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:48.537796Z digest=sha256:65d77a221580216ad760d5e6d894a90da867ee90073efb9966540b6f3aeafd28

Observation 2a484005-937b-47b3-912e-9e815524a744 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:48.595108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:48.595108Z digest=sha256:55cb5e5ad91c4fffd84ab9aa6e0e53b14ad0181385d844b6826f979c5fc938dc

Observation 3e219ea3-4a27-42f1-9288-8bbe7e7f46c3 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Learning transferable visual models from natural language supervision,

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:50.332408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:15:48.673221Z digest=sha256:98a1ff201a963f7bd68cf82a0bb573f3e1e3fd6479a2cb91ce8415a2234eb600

Observation dba5ab64-a687-4600-bbb6-142a5c5a19b2 · outbound

This paper cites Unmasked teacher: Towards training- efficient video foundation models,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Unmasked teacher: Towards training- efficient video foundation models,

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:50.106302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:15:48.732449Z digest=sha256:8f258fbc06e39b6c342c8c67f101b9890797cc86381e38e7ca954d1047c0a3be

Observation 7e7a514f-68bf-41ce-aea6-7b24568ecc26 · outbound

This paper cites Stablevqa: A deep no-reference quality assessment model for video stability,.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Stablevqa: A deep no-reference quality assessment model for video stability,

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:49.870172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:15:48.810810Z digest=sha256:f3f71e9695fc6c751075437ccd52674fa8c9f755cdba964f3ebb4cd406644202

Observation f8d41227-00e3-4ca1-8fdf-781525ee55a7 · outbound

This paper cites Q-Align: Teaching LMMs for Visual Scoring via Discrete Text-Defined Levels.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Q-Align: Teaching LMMs for Visual Scoring via Discrete Text-Defined Levels

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:48.881635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:48.881635Z digest=sha256:890877e70d34d85c228715529e3468285ac8fb737140882fd06684f8dc5faf1e

Observation b589e717-2e61-441e-bf33-a1400fd02fa0 · outbound

This paper cites Excellent.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Excellent

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:49.655959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:15:48.932160Z digest=sha256:f0be0f290945ba1ec56b60629f33337c4a20125b8942959981e2eed6093eec2e

Observation 4d17d73d-ae3a-467f-bceb-a0c55182ad0f · outbound

This paper cites Excellent.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Excellent

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:49.409764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:15:48.995779Z digest=sha256:098e49bffc2dc7fe12044e98e2bb2dd03dcd49e8a2fb8a3b47c534dfd391343d

Pith citing papers

No inbound Pith citation observations are available.