Pith. sign in

Paper Citation Record · LEDGER

VGA-Bench: A Unified Benchmark and Multi-Model Framework for Video Aesthetics and Generation Quality Evaluation

As of 15 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 1 inbound Pith citation observation for arXiv:2604.10127.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.10127 v1

Coverage vector

measured 44 of 44 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T15:44:02.709823Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-09T22:50:19.910504Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-09T22:56:37.668187Z

Reference resolution

44 of 44 outbound references displayed

  • verified exact15
  • verified fuzzy27
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5cd1592e-c817-4749-9c02-3a0e60639e58 · outbound

This paper cites Vivit: A video vision transformer.

VGA-Bench: A Unified Benchmark and Multi-Model Framework for Video Aesthetics and Generation Quality Evaluation Vivit: A video vision transformer

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:00:07.223853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T15:44:02.709823Z digest=sha256:cf4d2da1a0d48ce33f3c0add80dc68f8f2ef1367fa293cea802f3a745fcee2d9

Observation 2052df63-a9ef-4b86-b837-f5097c252c7c · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

VGA-Bench: A Unified Benchmark and Multi-Model Framework for Video Aesthetics and Generation Quality Evaluation Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-11T09:56:04.400139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T15:44:02.709823Z digest=sha256:476749d0742edec90acbecaa74fdde5562057ef4a554ea3f27ee5d43ffb72548

Observation 418dfc18-4ce5-4868-bf82-e384f8092c66 · outbound

This paper cites Align your latents: High-resolution video synthesis with la- tent diffusion models.

VGA-Bench: A Unified Benchmark and Multi-Model Framework for Video Aesthetics and Generation Quality Evaluation Align your latents: High-resolution video synthesis with la- tent diffusion models

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:00:07.250894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T15:44:02.709823Z digest=sha256:c25c59f7306165d54e0366eb5a3e843dcd45af95daa5b3f6a3bc5a7106c7dd77

Observation ce9155aa-2eae-492e-8751-ee4a8a65b054 · outbound

This paper cites Routledge.

VGA-Bench: A Unified Benchmark and Multi-Model Framework for Video Aesthetics and Generation Quality Evaluation Routledge

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:00:07.257428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T15:44:02.709823Z digest=sha256:19d645cfe50c1c4f19d7c32730a679bfd42091f35a56454f878d6e8b59a219e0

Observation 5b8f2416-d163-4790-80cf-d11411a243d4 · outbound

This paper cites Emerg- ing properties in self-supervised vision transformers.

VGA-Bench: A Unified Benchmark and Multi-Model Framework for Video Aesthetics and Generation Quality Evaluation Emerg- ing properties in self-supervised vision transformers

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:00:07.213941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T15:44:02.709823Z digest=sha256:e4dd134171e187ac6a4e8c4c21a4bd11009704f411d2215db53615dd4c34f596

Observation baaafa21-3a94-4569-b7b0-7a0cb5d0ac04 · outbound

This paper cites Vlp: A survey on vision-language pre-training.Machine Intelligence Re- search, 20(1):38–56.

VGA-Bench: A Unified Benchmark and Multi-Model Framework for Video Aesthetics and Generation Quality Evaluation Vlp: A survey on vision-language pre-training.Machine Intelligence Re- search, 20(1):38–56

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:00:07.241212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T15:44:02.709823Z digest=sha256:5d0a630275fd9d12cb17122f6645a9001fdc775e69a5120d7ecfa0cb5f456b4f

Observation 26acd60d-1786-483d-82a4-8b1ae874494d · outbound

This paper cites Cinematography: the creative use of reality.

VGA-Bench: A Unified Benchmark and Multi-Model Framework for Video Aesthetics and Generation Quality Evaluation Cinematography: the creative use of reality

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:00:07.279292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T15:44:02.709823Z digest=sha256:6cd7a1d437f1ee8c5123b79bb44d3fba54b3c3869a8951e270f296f1dcf86266

Observation 40feeb99-cead-4061-9298-823f4be60d9f · outbound

This paper cites Coarse-to-fine vision-language pre-training with fusion in the backbone.Advances in neural information processing systems, 35:32942–32956.

VGA-Bench: A Unified Benchmark and Multi-Model Framework for Video Aesthetics and Generation Quality Evaluation Coarse-to-fine vision-language pre-training with fusion in the backbone.Advances in neural information processing systems, 35:32942–32956

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:00:07.234927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T15:44:02.709823Z digest=sha256:7d2ba111aaa470ff0081d49d59c9d21041dda712f6f28e3b4443527a5fe2a538

Observation e904f042-1628-4e14-a2fd-8e03a48c0556 · outbound

This paper cites Vision-language pre-training: Basics, re- cent advances, and future trends.Foundations and Trends® in Computer Graphics and Vision, 14(3–4):163–352.

VGA-Bench: A Unified Benchmark and Multi-Model Framework for Video Aesthetics and Generation Quality Evaluation Vision-language pre-training: Basics, re- cent advances, and future trends.Foundations and Trends® in Computer Graphics and Vision, 14(3–4):163–352

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:00:07.262503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T15:44:02.709823Z digest=sha256:bf6e0d90897479455cff8fed07afb81226dced756e478100278144af9b356495

Observation 87f8e292-022d-49b9-bf66-f4138d8e9778 · outbound

This paper cites AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning.

VGA-Bench: A Unified Benchmark and Multi-Model Framework for Video Aesthetics and Generation Quality Evaluation AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-11T09:56:04.194692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T15:44:02.709823Z digest=sha256:f766ebf3a37e8d87cc55af16baa59abb1f579892a9a53fd235816a67884b28af

Observation 37174476-cb64-4463-8992-f5eec209e928 · outbound

This paper cites LTX-Video: Realtime Video Latent Diffusion.

VGA-Bench: A Unified Benchmark and Multi-Model Framework for Video Aesthetics and Generation Quality Evaluation LTX-Video: Realtime Video Latent Diffusion

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:36:12.849504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T15:44:02.709823Z digest=sha256:80eb68f94da222a43c4e5eacb9e2b1eb91558fd395c9b3384af0fedbc3f12ef3

Observation 52b0e74e-1bd0-434c-9d45-69d4d481d760 · outbound

This paper cites Clipscore: A reference-free evaluation met- ric for image captioning.

VGA-Bench: A Unified Benchmark and Multi-Model Framework for Video Aesthetics and Generation Quality Evaluation Clipscore: A reference-free evaluation met- ric for image captioning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:00:07.254265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T15:44:02.709823Z digest=sha256:b561546270940b38950c6e9c851fd92d7b0008bd6d4052f0c7032c0534a31b32

Observation e70f44df-7637-4e66-ac39-67b48e8e16e6 · outbound

This paper cites Denoising dif- fusion probabilistic models.Advances in neural information processing systems, 33:6840–6851.

VGA-Bench: A Unified Benchmark and Multi-Model Framework for Video Aesthetics and Generation Quality Evaluation Denoising dif- fusion probabilistic models.Advances in neural information processing systems, 33:6840–6851

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:00:07.266577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T15:44:02.709823Z digest=sha256:626aea40c58c2a5026caa76e52a4958872d2a4d640b63439f2e3bd29c95891e2

Observation db2091d2-1935-467a-bd86-888bd5926bd4 · outbound

This paper cites Imagen Video: High Definition Video Generation with Diffusion Models.

VGA-Bench: A Unified Benchmark and Multi-Model Framework for Video Aesthetics and Generation Quality Evaluation Imagen Video: High Definition Video Generation with Diffusion Models

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-11T09:56:04.306752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T15:44:02.709823Z digest=sha256:1b17eecd7ce9da1f5a4b7cce1f263596a64817e245483ac82b2ce8e8423d0270

Observation 78d66841-5e30-4077-b0e5-0971476d6840 · outbound

This paper cites Vbench: Comprehensive bench- mark suite for video generative models.

VGA-Bench: A Unified Benchmark and Multi-Model Framework for Video Aesthetics and Generation Quality Evaluation Vbench: Comprehensive bench- mark suite for video generative models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:00:07.282670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T15:44:02.709823Z digest=sha256:ae4448c5722029aaabdf1d86abe31b6f60284fd110fc1026bbe38dc4bcdb187a

Observation 1aa89c0c-d43b-4fb3-b28a-cedbdcc742a5 · outbound

This paper cites Musiq: Multi-scale image quality transformer.

VGA-Bench: A Unified Benchmark and Multi-Model Framework for Video Aesthetics and Generation Quality Evaluation Musiq: Multi-scale image quality transformer

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:00:07.220339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T15:44:02.709823Z digest=sha256:40062d081a07c18768d5e4ecd5db428759efadfbce88533a083ab18a9d39f18d

Observation 67e6148b-8642-4073-a109-5741278131b3 · outbound

This paper cites Text2video-zero: Text- to-image diffusion models are zero-shot video generators.

VGA-Bench: A Unified Benchmark and Multi-Model Framework for Video Aesthetics and Generation Quality Evaluation Text2video-zero: Text- to-image diffusion models are zero-shot video generators

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:00:07.217172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T15:44:02.709823Z digest=sha256:ff567bdab5c42ce392d6ff7a6eff350a169abda68b079a454d70d2fee3bac935

Observation 7ffbe592-bb58-4e18-b186-f29f1bc35cdc · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

VGA-Bench: A Unified Benchmark and Multi-Model Framework for Video Aesthetics and Generation Quality Evaluation HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-11T09:56:04.201536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T15:44:02.709823Z digest=sha256:c20fb9c654a18df885f22503f613f42c616e2ff6508f585e5822dbefe2cfd01c

Observation 790860a8-65b7-48c0-beb2-1a438e60a970 · outbound

This paper cites Dit: Self-supervised pre-training for docu- ment image transformer.

VGA-Bench: A Unified Benchmark and Multi-Model Framework for Video Aesthetics and Generation Quality Evaluation Dit: Self-supervised pre-training for docu- ment image transformer

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:00:07.238122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T15:44:02.709823Z digest=sha256:43c81ca7e351c0a896ee27a912154bc035c91298125bb7a0657dec51b3ec477a

Observation 12fb8910-0c8c-41f7-ba62-1ef463557e78 · outbound

This paper cites Fetv: A bench- mark for fine-grained evaluation of open-domain text-to- video generation.Advances in Neural Information Process- ing Systems, 36:62352–62387.

VGA-Bench: A Unified Benchmark and Multi-Model Framework for Video Aesthetics and Generation Quality Evaluation Fetv: A bench- mark for fine-grained evaluation of open-domain text-to- video generation.Advances in Neural Information Process- ing Systems, 36:62352–62387

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:00:07.270107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T15:44:02.709823Z digest=sha256:84077861575dd6623f99e0995bdbe71ab591294df12fcedbea89e1e4c6e84eb2

Observation faddecd5-2b36-42b3-a0cd-8a9fa2401e6c · outbound

This paper cites Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models.

VGA-Bench: A Unified Benchmark and Multi-Model Framework for Video Aesthetics and Generation Quality Evaluation Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-13T13:43:11.511714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T15:44:02.709823Z digest=sha256:f37a2c290e265c5cfb2877491e1c2eac80c008d7dab679b41b1e00f9c96cb8e3

Observation f5710549-4471-4ff3-bdb6-c5a262d95b1d · outbound

This paper cites Video swin transformer.

VGA-Bench: A Unified Benchmark and Multi-Model Framework for Video Aesthetics and Generation Quality Evaluation Video swin transformer

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:00:07.247853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T15:44:02.709823Z digest=sha256:dea8b99d5b2eb8e92a2cc192c06873b8f8631713e2e6a30b1bf6867f02645ca4

Observation a2a1123d-922c-43f3-a1f0-c1291002f8b1 · outbound

This paper cites VideoFusion: Decomposed Diffusion Models for High-Quality Video Generation.

VGA-Bench: A Unified Benchmark and Multi-Model Framework for Video Aesthetics and Generation Quality Evaluation VideoFusion: Decomposed Diffusion Models for High-Quality Video Generation

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:56:04.349804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T15:44:02.709823Z digest=sha256:e4da7ff3bebc79f00009bd68f0dea072d3070e08d2257229804a41a762785d92

Observation b871c440-51ff-42d3-a484-971a3f62b589 · outbound

This paper cites Latte: Latent Diffusion Transformer for Video Generation.

VGA-Bench: A Unified Benchmark and Multi-Model Framework for Video Aesthetics and Generation Quality Evaluation Latte: Latent Diffusion Transformer for Video Generation

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-13T21:45:35.835056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T15:44:02.709823Z digest=sha256:7961d5a760d740709948f6d75c3f5f59843daaaf09f56c7aa32e017cddbcc923

Observation 85075812-f637-4202-b862-8a26f2f21723 · outbound

This paper cites an unresolved cited work.

VGA-Bench: A Unified Benchmark and Multi-Model Framework for Video Aesthetics and Generation Quality Evaluation Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-05-17T19:00:07.244632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T15:44:02.709823Z digest=sha256:d2c2a36171ae96f5c24c0c1eac0e92adeffeddcf1abffe98614496203104c66f

Observation 14018c56-0b63-4e72-93be-da25f77fb2fd · outbound

This paper cites Qianqian Qiao, DanDan Zheng, Yihang Bo, Bao Peng, Heng Huang, Longteng Jiang, Huaye Wang, Jingdong Chen, Jun Zhou, and Xin Jin.

VGA-Bench: A Unified Benchmark and Multi-Model Framework for Video Aesthetics and Generation Quality Evaluation Qianqian Qiao, DanDan Zheng, Yihang Bo, Bao Peng, Heng Huang, Longteng Jiang, Huaye Wang, Jingdong Chen, Jun Zhou, and Xin Jin

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T09:56:04.239420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T15:44:02.709823Z digest=sha256:66ffb6310a9dab2f3afd6cabe56015ac3ffdcd94aeddb8463ebf4772eecba30f

Observation dfecb999-6bac-4bb7-b2cf-ee8558909ef8 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

VGA-Bench: A Unified Benchmark and Multi-Model Framework for Video Aesthetics and Generation Quality Evaluation Learning transferable visual models from natural language supervi- sion

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:00:07.286336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T15:44:02.709823Z digest=sha256:e674c8d040c92b8eaef5943542cbd549f425048b0bfb63e25a34ffa973267af8

Observation 22a320d7-b28c-4cf0-8749-4a3baa32ce47 · outbound

This paper cites Video transformers: A survey.IEEE Transactions on Pattern Anal- ysis and Machine Intelligence, 45(11):12922–12943.

VGA-Bench: A Unified Benchmark and Multi-Model Framework for Video Aesthetics and Generation Quality Evaluation Video transformers: A survey.IEEE Transactions on Pattern Anal- ysis and Machine Intelligence, 45(11):12922–12943

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:00:07.295548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T15:44:02.709823Z digest=sha256:1a51e8a55b6d33c1ccb0769fbbd9b405701205b22ab12758f394611a4da2aaff

Observation bd1106fe-a096-4951-9ec4-fa49a8765e4a · outbound

This paper cites Make-A-Video: Text-to-Video Generation without Text-Video Data.

VGA-Bench: A Unified Benchmark and Multi-Model Framework for Video Aesthetics and Generation Quality Evaluation Make-A-Video: Text-to-Video Generation without Text-Video Data

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-11T09:56:04.222407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T15:44:02.709823Z digest=sha256:3fc695ddc0385f48b154590fa9b401d0a97c73dc97976257034621ba84621f84

Observation 9549349f-09a3-40de-bc7c-d055cc52871a · outbound

This paper cites Score-Based Generative Modeling through Stochastic Differential Equations.

VGA-Bench: A Unified Benchmark and Multi-Model Framework for Video Aesthetics and Generation Quality Evaluation Score-Based Generative Modeling through Stochastic Differential Equations

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-11T09:56:04.363702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T15:44:02.709823Z digest=sha256:dcc70d5707b22fe49052d85ccfbf3b1eeb324058a6f9b7c8d577dd42a2b37199

Observation 8463eb7c-70e3-4208-b764-2b0f51df1934 · outbound

This paper cites T2v-compbench: A comprehen- sive benchmark for compositional text-to-video generation.

VGA-Bench: A Unified Benchmark and Multi-Model Framework for Video Aesthetics and Generation Quality Evaluation T2v-compbench: A comprehen- sive benchmark for compositional text-to-video generation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:00:07.299014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T15:44:02.709823Z digest=sha256:554fb9f2ac543f154b97b0b55695ed68337dab27d3ab438a23ac0f17c4c5f2c6

Observation ecec6641-0f10-4b85-9e6b-078593bce66c · outbound

This paper cites Human-Centric Foundation Models: Perception, Generation and Agentic Modeling.

VGA-Bench: A Unified Benchmark and Multi-Model Framework for Video Aesthetics and Generation Quality Evaluation Human-Centric Foundation Models: Perception, Generation and Agentic Modeling

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:56:04.291715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T15:44:02.709823Z digest=sha256:df830671f2e93ba282088268cf24d0298b3af315c71e2be3d49421603272371f

Observation 10c5522a-96a5-44ee-a949-37a09cc3dd18 · outbound

This paper cites Mochi 1.https :/ /github .com/ genmoai/models.

VGA-Bench: A Unified Benchmark and Multi-Model Framework for Video Aesthetics and Generation Quality Evaluation Mochi 1.https :/ /github .com/ genmoai/models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:00:07.289840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T15:44:02.709823Z digest=sha256:223eedb2c37cb083bb226fe44cbd0980f9bb9358398f7fe582a58b67793c5712

Observation e9e5e516-09d9-4c14-ad71-102e0b4b142c · outbound

This paper cites Fvd: A new metric for video generation.

VGA-Bench: A Unified Benchmark and Multi-Model Framework for Video Aesthetics and Generation Quality Evaluation Fvd: A new metric for video generation

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:00:07.292452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T15:44:02.709823Z digest=sha256:7a5d532f7a5a4a63a912574c948c6051de321987e0555f5395c540001e2c300f

Observation 96a76333-4039-40a6-8be8-a661adf831c0 · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

VGA-Bench: A Unified Benchmark and Multi-Model Framework for Video Aesthetics and Generation Quality Evaluation Wan: Open and Advanced Large-Scale Video Generative Models

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-11T09:56:04.332815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T15:44:02.709823Z digest=sha256:77964b7d7966a02aa4eb6b590bc1df45aa20733e525aebc3cdb92f6c141dd192

Observation 2fa36a80-8dbb-40a1-b1ce-ab5a4886134e · outbound

This paper cites ModelScope Text-to-Video Technical Report.

VGA-Bench: A Unified Benchmark and Multi-Model Framework for Video Aesthetics and Generation Quality Evaluation ModelScope Text-to-Video Technical Report

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-12T19:47:29.701736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T15:44:02.709823Z digest=sha256:4b71cb3c1e7a5fe720afa2c45566e64ae97e54c094be1db4b0a579dcd4681f8c

Observation a912fea2-09af-4060-9495-478091b24005 · outbound

This paper cites Image as a foreign language: Beit pretraining for vision and vision- language tasks.

VGA-Bench: A Unified Benchmark and Multi-Model Framework for Video Aesthetics and Generation Quality Evaluation Image as a foreign language: Beit pretraining for vision and vision- language tasks

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:00:07.302338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T15:44:02.709823Z digest=sha256:928c920bd0694c6d0e84f87e2eeb9a1994d0d45eaadc2f29cac516487a90da07

Observation 85b0ddfa-802e-40b1-829a-39a90a36a227 · outbound

This paper cites Lavie: High-quality video generation with cascaded latent diffusion models.International Journal of Computer Vision, 133(5):3059–3078.

VGA-Bench: A Unified Benchmark and Multi-Model Framework for Video Aesthetics and Generation Quality Evaluation Lavie: High-quality video generation with cascaded latent diffusion models.International Journal of Computer Vision, 133(5):3059–3078

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:00:07.305485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T15:44:02.709823Z digest=sha256:1d2d601b174f157e912ef281df1f56f5ac844ebdb7e95dd590db91b3407669a7

Observation 3a97ed02-abad-49cf-b48b-d11fe8595dee · outbound

This paper cites Is your world simulator a good story presenter? a consecu- tive events-based benchmark for future long video genera- tion.

VGA-Bench: A Unified Benchmark and Multi-Model Framework for Video Aesthetics and Generation Quality Evaluation Is your world simulator a good story presenter? a consecu- tive events-based benchmark for future long video genera- tion

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:00:07.276224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T15:44:02.709823Z digest=sha256:01eeb45f38bef517fa7401d5b6edf4b7c1826016cb8033d660146dc9a60b3e2a

Observation 4de2634f-3b2e-416d-825d-56f5761e7210 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

VGA-Bench: A Unified Benchmark and Multi-Model Framework for Video Aesthetics and Generation Quality Evaluation CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-11T09:56:04.207177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T15:44:02.709823Z digest=sha256:8d60ae91cf361e73b39b99843ba22ee6cedc7ec92b8869b8459e4f4444e9c37d

Observation dae9b72d-a58e-44be-9e54-faa4665804ef · outbound

This paper cites Chronomagic-bench: A bench- mark for metamorphic evaluation of text-to-time-lapse video generation.Advances in Neural Information Processing Sys- tems, 37:21236–21270.

VGA-Bench: A Unified Benchmark and Multi-Model Framework for Video Aesthetics and Generation Quality Evaluation Chronomagic-bench: A bench- mark for metamorphic evaluation of text-to-time-lapse video generation.Advances in Neural Information Processing Sys- tems, 37:21236–21270

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:00:07.228003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T15:44:02.709823Z digest=sha256:96036c7586573f6d2789a64bf94cf44f3e916054e4ade9f66bd630a7ea9d12f3

Observation f7a13e18-dde9-4921-82f6-6f1397bf8c2b · outbound

This paper cites Show-1: Marrying pixel and latent diffusion models for text-to-video generation.International Journal of Com- puter Vision, 133(4):1879–1893.

VGA-Bench: A Unified Benchmark and Multi-Model Framework for Video Aesthetics and Generation Quality Evaluation Show-1: Marrying pixel and latent diffusion models for text-to-video generation.International Journal of Com- puter Vision, 133(4):1879–1893

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:00:07.273251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T15:44:02.709823Z digest=sha256:3d7ae652a5853538b818130953e330ead4de9f9c76fcd178978072ad66119ea7

Observation 296ebec5-19c5-4325-a6e7-6fc5dbda2b7f · outbound

This paper cites Adding conditional control to text-to-image diffusion models.

VGA-Bench: A Unified Benchmark and Multi-Model Framework for Video Aesthetics and Generation Quality Evaluation Adding conditional control to text-to-image diffusion models

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:00:07.231615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T15:44:02.709823Z digest=sha256:a1c6459ed72f189ed0d935bcbcd8a48428beab609237c873ddced3f8751b30cd

Observation c44fc420-4c93-4dcc-9607-6c66d6cea954 · outbound

This paper cites VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness.

VGA-Bench: A Unified Benchmark and Multi-Model Framework for Video Aesthetics and Generation Quality Evaluation VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:42:03.548540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T15:44:02.709823Z digest=sha256:1b095474e3272e3ec40276bbb63de5b713114ec57b2c23bc3ebb11175cdb786e

Pith citing papers

Observation f846dc8d-16f9-4d8f-af17-f167e909f6d7 · inbound

MADB: A Large-Scale Music Aesthetics Dataset with Professional and Multi-Dimensional Annotations cites this paper.

MADB: A Large-Scale Music Aesthetics Dataset with Professional and Multi-Dimensional Annotations VGA-Bench: A Unified Benchmark and Multi-Model Framework for Video Aesthetics and Generation Quality Evaluation

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-07-09T22:56:37.669414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-09T22:50:19.910504Z digest=sha256:0236f92a3f6ee25c2e0448ceabab40f19fab30067b4273058988eb473e9ddb90