Pith. sign in

Paper Citation Record · LEDGER

T2VTextBench: A Human Evaluation Benchmark for Textual Control in Video Generation Models

As of 17 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 2 inbound Pith citation observations for arXiv:2505.04946.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.04946 v1

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:19:46.019472Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:10:59.330611Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T00:07:28.151343Z

Reference resolution

23 of 23 outbound references displayed

  • verified exact3
  • verified fuzzy4
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ff374fb2-fcd8-4be1-90ab-cd42505ff5aa · outbound

This paper cites Text-to-image diffusion models cannot count, and prompt refinement cannot help.

T2VTextBench: A Human Evaluation Benchmark for Textual Control in Video Generation Models Text-to-image diffusion models cannot count, and prompt refinement cannot help

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T23:19:45.939265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:19:45.939265Z digest=sha256:eb785a8d1512289644a47f8d6be48b30edf7a537f485e6e36229831c19d2d66c

Observation bb76a354-c470-4734-9513-319da0526a52 · outbound

This paper cites High-Order Matching for One-Step Shortcut Diffusion Models.

T2VTextBench: A Human Evaluation Benchmark for Textual Control in Video Generation Models High-Order Matching for One-Step Shortcut Diffusion Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T23:19:45.943262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:19:45.943262Z digest=sha256:d9a6725bc2154f07423385a5093fc931f8a0960b0e2f198bcae32534f31bf226

Observation 59f134ff-d6f4-42a4-89db-921c29027afc · outbound

This paper cites Provable Failure of Language Models in Learning Majority Boolean Logic via Gradient Descent.

T2VTextBench: A Human Evaluation Benchmark for Textual Control in Video Generation Models Provable Failure of Language Models in Learning Majority Boolean Logic via Gradient Descent

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T23:19:45.947572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:19:45.947572Z digest=sha256:e8dcfbd4bd28605e53ca43823c36ea3f1a11256542d89a45465a27b4b472f51a

Observation 631b5329-b302-454a-bf18-1270a1c42b94 · outbound

This paper cites TC-Bench: Benchmarking Temporal Compositionality in Text-to-Video and Image-to-Video Generation.

T2VTextBench: A Human Evaluation Benchmark for Textual Control in Video Generation Models TC-Bench: Benchmarking Temporal Compositionality in Text-to-Video and Image-to-Video Generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T23:19:45.952770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:19:45.952770Z digest=sha256:efb634c6f1cd3018685a9fc7933ed12295a13a851290a837719df74136f5f9fd

Observation 56dee2c6-9eb8-416d-a47b-4aef9609a5cc · outbound

This paper cites Can You Count to Nine? A Human Evaluation Benchmark for Counting Limits in Modern Text-to-Video Models.

T2VTextBench: A Human Evaluation Benchmark for Textual Control in Video Generation Models Can You Count to Nine? A Human Evaluation Benchmark for Counting Limits in Modern Text-to-Video Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T23:19:45.957133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:19:45.957133Z digest=sha256:82269f860932ef3197df92b67fdff300456ffea006f05f5982b9e0bebc3b7fe6

Observation 9ff221b9-cba9-4a39-abba-aad082528da1 · outbound

This paper cites LTX-Video: Realtime Video Latent Diffusion.

T2VTextBench: A Human Evaluation Benchmark for Textual Control in Video Generation Models LTX-Video: Realtime Video Latent Diffusion

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T23:19:45.960976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:19:45.960976Z digest=sha256:bbabb221d61045872371b65b9b979bc0dcd5bad283b5155b98c8059964a31247

Observation b1636a2b-8414-46e6-8da9-bf50daf6f16f · outbound

This paper cites Wordart designer: User- driven artistic typography synthesis using large language models.

T2VTextBench: A Human Evaluation Benchmark for Textual Control in Video Generation Models Wordart designer: User- driven artistic typography synthesis using large language models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:19:46.489612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:19:45.964659Z digest=sha256:5517446a5535f543dd8214d4c15ddc9d781f24e9d51d1f24b2be992a539c9d52

Observation b8b4c785-cf5f-408e-860d-7161fff3e218 · outbound

This paper cites VBench++: Comprehensive and Versatile Benchmark Suite for Video Generative Models.

T2VTextBench: A Human Evaluation Benchmark for Textual Control in Video Generation Models VBench++: Comprehensive and Versatile Benchmark Suite for Video Generative Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T23:19:45.968548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:19:45.968548Z digest=sha256:18adb2687be3c40b7a2bab6e10b4f59292e77b1057adf5ebd5a130e0bf10c052

Observation f4929530-d18c-4c7c-a04e-2b7c74f51a4b · outbound

This paper cites On the expressive power of modern hopfield networks.

T2VTextBench: A Human Evaluation Benchmark for Textual Control in Video Generation Models On the expressive power of modern hopfield networks

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T23:19:45.972705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:19:45.972705Z digest=sha256:d3582aab08cc8b5e4f0506d82cf29e44f9ef64d5504723840a6d48b4b6647007

Observation 7b76c554-49a7-44b0-8173-f4ead2ec32f0 · outbound

This paper cites Text-Animator: Controllable Visual Text Video Generation.

T2VTextBench: A Human Evaluation Benchmark for Textual Control in Video Generation Models Text-Animator: Controllable Visual Text Video Generation

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-15T23:19:46.219317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:19:45.976269Z digest=sha256:7df86d8fb0fdf97ea7121de41a711379da4ee86ebfda14f774a6c2b112699b77

Observation 3ce2210d-bba0-4c29-af49-5cb70a629055 · outbound

This paper cites On the Computational Capability of Graph Neural Networks: A Circuit Complexity Bound Perspective.

T2VTextBench: A Human Evaluation Benchmark for Textual Control in Video Generation Models On the Computational Capability of Graph Neural Networks: A Circuit Complexity Bound Perspective

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T23:19:45.980222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:19:45.980222Z digest=sha256:8cd98a7893524ac878326a37d71dd0978f2f61b46b4691da3b6bff9526583055

Observation e1af7c76-f89a-49cd-bc2a-5480e0df4f2d · outbound

This paper cites PhyBench: A Physical Commonsense Benchmark for Evaluating Text-to-Image Models.

T2VTextBench: A Human Evaluation Benchmark for Textual Control in Video Generation Models PhyBench: A Physical Commonsense Benchmark for Evaluating Text-to-Image Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T23:19:45.987881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:19:45.987881Z digest=sha256:ba1c9f12923ccb996c6541c611e5ff545d3e0067d4574ab116de008cb148c3f3

Observation b84d62b2-cfd6-4861-a432-255d234cea72 · outbound

This paper cites BizGen: Advancing Article-level Visual Text Rendering for Infographics Generation.

T2VTextBench: A Human Evaluation Benchmark for Textual Control in Video Generation Models BizGen: Advancing Article-level Visual Text Rendering for Infographics Generation

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-15T23:19:46.164815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:19:45.992395Z digest=sha256:2bea094171e4991238de91e1e796311240d1113c38eb8632e029e4f6cd3e0715

Observation 047b4ec9-ce8a-4a93-b413-0b4091a44877 · outbound

This paper cites T2V-CompBench: A Comprehensive Benchmark for Compositional Text-to-video Generation.

T2VTextBench: A Human Evaluation Benchmark for Textual Control in Video Generation Models T2V-CompBench: A Comprehensive Benchmark for Compositional Text-to-video Generation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T23:19:46.000639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:19:46.000639Z digest=sha256:e35ad0068105c5229c9e5e5ccd9353e7f0a14d95e2fd384cba454f8871378ae0

Observation 9ad058c4-f0c3-4871-bb06-c41493397b9a · outbound

This paper cites Dolfin: Diffusion Layout Transformers without Autoencoder.

T2VTextBench: A Human Evaluation Benchmark for Textual Control in Video Generation Models Dolfin: Diffusion Layout Transformers without Autoencoder

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T23:19:46.004346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:19:46.004346Z digest=sha256:d6025f33c5624a59f896de775c0f716b4851fefa6a4790e8136214c94f71f453

Observation ec9180d6-b003-482b-b97d-6d1d5d894570 · outbound

This paper cites Alignab: Pareto- optimal energy alignment for designing nature-like antibodies.

T2VTextBench: A Human Evaluation Benchmark for Textual Control in Video Generation Models Alignab: Pareto- optimal energy alignment for designing nature-like antibodies

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T23:19:46.008400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:19:46.008400Z digest=sha256:7f2f1b214dc99f99e43c28e741c30d8260f64d8fad3eaa2c7884966545fcc276

Observation 7a42ca0a-b428-4bfb-988a-b485017e4892 · outbound

This paper cites Typedance: Creating semantic typographic logos from image through personalized generation.

T2VTextBench: A Human Evaluation Benchmark for Textual Control in Video Generation Models Typedance: Creating semantic typographic logos from image through personalized generation

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:19:46.477474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:19:46.012061Z digest=sha256:d6e51c1bbe374d2c0c6e55ba17c6e458bc9a05ae730ce3e401a1729f434c03b3

Observation 23400afb-78ac-427b-927c-309011a3091a · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

T2VTextBench: A Human Evaluation Benchmark for Textual Control in Video Generation Models CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T23:19:46.015668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:19:46.015668Z digest=sha256:21ef6a132b5931d011b7fdedafe7c8981f1ddd1a7b1cc4a4b75f5cdc45aa3a3a

Observation 9c3cb00f-e799-4bde-b641-efcc3d4ea1fb · outbound

This paper cites The dragon’s weakness?.

T2VTextBench: A Human Evaluation Benchmark for Textual Control in Video Generation Models The dragon’s weakness?

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:19:46.465012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:19:46.019472Z digest=sha256:0a4a35699d92382734e052d78bdd46dfc91c3a29ccd6a79fb6f8f99ada4d1a51

Observation df0ca87a-ac19-4919-863d-fc1e00ae2568 · outbound

This paper cites HOFAR: High-Order Augmentation of Flow Autoregressive Transformers.

T2VTextBench: A Human Evaluation Benchmark for Textual Control in Video Generation Models HOFAR: High-Order Augmentation of Flow Autoregressive Transformers

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-15T23:19:45.984087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:19:45.984087Z digest=sha256:3e2f361f5a1e72ddb672af7e31369fc21eb9ef9cdeb355fd0377a5f1e7974c48

Observation 3be70a02-f180-4c9e-b8e9-322690d25b05 · outbound

This paper cites Videophy: Eval- uating physical commonsense for video generation.

T2VTextBench: A Human Evaluation Benchmark for Textual Control in Video Generation Models Videophy: Eval- uating physical commonsense for video generation

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:19:46.503929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:19:45.931000Z digest=sha256:c03be6e8663b5edf0428c0231693f11198115b4dd02894d95e75534dbc7442a2

Observation 30b3f5d9-030c-4d11-bd7f-8c9de14228f9 · outbound

This paper cites Force Matching with Relativistic Constraints: A Physics-Inspired Approach to Stable and Efficient Generative Modeling.

T2VTextBench: A Human Evaluation Benchmark for Textual Control in Video Generation Models Force Matching with Relativistic Constraints: A Physics-Inspired Approach to Stable and Efficient Generative Modeling

Reference 2024

Resolution
verified exact
local_arxiv, observed 2026-08-15T23:19:46.442642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:19:45.934862Z digest=sha256:93679c5258576c13e1c834075e7fba43d71c9517b994c395f644d0e78c45e527

Observation 215aacc4-35ff-45d3-b38b-e8e278159f07 · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

T2VTextBench: A Human Evaluation Benchmark for Textual Control in Video Generation Models Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-15T23:19:45.925985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:19:45.925985Z digest=sha256:5f3a064462f67a9ee203495d4636a055e76012cef17c346de858bcb9c03357b8

Pith citing papers

Observation 7b4ff729-256f-4dc6-98e9-7ba0a5ad3299 · inbound

Only Large Weights (And Not Skip Connections) Can Prevent the Perils of Rank Collapse cites this paper.

Only Large Weights (And Not Skip Connections) Can Prevent the Perils of Rank Collapse T2VTextBench: A Human Evaluation Benchmark for Textual Control in Video Generation Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:59.330611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:59.330611Z digest=sha256:0891b89ff974407dcc0ff5ddc38099a0e035b8a948a22baeba3578eaaf069cc0

Observation 352aabec-db72-4833-8c91-ef0e4d5d7bd4 · inbound

CineDance: Towards Next-Generation Multi-Shot Long-Form Cinematic Audio-Video Generation cites this paper.

CineDance: Towards Next-Generation Multi-Shot Long-Form Cinematic Audio-Video Generation T2VTextBench: A Human Evaluation Benchmark for Textual Control in Video Generation Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:07:28.152664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T17:30:25.371658Z digest=sha256:03db15b30c92fc0bf47c15834072b4c97cfc33bdb557c5181562cf74cffa9785