Pith. sign in

Paper Citation Record · LEDGER

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations

As of 15 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 0 inbound Pith citation observations for arXiv:2501.07647.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.07647 v1

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T20:41:46.943342Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

55 of 55 outbound references displayed

  • verified exact1
  • verified fuzzy26
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4b6e5f49-3636-464c-85e6-6c87130acc2b · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T20:41:46.661593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:41:46.661593Z digest=sha256:1dd3241a0d2fb2945d6543ac8f7ce442dc6429c1de9fd92aa2d2d9f55381d308

Observation 98b22be7-b64b-49b6-bb56-d74b7c0df6e8 · outbound

This paper cites Align your latents: High-resolution video synthesis with la- tent diffusion models.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Align your latents: High-resolution video synthesis with la- tent diffusion models

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:41:47.940970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:41:46.668067Z digest=sha256:effb1cb6746285cc9f96174567159f0ed38ed64e51e05a6a3ba1b4594f91a774

Observation 631277c9-2b1b-4913-9743-51ef281ec728 · outbound

This paper cites VideoCrafter1: Open Diffusion Models for High-Quality Video Generation.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations VideoCrafter1: Open Diffusion Models for High-Quality Video Generation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T20:41:46.673882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:41:46.673882Z digest=sha256:4108c0e7e68efca914a55853dc391b1d4c24fbf4af84d1558ca077d644466816

Observation 982eb44b-35c5-4c2f-bacf-68319494733f · outbound

This paper cites Videocrafter2: Overcoming data limitations for high-quality video diffu- sion models.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Videocrafter2: Overcoming data limitations for high-quality video diffu- sion models

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:41:47.924160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:41:46.679977Z digest=sha256:3caca8ff67ddd4e4c01d56b3e073a07a0f763d6e592a64c0a860ab274f9dd9bd

Observation 1c833fd9-240e-4cdc-b39b-776bf5163bc5 · outbound

This paper cites Training-free layout control with cross-attention guidance.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Training-free layout control with cross-attention guidance

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T20:41:46.685336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:41:46.685336Z digest=sha256:0dc4de42eb724e68465cd09441d0b3af22b3d797bb32fee5a9411474fd47d249

Observation 3db8ca74-d25a-4698-b9b4-01745dd7cbac · outbound

This paper cites Panda-70m: Captioning 70m videos with multiple cross-modality teachers.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Panda-70m: Captioning 70m videos with multiple cross-modality teachers

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:41:47.894451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:41:46.690463Z digest=sha256:c76685d7e186b8214710da4030ee16f91f435645540227a3881b85b934d813f8

Observation 62167b68-ea64-405b-95cd-b3d5386b4fcd · outbound

This paper cites Scaling recti- fied flow transformers for high-resolution image synthesis.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Scaling recti- fied flow transformers for high-resolution image synthesis

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T20:41:46.695844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:41:46.695844Z digest=sha256:c8d51bfbacb41940173a097e3d9ab3e202696d6187fe36496bded98741f02945

Observation 491fb305-9892-4bae-a1d2-045effcab649 · outbound

This paper cites TC-Bench: Benchmarking Temporal Compositionality in Text-to-Video and Image-to-Video Generation.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations TC-Bench: Benchmarking Temporal Compositionality in Text-to-Video and Image-to-Video Generation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T20:41:46.702000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:41:46.702000Z digest=sha256:d7afa56df5b559cf9c989007b0d527c1ff0b5b2ea65a694fbe18c1df8d935b27

Observation aef14e19-ec4e-4cf6-a00e-429c19d20942 · outbound

This paper cites Layoutgpt: Compositional visual plan- ning and generation with large language models.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Layoutgpt: Compositional visual plan- ning and generation with large language models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:41:47.865422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:41:46.707689Z digest=sha256:8270afe704aec1f9c5d9287a220bd3a8360209b4496ef98bccf229bff7bc956c

Observation 6f1c1714-89cf-4df0-a3f1-5849036a0d2a · outbound

This paper cites On the content bias in fr ´echet video distance.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations On the content bias in fr ´echet video distance

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:41:47.847242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:41:46.712470Z digest=sha256:edf3e2b60173194e6665e8ba4a9593eec1ef57dbba402d30397481205f943cbb

Observation 6e6a0b24-f824-4afb-b223-96d9176a029d · outbound

This paper cites Prompt-to-Prompt Image Editing with Cross Attention Control.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Prompt-to-Prompt Image Editing with Cross Attention Control

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T20:41:46.717039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:41:46.717039Z digest=sha256:7ead31fe9c21d038958f974156c41dd934d8b55a8d7ccc7f54b79fd8830dd95c

Observation f14832ef-e18d-4d32-a077-8c94087197b6 · outbound

This paper cites Denoising dif- fusion probabilistic models.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Denoising dif- fusion probabilistic models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T20:41:46.722487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:41:46.722487Z digest=sha256:aa29b7938d7779e6dbe87566bd913196544a90af2f79fd7ecf482845a041d732

Observation a509f378-ab9a-4a66-b701-430e417f8341 · outbound

This paper cites Cogvideo: Large-scale pretraining for text-to-video generation via transformers.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Cogvideo: Large-scale pretraining for text-to-video generation via transformers

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:41:47.818003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:41:46.727435Z digest=sha256:55dcb6540d02e4245b26ef73493cc6f099fddf2dc2f7a5adba04db0af5695735

Observation b512f7f3-1b04-4d41-99d7-6e620314bd92 · outbound

This paper cites Perceiver IO: A General Architecture for Structured Inputs & Outputs.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Perceiver IO: A General Architecture for Structured Inputs & Outputs

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T20:41:46.732481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:41:46.732481Z digest=sha256:7b20c1217657457d2b7a994807deeb0fd602245facb25b70f5599ddeaa64f053

Observation c315cf22-2268-4589-8112-96615752f09f · outbound

This paper cites A style-based generator architecture for generative adversarial networks.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations A style-based generator architecture for generative adversarial networks

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:41:47.801856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:41:46.737703Z digest=sha256:343e52592e7af6b8c254402e1951768fca8b4e30c2f2b96346aa2f87b07cb710

Observation 0ea4ed25-7e74-4691-ac5a-8e870a1ef29e · outbound

This paper cites Open-sora-plan, 2024.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Open-sora-plan, 2024

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:41:47.785045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:41:46.743338Z digest=sha256:ce90ea046e3918c4af6e37d6b806f193621055552a2da637708c60db6508229a

Observation f326661a-2d8c-4cec-935a-a19fc11da731 · outbound

This paper cites Dense optical tracking: Connecting the dots.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Dense optical tracking: Connecting the dots

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:41:47.768539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:41:46.748409Z digest=sha256:a99de7b593c3c31c1c72d144641630fe2b3d4f67d5ed17f64b288c0892a021a7

Observation 4f48b987-b94a-4382-80d7-4a9ce3047b56 · outbound

This paper cites T2V-Turbo: Breaking the Quality Bottleneck of Video Consistency Model with Mixed Reward Feedback.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations T2V-Turbo: Breaking the Quality Bottleneck of Video Consistency Model with Mixed Reward Feedback

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T20:41:46.753046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:41:46.753046Z digest=sha256:c1af7e194495c186378c27a3f38f91cb8073dcbfdd579f59b4c08c57dc577b8e

Observation 28659284-8aa8-4c06-ab70-20953d6fe2ce · outbound

This paper cites TrackDiffusion: Tracklet-Conditioned Video Generation via Diffusion Models.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations TrackDiffusion: Tracklet-Conditioned Video Generation via Diffusion Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T20:41:46.757694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:41:46.757694Z digest=sha256:fd817969cfa9a492bd89c4fbcc7c4ea063a4b5b23549c0853085888a1a813086

Observation 2205b1eb-5f02-41b9-bafa-78dd3044cfba · outbound

This paper cites Gligen: Open-set grounded text-to-image generation.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Gligen: Open-set grounded text-to-image generation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:41:47.751990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:41:46.763002Z digest=sha256:743a3c02e2314e8a4b645eb7211b8f8b8a5a4912de6301dc7e271f2f94cbacd2

Observation 63ba12ee-8134-4f6c-821f-c4423f069a4f · outbound

This paper cites Llm- grounded diffusion: Enhancing prompt understanding of text-to-image diffusion models with large language models.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Llm- grounded diffusion: Enhancing prompt understanding of text-to-image diffusion models with large language models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:41:47.735277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:41:46.767551Z digest=sha256:ea42f1733d948fd90446bd1ba771c15dcb67af4f467a1d8293d616f3442aa17e

Observation 98d54c78-f21b-462e-af0b-3d79843d08ec · outbound

This paper cites Llm-grounded video diffusion models.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Llm-grounded video diffusion models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:41:47.718723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:41:46.772368Z digest=sha256:37d8d9979b17273ee76d41a139954fcddf7a447793faa3ae061c49a33aae9678

Observation 79d97a22-8d0e-4683-a42c-1cbf14c23562 · outbound

This paper cites VideoDirectorGPT: Consistent Multi-scene Video Generation via LLM-Guided Planning.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations VideoDirectorGPT: Consistent Multi-scene Video Generation via LLM-Guided Planning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T20:41:46.777075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:41:46.777075Z digest=sha256:5f3c7d44d5325dd0422bfbfc94a9ee9f377d977e1b4fd0792daa1085794cab8e

Observation 1ca446ca-1ebd-4616-90ab-485c5903ad16 · outbound

This paper cites MotionClone: Training-Free Motion Cloning for Controllable Video Generation.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations MotionClone: Training-Free Motion Cloning for Controllable Video Generation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T20:41:46.782104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:41:46.782104Z digest=sha256:d98c3beff3b9ac1e809b79734ace4e4c0186514edfcfc2ec5674e30f3d5168f2

Observation 9cd566d9-8960-433c-84e1-b57251091fef · outbound

This paper cites 9 Blobgen-3d: Compositional 3d-consistent freeview image generation with 3d blobs.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations 9 Blobgen-3d: Compositional 3d-consistent freeview image generation with 3d blobs

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:41:47.698845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:41:46.786747Z digest=sha256:3c13cd6310249bd24eb4c15141bdaccfd254ab26b9886ec78e4e25ff9f13c073

Observation e8486c0f-a363-4208-a11c-460b454bed70 · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:41:47.679691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:41:46.791106Z digest=sha256:5b29e49665eaf558c30222a71bd6b95a176d266a6eeacde59afa142547867a70

Observation 0e5f15a6-bbf0-4b47-9d48-ff02b26a31f0 · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T20:41:46.795590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:41:46.795590Z digest=sha256:71dbd0ab6f5690cc94f45f65a4b7f11bdab8a17f9034f14d29bac2c96f7a9e5a

Observation 725a678a-e2aa-42c1-bcc8-e348297e1adb · outbound

This paper cites Evalcrafter: Benchmarking and eval- uating large video generation models.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Evalcrafter: Benchmarking and eval- uating large video generation models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:41:47.662749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:41:46.801884Z digest=sha256:a7c0743fdea80fdb555b64cbb7b761607666acc7c076a6ecc6ab1a96d7f42dc0

Observation 935fbd6d-8b40-454a-9c1a-3a63c044764a · outbound

This paper cites Fetv: A bench- mark for fine-grained evaluation of open-domain text-to- video generation.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Fetv: A bench- mark for fine-grained evaluation of open-domain text-to- video generation

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:41:47.646108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:41:46.807182Z digest=sha256:35b0d90e2cb935056fe46dab029322674f14b91ecadf70554c1f6cbb098e340f

Observation 60011944-a7c1-480e-9db9-e76d8c7bf9bd · outbound

This paper cites OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T20:41:46.812519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:41:46.812519Z digest=sha256:b4e21e1d004e9b8aac547fcf2db08ac362bc47f12929c915926b983db5205f6b

Observation 05fafea7-0cf5-4277-9f27-4771e99299a0 · outbound

This paper cites Compositional text-to-image gen- eration with dense blob representations.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Compositional text-to-image gen- eration with dense blob representations

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:41:47.628927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:41:46.817879Z digest=sha256:b8a127e48f266c72c6927399620d78eedd3cae1e9820256ab319ed91d15d037f

Observation 2d717236-49e6-4c5e-ba08-857d2cbaa3eb · outbound

This paper cites Scalable diffusion models with transformers.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Scalable diffusion models with transformers

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T20:41:46.822782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:41:46.822782Z digest=sha256:85accf3ca6ced9849aa0922f35313d5f44d1285c6f2c26bbac6882f7db2b3cd7

Observation 13761c41-a74a-434e-8e8f-72247be31bd4 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations SAM 2: Segment Anything in Images and Videos

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T20:41:46.827788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:41:46.827788Z digest=sha256:60907a2aaa4a5eda1b09d49d10e15fdb3ca6ca49ba7f5bd6c2bb18b099d62e95

Observation 6873bc93-5efe-403c-87a4-62149957d438 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations High-resolution image synthesis with latent diffusion models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T20:41:46.833115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:41:46.833115Z digest=sha256:8915c692b0f9f603cd6511f123713502cbd2e89eebe8628146c13edfd918866e

Observation 62089b14-6a46-4765-b1c9-f1f6b602a0d2 · outbound

This paper cites T2V-CompBench: A Comprehensive Benchmark for Compositional Text-to-video Generation.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations T2V-CompBench: A Comprehensive Benchmark for Compositional Text-to-video Generation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T20:41:46.838351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:41:46.838351Z digest=sha256:aa56ad92eba656eb3344f8ae48b17ade4ae4d623a5fd03e2cffd27d771ac73fe

Observation 5aa572ee-3d91-417c-9f34-4d0026793aa7 · outbound

This paper cites VidGen-1M: A Large-Scale Dataset for Text-to-video Generation.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations VidGen-1M: A Large-Scale Dataset for Text-to-video Generation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T20:41:46.844463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:41:46.844463Z digest=sha256:19e8bcdf366bb969f6f8b3a5e445f698568f79bd241eab7bd3eac3fc0bfb1e52

Observation 14c488e3-40ff-4233-a5a9-a8794e84021b · outbound

This paper cites Fourier features let networks learn high frequency functions in low dimen- sional domains.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Fourier features let networks learn high frequency functions in low dimen- sional domains

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T20:41:46.850622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:41:46.850622Z digest=sha256:2b7d2859e31ecdccfd67d240f817e5e401d94708e61d10b348b73642ee697dcb

Observation ac317e7e-1472-4e97-b502-6692c27daa47 · outbound

This paper cites VideoTetris: Towards Compositional Text-to-Video Generation.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations VideoTetris: Towards Compositional Text-to-Video Generation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T20:41:46.855840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:41:46.855840Z digest=sha256:3d1e00e5a12b42620e8171153f3cc730c2e0dc09d73fbdaa3d0ea97eaca1a7f9

Observation 6fa053de-4929-4f99-af7b-6c46cfe7c4d7 · outbound

This paper cites Towards Accurate Generative Models of Video: A New Metric & Challenges.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Towards Accurate Generative Models of Video: A New Metric & Challenges

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T20:41:46.861690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:41:46.861690Z digest=sha256:c2f37e28addb26611f9bf417462cc0f7f71730b6d1a516caa7a5e56f080e8a04

Observation 42f5a13b-4a3e-41a6-840e-c7eb99a4043d · outbound

This paper cites Score-based generative modeling in latent space.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Score-based generative modeling in latent space

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:41:47.578327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:41:46.867718Z digest=sha256:a9509c2d57fa81dc2481d2efb2cdda5bacc749d9f24fa1c259335ace4292cdeb

Observation 8e568363-7a99-4ff4-88c9-4ab15332b295 · outbound

This paper cites ModelScope Text-to-Video Technical Report.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations ModelScope Text-to-Video Technical Report

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T20:41:46.872600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:41:46.872600Z digest=sha256:34e7829689036d45a558f4c005bb085b23e97afe9268c890d028c0d60f0b7470

Observation c1ce706c-3b10-4e44-bc29-965b62ff289c · outbound

This paper cites Boximator: Gener- ating rich and controllable motions for video synthesis.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Boximator: Gener- ating rich and controllable motions for video synthesis

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:41:47.562145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:41:46.877786Z digest=sha256:1a401ec8d75209aff71cc28ab8db6b7573837a8e61c0735c2b791649aa80893d

Observation 237149c0-cfe4-4a60-b919-c6db8fe2eca8 · outbound

This paper cites LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T20:41:46.883364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:41:46.883364Z digest=sha256:bdda3366efbad9b135c6d1452902e32e5b18f5339f506f47be6552e9c1532857

Observation f9b1cfec-798c-4e18-b41d-b42376453dfd · outbound

This paper cites Msr-vtt: A large video description dataset for bridging video and language.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Msr-vtt: A large video description dataset for bridging video and language

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T20:41:46.888799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:41:46.888799Z digest=sha256:14c164f7bb197e78b95f46287cbe62e4711c12c7f41f276154cf43cc96321571

Observation a6d7aa9d-038b-4415-9fa3-c9886b3c3a7f · outbound

This paper cites Open-vocabulary panop- tic segmentation with text-to-image diffusion models.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Open-vocabulary panop- tic segmentation with text-to-image diffusion models

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:41:47.535970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:41:46.893683Z digest=sha256:7d82e335a7c28aea3d85fe13a8c28147cff91039a680bb239b3124c540cb861a

Observation 29c53d94-d326-43b2-bc10-7c6520361e72 · outbound

This paper cites Ad- vancing high-resolution video-language representation with large-scale video transcriptions.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Ad- vancing high-resolution video-language representation with large-scale video transcriptions

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:41:47.519590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:41:46.898852Z digest=sha256:48141b578976bdc31e2d279b6a853507a902bab0683ea50d06180643f2f59c62

Observation ff31b245-ad15-466e-ac4c-245d8d1ddc02 · outbound

This paper cites The 3rd large-scale video object segmentation challenge - video in- stance segmentation track, 2021.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations The 3rd large-scale video object segmentation challenge - video in- stance segmentation track, 2021

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:41:47.503207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:41:46.903448Z digest=sha256:42f43e57bcc53699acd4bb1050e891112b8b347a24d83826518915523f42555e

Observation 1c8caf67-ceb5-4d66-87b9-2a5e3a4fed17 · outbound

This paper cites Compositional Video Generation as Flow Equalization.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Compositional Video Generation as Flow Equalization

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T20:41:46.907745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:41:46.907745Z digest=sha256:bdfbc1ed7fc5a5d098d1654a0955600549076504f87b5f0e6afb3db63c227d33

Observation b362f654-ab1f-418b-a79a-7247e9727bb1 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T20:41:46.912771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:41:46.912771Z digest=sha256:50a4edc01be3c34f4eea5ed9e3d9acc56e0c8fe8903cccb40c9d727bc857f735

Observation d215340d-085f-45be-b3e4-e7d42fc77633 · outbound

This paper cites Scannet++: A high-fidelity dataset of 3d in- door scenes.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Scannet++: A high-fidelity dataset of 3d in- door scenes

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:41:47.486813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:41:46.917408Z digest=sha256:a78a259d405cb3622683d8a0e0d189c0173b67333d63e3f4d1103ee82a3f0810

Observation 6ae457b4-859d-4308-aebf-f18053923296 · outbound

This paper cites Show-1: Marrying pixel and latent diffusion models for text-to-video generation.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Show-1: Marrying pixel and latent diffusion models for text-to-video generation

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T20:41:46.921729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:41:46.921729Z digest=sha256:a083d20091c5a3da9df540b1bed63f8a22f490f0bff56bf22e8e59bd46542340

Observation 127e9784-ed54-430e-b1b0-ff3e5fc55929 · outbound

This paper cites Adding conditional control to text-to-image diffusion models.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Adding conditional control to text-to-image diffusion models

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:41:47.458369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:41:46.926761Z digest=sha256:a96243d98943d97686c9f1dd2232bd4dce6808dbf4b49be3d8af53df1762b754

Observation 51f1e579-ef71-4ae9-9ec9-8cbd61e9839d · outbound

This paper cites Llava- next: A strong zero-shot video understanding model, 2024.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Llava- next: A strong zero-shot video understanding model, 2024

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:41:47.442037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:41:46.931591Z digest=sha256:8df430b10d27fd707ee7a40884fe389f19ebe1b93f33c1c4e4904db1d860f799

Observation 5c3d4a12-2050-450f-81be-34389c94427f · outbound

This paper cites both foreground and background.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations both foreground and background

Reference 54

Resolution
verified exact
raw_fallback, observed 2026-08-10T20:41:47.062580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:41:46.937177Z digest=sha256:cd011e18c8ea2da8971c8a6f5e98de4bb0a42ab473fdc2d5d9643876f49103fc

Observation 87f0d80a-6b97-42d8-b5aa-e73017baeb4f · outbound

This paper cites Frame0”: “Object2.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Frame0”: “Object2

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:41:47.424318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:41:46.943342Z digest=sha256:43668a9b52fda8ea2e0f59492991d7f6c55154464e5d3c7486e7b97ceb9fbaaf

Pith citing papers

No inbound Pith citation observations are available.