Pith. sign in

Paper Citation Record · LEDGER

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement

As of 17 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 0 inbound Pith citation observations for arXiv:2412.18966.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.18966 v1

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T01:04:35.416676Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

57 of 57 outbound references displayed

  • verified exact0
  • verified fuzzy23
  • unresolved34
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4f93188e-2dfe-400a-946b-8444a5f851d3 · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to-end retrieval.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Frozen in time: A joint video and image encoder for end-to-end retrieval

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:34.891030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:34.891030Z digest=sha256:ef7f28a80a226ab24571bcf8f24126a3518f8c26f1996cafde9d3d64bf233c46

Observation 2d506507-06d0-4fd2-8408-ee526a2c4173 · outbound

This paper cites Lumiere: A space-time dif- fusion model for video generation.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Lumiere: A space-time dif- fusion model for video generation

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:37.424442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T01:04:34.899673Z digest=sha256:20ca3c2b64efd6abcf92e84ef8c5c9300172ab0a4fc90d427944c2dceeb67204

Observation 65264e48-4f78-415f-93cf-24f23abd78c0 · outbound

This paper cites Improving image generation with better captions.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Improving image generation with better captions

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:37.391062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T01:04:34.909158Z digest=sha256:452867454f2fd49d5cc76b6a9d44d9328b483138a0049c013fdc0fb6b7ff1aee

Observation 7acdce46-e9ff-4954-84f5-bb8420c2d8ac · outbound

This paper cites Align your latents: High-resolution video synthesis with la- tent diffusion models.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Align your latents: High-resolution video synthesis with la- tent diffusion models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:34.916707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:34.916707Z digest=sha256:7a4a42e54cd82d32be674f3d2e3b2c160367dcc56bacd690944c019b18a1835a

Observation 1071c5b4-43d3-4666-ab46-58dc859088b2 · outbound

This paper cites In- structpix2pix: Learning to follow image editing instructions.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement In- structpix2pix: Learning to follow image editing instructions

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:34.924557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:34.924557Z digest=sha256:0b7b67915a0bdbfe59420124792ac231aa0e7866a56c4ae0ee7ec88bbc398631

Observation 7f99b4b2-d725-4e7b-bdcc-969ce323ba1b · outbound

This paper cites Video generation models as world simulators.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Video generation models as world simulators

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:34.934310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:34.934310Z digest=sha256:3fb51209c9d45e0d1e91c2b9203207f3b3c57e2060104e8480413fccb117f2fb

Observation 86ddf32c-274e-4583-bff1-365c83fb2e51 · outbound

This paper cites VideoCrafter1: Open Diffusion Models for High-Quality Video Generation.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement VideoCrafter1: Open Diffusion Models for High-Quality Video Generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:34.943735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:34.943735Z digest=sha256:5a67062d90853d0ec22ccae54c0bd03a4b7fd7bf098739bb9c1fc71c66661998

Observation 9afefb62-49e9-4ccb-be21-a32ee5563526 · outbound

This paper cites Videocrafter2: Overcoming data limitations for high-quality video diffu- sion models.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Videocrafter2: Overcoming data limitations for high-quality video diffu- sion models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:34.954429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:34.954429Z digest=sha256:be86e3fe3c0d26da4c60d1b61c4994efd7aa21b0ec230a2b1cff038550d1907d

Observation cf650a6b-a6d3-4881-a51a-48b5312e9fff · outbound

This paper cites Pixart-alpha: Fast training of diffusion transformer for photorealistic text-to-image syn- thesis.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Pixart-alpha: Fast training of diffusion transformer for photorealistic text-to-image syn- thesis

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:37.226592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T01:04:34.961354Z digest=sha256:67aa5a2799e937aa122f55d0f0c0904ef3f190aaed9a5b0a0389db450077966c

Observation 5e5f47e3-c7d1-43b0-a32b-271cabe9a253 · outbound

This paper cites Panda-70m: Captioning 70m videos with multiple cross-modality teachers.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Panda-70m: Captioning 70m videos with multiple cross-modality teachers

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:37.198557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T01:04:34.968168Z digest=sha256:9e76d3b7610aef72f40ecc7db27af6d38a17d94b0e6396c25e58aa847ae7fe11

Observation 1857c1e7-f202-4dc7-b3bc-572ceb3da728 · outbound

This paper cites Contin- ual pre-training mitigates forgetting in language and vision.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Contin- ual pre-training mitigates forgetting in language and vision

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:37.141649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T01:04:34.977168Z digest=sha256:06dfaf19a8565f11796abf3cac6d0dd4faadd42c62e59b88292c2178a25993b9

Observation d5edb5ab-538e-4b8e-ae38-07867dbf59e8 · outbound

This paper cites A continual learning survey: Defying for- getting in classification tasks.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement A continual learning survey: Defying for- getting in classification tasks

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:37.090386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T01:04:34.986248Z digest=sha256:0ad31265c57e030232c0529640217b54bafdb4c44ec998b872a0c6fde0af4ce1

Observation 6115f435-d074-4e1d-a2e9-6d24a305a36d · outbound

This paper cites Irc- gan: Introspective recurrent convolutional gan for text-to- video generation.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Irc- gan: Introspective recurrent convolutional gan for text-to- video generation

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:37.054938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T01:04:34.999548Z digest=sha256:ed428a9eb2302f145691085403d2f7a08142d5fec74b6d975103a0bb990033e0

Observation 8b915457-2b3e-463a-8ebf-507109f939a3 · outbound

This paper cites Catastrophic forgetting in connectionist networks.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Catastrophic forgetting in connectionist networks

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:37.020898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T01:04:35.009018Z digest=sha256:4adadb9e3f4f9af1b39945bc22cff742915bb3ef0a3bb45f283f803880b9175f

Observation eb0f6510-8aa8-4672-9a72-3c07a375f1bb · outbound

This paper cites Videostu- dio: Generating consistent-content and multi-scene videos.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Videostu- dio: Generating consistent-content and multi-scene videos

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:36.979001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T01:04:35.020241Z digest=sha256:7dbeb9cd538ba99659736fe3663e85910c848470bd6668f3e5d51e4193551aae

Observation c5a220a6-cc78-4066-bc0d-e24c36358e57 · outbound

This paper cites Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.030224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.030224Z digest=sha256:5ad83407241de3bb9dc238ab3169cd98dde44c7ab8d60d1f6ec177d4bc8efef8

Observation 9b0f21a0-c469-4180-8caf-89b54035d244 · outbound

This paper cites Preserve your own correlation: A noise prior for video diffusion models.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Preserve your own correlation: A noise prior for video diffusion models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:36.951409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T01:04:35.038115Z digest=sha256:b38a844ec238a681ba6c56890d5fd9e08a7435491237a1403a5a90c314df8f9f

Observation dbf4ee18-95f0-4d83-9c83-80c374b26fb5 · outbound

This paper cites SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.050124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.050124Z digest=sha256:4db25c9c97b3251d0e73a153c5d131c7bc3b12277c7a01a3f68eb2816d0d33a2

Observation 6ae13927-af9e-4072-9a43-94d2677e0788 · outbound

This paper cites Animatediff: Animate your personalized text-to- image diffusion models without specific tuning.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Animatediff: Animate your personalized text-to- image diffusion models without specific tuning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.056511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.056511Z digest=sha256:eb54021bc10c679cf44cd03e2313a36046bcade295e09c1613f4f1e95ca78e14

Observation e2eb497d-0d51-4551-bac0-1f4449f9b56d · outbound

This paper cites Photorealistic video generation with diffusion models.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Photorealistic video generation with diffusion models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:36.901018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T01:04:35.068618Z digest=sha256:aae158f4439f00fa421497348ff916893e32d4aa3b21050c4d3760f77c69e34a

Observation 54c4ca97-239e-448c-8ed0-cd1d32be011a · outbound

This paper cites Latent Video Diffusion Models for High-Fidelity Long Video Generation.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Latent Video Diffusion Models for High-Fidelity Long Video Generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.077390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.077390Z digest=sha256:c29540fa1a76ddb3534a17db1793aa5d92184299c3af79c2ba8844bb242cfed3

Observation 65b639a6-8f91-43ad-8743-471ad106b4f3 · outbound

This paper cites Animate-A-Story: Storytelling with Retrieval-Augmented Video Generation.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Animate-A-Story: Storytelling with Retrieval-Augmented Video Generation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.082441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.082441Z digest=sha256:2d5b86a5e573a78c40608d66708ff1190884699de5deef29c8cc41fbcc176ec0

Observation 4c28b58e-12cf-4e74-b374-64c44e730325 · outbound

This paper cites LLMs Meet Multimodal Generation and Editing: A Survey.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement LLMs Meet Multimodal Generation and Editing: A Survey

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.095299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.095299Z digest=sha256:9d59c4bc67fa4b124aa6fc05740fb1d09c9fd6af7b4ef9bb1855ed1b7b49f1fc

Observation 0ef79add-b938-4221-b8c6-13b346135884 · outbound

This paper cites Imagen Video: High Definition Video Generation with Diffusion Models.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Imagen Video: High Definition Video Generation with Diffusion Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.105992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.105992Z digest=sha256:0da46d4ec0e68b732d76bdb4986f1dd08e938ae0f20dabbc766a180f7909a884

Observation 6502a7d2-95b9-479f-87be-87b053ff6552 · outbound

This paper cites Video dif- fusion models.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Video dif- fusion models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.115211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.115211Z digest=sha256:4c6765628c085e0d923f4a2b6eced200122298d06b6d185c74677fa0196d2245

Observation 6b1d7871-dc3a-42d5-9900-de618ddb4c2c · outbound

This paper cites DirecT2V: Large Language Models are Frame-Level Directors for Zero-Shot Text-to-Video Generation.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement DirecT2V: Large Language Models are Frame-Level Directors for Zero-Shot Text-to-Video Generation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.123565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.123565Z digest=sha256:1d401a3f36dac923aa6a2af9b69cb17033405c3b00af9eb6f260b056336052e6

Observation 6df7b4bc-64e9-4b8c-ad73-677bb985d473 · outbound

This paper cites Parameter-efficient transfer learning for nlp.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Parameter-efficient transfer learning for nlp

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.129333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.129333Z digest=sha256:0a95704721798f2e52dbce48801b1a98b887dc64e803db14ffcbb18953d8cacf

Observation cfa864c4-37bc-4e12-8618-5eb4d23d8d6f · outbound

This paper cites Lora: Low- rank adaptation of large language models.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Lora: Low- rank adaptation of large language models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:36.814147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T01:04:35.136837Z digest=sha256:799feacaea4478cb8e8f5afe0b8ac698cfa8ab89d5a19cf391fe6842807e965a

Observation e25e9399-74f3-4834-82c9-e216993040ef · outbound

This paper cites VBench: Com- prehensive benchmark suite for video generative models.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement VBench: Com- prehensive benchmark suite for video generative models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:36.774461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T01:04:35.142817Z digest=sha256:91e6d825bf26d3025470e12edbf9b7035d6f7e9fef0750e7625d4db886e38fab

Observation a122b23e-ac11-46a5-a1ea-d759b41a5827 · outbound

This paper cites Simple and Scalable Strategies to Continually Pre-train Large Language Models.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Simple and Scalable Strategies to Continually Pre-train Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.150244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.150244Z digest=sha256:414d867d82502ef5216312801e34ffe75b8dfa23c78768c3b874d2a229a77774

Observation e0065cca-0473-4b6e-826a-66aa850c447e · outbound

This paper cites Continual pre-training of language models.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Continual pre-training of language models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:36.738658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T01:04:35.170485Z digest=sha256:be14703fdfb2e4f9dac5d3a6359f000e79fd298066fdffba857976d9878b3b84

Observation 17de3f09-e731-4c82-8bc1-b9407b894cb5 · outbound

This paper cites Open-sora-plan, 2024.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Open-sora-plan, 2024

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.184878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.184878Z digest=sha256:f6203afb0a62e38e73c02f56f32144366bde13af2e51673b1f2c7960c4db0b92

Observation 83037e88-3364-4874-9c81-06457a5e886b · outbound

This paper cites Llava-next: Tack- ling multi-image, video, and 3d in large multimodal models,.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Llava-next: Tack- ling multi-image, video, and 3d in large multimodal models,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.197675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.197675Z digest=sha256:5971d516e86b9227ede2e9415faacdef3cf4eb489128fdf812d504e864017ee5

Observation 99fc34b5-5ff3-4507-b61a-ceae3cb6d0ce · outbound

This paper cites Video generation from text.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Video generation from text

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.206474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.206474Z digest=sha256:deffeac7a38f69b1f820e43c3816627d461bca994db6c875d310125e09cfe934

Observation b1a42560-6736-442a-9285-dbd368053cf5 · outbound

This paper cites Llm-grounded video diffusion models.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Llm-grounded video diffusion models

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:36.630561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T01:04:35.223224Z digest=sha256:46503097b108e624ae013f71411e13e970751b5aa41c44df8742783f20c44dbb

Observation f69a88bc-708c-4782-a945-7b7a06fc6934 · outbound

This paper cites Latte: Latent Diffusion Transformer for Video Generation.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Latte: Latent Diffusion Transformer for Video Generation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.230155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.230155Z digest=sha256:f37f0aa825574a724017ec9f53c3e06cd32bb791d558a40fd7c054e68b198e05

Observation e6ca6569-8f76-452c-8ac1-19b4e607729b · outbound

This paper cites Snap video: Scaled spatiotemporal transformers for text-to-video synthesis.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Snap video: Scaled spatiotemporal transformers for text-to-video synthesis

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.240530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.240530Z digest=sha256:c3781eaff2210628408f11e80d05506d9b452a0c26e3e00a0c130d2fe44ae58d

Observation b69fd12d-e3c1-4381-8baf-34053845f54d · outbound

This paper cites Sync-draw: Automatic video generation using deep recurrent attentive architectures.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Sync-draw: Automatic video generation using deep recurrent attentive architectures

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:36.572032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T01:04:35.256364Z digest=sha256:97b145fe31fb62b6f48cba29b8b8cf54c2fd44fc4d4a127f62159e676bb899e8

Observation 5128715b-3fb4-45e2-a9ef-496d73edf303 · outbound

This paper cites Jour- neydb: A benchmark for generative image understanding,.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Jour- neydb: A benchmark for generative image understanding,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.263136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.263136Z digest=sha256:1efd03f1538da4ad248e3aaf5d8bf2834410b7c39bcd27638dc01e0bc9c07799

Observation f19c022c-0c68-418c-b541-b8aee8e2c3c4 · outbound

This paper cites Scalable diffusion models with transformers.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Scalable diffusion models with transformers

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.270449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.270449Z digest=sha256:c121f136437b98ebdcb8367aabc8e4622993ce0aec5c9b2f46e7930e8d57c536

Observation 8a78e60a-17bf-43ac-8009-6d75cfc50cf4 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Learning transferable visual models from natural language supervi- sion

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.278274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.278274Z digest=sha256:eb74135f7bcef6cac887feedaf51dceb1651a50fd11350064cef1b243759c23f

Observation b467d429-bced-4f93-830d-1607b803caee · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:36.483662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T01:04:35.282958Z digest=sha256:0c10441111a2d1489e2508f6c622e7bd4d6604f54599c6a7475b8c59c53418d5

Observation 3da6630f-6637-41ae-a223-0035c167fc5e · outbound

This paper cites T2V-CompBench: A Comprehensive Benchmark for Compositional Text-to-video Generation.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement T2V-CompBench: A Comprehensive Benchmark for Compositional Text-to-video Generation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.287887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.287887Z digest=sha256:705fa144d3bdf83888ba81bdc6be65f24dafd49e4596848a77c5785aa18c3adf

Observation 19552679-5593-49c2-9c43-f928984dc853 · outbound

This paper cites Emu: Generative Pretraining in Multimodality.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Emu: Generative Pretraining in Multimodality

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.295977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.295977Z digest=sha256:f3ed3e82694b19d277ec70e59a8186a57c98c1ba67f8c405836023c8d4d13d29

Observation e4111d41-5972-45a6-9894-bab94533913a · outbound

This paper cites Generative multimodal mod- els are in-context learners.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Generative multimodal mod- els are in-context learners

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:36.435782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T01:04:35.304811Z digest=sha256:ecb9d6d206a2a4f0328d279aefd03c68e4ea522834ab0a9b2756cf6634002888

Observation c01c39a3-e640-4fb8-9888-02961b128f2f · outbound

This paper cites Neural discrete representation learning.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Neural discrete representation learning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.316897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.316897Z digest=sha256:68e18896896d09aeb4f4089ffd9b452021752fc80773832f24b9095290c4a77f

Observation c7c2b8c9-8e60-4fe8-81f9-03bfbb476481 · outbound

This paper cites Magicvideo-v2: Multi-stage high-aesthetic video generation, 2024.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Magicvideo-v2: Multi-stage high-aesthetic video generation, 2024

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:36.384581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T01:04:35.323886Z digest=sha256:42ee156afa92aaff84bab388b0150fe7b230791b451ed824bc0c784de607a094

Observation 31901252-b885-434a-9e13-e2a4b7f0fde7 · outbound

This paper cites TRACE: A Comprehensive Benchmark for Continual Learning in Large Language Models.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement TRACE: A Comprehensive Benchmark for Continual Learning in Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.334537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.334537Z digest=sha256:7944dd3eab5c1a4f6ea8b680b5dd094412bf1afbec1e4581af35ce65fa707101

Observation 1c90472d-1505-4b1b-8dbc-f0ea85c79723 · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Emu3: Next-Token Prediction is All You Need

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.342459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.342459Z digest=sha256:88ad34d39fc210499cc1238742aca8dcc088c537623abed29d8e8b2ea57f36d6

Observation ef30cbfb-6a67-4517-b548-ff8ac4c01afd · outbound

This paper cites LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.349663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.349663Z digest=sha256:aee42f4fd3691256b10ac1f0b9a4e3e4a2a0d56c979148d3bff321ff28d4a2c7

Observation 1d6f23ce-6dc9-4cd2-a31e-5259de3f0449 · outbound

This paper cites LLaMA pro: Progressive LLaMA with block expansion.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement LLaMA pro: Progressive LLaMA with block expansion

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:36.330890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T01:04:35.358531Z digest=sha256:7e63980e18aa7f0d4ffe7bbdda3d0fcca9c9036481faad9df827efda647cf3b6

Observation cf330d11-9aa7-4b63-85bd-0c51fcab8702 · outbound

This paper cites Vript: A video is worth thousands of words, 2024.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Vript: A video is worth thousands of words, 2024

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:36.288520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T01:04:35.367105Z digest=sha256:e47806352ecab6e5a74fffc66b8502ef819bd58cbbc1bc7f906f454c661a7f27

Observation c755604e-2b24-4701-8e21-fe7597d3bb59 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.377149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.377149Z digest=sha256:683f5d004534ff4b48648f819bc305bca61d785d2fad54b838d06d8d0b85430d

Observation 5cc165ec-c1aa-4c8b-8db4-733edfaaf277 · outbound

This paper cites Language model beats diffusion-tokenizer is key to visual generation.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Language model beats diffusion-tokenizer is key to visual generation

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:36.250826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T01:04:35.385041Z digest=sha256:a6018e9f69e80a9855fa9970bdf7f37f6d0497686f5ad85eda153f1e221b2485

Observation 8f96385c-1f72-4902-8aac-1f7e543c5e7c · outbound

This paper cites Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.396189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.396189Z digest=sha256:319b1947f176e418dc1418422a1ced09034b0e4699911808441b3d649e03d686

Observation cb8924cb-da26-4933-a0ca-c805c9a3a61e · outbound

This paper cites Open-sora: Democratizing efficient video production for all, 2024.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Open-sora: Democratizing efficient video production for all, 2024

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:36.218328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T01:04:35.406134Z digest=sha256:8581a4f40dda06d90c011d1c0024c373a18a86c7c7ed1ce39a9e4f976cc09e67

Observation e05f910e-1c09-4a7a-834a-311f74a36604 · outbound

This paper cites MagicVideo: Efficient Video Generation With Latent Diffusion Models.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement MagicVideo: Efficient Video Generation With Latent Diffusion Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.416676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.416676Z digest=sha256:7ecc1010b05b46e1a936d69611e00da3a903222c927fc787e61bc4ab4a22dcd0

Pith citing papers

No inbound Pith citation observations are available.