Pith. sign in

Paper Citation Record · LEDGER

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement

As of 13 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 0 inbound Pith citation observations for arXiv:2412.18966.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.18966 v1

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T01:04:35.416676Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

57 of 57 outbound references displayed

  • verified exact0
  • verified fuzzy23
  • unresolved34
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4f93188e-2dfe-400a-946b-8444a5f851d3 · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to-end retrieval.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Frozen in time: A joint video and image encoder for end-to-end retrieval

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:34.891030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:34.891030Z digest=sha256:0b8e42dd00712cc05b182f472cfe3de66c84c7f7d843a07912e89de677d196a0

Observation 2d506507-06d0-4fd2-8408-ee526a2c4173 · outbound

This paper cites Lumiere: A space-time dif- fusion model for video generation.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Lumiere: A space-time dif- fusion model for video generation

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:37.424442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T01:04:34.899673Z digest=sha256:17caf6c6fb16b89d09316e0fafb8f44be3baca730cfbe3532a5672fbf0605140

Observation 65264e48-4f78-415f-93cf-24f23abd78c0 · outbound

This paper cites Improving image generation with better captions.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Improving image generation with better captions

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:37.391062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T01:04:34.909158Z digest=sha256:6f6f005ef9c9f7fe8725c87f2ddf5420a1a17064f16c51c69422dff1216212a4

Observation 7acdce46-e9ff-4954-84f5-bb8420c2d8ac · outbound

This paper cites Align your latents: High-resolution video synthesis with la- tent diffusion models.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Align your latents: High-resolution video synthesis with la- tent diffusion models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:34.916707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:34.916707Z digest=sha256:9982ae960ed2f85e26706c15504be089bdef1a10d33bb86139c33c38c1085aec

Observation 1071c5b4-43d3-4666-ab46-58dc859088b2 · outbound

This paper cites In- structpix2pix: Learning to follow image editing instructions.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement In- structpix2pix: Learning to follow image editing instructions

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:34.924557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:34.924557Z digest=sha256:d94cd9756685daaa0798e23a9c91f9162c051bf63afa20fdebdda7c0364124a3

Observation 7f99b4b2-d725-4e7b-bdcc-969ce323ba1b · outbound

This paper cites Video generation models as world simulators.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Video generation models as world simulators

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:34.934310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:34.934310Z digest=sha256:275240f3aee1c555bc3ee236ea9165710d73ee6effdedfcabd26bf6ccd9618e5

Observation 86ddf32c-274e-4583-bff1-365c83fb2e51 · outbound

This paper cites VideoCrafter1: Open Diffusion Models for High-Quality Video Generation.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement VideoCrafter1: Open Diffusion Models for High-Quality Video Generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:34.943735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:34.943735Z digest=sha256:9a0a8da7d5c40c809a2f780473b1a43ac356370413e2e4cfddd247323c6e63cc

Observation 9afefb62-49e9-4ccb-be21-a32ee5563526 · outbound

This paper cites Videocrafter2: Overcoming data limitations for high-quality video diffu- sion models.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Videocrafter2: Overcoming data limitations for high-quality video diffu- sion models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:34.954429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:34.954429Z digest=sha256:e35ee6b80f7ded01e198ee0a2cd5ee58551e02877381477e83819f5f43686b15

Observation cf650a6b-a6d3-4881-a51a-48b5312e9fff · outbound

This paper cites Pixart-alpha: Fast training of diffusion transformer for photorealistic text-to-image syn- thesis.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Pixart-alpha: Fast training of diffusion transformer for photorealistic text-to-image syn- thesis

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:37.226592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T01:04:34.961354Z digest=sha256:fbc6eb03a5feeb2bd2739bda333960ec7e35c93ff21017aa12626c59de6dac15

Observation 5e5f47e3-c7d1-43b0-a32b-271cabe9a253 · outbound

This paper cites Panda-70m: Captioning 70m videos with multiple cross-modality teachers.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Panda-70m: Captioning 70m videos with multiple cross-modality teachers

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:37.198557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T01:04:34.968168Z digest=sha256:83570dc9d724ecf20c784695e2862d95fd95e14f1cf67c11ece38a81556b0b2f

Observation 1857c1e7-f202-4dc7-b3bc-572ceb3da728 · outbound

This paper cites Contin- ual pre-training mitigates forgetting in language and vision.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Contin- ual pre-training mitigates forgetting in language and vision

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:37.141649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T01:04:34.977168Z digest=sha256:880f95bb2612398eec463aaab7d3c59f5ac2d27ca4a47731e63dc5b35432cdeb

Observation d5edb5ab-538e-4b8e-ae38-07867dbf59e8 · outbound

This paper cites A continual learning survey: Defying for- getting in classification tasks.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement A continual learning survey: Defying for- getting in classification tasks

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:37.090386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T01:04:34.986248Z digest=sha256:abf86883af034d7dfd93a2f72b680eaf91fdccaa86c4fa2f4a85db0587e8fabb

Observation 6115f435-d074-4e1d-a2e9-6d24a305a36d · outbound

This paper cites Irc- gan: Introspective recurrent convolutional gan for text-to- video generation.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Irc- gan: Introspective recurrent convolutional gan for text-to- video generation

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:37.054938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T01:04:34.999548Z digest=sha256:b400a5792b383a9652d5466708c0be1beeba427347727eba4b808d6c330fd605

Observation 8b915457-2b3e-463a-8ebf-507109f939a3 · outbound

This paper cites Catastrophic forgetting in connectionist networks.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Catastrophic forgetting in connectionist networks

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:37.020898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T01:04:35.009018Z digest=sha256:37bea58df1550e955520e32b86d87cad9ebe07573fc307a75a920d53230ced23

Observation eb0f6510-8aa8-4672-9a72-3c07a375f1bb · outbound

This paper cites Videostu- dio: Generating consistent-content and multi-scene videos.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Videostu- dio: Generating consistent-content and multi-scene videos

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:36.979001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T01:04:35.020241Z digest=sha256:f7b5844a4bbafdc54c564bba16ef4c16dbe1071e1ec7e3189e66f75f741d30a5

Observation c5a220a6-cc78-4066-bc0d-e24c36358e57 · outbound

This paper cites Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.030224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.030224Z digest=sha256:cbaeb817f8a0ebbf7941d4717f8edafbe9c883ae7dbba1003a97c7f57612dbfb

Observation 9b0f21a0-c469-4180-8caf-89b54035d244 · outbound

This paper cites Preserve your own correlation: A noise prior for video diffusion models.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Preserve your own correlation: A noise prior for video diffusion models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:36.951409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T01:04:35.038115Z digest=sha256:f6b6d61e4c9d67e080fc9549dd31d2059c5e66175b6587db4e31a961010079d6

Observation dbf4ee18-95f0-4d83-9c83-80c374b26fb5 · outbound

This paper cites SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.050124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.050124Z digest=sha256:cf6c5a1cad1a33d80b2fdad8db9bf65c32ef040dd006454c840084a4c439583d

Observation 6ae13927-af9e-4072-9a43-94d2677e0788 · outbound

This paper cites Animatediff: Animate your personalized text-to- image diffusion models without specific tuning.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Animatediff: Animate your personalized text-to- image diffusion models without specific tuning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.056511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.056511Z digest=sha256:deb33aff0846ce87dabbf4e6c65fdecc40a2ff92ec71f025b09542fb1b108694

Observation e2eb497d-0d51-4551-bac0-1f4449f9b56d · outbound

This paper cites Photorealistic video generation with diffusion models.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Photorealistic video generation with diffusion models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:36.901018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T01:04:35.068618Z digest=sha256:461d02a96a9e5ea78193fce71a117d30e83733e26a3be5c1a4243af19f94909f

Observation 54c4ca97-239e-448c-8ed0-cd1d32be011a · outbound

This paper cites Latent Video Diffusion Models for High-Fidelity Long Video Generation.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Latent Video Diffusion Models for High-Fidelity Long Video Generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.077390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.077390Z digest=sha256:c644da314638dc74fa9058511c5e5089304cd200ff85c178f18bb56fdbc875a5

Observation 65b639a6-8f91-43ad-8743-471ad106b4f3 · outbound

This paper cites Animate-A-Story: Storytelling with Retrieval-Augmented Video Generation.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Animate-A-Story: Storytelling with Retrieval-Augmented Video Generation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.082441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.082441Z digest=sha256:ab9bfd99cd57b65c885a3f5104395b9f0c9acb43bbe0943bdfdcf479cb8a3f40

Observation 4c28b58e-12cf-4e74-b374-64c44e730325 · outbound

This paper cites LLMs Meet Multimodal Generation and Editing: A Survey.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement LLMs Meet Multimodal Generation and Editing: A Survey

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.095299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.095299Z digest=sha256:080da1bc4b35219bc22dd732115168399b16b62082dee5439d68a04b1759d80b

Observation 0ef79add-b938-4221-b8c6-13b346135884 · outbound

This paper cites Imagen Video: High Definition Video Generation with Diffusion Models.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Imagen Video: High Definition Video Generation with Diffusion Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.105992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.105992Z digest=sha256:b6b88a37575d26dd00dc330c2ffa0dbf562bc46539fee068eefdf9e16aa002cf

Observation 6502a7d2-95b9-479f-87be-87b053ff6552 · outbound

This paper cites Video dif- fusion models.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Video dif- fusion models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.115211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.115211Z digest=sha256:d3e21400fda3556990ca9f9e49893fcb09e1b71859b1c30ecf0931dc83f3198f

Observation 6b1d7871-dc3a-42d5-9900-de618ddb4c2c · outbound

This paper cites DirecT2V: Large Language Models are Frame-Level Directors for Zero-Shot Text-to-Video Generation.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement DirecT2V: Large Language Models are Frame-Level Directors for Zero-Shot Text-to-Video Generation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.123565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.123565Z digest=sha256:1d3d111aeb037fb4580715bce27568da8991e92a985e3acdde8beddb85e7dfc9

Observation 6df7b4bc-64e9-4b8c-ad73-677bb985d473 · outbound

This paper cites Parameter-efficient transfer learning for nlp.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Parameter-efficient transfer learning for nlp

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.129333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.129333Z digest=sha256:03555a0e55a942645b74a954a824b02ca871391f2f5bb3f0329d1bcd4ddc09b0

Observation cfa864c4-37bc-4e12-8618-5eb4d23d8d6f · outbound

This paper cites Lora: Low- rank adaptation of large language models.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Lora: Low- rank adaptation of large language models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:36.814147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T01:04:35.136837Z digest=sha256:ea62dae68b267ba2b9429ea7c2110bcbadd7290e28d5461350ff3eab7f57338b

Observation e25e9399-74f3-4834-82c9-e216993040ef · outbound

This paper cites VBench: Com- prehensive benchmark suite for video generative models.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement VBench: Com- prehensive benchmark suite for video generative models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:36.774461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T01:04:35.142817Z digest=sha256:7523acf7bbb5082c5c8976510e9f5d229d670d6d126b999649d80f9d11f4c65e

Observation a122b23e-ac11-46a5-a1ea-d759b41a5827 · outbound

This paper cites Simple and Scalable Strategies to Continually Pre-train Large Language Models.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Simple and Scalable Strategies to Continually Pre-train Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.150244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.150244Z digest=sha256:058a3571d3e883325d43318f5dc437a7345479b9eadc9bb6d71813c18993537f

Observation e0065cca-0473-4b6e-826a-66aa850c447e · outbound

This paper cites Continual pre-training of language models.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Continual pre-training of language models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:36.738658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T01:04:35.170485Z digest=sha256:cbec5618287dfce66263e4b68764f5f502432f0d1ef64cd9417d3201d7732fd6

Observation 17de3f09-e731-4c82-8bc1-b9407b894cb5 · outbound

This paper cites Open-sora-plan, 2024.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Open-sora-plan, 2024

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.184878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.184878Z digest=sha256:9461bd3d9676ea92ec7a7d623c7e39a49bf0e47a45d99cf63f333c169d7310de

Observation 83037e88-3364-4874-9c81-06457a5e886b · outbound

This paper cites Llava-next: Tack- ling multi-image, video, and 3d in large multimodal models,.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Llava-next: Tack- ling multi-image, video, and 3d in large multimodal models,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.197675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.197675Z digest=sha256:b8f636a7339f112605f9e9597ef6eee1de34c31bc361bcff795c17e8e77d632e

Observation 99fc34b5-5ff3-4507-b61a-ceae3cb6d0ce · outbound

This paper cites Video generation from text.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Video generation from text

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.206474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.206474Z digest=sha256:38488393a2e2f4adfb859779e27cc5f973425d556485ab5dca10b7b874cecb3b

Observation b1a42560-6736-442a-9285-dbd368053cf5 · outbound

This paper cites Llm-grounded video diffusion models.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Llm-grounded video diffusion models

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:36.630561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T01:04:35.223224Z digest=sha256:1d82b95cee580025c03db00680e089653fee26259d3e6231f7073dfaef85dcf5

Observation f69a88bc-708c-4782-a945-7b7a06fc6934 · outbound

This paper cites Latte: Latent Diffusion Transformer for Video Generation.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Latte: Latent Diffusion Transformer for Video Generation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.230155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.230155Z digest=sha256:4abd5ed14d564c78ccfdeb46a122b7c88be4787e636a9110a3d32f7773522cff

Observation e6ca6569-8f76-452c-8ac1-19b4e607729b · outbound

This paper cites Snap video: Scaled spatiotemporal transformers for text-to-video synthesis.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Snap video: Scaled spatiotemporal transformers for text-to-video synthesis

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.240530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.240530Z digest=sha256:5030a380664f8dbc2da6bebcb187c6aedc7c842cee381c7dc49b537f8e6eb1c0

Observation b69fd12d-e3c1-4381-8baf-34053845f54d · outbound

This paper cites Sync-draw: Automatic video generation using deep recurrent attentive architectures.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Sync-draw: Automatic video generation using deep recurrent attentive architectures

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:36.572032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T01:04:35.256364Z digest=sha256:882e426306cc04e3768da4fcf1bf84521f278d07358d51f745ccdef7305195a5

Observation 5128715b-3fb4-45e2-a9ef-496d73edf303 · outbound

This paper cites Jour- neydb: A benchmark for generative image understanding,.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Jour- neydb: A benchmark for generative image understanding,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.263136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.263136Z digest=sha256:53ddfa2d2dc0d6065d552e4dac2a6dda46e30c6a902e659715ec6bfe0ee03ac6

Observation f19c022c-0c68-418c-b541-b8aee8e2c3c4 · outbound

This paper cites Scalable diffusion models with transformers.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Scalable diffusion models with transformers

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.270449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.270449Z digest=sha256:c52309936666b91819d1a45080a8904c54a464ff00173b2011a52b06364bd180

Observation 8a78e60a-17bf-43ac-8009-6d75cfc50cf4 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Learning transferable visual models from natural language supervi- sion

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.278274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.278274Z digest=sha256:10b56541dae833bbd7f29fce1275122fa59f4d47d35e9e34de28f81ab8d19be5

Observation b467d429-bced-4f93-830d-1607b803caee · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:36.483662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T01:04:35.282958Z digest=sha256:4be1e829311ee52593ab937cccbb4e69218626469bc917b0d3912555c80ed4bf

Observation 3da6630f-6637-41ae-a223-0035c167fc5e · outbound

This paper cites T2V-CompBench: A Comprehensive Benchmark for Compositional Text-to-video Generation.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement T2V-CompBench: A Comprehensive Benchmark for Compositional Text-to-video Generation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.287887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.287887Z digest=sha256:ad918144a170ce2cd5fb4eeb4db3ab7dac33368f4946f6b7fcf7210ff81220b8

Observation 19552679-5593-49c2-9c43-f928984dc853 · outbound

This paper cites Emu: Generative Pretraining in Multimodality.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Emu: Generative Pretraining in Multimodality

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.295977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.295977Z digest=sha256:22ed437b37cb321f0c3cba7cbed67b15bb2fd955a4a251b2a97734c0ca026e8a

Observation e4111d41-5972-45a6-9894-bab94533913a · outbound

This paper cites Generative multimodal mod- els are in-context learners.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Generative multimodal mod- els are in-context learners

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:36.435782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T01:04:35.304811Z digest=sha256:9cbd2bd2ec47cf4408d63de22de72d5a23bcecf023e38d05ef2d0313e228bfaf

Observation c01c39a3-e640-4fb8-9888-02961b128f2f · outbound

This paper cites Neural discrete representation learning.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Neural discrete representation learning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.316897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.316897Z digest=sha256:d61defd65726ff78303aa73856d3c84d9d8e2fa34aceb7af8fd4e0d6e772a9c9

Observation c7c2b8c9-8e60-4fe8-81f9-03bfbb476481 · outbound

This paper cites Magicvideo-v2: Multi-stage high-aesthetic video generation, 2024.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Magicvideo-v2: Multi-stage high-aesthetic video generation, 2024

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:36.384581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T01:04:35.323886Z digest=sha256:2ea23e579133b2deec907c2787c58ba8f7c9c7b87578108e2aaccea6ce7fa01c

Observation 31901252-b885-434a-9e13-e2a4b7f0fde7 · outbound

This paper cites TRACE: A Comprehensive Benchmark for Continual Learning in Large Language Models.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement TRACE: A Comprehensive Benchmark for Continual Learning in Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.334537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.334537Z digest=sha256:bc194de8b6b83340f6af4fbb21d49d60e8696c7d2d6d13f453678c958ec883ff

Observation 1c90472d-1505-4b1b-8dbc-f0ea85c79723 · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Emu3: Next-Token Prediction is All You Need

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.342459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.342459Z digest=sha256:58dae00e2b5231aec59686d1141179ecffd7d1366632ecb3dbeff1164dad75a9

Observation ef30cbfb-6a67-4517-b548-ff8ac4c01afd · outbound

This paper cites LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.349663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.349663Z digest=sha256:8cac0c571921660b44112e0457f4b40e68d3460562d12e99a24e4a192f14cfd6

Observation 1d6f23ce-6dc9-4cd2-a31e-5259de3f0449 · outbound

This paper cites LLaMA pro: Progressive LLaMA with block expansion.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement LLaMA pro: Progressive LLaMA with block expansion

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:36.330890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T01:04:35.358531Z digest=sha256:5949a380c033da97c2a5c46bdb09de469519f074af3b873500dd42ac84a28974

Observation cf330d11-9aa7-4b63-85bd-0c51fcab8702 · outbound

This paper cites Vript: A video is worth thousands of words, 2024.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Vript: A video is worth thousands of words, 2024

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:36.288520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T01:04:35.367105Z digest=sha256:be3a76f6b890feccf2740ba0036fc2c877f8c7546af755b75cbc5953de0c71a1

Observation c755604e-2b24-4701-8e21-fe7597d3bb59 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.377149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.377149Z digest=sha256:0b52b8afb67ca5c435180bab6e62ca52db8d920353f6cab42648cef013806482

Observation 5cc165ec-c1aa-4c8b-8db4-733edfaaf277 · outbound

This paper cites Language model beats diffusion-tokenizer is key to visual generation.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Language model beats diffusion-tokenizer is key to visual generation

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:36.250826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T01:04:35.385041Z digest=sha256:6c933630f01a0605fab28eb2efe3ff1a32fa0ee8fe2c3393411df221e0f7e04a

Observation 8f96385c-1f72-4902-8aac-1f7e543c5e7c · outbound

This paper cites Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.396189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.396189Z digest=sha256:051c4fa8ccee0b91068dc241ea84ca7f4af57c8d3413ce36c6a934a4861b46da

Observation cb8924cb-da26-4933-a0ca-c805c9a3a61e · outbound

This paper cites Open-sora: Democratizing efficient video production for all, 2024.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Open-sora: Democratizing efficient video production for all, 2024

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:36.218328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T01:04:35.406134Z digest=sha256:0e8d1a5ce578fec4d33ca7b37d272ff36dfe4abca1b901898bb2280f87a5269d

Observation e05f910e-1c09-4a7a-834a-311f74a36604 · outbound

This paper cites MagicVideo: Efficient Video Generation With Latent Diffusion Models.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement MagicVideo: Efficient Video Generation With Latent Diffusion Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.416676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.416676Z digest=sha256:bcfcb767dc1156b080e944ee01f36cda1908f645c8e43e87acc0acc6240c32fd

Pith citing papers

No inbound Pith citation observations are available.