Pith. sign in

Paper Citation Record · LEDGER

Mimir: Improving Video Diffusion Models for Precise Text Understanding

As of 13 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 3 inbound Pith citation observations for arXiv:2412.03085.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.03085 v1

Coverage vector

measured 68 of 68 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T22:51:11.861299Z

measured 71 of 71 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T21:09:07.991313Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-15T01:53:28.925576Z

Reference resolution

68 of 68 outbound references displayed

  • verified exact0
  • verified fuzzy23
  • unresolved45
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fa3d78dd-059b-486f-854a-f07c0d1cf60d · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.612873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.612873Z digest=sha256:52ec2daaba9b53e258edf6ae703b6540c6fa6ac828833ce2148fa301fd591a04

Observation 174ce6c1-871c-49fe-be62-2eee06cdd298 · outbound

This paper cites Qwen Technical Report.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Qwen Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.617764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.617764Z digest=sha256:668a4f7b2ac6c40e39d2bfa75d1bb9fbf19f8b0f51faf85a1d5e94e0b864e02a

Observation 93c2fe9c-38a0-42f0-b0c2-532a69bedd52 · outbound

This paper cites eDiff-I: Text-to-Image Diffusion Models with an Ensemble of Expert Denoisers.

Mimir: Improving Video Diffusion Models for Precise Text Understanding eDiff-I: Text-to-Image Diffusion Models with an Ensemble of Expert Denoisers

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.621761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.621761Z digest=sha256:bcd1b857e43348c7c6752673946e2600770b0ecb0741fff0f0170679aebe4472

Observation 31fbc9c4-a06d-4694-896e-45314cf6334a · outbound

This paper cites LLM2Vec: Large Language Models Are Secretly Powerful Text Encoders.

Mimir: Improving Video Diffusion Models for Precise Text Understanding LLM2Vec: Large Language Models Are Secretly Powerful Text Encoders

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.626353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.626353Z digest=sha256:c3ef01dd4535f0c254c8dc053e05adc341c622effb706e3f7b868b24e59fc988

Observation a5492bb5-849a-4f41-a4fb-73ab32b8f523 · outbound

This paper cites Improving image generation with better captions.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Improving image generation with better captions

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.630859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.630859Z digest=sha256:8896424a27ab9b5b5b1ab4e0a5a8e34181606409dddb51e36e4fed5d8d26de26

Observation dc2703c4-4df9-446b-ae3f-c6d6f1b7fdb4 · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.634791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.634791Z digest=sha256:a14105626b616fe65f58314e1aa00c4552c4cc952ee6f2ede8f2c0165d9c16d6

Observation f009921e-e3c9-4f5c-aa70-1002c1168767 · outbound

This paper cites PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis.

Mimir: Improving Video Diffusion Models for Precise Text Understanding PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.639492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.639492Z digest=sha256:2ae428c9dc87613280faf5a2a97dbbd72dfd460863ec69238a826946a0c444ff

Observation 5b167563-b1c0-4aff-a3e9-c7d1ffa9c571 · outbound

This paper cites PixArt-\Sigma: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation.

Mimir: Improving Video Diffusion Models for Precise Text Understanding PixArt-\Sigma: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.643736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.643736Z digest=sha256:b9907bb0a2ba4a701ce579771cf393831ec277dc08d2967f8a4b4aca2edbc26c

Observation de3a697d-3f46-4bb7-ad16-68da13c74268 · outbound

This paper cites PIXART-{\delta}: Fast and Controllable Image Generation with Latent Consistency Models.

Mimir: Improving Video Diffusion Models for Precise Text Understanding PIXART-{\delta}: Fast and Controllable Image Generation with Latent Consistency Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.648408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.648408Z digest=sha256:d13ea4729a8b8175f3d1b3449c836be2616022b26eb973a0991b55a8e7e8c09c

Observation 66913913-901d-4840-89c9-15821d23ebca · outbound

This paper cites OD-VAE: An Omni-dimensional Video Compressor for Improving Latent Video Diffusion Model.

Mimir: Improving Video Diffusion Models for Precise Text Understanding OD-VAE: An Omni-dimensional Video Compressor for Improving Latent Video Diffusion Model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.651909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.651909Z digest=sha256:52d6b2517141e97b9bb0976c4e89a8f1edd3768214a6f051d2c821c8ff046027

Observation b244d4ae-c421-41b0-a980-c1150df05139 · outbound

This paper cites Taming transformers for high-resolution image synthesis.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Taming transformers for high-resolution image synthesis

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.656089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.656089Z digest=sha256:e718656b410e150bca0b41af032f7fe4220df9d66a4f92c9e030162088f71b68

Observation 62d3c6c2-d215-4a35-b549-8674f9f8b614 · outbound

This paper cites Scaling rectified flow transformers for high-resolution image synthesis.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Scaling rectified flow transformers for high-resolution image synthesis

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:51:12.504267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T22:51:11.659739Z digest=sha256:e3c6ff49206f3a0ad36868c0dd72d7a8c4fafd68ec072202a2c0fd4b27fc159f

Observation a619a1c5-5332-4ef9-ba53-7b5841425048 · outbound

This paper cites Perceptual quality assessment of smartphone photography.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Perceptual quality assessment of smartphone photography

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.663489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.663489Z digest=sha256:07b5aa7f9e253f7a887ee186a47dfe0aab6d82cf1d863d2d6316c5f524b4da80

Observation e2a60417-8b0d-4c08-9027-32ec10e61a03 · outbound

This paper cites Ranni: Taming text-to-image diffusion for accurate instruction following.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Ranni: Taming text-to-image diffusion for accurate instruction following

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:51:12.488580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T22:51:11.667579Z digest=sha256:73d03004ad914cd2a74625cd3e11fe7e5e7cd0c5d549858bbbdbc98defd464fc

Observation 45b05fbf-8760-4267-bc57-45d5b7dfed17 · outbound

This paper cites Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.670823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.670823Z digest=sha256:0fd5fb21268be6ac3547e20a6490e0d73aa921613c4eff50ed5bda36ec025806

Observation c3cf3bea-0787-4c18-a181-d61f21012c25 · outbound

This paper cites Preserve your own correlation: A noise prior for video diffusion models.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Preserve your own correlation: A noise prior for video diffusion models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:51:12.477864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T22:51:11.674621Z digest=sha256:485a251add067a117551e1bd54d3fdbd1267c387d4ede777c53b2d01fc588689

Observation 8db5b61a-cb49-4097-bf14-00676a7ecaf5 · outbound

This paper cites Check locate rectify: A training- free layout calibration system for text-to-image generation.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Check locate rectify: A training- free layout calibration system for text-to-image generation

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:51:12.466479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T22:51:11.678765Z digest=sha256:88722d6d89df85e92873f69fd82c032377bd4f3ffe309a3d7202279940c83d0b

Observation 7a1b328f-6ac9-4d11-a578-f0ca4f9f0ed1 · outbound

This paper cites AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning.

Mimir: Improving Video Diffusion Models for Precise Text Understanding AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.682448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.682448Z digest=sha256:298d7a5e70fac445e0848ee4c0d2bb4d663f9fced468f3da0492cc6c7e3af592

Observation 7d66e66d-b89d-4fc4-8936-e018f8b2aada · outbound

This paper cites Denoising dif- fusion probabilistic models.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Denoising dif- fusion probabilistic models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.686779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.686779Z digest=sha256:f1f9e4eba682fe585adc3e327b70b9212233f6c1ba994aeb6a0268606619e6e6

Observation d8128cdc-9f30-49be-9d2c-1ecf7b7ddc93 · outbound

This paper cites Video dif- fusion models.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Video dif- fusion models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.690080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.690080Z digest=sha256:65d2a47a72029780005c3243f51d707b1521def5bef96d10b154a2b0cf4f92fe

Observation 72645895-e98b-4a7f-a8a6-0a3b48f4dc0d · outbound

This paper cites CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers.

Mimir: Improving Video Diffusion Models for Precise Text Understanding CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.693552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.693552Z digest=sha256:b76917a2e3820b80cc55134ec0b18f3cac0ed3a986f2f408c4dacd88e797726d

Observation a21b01a0-49ff-4cb7-b178-40b54e03990d · outbound

This paper cites CogVLM2: Visual Language Models for Image and Video Understanding.

Mimir: Improving Video Diffusion Models for Precise Text Understanding CogVLM2: Visual Language Models for Image and Video Understanding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.697842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.697842Z digest=sha256:cc8fc0ac6ae5409d57def19de455dbfab16a6d1f3e03240b97ea674e2beb34b2

Observation 22e6d8b2-de04-450e-83c8-346d24e7b021 · outbound

This paper cites ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment.

Mimir: Improving Video Diffusion Models for Precise Text Understanding ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.700744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.700744Z digest=sha256:bb2af417fe5448d9a7f009055aec203b3fefd5ae7d22aef9156d1bbbe84edf7c

Observation 6baa9656-9778-4f21-8fb8-8ab3caa37b8b · outbound

This paper cites T2i-compbench: A comprehensive bench- mark for open-world compositional text-to-image genera- 9 tion.

Mimir: Improving Video Diffusion Models for Precise Text Understanding T2i-compbench: A comprehensive bench- mark for open-world compositional text-to-image genera- 9 tion

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:51:12.444116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T22:51:11.703692Z digest=sha256:2d06174f410a87516013e1f92dbbf46a748a7ba4a7b0f150e57d9e4ac2d3ede4

Observation ebc51882-8616-4c0c-bf7d-49852386588f · outbound

This paper cites VBench: Com- prehensive benchmark suite for video generative models.

Mimir: Improving Video Diffusion Models for Precise Text Understanding VBench: Com- prehensive benchmark suite for video generative models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:51:12.431850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T22:51:11.706469Z digest=sha256:6b1722c91959347179123161d4dbc319ed4d85366d7f54b18d557b92be378fcb

Observation 6239cb00-0320-49f7-9520-3e90afdea811 · outbound

This paper cites Musiq: Multi-scale image quality transformer.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Musiq: Multi-scale image quality transformer

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:51:12.420034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T22:51:11.709085Z digest=sha256:3e06c99c6ad7b2733f01fadbb54a109594c7a2906ed983419f0ee87538a942b0

Observation b88ade9b-eb5f-425a-a074-e21a22a4ae57 · outbound

This paper cites VideoPoet: A Large Language Model for Zero-Shot Video Generation.

Mimir: Improving Video Diffusion Models for Precise Text Understanding VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.711501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.711501Z digest=sha256:7885f685ffbf220ef2193f41c5f3c88089730055f90e5fbe26aebc337b990d05

Observation 397fdc8b-b211-4ac7-a6b7-8bfe676f77f5 · outbound

This paper cites Open-sora-plan, 2024.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Open-sora-plan, 2024

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:51:12.409609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T22:51:11.714777Z digest=sha256:3d0b24dd260e234af564db8ee7b51770a55f920324137777d31bef716e8add03

Observation 2220ec05-9a82-44ac-a17d-68ab422d66f6 · outbound

This paper cites Common diffusion noise schedules and sample steps are flawed.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Common diffusion noise schedules and sample steps are flawed

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.718305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.718305Z digest=sha256:01e48cf93d93a1a039ea1ef6df7c31c96edfd2591b40613004d909180e019ab6

Observation d22d68cb-0418-4f31-a267-d067f09854cd · outbound

This paper cites Videofusion: Decomposed diffusion models for high-quality video generation.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Videofusion: Decomposed diffusion models for high-quality video generation

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:51:12.394518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T22:51:11.721349Z digest=sha256:5dd6a49d3da437b2862bf1af61c501940bded5dc41cfaa4460ba20ffadba5c5a

Observation f1f1dd64-9159-4d62-bedc-24cca4f2eeb3 · outbound

This paper cites Exploring the Role of Large Language Models in Prompt Encoding for Diffusion Models.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Exploring the Role of Large Language Models in Prompt Encoding for Diffusion Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.725237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.725237Z digest=sha256:38ead558561b44ac5968f98ed44f88f139c340cc5b5abab803f3578ff2cc6b56

Observation 94af8bdb-1df9-49ce-be09-ab6e4bf4e9aa · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Mimir: Improving Video Diffusion Models for Precise Text Understanding SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.729407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.729407Z digest=sha256:f2202dd78c061c27154e7a857c0e6d103066cd684025c33b06fd26d262e9c2ef

Observation caf6a1f7-5e7c-418c-b5eb-ce2e29ce478f · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Learning transferable visual models from natural language supervi- sion

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:51:12.384419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T22:51:11.733125Z digest=sha256:5d82bd1d24d8bbfde731cb769049f0fd8f4dae1d39a055bdb3eb2f8c5a4d4caa

Observation 9cdd973b-fe4b-4b4c-ba09-61c8c29dd0c8 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.736887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.736887Z digest=sha256:7cadb300205e741eb7d35536560b07cd01e873485a163e8fb2a6e3761ab81f7b

Observation ca035896-122e-4992-b697-583963c91b40 · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.740143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.740143Z digest=sha256:2443466890675ce7e8167b9c2c71dd9621d20dba1f0616354afa4550c5d205ac

Observation f3ae12f5-5f68-42c8-b040-8a8916b2c07b · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Mimir: Improving Video Diffusion Models for Precise Text Understanding High-resolution image synthesis with latent diffusion models

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:51:12.367811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T22:51:11.743465Z digest=sha256:7c62197d520f72a1af80ec22052f44b360912e1333d8e20f7315b71fcdcbbdb2

Observation 3cc13530-d148-47f9-a88c-90caab8fef4e · outbound

This paper cites Photorealistic text-to-image diffusion models with deep language understanding.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Photorealistic text-to-image diffusion models with deep language understanding

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:51:12.358061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T22:51:11.746344Z digest=sha256:2213459a6b513d95bf94dc5767bddaa750108bfcd583903615aab9d3f96f88b4

Observation 604fe72a-990a-44bf-a5b3-5b2f919b16ff · outbound

This paper cites Progressive Distillation for Fast Sampling of Diffusion Models.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Progressive Distillation for Fast Sampling of Diffusion Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.749400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.749400Z digest=sha256:4e6aab17f67b90ff2e77fb7e4d87d48c038dd8bf8983600b7a51e957db18d596

Observation 51edff70-7685-45e7-81c3-482dcdb6b7cb · outbound

This paper cites Deep unsupervised learning using nonequilibrium thermodynamics.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Deep unsupervised learning using nonequilibrium thermodynamics

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.753568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.753568Z digest=sha256:ed1f47ab9abd711855d079ad24a6d0358c7a50be73c7479c7d90afddeb7f112f

Observation 2d940805-e791-4ff3-8f01-eb3f97354397 · outbound

This paper cites Stable Diffusion 2.0 Release, 2022.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Stable Diffusion 2.0 Release, 2022

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:51:12.341114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T22:51:11.757006Z digest=sha256:7cb2c2f4ac0fd18902190476ffa508cc856e924d344f3b0ab883bda8921e73a0

Observation 8c4b294e-7380-4d32-bce9-ee3cf89529b9 · outbound

This paper cites Galip: Generative adversarial clips for text-to-image synthesis.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Galip: Generative adversarial clips for text-to-image synthesis

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:51:12.328642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T22:51:11.760544Z digest=sha256:820a740f7ee1fdc2b3ccb1d7188a8f9dce7c8d174bee1994836a06cf78c0c128

Observation 9b869e0b-9656-42a2-a5a3-2275c644060a · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Gemini: A Family of Highly Capable Multimodal Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.763823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.763823Z digest=sha256:23273e3c4c3d5e2ac2548a302d0a9080e8b918f7c8d453d8d1bd9c04f62b177d

Observation 14280fb3-5cc0-48cf-b81b-a137f4750877 · outbound

This paper cites Internlm: A multilingual language model with progressively enhanced capabilities, 2023.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Internlm: A multilingual language model with progressively enhanced capabilities, 2023

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.767558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.767558Z digest=sha256:a3cd809af08e36a2ca5174b71bcc06046d9159384662dc20890209ab4dd48f7f

Observation 8c8393dd-ca0a-4920-b893-cfee5060a49e · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Mimir: Improving Video Diffusion Models for Precise Text Understanding LLaMA: Open and Efficient Foundation Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.770731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.770731Z digest=sha256:e1ed8f3aae73eaada60c15b91503bdcd89e3738eb3f3a31175aa900c477196e1

Observation ebbd1558-1adf-4a4b-8507-97a12bac64dc · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.774507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.774507Z digest=sha256:a6684c7e92945403caf239f8ff64859d13aeac97f9c3e69ae21743c2d89d7a48

Observation d83c2f09-c7e3-4a8f-ac8b-42d26310f4ac · outbound

This paper cites Visualizing data using t-sne.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Visualizing data using t-sne

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.778460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.778460Z digest=sha256:f215058e4dcecef3ab762d7351cc73ad2fa43873ff0d0cbb89e5a38979fa6f39

Observation 370095dc-e1e5-452a-ba97-170d31db0697 · outbound

This paper cites Cogvlm: Visual expert for pretrained language models, 2023.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Cogvlm: Visual expert for pretrained language models, 2023

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:51:12.305573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T22:51:11.781986Z digest=sha256:bce07c157cb072256391798773bee8109b8407d7010a7a08250c70024101b964

Observation b4c52a21-8243-4808-973b-25604af2bee3 · outbound

This paper cites InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation.

Mimir: Improving Video Diffusion Models for Precise Text Understanding InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.786109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.786109Z digest=sha256:172edba7c84fac86e21f2ec221bfd614acf531f7f003bec1c964c33e9837a5f3

Observation a17778bf-b914-4a82-96f6-5f13fd06f7bf · outbound

This paper cites Grit: A gener- ative region-to-text transformer for object understanding.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Grit: A gener- ative region-to-text transformer for object understanding

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:51:12.295080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T22:51:11.790031Z digest=sha256:efc8ad3e4ebd8adb64bf94be9f6a789a7077f169a691b6936e8ac74f48cdc3c3

Observation 078ce21b-e4e4-407b-9b82-e25718a751a9 · outbound

This paper cites Paragraph-to-Image Generation with Information-Enriched Diffusion Model.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Paragraph-to-Image Generation with Information-Enriched Diffusion Model

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.793880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.793880Z digest=sha256:6a3915034d2691a017b2487c08dfbbf6488e61895ec23af52c00236af967d342

Observation 02eba41d-9f32-41c5-9774-7c3e884688df · outbound

This paper cites SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformers.

Mimir: Improving Video Diffusion Models for Precise Text Understanding SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformers

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.797701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.797701Z digest=sha256:d117f152c149247e0849ee325224986a8cd03b820c013ede7742ec2ccba36e93

Observation b644eab3-2884-4903-a021-e10675749e39 · outbound

This paper cites Baichuan 2: Open Large-scale Language Models.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Baichuan 2: Open Large-scale Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.800770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.800770Z digest=sha256:3f757e46e3a685eed65d669663128319b8ad5f7bc13b61a26af3f9c811763976

Observation c6430e2a-71bf-494e-8f17-c1d6284c3712 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

Mimir: Improving Video Diffusion Models for Precise Text Understanding CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.804287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.804287Z digest=sha256:96f60efda31a51242f36e9572642e56e6bd17d4c83d241c4caef964f45eb4ba0

Observation f7e04b10-0996-4aa7-9f3e-29316a62060f · outbound

This paper cites Yi: Open Foundation Models by 01.AI.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Yi: Open Foundation Models by 01.AI

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.807591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.807591Z digest=sha256:939822c1afb36128f97c48f57b6e9e3350362f11acd9c405c640962d767589da

Observation 2a3bf3b5-821a-428c-8285-f8604899fcdb · outbound

This paper cites Magvit: Masked generative video transformer.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Magvit: Masked generative video transformer

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:51:12.284845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T22:51:11.810768Z digest=sha256:2d1eed8ace927f7a615f4047675d7e29a8e9d24e22aa90f9812d0513f833dcb4

Observation d22537aa-a24e-42e8-bc1b-85ac610b88c6 · outbound

This paper cites Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.813821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.813821Z digest=sha256:aa4e330c6de86069951d5e7dd4a571b7935dfd9c6ad7a426d2b6cfd7f2698f2b

Observation 68908d34-bc99-438d-bc82-1be8c17d0042 · outbound

This paper cites Spae: Semantic pyramid autoencoder for multimodal generation with frozen llms.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Spae: Semantic pyramid autoencoder for multimodal generation with frozen llms

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:51:12.274098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T22:51:11.818765Z digest=sha256:aa1e9e863d0a0726ee49d96d8b817251a1b02df08674b5325530b64241c1e26b

Observation 560ddeec-0fa3-42e1-8d32-b12b9fe9caf4 · outbound

This paper cites Bridging Different Language Models and Generative Vision Models for Text-to-Image Generation.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Bridging Different Language Models and Generative Vision Models for Text-to-Image Generation

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.822595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.822595Z digest=sha256:b86b328c7dba694410bb246792c3c76da5e27c7dc43f5fa799d61daefe75a3b1

Observation f36d674f-0dae-4596-95b0-33d22e2afff5 · outbound

This paper cites CV-VAE: A Compatible Video VAE for Latent Generative Video Models.

Mimir: Improving Video Diffusion Models for Precise Text Understanding CV-VAE: A Compatible Video VAE for Latent Generative Video Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.827049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.827049Z digest=sha256:86da723a41c45d8e18b1ab51eab7b213b55ee4fbe0327fd45b19467fba0db89a

Observation 1199afe8-f167-4d7b-abc3-248dd0a6b825 · outbound

This paper cites Open-sora: Democratizing efficient video production for all, march 2024.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Open-sora: Democratizing efficient video production for all, march 2024

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:51:12.263393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T22:51:11.831286Z digest=sha256:079ae8e76d604d4aecf49d67e577a65ca078adeaf2fdd8d314c57b8efdf77264

Observation 2fde7757-3226-4491-89de-2b0d958a469c · outbound

This paper cites an unresolved cited work.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-08-11T22:51:12.254031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T22:51:11.835420Z digest=sha256:b3cd5bf8ce08b6095a35f1235cf4d629ff83947cc45bbdce96c0506c2cfaf1fb

Observation f956b5f7-5c53-49a0-8f7e-41a835d0d893 · outbound

This paper cites • Videos with a motion score of 0, determined using optical flow, are excluded.

Mimir: Improving Video Diffusion Models for Precise Text Understanding • Videos with a motion score of 0, determined using optical flow, are excluded

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:51:12.243362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T22:51:11.839179Z digest=sha256:3e6c8b48c1a15763debeca9b12289b169be1b4880d5464e93d10c4bae5557793

Observation 8515fc69-f386-4409-8cea-ba987206a161 · outbound

This paper cites an unresolved cited work.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-11T22:51:12.231669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T22:51:11.843069Z digest=sha256:61dafb7bb34b1e05a92e07a399bfe0006a9a7ee3cfa9b66749e64ddc5fcfe06d

Observation ed2bb0e4-2c56-40f2-a982-cc8cd9bf3c35 · outbound

This paper cites Input text prompt.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Input text prompt

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:51:12.217479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T22:51:11.847314Z digest=sha256:b6baffd7e196c697ca4b02b7e26e79ea38c98f56990cc3224f1385b3a7cd655b

Observation 537e7d6e-a8cd-4598-83db-03bae41de380 · outbound

This paper cites an unresolved cited work.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Unresolved cited work

Reference 65

Resolution
unresolved
raw_fallback, observed 2026-08-11T22:51:12.206404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T22:51:11.850609Z digest=sha256:93ae4185b633337b98fb4c98ccd6afeef644f6fada9ba875ad3df0c6d6c3311a

Observation a4417213-4970-4ef3-9ecb-4ba35cbc2b33 · outbound

This paper cites containing watermarks.

Mimir: Improving Video Diffusion Models for Precise Text Understanding containing watermarks

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:51:12.196304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T22:51:11.854420Z digest=sha256:08b5614cd7155062e432496a45a29ed1c672cf96518a3ce8f9b2d7d0b8b1c24d

Observation b87fc7b9-720d-48af-8501-26cd4788d9ce · outbound

This paper cites an unresolved cited work.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-08-11T22:51:12.186046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T22:51:11.858303Z digest=sha256:d60e8d902b9da44a5703e1ddd52ee7b57f7066677fadde3c2b676bd03ba5e267

Observation a99ad1bb-e85d-4473-823a-b87f226ba4cb · outbound

This paper cites top”, “ below.

Mimir: Improving Video Diffusion Models for Precise Text Understanding top”, “ below

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:51:12.173566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T22:51:11.861299Z digest=sha256:5609cda10c0b243f2c787d0771971fff7faa17c62a221b4414d0ac9ccf637bee

Pith citing papers

Observation 71f99619-f953-4740-9ac9-36f852f104e7 · inbound

Animate-X++: Universal Character Image Animation with Dynamic Backgrounds cites this paper.

Animate-X++: Universal Character Image Animation with Dynamic Backgrounds Mimir: Improving Video Diffusion Models for Precise Text Understanding

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T21:09:07.991313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:09:07.991313Z digest=sha256:e68dd911a167759afc04c9059dd0b68be73e43f73717c1843aebbc7ee21c8dbd

Observation ed581e4b-c3b5-4117-9062-2b5d37bed3de · inbound

EDTalk++: Full Disentanglement for Controllable Talking Head Synthesis cites this paper.

EDTalk++: Full Disentanglement for Controllable Talking Head Synthesis Mimir: Improving Video Diffusion Models for Precise Text Understanding

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T19:07:34.083942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:07:34.083942Z digest=sha256:2a685fee509b30aff392a3d40312da67af71e079e1895c7c57b963641b56f872

Observation 8513a963-5d84-4bb1-b70d-40f6565b1a0b · inbound

Bridging Brain and Semantics: A Hierarchical Framework for Semantically Enhanced fMRI-to-Video Reconstruction cites this paper.

Bridging Brain and Semantics: A Hierarchical Framework for Semantically Enhanced fMRI-to-Video Reconstruction Mimir: Improving Video Diffusion Models for Precise Text Understanding

Reference 86

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:53:28.927343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T01:51:57.018809Z digest=sha256:f8fed50aa07ceed858736f19f890baecf8bbfa7115041b6f3abb126ccca60d28