Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T22:51:11.861299Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 3 inbound Pith citation observations for arXiv:2412.03085.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T22:51:11.861299Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-05T21:09:07.991313Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-15T01:53:28.925576Z
68 of 68 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation fa3d78dd-059b-486f-854a-f07c0d1cf60d · outbound
Mimir: Improving Video Diffusion Models for Precise Text Understanding Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 174ce6c1-871c-49fe-be62-2eee06cdd298 · outbound
Mimir: Improving Video Diffusion Models for Precise Text Understanding Qwen Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93c2fe9c-38a0-42f0-b0c2-532a69bedd52 · outbound
Mimir: Improving Video Diffusion Models for Precise Text Understanding eDiff-I: Text-to-Image Diffusion Models with an Ensemble of Expert Denoisers
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31fbc9c4-a06d-4694-896e-45314cf6334a · outbound
Mimir: Improving Video Diffusion Models for Precise Text Understanding LLM2Vec: Large Language Models Are Secretly Powerful Text Encoders
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5492bb5-849a-4f41-a4fb-73ab32b8f523 · outbound
Mimir: Improving Video Diffusion Models for Precise Text Understanding Improving image generation with better captions
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc2703c4-4df9-446b-ae3f-c6d6f1b7fdb4 · outbound
Mimir: Improving Video Diffusion Models for Precise Text Understanding Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f009921e-e3c9-4f5c-aa70-1002c1168767 · outbound
Mimir: Improving Video Diffusion Models for Precise Text Understanding PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b167563-b1c0-4aff-a3e9-c7d1ffa9c571 · outbound
Mimir: Improving Video Diffusion Models for Precise Text Understanding PixArt-\Sigma: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de3a697d-3f46-4bb7-ad16-68da13c74268 · outbound
Mimir: Improving Video Diffusion Models for Precise Text Understanding PIXART-{\delta}: Fast and Controllable Image Generation with Latent Consistency Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66913913-901d-4840-89c9-15821d23ebca · outbound
Mimir: Improving Video Diffusion Models for Precise Text Understanding OD-VAE: An Omni-dimensional Video Compressor for Improving Latent Video Diffusion Model
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b244d4ae-c421-41b0-a980-c1150df05139 · outbound
Mimir: Improving Video Diffusion Models for Precise Text Understanding Taming transformers for high-resolution image synthesis
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62d3c6c2-d215-4a35-b549-8674f9f8b614 · outbound
Mimir: Improving Video Diffusion Models for Precise Text Understanding Scaling rectified flow transformers for high-resolution image synthesis
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a619a1c5-5332-4ef9-ba53-7b5841425048 · outbound
Mimir: Improving Video Diffusion Models for Precise Text Understanding Perceptual quality assessment of smartphone photography
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2a60417-8b0d-4c08-9027-32ec10e61a03 · outbound
Mimir: Improving Video Diffusion Models for Precise Text Understanding Ranni: Taming text-to-image diffusion for accurate instruction following
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 45b05fbf-8760-4267-bc57-45d5b7dfed17 · outbound
Mimir: Improving Video Diffusion Models for Precise Text Understanding Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3cf3bea-0787-4c18-a181-d61f21012c25 · outbound
Mimir: Improving Video Diffusion Models for Precise Text Understanding Preserve your own correlation: A noise prior for video diffusion models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8db5b61a-cb49-4097-bf14-00676a7ecaf5 · outbound
Mimir: Improving Video Diffusion Models for Precise Text Understanding Check locate rectify: A training- free layout calibration system for text-to-image generation
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7a1b328f-6ac9-4d11-a578-f0ca4f9f0ed1 · outbound
Mimir: Improving Video Diffusion Models for Precise Text Understanding AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d66e66d-b89d-4fc4-8936-e018f8b2aada · outbound
Mimir: Improving Video Diffusion Models for Precise Text Understanding Denoising dif- fusion probabilistic models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8128cdc-9f30-49be-9d2c-1ecf7b7ddc93 · outbound
Mimir: Improving Video Diffusion Models for Precise Text Understanding Video dif- fusion models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72645895-e98b-4a7f-a8a6-0a3b48f4dc0d · outbound
Mimir: Improving Video Diffusion Models for Precise Text Understanding CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a21b01a0-49ff-4cb7-b178-40b54e03990d · outbound
Mimir: Improving Video Diffusion Models for Precise Text Understanding CogVLM2: Visual Language Models for Image and Video Understanding
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22e6d8b2-de04-450e-83c8-346d24e7b021 · outbound
Mimir: Improving Video Diffusion Models for Precise Text Understanding ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6baa9656-9778-4f21-8fb8-8ab3caa37b8b · outbound
Mimir: Improving Video Diffusion Models for Precise Text Understanding T2i-compbench: A comprehensive bench- mark for open-world compositional text-to-image genera- 9 tion
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ebc51882-8616-4c0c-bf7d-49852386588f · outbound
Mimir: Improving Video Diffusion Models for Precise Text Understanding VBench: Com- prehensive benchmark suite for video generative models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6239cb00-0320-49f7-9520-3e90afdea811 · outbound
Mimir: Improving Video Diffusion Models for Precise Text Understanding Musiq: Multi-scale image quality transformer
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b88ade9b-eb5f-425a-a074-e21a22a4ae57 · outbound
Mimir: Improving Video Diffusion Models for Precise Text Understanding VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 397fdc8b-b211-4ac7-a6b7-8bfe676f77f5 · outbound
Mimir: Improving Video Diffusion Models for Precise Text Understanding Open-sora-plan, 2024
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2220ec05-9a82-44ac-a17d-68ab422d66f6 · outbound
Mimir: Improving Video Diffusion Models for Precise Text Understanding Common diffusion noise schedules and sample steps are flawed
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d22d68cb-0418-4f31-a267-d067f09854cd · outbound
Mimir: Improving Video Diffusion Models for Precise Text Understanding Videofusion: Decomposed diffusion models for high-quality video generation
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f1f1dd64-9159-4d62-bedc-24cca4f2eeb3 · outbound
Mimir: Improving Video Diffusion Models for Precise Text Understanding Exploring the Role of Large Language Models in Prompt Encoding for Diffusion Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94af8bdb-1df9-49ce-be09-ab6e4bf4e9aa · outbound
Mimir: Improving Video Diffusion Models for Precise Text Understanding SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation caf6a1f7-5e7c-418c-b5eb-ce2e29ce478f · outbound
Mimir: Improving Video Diffusion Models for Precise Text Understanding Learning transferable visual models from natural language supervi- sion
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9cdd973b-fe4b-4b4c-ba09-61c8c29dd0c8 · outbound
Mimir: Improving Video Diffusion Models for Precise Text Understanding Exploring the limits of transfer learning with a unified text-to-text transformer
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca035896-122e-4992-b697-583963c91b40 · outbound
Mimir: Improving Video Diffusion Models for Precise Text Understanding Hierarchical Text-Conditional Image Generation with CLIP Latents
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3ae12f5-5f68-42c8-b040-8a8916b2c07b · outbound
Mimir: Improving Video Diffusion Models for Precise Text Understanding High-resolution image synthesis with latent diffusion models
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 3cc13530-d148-47f9-a88c-90caab8fef4e · outbound
Mimir: Improving Video Diffusion Models for Precise Text Understanding Photorealistic text-to-image diffusion models with deep language understanding
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 604fe72a-990a-44bf-a5b3-5b2f919b16ff · outbound
Mimir: Improving Video Diffusion Models for Precise Text Understanding Progressive Distillation for Fast Sampling of Diffusion Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51edff70-7685-45e7-81c3-482dcdb6b7cb · outbound
Mimir: Improving Video Diffusion Models for Precise Text Understanding Deep unsupervised learning using nonequilibrium thermodynamics
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d940805-e791-4ff3-8f01-eb3f97354397 · outbound
Mimir: Improving Video Diffusion Models for Precise Text Understanding Stable Diffusion 2.0 Release, 2022
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8c4b294e-7380-4d32-bce9-ee3cf89529b9 · outbound
Mimir: Improving Video Diffusion Models for Precise Text Understanding Galip: Generative adversarial clips for text-to-image synthesis
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9b869e0b-9656-42a2-a5a3-2275c644060a · outbound
Mimir: Improving Video Diffusion Models for Precise Text Understanding Gemini: A Family of Highly Capable Multimodal Models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14280fb3-5cc0-48cf-b81b-a137f4750877 · outbound
Mimir: Improving Video Diffusion Models for Precise Text Understanding Internlm: A multilingual language model with progressively enhanced capabilities, 2023
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c8393dd-ca0a-4920-b893-cfee5060a49e · outbound
Mimir: Improving Video Diffusion Models for Precise Text Understanding LLaMA: Open and Efficient Foundation Language Models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ebbd1558-1adf-4a4b-8507-97a12bac64dc · outbound
Mimir: Improving Video Diffusion Models for Precise Text Understanding Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d83c2f09-c7e3-4a8f-ac8b-42d26310f4ac · outbound
Mimir: Improving Video Diffusion Models for Precise Text Understanding Visualizing data using t-sne
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 370095dc-e1e5-452a-ba97-170d31db0697 · outbound
Mimir: Improving Video Diffusion Models for Precise Text Understanding Cogvlm: Visual expert for pretrained language models, 2023
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b4c52a21-8243-4808-973b-25604af2bee3 · outbound
Mimir: Improving Video Diffusion Models for Precise Text Understanding InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a17778bf-b914-4a82-96f6-5f13fd06f7bf · outbound
Mimir: Improving Video Diffusion Models for Precise Text Understanding Grit: A gener- ative region-to-text transformer for object understanding
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 078ce21b-e4e4-407b-9b82-e25718a751a9 · outbound
Mimir: Improving Video Diffusion Models for Precise Text Understanding Paragraph-to-Image Generation with Information-Enriched Diffusion Model
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02eba41d-9f32-41c5-9774-7c3e884688df · outbound
Mimir: Improving Video Diffusion Models for Precise Text Understanding SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformers
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b644eab3-2884-4903-a021-e10675749e39 · outbound
Mimir: Improving Video Diffusion Models for Precise Text Understanding Baichuan 2: Open Large-scale Language Models
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6430e2a-71bf-494e-8f17-c1d6284c3712 · outbound
Mimir: Improving Video Diffusion Models for Precise Text Understanding CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7e04b10-0996-4aa7-9f3e-29316a62060f · outbound
Mimir: Improving Video Diffusion Models for Precise Text Understanding Yi: Open Foundation Models by 01.AI
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a3bf3b5-821a-428c-8285-f8604899fcdb · outbound
Mimir: Improving Video Diffusion Models for Precise Text Understanding Magvit: Masked generative video transformer
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d22537aa-a24e-42e8-bc1b-85ac610b88c6 · outbound
Mimir: Improving Video Diffusion Models for Precise Text Understanding Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68908d34-bc99-438d-bc82-1be8c17d0042 · outbound
Mimir: Improving Video Diffusion Models for Precise Text Understanding Spae: Semantic pyramid autoencoder for multimodal generation with frozen llms
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 560ddeec-0fa3-42e1-8d32-b12b9fe9caf4 · outbound
Mimir: Improving Video Diffusion Models for Precise Text Understanding Bridging Different Language Models and Generative Vision Models for Text-to-Image Generation
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f36d674f-0dae-4596-95b0-33d22e2afff5 · outbound
Mimir: Improving Video Diffusion Models for Precise Text Understanding CV-VAE: A Compatible Video VAE for Latent Generative Video Models
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1199afe8-f167-4d7b-abc3-248dd0a6b825 · outbound
Mimir: Improving Video Diffusion Models for Precise Text Understanding Open-sora: Democratizing efficient video production for all, march 2024
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2fde7757-3226-4491-89de-2b0d958a469c · outbound
Mimir: Improving Video Diffusion Models for Precise Text Understanding Unresolved cited work
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f956b5f7-5c53-49a0-8f7e-41a835d0d893 · outbound
Mimir: Improving Video Diffusion Models for Precise Text Understanding • Videos with a motion score of 0, determined using optical flow, are excluded
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8515fc69-f386-4409-8cea-ba987206a161 · outbound
Mimir: Improving Video Diffusion Models for Precise Text Understanding Unresolved cited work
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ed2bb0e4-2c56-40f2-a982-cc8cd9bf3c35 · outbound
Mimir: Improving Video Diffusion Models for Precise Text Understanding Input text prompt
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 537e7d6e-a8cd-4598-83db-03bae41de380 · outbound
Mimir: Improving Video Diffusion Models for Precise Text Understanding Unresolved cited work
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a4417213-4970-4ef3-9ecb-4ba35cbc2b33 · outbound
Mimir: Improving Video Diffusion Models for Precise Text Understanding containing watermarks
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b87fc7b9-720d-48af-8501-26cd4788d9ce · outbound
Mimir: Improving Video Diffusion Models for Precise Text Understanding Unresolved cited work
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a99ad1bb-e85d-4473-823a-b87f226ba4cb · outbound
Mimir: Improving Video Diffusion Models for Precise Text Understanding top”, “ below
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 71f99619-f953-4740-9ac9-36f852f104e7 · inbound
Animate-X++: Universal Character Image Animation with Dynamic Backgrounds Mimir: Improving Video Diffusion Models for Precise Text Understanding
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed581e4b-c3b5-4117-9062-2b5d37bed3de · inbound
EDTalk++: Full Disentanglement for Controllable Talking Head Synthesis Mimir: Improving Video Diffusion Models for Precise Text Understanding
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8513a963-5d84-4bb1-b70d-40f6565b1a0b · inbound
Bridging Brain and Semantics: A Hierarchical Framework for Semantically Enhanced fMRI-to-Video Reconstruction Mimir: Improving Video Diffusion Models for Precise Text Understanding
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.