Pith. sign in

Paper Citation Record · LEDGER

Early Estimation of Language to Latent Alignment in Diffusion Models

As of 21 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 1 inbound Pith citation observation for arXiv:2512.08505.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2512.08505 v2

Coverage vector

measured 64 of 64 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T17:42:47.519295Z

measured 65 of 65 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T17:21:37.570778Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-03T00:17:29.054872Z

Reference resolution

64 of 64 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved64
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0395128c-1357-4672-861e-97a360a0a8e9 · outbound

This paper cites an unresolved cited work.

Early Estimation of Language to Latent Alignment in Diffusion Models Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T17:42:41.538992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:42:41.538992Z digest=sha256:d83bc90b536b902d5819834494b42a6b0606cd96fd632708463f4799589b6959

Observation 95dd2ad9-901d-4648-9c4c-9f96e03f0443 · outbound

This paper cites Imagen 3.

Early Estimation of Language to Latent Alignment in Diffusion Models Imagen 3

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T17:42:41.686457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:42:41.686457Z digest=sha256:bcf1ed8e86a707cc8247956cecec66a4b35febb9d555d68355a04a989a165d4c

Observation e641d3c0-800b-4ac3-9c67-159cee1b7612 · outbound

This paper cites Controlling Latent Diffusion Using Latent CLIP.

Early Estimation of Language to Latent Alignment in Diffusion Models Controlling Latent Diffusion Using Latent CLIP

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T17:42:41.792537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:42:41.792537Z digest=sha256:dab4facbaa321d32392f0ffa7007f6e45900dc46dee4ff352e1ff34736f57ced

Observation d18a8c6a-0ad5-4106-b234-d63bfaee54fd · outbound

This paper cites Natural language in- ference improves compositionality in vision-language mod- els.

Early Estimation of Language to Latent Alignment in Diffusion Models Natural language in- ference improves compositionality in vision-language mod- els

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T17:42:41.883627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:42:41.883627Z digest=sha256:ac269b880d5b2e05bd9aecaa3262ddf931338af106b407dbe0d907e7f9f38e47

Observation d1c07e8f-c5cc-450b-9232-388adc718de1 · outbound

This paper cites Conceptual 12m: Pushing web-scale image-text pre- training to recognize long-tail visual concepts.

Early Estimation of Language to Latent Alignment in Diffusion Models Conceptual 12m: Pushing web-scale image-text pre- training to recognize long-tail visual concepts

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T17:42:41.995177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:42:41.995177Z digest=sha256:6bcb4fd020ccd626676486a21684cac6c6d74df5441af55f03034d90bc24ea48

Observation d93c754c-9528-4284-b1e0-b1fa696d4bb1 · outbound

This paper cites The hidden language of diffusion models.

Early Estimation of Language to Latent Alignment in Diffusion Models The hidden language of diffusion models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T17:42:42.085964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:42:42.085964Z digest=sha256:da88cc19db9f1b253949d2783820ac52982b1dbc93453eed659b5d95a6e3a41a

Observation 3dc468c2-5ef7-4a17-9af7-f7a0579944ef · outbound

This paper cites Visual pro- gramming for step-by-step text-to-image generation and evaluation.

Early Estimation of Language to Latent Alignment in Diffusion Models Visual pro- gramming for step-by-step text-to-image generation and evaluation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T17:42:42.179451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:42:42.179451Z digest=sha256:353af632c53e29cc8cdefac176d6fe3c48bac0be3f8a2d12d729afbe77d4d687

Observation 7067c17b-2878-4b86-a5e5-4c2084c7f7d3 · outbound

This paper cites Prompt tuning inversion for text-driven image editing using diffusion models.

Early Estimation of Language to Latent Alignment in Diffusion Models Prompt tuning inversion for text-driven image editing using diffusion models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T17:42:42.239832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:42:42.239832Z digest=sha256:e7b3dc56b4cbde79122aadd118fb7001e4b79cb32a04baaf0becfd6bd099f85b

Observation 76c8c5cc-00cb-48fa-8357-a9af1030f2bb · outbound

This paper cites conceptual-captions-cc12m- llavanext.https://huggingface.co/datasets/ CaptionEmporium / conceptual - captions - cc12m-llavanext, 2024.

Early Estimation of Language to Latent Alignment in Diffusion Models conceptual-captions-cc12m- llavanext.https://huggingface.co/datasets/ CaptionEmporium / conceptual - captions - cc12m-llavanext, 2024

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T17:42:42.316925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:42:42.316925Z digest=sha256:ec26a566c37f7dfc784e321c2ab6bf917cfc1db6061b49aed293c57da19445ca

Observation 047f19df-285a-4f12-b84e-d65f5988be04 · outbound

This paper cites Scaling rec- tified flow transformers for high-resolution image synthesis.

Early Estimation of Language to Latent Alignment in Diffusion Models Scaling rec- tified flow transformers for high-resolution image synthesis

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T17:42:42.385164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:42:42.385164Z digest=sha256:c05f5499292bcb2b6f9ce6ac5943e5e45380699792249ac8adb35662a837a72a

Observation a2a87ed1-974c-490c-8ec5-23e61722f996 · outbound

This paper cites W AFFLE: multimodal floorplan understand- ing in the wild.

Early Estimation of Language to Latent Alignment in Diffusion Models W AFFLE: multimodal floorplan understand- ing in the wild

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T17:42:42.462531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:42:42.462531Z digest=sha256:ae4b93a3020ef6eb90a0ecb14533997f3f33b04bc1dd3ed0bbd61864d8103ffd

Observation b20d284a-67ce-4b45-9d46-47f1bf61cd31 · outbound

This paper cites A large scale analysis of gender biases in text- to-image generative models.CoRR, abs/2503.23398, 2025.

Early Estimation of Language to Latent Alignment in Diffusion Models A large scale analysis of gender biases in text- to-image generative models.CoRR, abs/2503.23398, 2025

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T17:42:42.543822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:42:42.543822Z digest=sha256:4fd3fbd3584db454dddc67204778d54fbc69285a4fb7b9a555892aae2b22d30a

Observation e75551d3-a408-44c6-b93c-1c017d0a1f53 · outbound

This paper cites Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio.

Early Estimation of Language to Latent Alignment in Diffusion Models Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T17:42:42.679901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:42:42.679901Z digest=sha256:de523a323012e117a0b917f4151b9e01535199e51ca72a4970ea22547041efbc

Observation caba42ec-6237-4680-8097-afb840c3f172 · outbound

This paper cites Clipscore: A reference-free evaluation met- ric for image captioning.

Early Estimation of Language to Latent Alignment in Diffusion Models Clipscore: A reference-free evaluation met- ric for image captioning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T17:42:42.796165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:42:42.796165Z digest=sha256:05b03d671469dc072dca95a5b3e1df9355ada1bd53c24b3683a8a85e62661cb4

Observation 751a7e2f-28b5-4050-a817-8220e971f0c1 · outbound

This paper cites Classifier-free diffusion guidance.

Early Estimation of Language to Latent Alignment in Diffusion Models Classifier-free diffusion guidance

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T17:42:42.867695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:42:42.867695Z digest=sha256:770e4cba44107973295df14c4e4597b20dff3b108b08698f6e8aabd8124c22d8

Observation 96922295-a8d6-4e52-8d30-21bfadcefc5a · outbound

This paper cites Denoising Dif- fusion Probabilistic Models.

Early Estimation of Language to Latent Alignment in Diffusion Models Denoising Dif- fusion Probabilistic Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T17:42:43.017973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:42:43.017973Z digest=sha256:f1931999051e5d12d9fe8b8695dc8416a23c6fdfdf277100491d9722c8ff245b

Observation 0c71f37b-6b80-4140-9174-38bb78ac7d80 · outbound

This paper cites an unresolved cited work.

Early Estimation of Language to Latent Alignment in Diffusion Models Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T17:42:43.101096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:42:43.101096Z digest=sha256:3f45b4968611eec835a58ebff75bf558fa0a569a84afe982a610ea8832c3962f

Observation 1383ef8d-d2b1-47c7-ae06-e603f09bdf19 · outbound

This paper cites T2i-compbench: A comprehensive benchmark for open-world compositional text-to-image generation.

Early Estimation of Language to Latent Alignment in Diffusion Models T2i-compbench: A comprehensive benchmark for open-world compositional text-to-image generation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T17:42:43.165022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:42:43.165022Z digest=sha256:d29efe95fca1534ffc77730e3fdeba3accf591d6eb5082d8a74cdde5942159c9

Observation 4a359c2e-82ac-41be-991b-50ce1fd44617 · outbound

This paper cites Vi- sual hallucinations of multi-modal large language models.

Early Estimation of Language to Latent Alignment in Diffusion Models Vi- sual hallucinations of multi-modal large language models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T17:42:43.273452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:42:43.273452Z digest=sha256:3cdb3b7c2b847c13f7b3e6347f38473abbd624c7d91f23181c3ff0cc5aaddc7c

Observation c3f43404-8aa9-42c4-824a-42c258e14c35 · outbound

This paper cites Perceiver: General perception with iterative attention.

Early Estimation of Language to Latent Alignment in Diffusion Models Perceiver: General perception with iterative attention

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T17:42:43.387476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:42:43.387476Z digest=sha256:af2ed4479b63849c1ffbf7ccd72f167a2524175fcc1823279a169114c27fec57

Observation 8e18b373-1a08-4bbd-ad84-10ff5e7d8a89 · outbound

This paper cites Gemma 3 Technical Report.

Early Estimation of Language to Latent Alignment in Diffusion Models Gemma 3 Technical Report

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T17:42:43.495579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:42:43.495579Z digest=sha256:de07d2c0d41d50be6e3cb128ad6da58e27cdd904dbcf49cdab934d6aed60c593

Observation 49174265-638a-43a8-bb78-12e567db0ee4 · outbound

This paper cites Auto-encoding varia- tional bayes, 2022.

Early Estimation of Language to Latent Alignment in Diffusion Models Auto-encoding varia- tional bayes, 2022

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T17:42:43.639090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:42:43.639090Z digest=sha256:36230dd923428c92127d84e4bb40179f438557668c56154acfa4d83e8f9dc241

Observation 399e3402-6d56-4ec0-b8bf-b33a71a54e9a · outbound

This paper cites Pick-a-pic: An open dataset of user preferences for text-to-image generation.

Early Estimation of Language to Latent Alignment in Diffusion Models Pick-a-pic: An open dataset of user preferences for text-to-image generation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T17:42:43.758276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:42:43.758276Z digest=sha256:57330eb8375af9bc9647d9ffa4be9fe208506e09f0c1100269845e274a2906b8

Observation 9e2288a9-e8a4-4b19-85a0-1741d6505915 · outbound

This paper cites GenAI-Bench: Evaluating and Improving Compositional Text-to-Visual Generation.

Early Estimation of Language to Latent Alignment in Diffusion Models GenAI-Bench: Evaluating and Improving Compositional Text-to-Visual Generation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T17:42:43.835726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:42:43.835726Z digest=sha256:3fd0a5ad5d7abc383f8e1646d789bafc2f79393e1bd67d6e1ec6b105b0cbfdc1

Observation a0197877-0fe8-44a2-b3b1-eb8d4691a980 · outbound

This paper cites Crowdclip: Unsupervised crowd counting via vision-language model.

Early Estimation of Language to Latent Alignment in Diffusion Models Crowdclip: Unsupervised crowd counting via vision-language model

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T17:42:43.942840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:42:43.942840Z digest=sha256:0787fa499e88ae70f57d4828c1006ce38f355765ae37aaea838b74cf866b7089

Observation 231b8c3b-48fd-44cd-9ec0-42b48704ff58 · outbound

This paper cites Evaluating text-to-visual generation with image-to-text gen- eration.

Early Estimation of Language to Latent Alignment in Diffusion Models Evaluating text-to-visual generation with image-to-text gen- eration

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T17:42:44.059757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:42:44.059757Z digest=sha256:31d4db8a2f4718013ee1a7fbd2a356b7056eb4990cbf0ab484c57e6a8c7393cb

Observation 31b088cf-2787-4f5b-9b81-c8fe2ac797d6 · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

Early Estimation of Language to Latent Alignment in Diffusion Models Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T17:42:44.145308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:42:44.145308Z digest=sha256:a4407c78c45ec5e1f95774835158f702531ff740775920eca7a5af5501f3e1c5

Observation 9ff244e6-d051-4706-9700-4f989150a652 · outbound

This paper cites Jaakkola, Xuhui Jia, and Saining Xie.

Early Estimation of Language to Latent Alignment in Diffusion Models Jaakkola, Xuhui Jia, and Saining Xie

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T17:42:44.216055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:42:44.216055Z digest=sha256:ca376b161afb2c2c9695b2c31c3f3436b7335ec09ce48b02fd88425e8e7cfcbc

Observation 421cb7b0-11b4-4bf7-8802-ca588f841bd7 · outbound

This paper cites Teaching CLIP to count to ten.

Early Estimation of Language to Latent Alignment in Diffusion Models Teaching CLIP to count to ten

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T17:42:44.320322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:42:44.320322Z digest=sha256:be6e4481b3a02b6cc284c6b3c6621028120c49b3ff6dbb2d11b6bcaf831b71d8

Observation f3c69b6a-0519-4606-97e1-64f34e41531f · outbound

This paper cites Dynamic classifier-free diffusion guidance via online feedback.CoRR, abs/2509.16131, 2025.

Early Estimation of Language to Latent Alignment in Diffusion Models Dynamic classifier-free diffusion guidance via online feedback.CoRR, abs/2509.16131, 2025

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T17:42:44.396799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:42:44.396799Z digest=sha256:678cbb9902db96f0cd175b8ab1a8d904b35d73ced52af11e4e0cd546d5e40c61

Observation 3a35a985-9a2c-4f3d-8c30-141be8145d57 · outbound

This paper cites Know ”no” better: A data- driven approach for enhancing negation awareness in clip.

Early Estimation of Language to Latent Alignment in Diffusion Models Know ”no” better: A data- driven approach for enhancing negation awareness in clip

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T17:42:44.477545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:42:44.477545Z digest=sha256:ba9fc7351a6660279c66adc2eaa74abce14d8a4d0a2d464958615003a68a15a7

Observation dadeb2c0-1ee5-44c6-9cac-0af519e77692 · outbound

This paper cites Understanding the latent space of dif- fusion models through the lens of riemannian geometry.

Early Estimation of Language to Latent Alignment in Diffusion Models Understanding the latent space of dif- fusion models through the lens of riemannian geometry

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T17:42:44.581515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:42:44.581515Z digest=sha256:3497b4259418c2df7916b24a6b191c00140cd2fcb35f349c2f103bbd8ed4e231

Observation 0763b12a-7f0c-4064-8b3a-7ed3b413be86 · outbound

This paper cites Scalable diffusion models with transformers.

Early Estimation of Language to Latent Alignment in Diffusion Models Scalable diffusion models with transformers

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T17:42:44.664646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:42:44.664646Z digest=sha256:9ba53fc98605d07316376c7a189b847c98cb2842773270c08ebcf36477f5e17e

Observation db2d023f-35f5-4b32-990b-d8c11bf1d091 · outbound

This paper cites Seeing what mat- ters: Empowering CLIP with patch generation-to-selection.

Early Estimation of Language to Latent Alignment in Diffusion Models Seeing what mat- ters: Empowering CLIP with patch generation-to-selection

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T17:42:44.744314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:42:44.744314Z digest=sha256:78f0998b1ab2c07b96a227cf570560a5c9b6ae9f293f1633a5a68501cea980ca

Observation be42a573-1fd4-4c83-85a2-5007808bfd0d · outbound

This paper cites SDXL: Improving Latent Diffusion Mod- els for High-Resolution Image Synthesis.

Early Estimation of Language to Latent Alignment in Diffusion Models SDXL: Improving Latent Diffusion Mod- els for High-Resolution Image Synthesis

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T17:42:44.805647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:42:44.805647Z digest=sha256:979cf66a5ac517d646347106e8973657e49d4ac8aed46a9f48c93ed738bad540

Observation abb01e2e-3132-45d5-9d30-5032d9405dbb · outbound

This paper cites Learning transferable visual models from natural language supervision.

Early Estimation of Language to Latent Alignment in Diffusion Models Learning transferable visual models from natural language supervision

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T17:42:44.921513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:42:44.921513Z digest=sha256:28f0a23e89367cb80e1eac681e5c09676d65d3c041382eb36af561862ec23d55

Observation f3b2caac-2d11-4db7-b48e-45e8da855029 · outbound

This paper cites an unresolved cited work.

Early Estimation of Language to Latent Alignment in Diffusion Models Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T17:42:45.034164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:42:45.034164Z digest=sha256:b7da87f4539f5eb531cbc0a716f575f09b4b2cf0d715df41f7cea5188d14b588

Observation 6ff6ad89-c252-4bd2-9eba-60beae412ffb · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

Early Estimation of Language to Latent Alignment in Diffusion Models Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T17:42:45.118237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:42:45.118237Z digest=sha256:fda5788f0823e3e3a0f9c99f90a28c4d24809a5139a9a5613e45bf8f9bf4c5da

Observation 681b55b6-b14d-4277-895f-17966a2200fb · outbound

This paper cites Contrastive sequential-diffusion learn- ing: Non-linear and multi-scene instructional video synthe- sis.

Early Estimation of Language to Latent Alignment in Diffusion Models Contrastive sequential-diffusion learn- ing: Non-linear and multi-scene instructional video synthe- sis

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T17:42:45.178558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:42:45.178558Z digest=sha256:a38aa2cd3b2e8f90634d93b2a86834fa68aab36d567da74a66235f2f3c80eeaa

Observation 5d532076-0834-4212-b118-922c42aa6257 · outbound

This paper cites High-resolution image syn- thesis with latent diffusion models.

Early Estimation of Language to Latent Alignment in Diffusion Models High-resolution image syn- thesis with latent diffusion models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-03T17:42:45.302760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:42:45.302760Z digest=sha256:7577fdea9d1e08e9f4097285d761270cc9e01c1ed5b126ec1f5eb867a3586cf4

Observation c202dba4-bd51-4a02-be43-2888b052a6a8 · outbound

This paper cites High-Resolution Image Synthesis with Latent Diffusion Models.

Early Estimation of Language to Latent Alignment in Diffusion Models High-Resolution Image Synthesis with Latent Diffusion Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T17:42:45.378085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:42:45.378085Z digest=sha256:cc2d1c4da2c0aa4808ae9601d57a4976878ac8652daca2a2422834eb4c09a138

Observation 8001f49e-2a7f-42cc-a2e9-be7acb5a3311 · outbound

This paper cites Denton, Seyed Kamyar Seyed Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, Jonathan Ho, David J.

Early Estimation of Language to Latent Alignment in Diffusion Models Denton, Seyed Kamyar Seyed Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, Jonathan Ho, David J

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-03T17:42:45.431054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:42:45.431054Z digest=sha256:4a7ef9701385d9261f4825e3a2faa9cfb7bfb242d625701b45e62d3ce86a1c70

Observation 995c11eb-de6c-44cc-bfab-e8b0307c0ade · outbound

This paper cites Fast high- resolution image synthesis with latent adversarial diffusion distillation.

Early Estimation of Language to Latent Alignment in Diffusion Models Fast high- resolution image synthesis with latent adversarial diffusion distillation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-03T17:42:45.494569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:42:45.494569Z digest=sha256:1454f2ef7548e9cc2abeb13081bf08772865f29fb8eeddcb5489e39a539f8a84

Observation 1d0e75c2-5898-4b07-bebc-a94fae7ab41e · outbound

This paper cites Proposalclip: Unsupervised open-category object proposal generation via exploiting CLIP cues.

Early Estimation of Language to Latent Alignment in Diffusion Models Proposalclip: Unsupervised open-category object proposal generation via exploiting CLIP cues

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-03T17:42:45.555854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:42:45.555854Z digest=sha256:a910af61ddaf0fc2ea65dedbaf9ea971b14086d117018fd13c31dfc9e56cd71c

Observation 9b4e1ed2-d1f0-42c3-81ea-59cb3b404f96 · outbound

This paper cites Generalizing alignment paradigm of text-to-image genera- tion with preferences through f-divergence minimization.

Early Estimation of Language to Latent Alignment in Diffusion Models Generalizing alignment paradigm of text-to-image genera- tion with preferences through f-divergence minimization

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-03T17:42:45.615933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:42:45.615933Z digest=sha256:307e76ec14117a667e7be4d3601dff56c5b17fafd7a7381e094c4191beb2b776

Observation cef10814-e137-4ea1-9f05-8eae77c4ed4e · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

Early Estimation of Language to Latent Alignment in Diffusion Models Representation Learning with Contrastive Predictive Coding

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-03T17:42:45.715373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:42:45.715373Z digest=sha256:280da33173c39d3730a4b1271df9043fd77f663b21891b16a43df45a969eb113

Observation 15def88c-675d-4a4f-afbe-7372394baa26 · outbound

This paper cites Explaining the SDXL la- tent space.https : / / huggingface.

Early Estimation of Language to Latent Alignment in Diffusion Models Explaining the SDXL la- tent space.https : / / huggingface

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-03T17:42:45.791008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:42:45.791008Z digest=sha256:2bc717ee1cee5315c6728ca60da961582aac4067039702f0c6373ccdd3b695d6

Observation f3b47dc0-083c-4d51-9fa2-ba81dc74d906 · outbound

This paper cites Wang, Songwei Ge, Tero Karras, Ming-Yu Liu, and Yogesh Balaji.

Early Estimation of Language to Latent Alignment in Diffusion Models Wang, Songwei Ge, Tero Karras, Ming-Yu Liu, and Yogesh Balaji

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-03T17:42:45.846834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:42:45.846834Z digest=sha256:c155c7eed90751780440f2685e7ce8e75453ae60f945f25af6cfa99e2ac4d31a

Observation 88f45ffb-8ef8-4d48-a899-14bb0e583a1a · outbound

This paper cites DIRE for diffusion-generated image detection.

Early Estimation of Language to Latent Alignment in Diffusion Models DIRE for diffusion-generated image detection

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-03T17:42:45.906111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:42:45.906111Z digest=sha256:a2a36ce16f3e42d342b85224833ed75590305809bd34829a7c3fc34ebf94bc7d

Observation 496e8ca8-7542-4cbe-be1e-1ba8961f9ac2 · outbound

This paper cites Show-o: One single transformer to unify multimodal understanding and generation.

Early Estimation of Language to Latent Alignment in Diffusion Models Show-o: One single transformer to unify multimodal understanding and generation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-03T17:42:46.000697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:42:46.000697Z digest=sha256:c7851e5b06f97a3eaee124f98a63065907d561f39c8f90ab96959d7cb10866c5

Observation 6afb50b1-ce2e-49c1-b527-b06f1bdda2f5 · outbound

This paper cites Good seed makes a good crop: Discovering secret seeds in text-to- image diffusion models.

Early Estimation of Language to Latent Alignment in Diffusion Models Good seed makes a good crop: Discovering secret seeds in text-to- image diffusion models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-03T17:42:46.018049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:42:46.018049Z digest=sha256:7edfb6b9442da86e85b206ae10005df4245eb9c9c6c7d5fee899f21fc6415221

Observation 45c92c83-0408-4789-82f4-6fc233b19171 · outbound

This paper cites Multimodal Inconsistency Reasoning (MMIR): A New Benchmark for Multimodal Reasoning Models.

Early Estimation of Language to Latent Alignment in Diffusion Models Multimodal Inconsistency Reasoning (MMIR): A New Benchmark for Multimodal Reasoning Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-03T17:42:46.128986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:42:46.128986Z digest=sha256:cb343a2c44d4794860cd79e3eec96a1e78dc193b6c12fb4016c5ed0552e5b5c7

Observation 922598df-8391-4fd1-abde-aab6b70b7716 · outbound

This paper cites Diffusion models: A comprehensive survey of methods and applications.ACM computing surveys, 56(4): 1–39, 2023.

Early Estimation of Language to Latent Alignment in Diffusion Models Diffusion models: A comprehensive survey of methods and applications.ACM computing surveys, 56(4): 1–39, 2023

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-03T17:42:46.241114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:42:46.241114Z digest=sha256:2d2ec362e26590de69fb520feaa1ce69e69d2651c192f9dff3ca31b197e093b6

Observation 4658c2ee-3d59-4420-afd5-65465441b7c7 · outbound

This paper cites Object-aware inversion and re- assembly for image editing.

Early Estimation of Language to Latent Alignment in Diffusion Models Object-aware inversion and re- assembly for image editing

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-03T17:42:46.342889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:42:46.342889Z digest=sha256:6df49d334fbe12b0b174fac89a8502ac09b84802ad085921b7a5764c83a85182

Observation 05c79a39-fed9-4d1d-8acc-ae3225567327 · outbound

This paper cites What you see is what you read? improving text- image alignment evaluation.

Early Estimation of Language to Latent Alignment in Diffusion Models What you see is what you read? improving text- image alignment evaluation

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-03T17:42:46.421924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:42:46.421924Z digest=sha256:3bbcfd8c4d23227495b0675afd6633087bfe37ac637a834a24d67e603cdd89dc

Observation db63a228-8628-43d1-b1a0-ee7a7a5bd8d6 · outbound

This paper cites Sigmoid loss for language image pre-training.

Early Estimation of Language to Latent Alignment in Diffusion Models Sigmoid loss for language image pre-training

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-03T17:42:46.584942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:42:46.584942Z digest=sha256:7968ecebc8701d30698dad6a726f3f6157cc9a548da5daefbd2c7902db7e156d

Observation 6e3403d7-df90-45ae-b584-b4d393019349 · outbound

This paper cites Long-clip: Unlocking the long-text capabil- ity of CLIP.

Early Estimation of Language to Latent Alignment in Diffusion Models Long-clip: Unlocking the long-text capabil- ity of CLIP

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-03T17:42:46.705248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:42:46.705248Z digest=sha256:fd3ea9ab3df55e7cbf045d38637aa57a659b73bc2c8f124a613dc7016023243d

Observation 2529d679-a6e7-4502-9442-45ed8355261e · outbound

This paper cites Let’s verify and reinforce image gener- ation step by step.

Early Estimation of Language to Latent Alignment in Diffusion Models Let’s verify and reinforce image gener- ation step by step

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-03T17:42:46.780282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:42:46.780282Z digest=sha256:b3d7d081fd1334c765ff4d363a6ec529d36ba9689f5b5dad30c1344603b9104f

Observation 3f6003db-be9c-4aef-adc5-d3f6714cc218 · outbound

This paper cites Dis- tinctive image captioning via CLIP guided group optimiza- tion.

Early Estimation of Language to Latent Alignment in Diffusion Models Dis- tinctive image captioning via CLIP guided group optimiza- tion

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-03T17:42:46.928724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:42:46.928724Z digest=sha256:eaf0d1bf723bca5a791acdb5245faff23321ffa9c504a0afa7c51f569c745b21

Observation 59ae0145-59d8-41d3-8c90-68bbe2fb9599 · outbound

This paper cites Investigating and mitigating the multimodal hallucination snowballing in large vision-language models.

Early Estimation of Language to Latent Alignment in Diffusion Models Investigating and mitigating the multimodal hallucination snowballing in large vision-language models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-03T17:42:47.054752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:42:47.054752Z digest=sha256:b91c011cd7f93eeb5578a2c9cbbd68dbd4519c961f1d6fdb4724ff8c34f124ea

Observation 07542c4f-819b-41bc-b4b5-c9517346f555 · outbound

This paper cites an unresolved cited work.

Early Estimation of Language to Latent Alignment in Diffusion Models Unresolved cited work

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-03T17:42:47.190226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:42:47.190226Z digest=sha256:6894da284261b3ee4f7d54fd94a7f3f88c618af3902cb2d32fccb1990d4aee77

Observation 0f158136-4bdb-44e1-82f2-0412013a93ce · outbound

This paper cites This modification should render that specific aspect non-factual or illogical within the context of the original sentence.

Early Estimation of Language to Latent Alignment in Diffusion Models This modification should render that specific aspect non-factual or illogical within the context of the original sentence

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-03T17:42:47.314365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:42:47.314365Z digest=sha256:60176053879b00dd56d69fe6847f95d521d2a0a18862113bc9f6805dc12d3054

Observation f52c40aa-f61b-49fb-9e61-0b431463d43a · outbound

This paper cites No additional text, explanations, or formatting.

Early Estimation of Language to Latent Alignment in Diffusion Models No additional text, explanations, or formatting

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-03T17:42:47.413249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:42:47.413249Z digest=sha256:69a074cc8a3f5f0b8462dd838ed41f6a2103535bc1db235ef95ace39c75ef86c

Observation 4cc92bc9-05e4-4a85-84c2-b5f79d1c88e9 · outbound

This paper cites Change the {ERROR_TYPE} in the following PROMPT: PROMPT: {PROMPT} B.

Early Estimation of Language to Latent Alignment in Diffusion Models Change the {ERROR_TYPE} in the following PROMPT: PROMPT: {PROMPT} B

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-03T17:42:47.519295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:42:47.519295Z digest=sha256:07bc01f7b6178539b23a8e3511fa4dc427b5536e04a07e5f6268543a950ea1d7

Pith citing papers

Observation 2a6d5fb0-c158-4713-81a5-0baf1e9c690f · inbound

Assessing Sample Quality in Conditional Generation under Compositional Shift cites this paper.

Assessing Sample Quality in Conditional Generation under Compositional Shift Early Estimation of Language to Latent Alignment in Diffusion Models

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-07-03T00:17:29.056144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T17:21:37.570778Z digest=sha256:59eca5633bfd9149f8bd21b35cbc41858aad8a11962faad90db6ef7a25f35712