Pith. sign in

Paper Citation Record · LEDGER

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs

As of 9 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 1 inbound Pith citation observation for arXiv:2507.19939.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.19939 v1

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T13:58:05.169969Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T09:52:01.768852Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

58 of 58 outbound references displayed

  • verified exact2
  • verified fuzzy35
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 07e9ae56-ddf1-447f-bc70-bf34fe543b0c · outbound

This paper cites Diffit: Diffusion vision transformers for im- age generation.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Diffit: Diffusion vision transformers for im- age generation

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:06.338680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:58:04.946931Z digest=sha256:8e279f3a6107a619104e438771d73f77464048b08359701e09d2354609e18f36

Observation ba332586-5090-4171-81ed-75404db30ac2 · outbound

This paper cites Sketch-guided text-to-image diffusion models.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Sketch-guided text-to-image diffusion models

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:06.325945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:58:04.951193Z digest=sha256:a0741562fbd5e20c657401c94e8ce2e1ca5a67a25883ba1b58313b978bf305a6

Observation c4ee3e43-266a-467d-95fb-8239ef05321d · outbound

This paper cites Language Models are Few-Shot Learners.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Language Models are Few-Shot Learners

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T13:58:04.955410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:58:04.955410Z digest=sha256:37716c57aa82036e176a8953fd41cfabfe67c1cb379e3ec623dea0e874e45084

Observation b7146fd2-fcec-4aa8-aa49-eb21f363e065 · outbound

This paper cites Emergent Abilities of Large Language Models.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Emergent Abilities of Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T13:58:04.959610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:58:04.959610Z digest=sha256:c3f3f5af3f68adf13536f3e6bb0c5dabe45e98e2b9b44b06fb9428cf8900a9d5

Observation 117e6c31-1f2c-4410-bbcc-8af04532b50f · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs LLaMA: Open and Efficient Foundation Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T13:58:04.964121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:58:04.964121Z digest=sha256:4f20acb7abc453be081705cc184142ba340f5739e30c83a8e06a472b6101e9dc

Observation 084b781f-ebe9-42a4-8549-550deea7ee0c · outbound

This paper cites Instruction Tuning with GPT-4.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Instruction Tuning with GPT-4

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T13:58:04.968184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:58:04.968184Z digest=sha256:1b97fe1c0a9eaadadd41291d0b6ce69b34fe31370184dc7370863eb7f88f6bcc

Observation e7c8796a-5a87-450c-9b6f-9242b661bb1f · outbound

This paper cites UniControl: A Unified Diffusion Model for Controllable Visual Generation In the Wild.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs UniControl: A Unified Diffusion Model for Controllable Visual Generation In the Wild

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T13:58:04.972222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:58:04.972222Z digest=sha256:a984d63c4f053788106da5de593264011bd25bddfbe9f789adba325c9f148de3

Observation 171efb84-11c5-48a5-b5d8-9d20abe572af · outbound

This paper cites T2i-adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs T2i-adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:06.313044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:58:04.976692Z digest=sha256:92cdd53210dae5aeeb370a638d477e5b7bb5c7e177be415af6fedfde936146ed

Observation bcd86d95-aa16-4441-ae02-c68103fd6549 · outbound

This paper cites Migc: Multi-instance generation controller for text-to-image synthesis.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Migc: Multi-instance generation controller for text-to-image synthesis

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:06.299741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:58:04.980592Z digest=sha256:c2d41d4cfed18728c43e0ced03801d7c46eb5e58d79b29e41aa8db3a8294f259

Observation 52a95384-79c8-48c5-9f2d-ef7c6610dcd7 · outbound

This paper cites Minigpt-4: Enhancing vision-language understanding with advanced large language models.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Minigpt-4: Enhancing vision-language understanding with advanced large language models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:06.285704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:58:04.984214Z digest=sha256:6fc43c0657c81fce19296582477c70feb39b8ba19059f96bf184ad0fd40ded97

Observation 6eaf496b-d8b5-457e-bbea-6eca8fcde988 · outbound

This paper cites Cogview: Mastering text-to-image generation via transformers.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Cogview: Mastering text-to-image generation via transformers

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:06.272597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:58:04.987777Z digest=sha256:9eff2b35fbbb3fe11dc826c09d5669db55d9c8485c30be4026e454b0b462a35d

Observation 6900dd29-b12c-4755-a336-471b5f0335be · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T13:58:04.991565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:58:04.991565Z digest=sha256:05d4f611472157c515f030a3b9cb1cab1c55bfec776e77e68b8019b2743144af

Observation 13a8f93a-9671-44cd-9f4c-7d768096f52f · outbound

This paper cites Freeman, Fr ´edo Durand, and Song Han.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Freeman, Fr ´edo Durand, and Song Han

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:06.259576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:58:04.995851Z digest=sha256:300d8471ec9aac170b5774e286cbade1317487a201cd827a176c7f935a87e07c

Observation e372bee8-fdf3-427a-a54e-9430c798aad7 · outbound

This paper cites Freeman, Fr ´edo Durand, and Song Han.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Freeman, Fr ´edo Durand, and Song Han

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:06.245961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:58:04.999510Z digest=sha256:03dcc0b422efbb53d01c5f9ea1ee00753d753831babfac7e9bdb1db3f7e18c89

Observation a026925c-8916-426c-af89-1f7a18dfca60 · outbound

This paper cites Cross-modal contrastive learning for text-to- image generation.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Cross-modal contrastive learning for text-to- image generation

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:06.231757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:58:05.003315Z digest=sha256:88624968b702e7092c770a5bd9c17ab53d8ef8c1585fa8dea34ae01367210f17

Observation d7ddf754-1a80-441a-9475-41b3333b530c · outbound

This paper cites IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T13:58:05.006991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:58:05.006991Z digest=sha256:7baf8273642362b94677c2fd3f9318e4258e4ccb856334f90946d2284301798d

Observation c672f31d-3987-41dd-ac11-5cb1fdbadc9d · outbound

This paper cites an unresolved cited work.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Unresolved cited work

Reference 17

Resolution
verified exact
raw_fallback, observed 2026-08-06T13:58:05.764736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:58:05.011442Z digest=sha256:0d8181782d0bd5f916ed717035d698e6f0072bf95d76a405fafcf80e25a7e1e0

Observation f49346e4-5003-42e9-b883-139052201eca · outbound

This paper cites Scaling autoregressive models for content-rich text-to-image genera- tion.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Scaling autoregressive models for content-rich text-to-image genera- tion

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T13:58:05.015533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:58:05.015533Z digest=sha256:af61d7d66075a5e7142f54738626fda8f6f31524a90c0ec6b55a702022a3a109

Observation 49cc29f7-ae85-47f5-9d8d-d0354b7ef22d · outbound

This paper cites Denoising Diffusion Implicit Models.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Denoising Diffusion Implicit Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T13:58:05.019334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:58:05.019334Z digest=sha256:1ed42a4efd4d4e3baaadb643c266b78da6a2eb96d58b94b8ce93bbf3e0322baf

Observation 387a71a9-9a1e-4380-965d-b4a3df88279d · outbound

This paper cites Open-vocabulary panop- tic segmentation with text-to-image diffusion models.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Open-vocabulary panop- tic segmentation with text-to-image diffusion models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:06.218527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:58:05.023001Z digest=sha256:8b3062d61587ae783d24a769a5c1fe8c00df17e9fa5d1741419a1241b620979e

Observation bb990f3b-b75d-4482-b143-6cdd2f770cd2 · outbound

This paper cites Classifier-Free Diffusion Guidance.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Classifier-Free Diffusion Guidance

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T13:58:05.026909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:58:05.026909Z digest=sha256:2084aab014b37b7f2fb1f262670da0acb3a699ce2af96aaf0d21476ca6e5a4ee

Observation 045cb68c-5a5e-4c06-8745-e09cd8738412 · outbound

This paper cites Perception priori- tized training of diffusion models.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Perception priori- tized training of diffusion models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:06.205937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:58:05.030582Z digest=sha256:eb99545a1772d77d3e85c8ffd6dfe38869d4c98e4e75a7cb8807afe3ed6dcf02

Observation 01ca8309-71cb-4dc4-9625-b066f51b68a3 · outbound

This paper cites Training Compute-Optimal Large Language Models.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Training Compute-Optimal Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T13:58:05.034238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:58:05.034238Z digest=sha256:32dca58687769b648baf2b60b85744381c47140fd03b0c6e588522f278b269d0

Observation ba033c7d-4d55-45e4-9218-fd745383be60 · outbound

This paper cites GPT-4 Technical Report.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs GPT-4 Technical Report

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T13:58:05.038403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:58:05.038403Z digest=sha256:dcbea364816b9c8ec11ae6e711f852bb46e34bfbe3c44c30a9ccdf492bf0f9ab

Observation 96d40ce9-fc3e-47cc-8936-c5206b6cfb6e · outbound

This paper cites Blip: Bootstrapping language-image pre-training for uni- fied vision-language understanding and generation.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Blip: Bootstrapping language-image pre-training for uni- fied vision-language understanding and generation

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:06.193167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:58:05.042052Z digest=sha256:6d24f2da1dbd21301fcef310399a38df7e0fa6870b0d208a169d25b32a4cf003

Observation dec473ca-5274-436a-8276-a27034d4da20 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:06.179252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:58:05.045817Z digest=sha256:5a1d276c4fe600279d8918e029c89f41d4d2d5ed7ff6050d90efa45a8d1401de

Observation aedc824d-b6f6-442a-a7bd-ae989798ff91 · outbound

This paper cites Composer: Creative and Controllable Image Synthesis with Composable Conditions.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Composer: Creative and Controllable Image Synthesis with Composable Conditions

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T13:58:05.049275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:58:05.049275Z digest=sha256:61449abbb9ea1f4522b8a85595fc98cfaa8c92909b64dff1b372195b7c997640

Observation 0a970c1c-d1f8-4434-9689-93710fb5c2ff · outbound

This paper cites LLM-grounded Diffusion: Enhancing Prompt Understanding of Text-to-Image Diffusion Models with Large Language Models.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs LLM-grounded Diffusion: Enhancing Prompt Understanding of Text-to-Image Diffusion Models with Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T13:58:05.053158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:58:05.053158Z digest=sha256:4b071f12d15da08ddb759b325c68ec297dfe6ee42985d3cb8ab603c88574085c

Observation fd0473fa-31b1-4309-9aeb-72ca7fb8cc91 · outbound

This paper cites Adding conditional control to text-to-image diffusion models.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Adding conditional control to text-to-image diffusion models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:06.165248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:58:05.057314Z digest=sha256:7ba21df57cfc5febd20d39d31220c229ad4101509b954bf6302e0f74a20d5243

Observation d0195274-c817-4e8b-a3ba-0f901497dd7d · outbound

This paper cites Fourier features let networks learn high frequency functions in low dimen- sional domains.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Fourier features let networks learn high frequency functions in low dimen- sional domains

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:06.152391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:58:05.061006Z digest=sha256:e3f5a73f072ba678662b86f933a1daf2745a652b777db15ea268b49ae27c19ec

Observation 9d934d97-73ad-40df-b623-a83d283a0493 · outbound

This paper cites Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:06.139406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:58:05.064778Z digest=sha256:3ecd7a1733835ceb57f230e2f94ec5bd9eaf174cad2f974bd3fa26b4b9bb875f

Observation 5dd8cdb2-bf75-499e-b609-e59e05b3f2c8 · outbound

This paper cites Hyperdreambooth: Hypernetworks for fast personalization of text-to-image models.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Hyperdreambooth: Hypernetworks for fast personalization of text-to-image models

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:06.127185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:58:05.068589Z digest=sha256:11ecae24eaa77915c08738d45be11c9aaa193ccd0bebe298fc173e08565304e7

Observation d3c4b848-ad70-490f-9bbc-fdcc948ddf11 · outbound

This paper cites Spatext: Spatio-textual representation for con- trollable image generation.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Spatext: Spatio-textual representation for con- trollable image generation

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:06.114259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:58:05.072153Z digest=sha256:bb4cb6a8baedab0ba2042b633e3af71dbaf3ba1a5de1d79e1cbec1a4a652bd44

Observation f00b8a3d-e113-493b-ae07-9bb2c96cdf12 · outbound

This paper cites Scaling Rectified Flow Transformers for High-Resolution Image Synthesis.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Scaling Rectified Flow Transformers for High-Resolution Image Synthesis

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T13:58:05.075777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:58:05.075777Z digest=sha256:70d3a77b8d51b86082918bf0d2a8cc496465c389fd057089e0222704fd000f7d

Observation 5f9b0f5c-a565-44a1-91c0-e3870f289e9c · outbound

This paper cites Diffusion models beat gans on image synthesis.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Diffusion models beat gans on image synthesis

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:06.101424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:58:05.080272Z digest=sha256:8e840d98f337c56ef6b9412fd614ea2ab3544770671c5378ec3c99accf85c778

Observation a820c14e-3831-45c8-b988-edf429f9c255 · outbound

This paper cites An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T13:58:05.084309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:58:05.084309Z digest=sha256:08d523110e41c745c4b63ea968f76ae4d5ba540c3808d31400dccad2114bb054

Observation e18275d2-9c57-48f6-a3e7-aca56999271f · outbound

This paper cites High-resolution image syn- thesis with latent diffusion models.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs High-resolution image syn- thesis with latent diffusion models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:06.089138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:58:05.088465Z digest=sha256:abe7bfdab3b384c7619bc72fe37072cb6abeddfd823a723d09c91af03e897fe1

Observation 1590a588-bddb-4a99-9aa1-282788ff9afe · outbound

This paper cites Generative ad- versarial text to image synthesis.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Generative ad- versarial text to image synthesis

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:06.076676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:58:05.091929Z digest=sha256:8fb3fc2f988453ac7ed6d2ace697b3f4da570adce7811d99286eadff0d2e2279

Observation 2882007e-6483-443f-b211-b9d621daf465 · outbound

This paper cites Younes Mirinezhad.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Younes Mirinezhad

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:06.063281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:58:05.095580Z digest=sha256:c904382a6a6f76b6034ebcd35dee64c5ec266827cc23c1449f960d935f8bd697

Observation aa69289a-f977-41b8-8bce-40f4f1818b70 · outbound

This paper cites an unresolved cited work.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:58:06.050460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:58:05.099809Z digest=sha256:39382e8c2a505a13e067620583672ea013adf0563326b1cd34b17bbae8da5007

Observation 8b607d78-345b-402f-9019-477285db138b · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T13:58:05.103489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:58:05.103489Z digest=sha256:9ede8f97c769e1c23b79f9de0d981ab0b3b4e850d3882d4dc7e4d7eaf6888265

Observation dff75b55-9d13-4164-81f5-79445e019411 · outbound

This paper cites Freecontrol: Training-free spatial control of any text-to-image diffusion model with any condition.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Freecontrol: Training-free spatial control of any text-to-image diffusion model with any condition

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:06.037096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:58:05.107430Z digest=sha256:45f36dd72a96198eeb7cbe7f4c03b011e0c09b11bb4f7674606fdb59ce437697

Observation 9f202afb-5277-4256-9229-d79520cce8a1 · outbound

This paper cites Deep unsupervised learning using nonequilibrium thermodynamics.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Deep unsupervised learning using nonequilibrium thermodynamics

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:06.024283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:58:05.111372Z digest=sha256:be2a561c868c8eefdaf8fd2f1fcf18638befa50687a54d71a4c6844f2b3eb1c8

Observation 365b3635-be53-442c-bcc9-5976687a29ed · outbound

This paper cites Score-based 9 generative modeling through stochastic differential equa- tions.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Score-based 9 generative modeling through stochastic differential equa- tions

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:06.011333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:58:05.115135Z digest=sha256:ac5ea01e5cb933fac7859692f4e16dc04e26b4c22005036bc5d0e54c4cea874a

Observation f4a4d41f-fe1c-4a7e-9c56-368377856526 · outbound

This paper cites It’s all about your sketch: Democratising sketch control in diffusion models.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs It’s all about your sketch: Democratising sketch control in diffusion models

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:05.997752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:58:05.119314Z digest=sha256:a39b5563b686295c6514fef755711b42188a2bf4ee2875fa9af1963e1bd4d4d0

Observation ad81617a-d9db-4dc4-867e-6a734b62bd54 · outbound

This paper cites Attngan: Fine- grained text to image generation with attentional generative adversarial networks.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Attngan: Fine- grained text to image generation with attentional generative adversarial networks

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:05.984472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:58:05.123447Z digest=sha256:1c5885bcebed2de3a83b640233b359d387102af3e545d72ba718272ca79513f6

Observation 042213c5-e384-4891-81bd-13bd131496f3 · outbound

This paper cites Plug-and-play diffusion features for text-driven image-to-image translation.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Plug-and-play diffusion features for text-driven image-to-image translation

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:05.971428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:58:05.127180Z digest=sha256:5ee993278bff40428e656fa9dad4036500b7bd5967b7a715ecd8c5c413b0901b

Observation 40174456-e7e8-4dc2-aaae-aec564f43fd0 · outbound

This paper cites Compositional Text-to-Image Generation with Dense Blob Representations.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Compositional Text-to-Image Generation with Dense Blob Representations

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T13:58:05.130972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:58:05.130972Z digest=sha256:0cb74229aa8e317fb55aea9c4e9b0202c6e05d54ce7798855fc9bd1d5d236616

Observation 85990893-f0a0-45e3-b0ac-bfd9e6f7ab99 · outbound

This paper cites Layoutgpt: Compositional visual plan- ning and generation with large language models.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Layoutgpt: Compositional visual plan- ning and generation with large language models

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:05.957221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:58:05.135424Z digest=sha256:5704c985db38206a6f105b43f36446a4cc7dea8694b84b855999d5100b2bc030

Observation 6625bbde-6465-4073-9aa9-78dc70909b7a · outbound

This paper cites ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T13:58:05.138991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:58:05.138991Z digest=sha256:f12c0a8701c2d2f8ac5ffa6fe86ae33802a7f351d29cffe07fbed1c0627ee36b

Observation dac461c6-ad6a-48a2-86ce-d515f6dd62ed · outbound

This paper cites Humansd: A native skeleton-guided diffusion model for human image generation.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Humansd: A native skeleton-guided diffusion model for human image generation

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:05.943248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:58:05.142871Z digest=sha256:2fac519655542f1adf9511dc72994473abfb0176d90a57a343d9ee2e1ccd8c04

Observation 2be2634a-9867-456f-97c1-6ceb237ba7e2 · outbound

This paper cites eDiff-I: Text-to-Image Diffusion Models with an Ensemble of Expert Denoisers.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs eDiff-I: Text-to-Image Diffusion Models with an Ensemble of Expert Denoisers

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T13:58:05.146539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:58:05.146539Z digest=sha256:9685223bddc1c5fec9d9e05c115531bb7e808ddf67d283abf2604327377586bf

Observation 8b6efc69-2ffd-4ae7-bccb-0251a4f3104e · outbound

This paper cites Lafite2: Few-shot Text-to-Image Generation.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Lafite2: Few-shot Text-to-Image Generation

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-08-06T13:58:05.211626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:58:05.150622Z digest=sha256:e17423769c9781d7235627dd38cf74d834777f84cdd6d124b0686f4e90a1a515

Observation b1d736c3-2c9e-4aed-b529-7dc8e21c151e · outbound

This paper cites Gligen: Open-set grounded text-to-image generation.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Gligen: Open-set grounded text-to-image generation

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:05.929027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:58:05.154643Z digest=sha256:b7f937b711d12a151a536417ffaa17c2e4b5f74b50501ffc66f6991f650e690f

Observation 64e4b6a5-2c47-43dc-998c-bf7fb7cd7e00 · outbound

This paper cites Ranni: Taming text-to-image diffusion for accurate instruction following.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Ranni: Taming text-to-image diffusion for accurate instruction following

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:05.915734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:58:05.158515Z digest=sha256:9b26b30fb02033d4757b37feece675a53b01c99d3dec7028050189c19a879490

Observation c05889d0-fa51-4631-a53c-603f36e5e8fb · outbound

This paper cites Smartedit: Ex- ploring complex instruction-based image editing with multi- modal large language models.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Smartedit: Ex- ploring complex instruction-based image editing with multi- modal large language models

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:05.902359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:58:05.162408Z digest=sha256:ac4381f5f1addfa01a0218e650d0a4cca27f7f95540bb04719b9c8fa62752bf8

Observation 0fbbccd9-48a4-463e-90e8-709a42ec2a94 · outbound

This paper cites Layoutdiffusion: Controllable diffu- sion model for layout-to-image generation.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Layoutdiffusion: Controllable diffu- sion model for layout-to-image generation

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:05.888090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:58:05.166417Z digest=sha256:635ea0648b3ece6eb8ac9455d8cc74c2f914ab08edf39783ae9d51fff3a84cc1

Observation aa617503-f3f4-4e57-a2f1-b75cbf59137a · outbound

This paper cites Reco: Region- controlled text-to-image generation.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Reco: Region- controlled text-to-image generation

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:05.875093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:58:05.169969Z digest=sha256:63b2f98f2a4dd3ce5988e780713978a4b81cf9b7f1d67cffce7b4c022b61497d

Pith citing papers

Observation 136e6b49-1b99-449a-b459-2878b5d978c5 · inbound

EventOD: Event-Aware OD Flow Generation via LLM-Guided Semantic Modulation cites this paper.

EventOD: Event-Aware OD Flow Generation via LLM-Guided Semantic Modulation LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T09:52:01.768852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:52:01.768852Z digest=sha256:603d669bb578ffcc2969125a0476b960ce36ea9dfd17b1bcebe81584a52c3b7e