Pith. sign in

Paper Citation Record · LEDGER

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs

As of 23 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 1 inbound Pith citation observation for arXiv:2507.19939.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.19939 v1

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T13:58:05.169969Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T09:52:01.768852Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

58 of 58 outbound references displayed

  • verified exact2
  • verified fuzzy35
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 07e9ae56-ddf1-447f-bc70-bf34fe543b0c · outbound

This paper cites Diffit: Diffusion vision transformers for im- age generation.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Diffit: Diffusion vision transformers for im- age generation

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:06.338680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T13:58:04.946931Z digest=sha256:ad0c0236e0304d00922d2c3c039538a2dc48696e2990708dadfe8c4cbe577b23

Observation ba332586-5090-4171-81ed-75404db30ac2 · outbound

This paper cites Sketch-guided text-to-image diffusion models.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Sketch-guided text-to-image diffusion models

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:06.325945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T13:58:04.951193Z digest=sha256:0d57cc54eb512c10eb9c4c1a6b784d449f4d9dc5b45db4295d9525049856cd08

Observation c4ee3e43-266a-467d-95fb-8239ef05321d · outbound

This paper cites Language Models are Few-Shot Learners.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Language Models are Few-Shot Learners

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T13:58:04.955410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:58:04.955410Z digest=sha256:32600a8ee07683489326cab1a360d69d82bcf592748e501ce1253d61d258cbb9

Observation b7146fd2-fcec-4aa8-aa49-eb21f363e065 · outbound

This paper cites Emergent Abilities of Large Language Models.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Emergent Abilities of Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T13:58:04.959610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:58:04.959610Z digest=sha256:194176eaddc337c9417c4e9a59c28203e200b9f449db6c765dbce401eac30697

Observation 117e6c31-1f2c-4410-bbcc-8af04532b50f · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs LLaMA: Open and Efficient Foundation Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T13:58:04.964121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:58:04.964121Z digest=sha256:21dd98c0e015cdcfc0a6fa94d85fecc571b5e8e4c7dfe710fe932f835e00310a

Observation 084b781f-ebe9-42a4-8549-550deea7ee0c · outbound

This paper cites Instruction Tuning with GPT-4.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Instruction Tuning with GPT-4

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T13:58:04.968184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:58:04.968184Z digest=sha256:97d3d8f4c90f73978c7c77d61d84a7329020aeaddd27790941f2ba7d2f5921d0

Observation e7c8796a-5a87-450c-9b6f-9242b661bb1f · outbound

This paper cites UniControl: A Unified Diffusion Model for Controllable Visual Generation In the Wild.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs UniControl: A Unified Diffusion Model for Controllable Visual Generation In the Wild

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T13:58:04.972222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:58:04.972222Z digest=sha256:485604ce7c53604ca5d41c65bb4a001f58dcce3883cde6d068bfb5502b09e30f

Observation 171efb84-11c5-48a5-b5d8-9d20abe572af · outbound

This paper cites T2i-adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs T2i-adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:06.313044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T13:58:04.976692Z digest=sha256:38b83d68b9f9cdecc714586097a30f2438a0918c580746f659d81c1b2d7a9488

Observation bcd86d95-aa16-4441-ae02-c68103fd6549 · outbound

This paper cites Migc: Multi-instance generation controller for text-to-image synthesis.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Migc: Multi-instance generation controller for text-to-image synthesis

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:06.299741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T13:58:04.980592Z digest=sha256:1287b2b65eb29867fee343da2e479dd64f2b181ffdecbabc0ab5a00d22a881fb

Observation 52a95384-79c8-48c5-9f2d-ef7c6610dcd7 · outbound

This paper cites Minigpt-4: Enhancing vision-language understanding with advanced large language models.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Minigpt-4: Enhancing vision-language understanding with advanced large language models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:06.285704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T13:58:04.984214Z digest=sha256:b6d1dbbee74a739197fee4d73777e7e0784530eb687c74b02cdbda0326df6376

Observation 6eaf496b-d8b5-457e-bbea-6eca8fcde988 · outbound

This paper cites Cogview: Mastering text-to-image generation via transformers.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Cogview: Mastering text-to-image generation via transformers

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:06.272597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T13:58:04.987777Z digest=sha256:7fcf458d367c0669a22cc63c551cfe6f1d077371852d8d021a971682a4a5f704

Observation 6900dd29-b12c-4755-a336-471b5f0335be · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T13:58:04.991565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:58:04.991565Z digest=sha256:40e54dadbbd1b5485293b25428b7a9680dc1586fb9201dc6dedfd7175bd6c703

Observation 13a8f93a-9671-44cd-9f4c-7d768096f52f · outbound

This paper cites Freeman, Fr ´edo Durand, and Song Han.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Freeman, Fr ´edo Durand, and Song Han

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:06.259576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T13:58:04.995851Z digest=sha256:d4400fc89c5f1cec5bf81f83a486a9d13b05fe516cc438c83d14e3a6a79105cd

Observation e372bee8-fdf3-427a-a54e-9430c798aad7 · outbound

This paper cites Freeman, Fr ´edo Durand, and Song Han.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Freeman, Fr ´edo Durand, and Song Han

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:06.245961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T13:58:04.999510Z digest=sha256:50775d05ffa71ae824516a1dd2b4afa65f1391d1ca5f8a665926226ccd58145a

Observation a026925c-8916-426c-af89-1f7a18dfca60 · outbound

This paper cites Cross-modal contrastive learning for text-to- image generation.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Cross-modal contrastive learning for text-to- image generation

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:06.231757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T13:58:05.003315Z digest=sha256:0b3ae738dd86e542c75c206c5d7649fff97cabf69eacd2c141fb7e0514e625ad

Observation d7ddf754-1a80-441a-9475-41b3333b530c · outbound

This paper cites IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T13:58:05.006991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:58:05.006991Z digest=sha256:695483e2df7a33368e5790466c3d09dba30cd10193bc7b524201835842b63d14

Observation c672f31d-3987-41dd-ac11-5cb1fdbadc9d · outbound

This paper cites an unresolved cited work.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Unresolved cited work

Reference 17

Resolution
verified exact
raw_fallback, observed 2026-08-06T13:58:05.764736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T13:58:05.011442Z digest=sha256:cec5d8b6f72420b757fc9b9caa781abc6d29e8c1d14d13c8b1412f13a4658fd8

Observation f49346e4-5003-42e9-b883-139052201eca · outbound

This paper cites Scaling autoregressive models for content-rich text-to-image genera- tion.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Scaling autoregressive models for content-rich text-to-image genera- tion

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T13:58:05.015533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:58:05.015533Z digest=sha256:d8c002e1cc5bf6907d98afbbf0e12754202279413b09d2acc816cc98d7a20106

Observation 49cc29f7-ae85-47f5-9d8d-d0354b7ef22d · outbound

This paper cites Denoising Diffusion Implicit Models.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Denoising Diffusion Implicit Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T13:58:05.019334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:58:05.019334Z digest=sha256:b9ed9533a784e7d4dfd9b29404a00ff58c86cf603d00253dc336bc80ef0150b4

Observation 387a71a9-9a1e-4380-965d-b4a3df88279d · outbound

This paper cites Open-vocabulary panop- tic segmentation with text-to-image diffusion models.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Open-vocabulary panop- tic segmentation with text-to-image diffusion models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:06.218527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T13:58:05.023001Z digest=sha256:7c8a4d6b2860c4abd9f41df7ada1266128df98ecdda0f44f82237de874fc5852

Observation bb990f3b-b75d-4482-b143-6cdd2f770cd2 · outbound

This paper cites Classifier-Free Diffusion Guidance.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Classifier-Free Diffusion Guidance

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T13:58:05.026909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:58:05.026909Z digest=sha256:ba2cb606578a1f34308c42292c2a61d378b10951601db29a43739d8a2770a0b6

Observation 045cb68c-5a5e-4c06-8745-e09cd8738412 · outbound

This paper cites Perception priori- tized training of diffusion models.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Perception priori- tized training of diffusion models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:06.205937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T13:58:05.030582Z digest=sha256:8ca79ca297df270aa1195c86af2c0a518f2e7140f7f8826bca67f9b67aeb1f83

Observation 01ca8309-71cb-4dc4-9625-b066f51b68a3 · outbound

This paper cites Training Compute-Optimal Large Language Models.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Training Compute-Optimal Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T13:58:05.034238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:58:05.034238Z digest=sha256:4503d1e8e36ebce6fa4d0c880949a8b3030305b9d38c9c1b72c839bd753d1f66

Observation ba033c7d-4d55-45e4-9218-fd745383be60 · outbound

This paper cites GPT-4 Technical Report.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs GPT-4 Technical Report

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T13:58:05.038403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:58:05.038403Z digest=sha256:4c1c21f47e85c2b155290e50c70e68af8b19d951597a973c33249973ae3e15a3

Observation 96d40ce9-fc3e-47cc-8936-c5206b6cfb6e · outbound

This paper cites Blip: Bootstrapping language-image pre-training for uni- fied vision-language understanding and generation.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Blip: Bootstrapping language-image pre-training for uni- fied vision-language understanding and generation

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:06.193167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T13:58:05.042052Z digest=sha256:ce85082223b879b3095ef5805a4bcd5ae5fc4996bb2b8c60f9c8e8f9f6914c88

Observation dec473ca-5274-436a-8276-a27034d4da20 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:06.179252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T13:58:05.045817Z digest=sha256:bd32ab1758dc39e83d30c7ce7e267c36f03e709528419b6ae0741da1c132a4b3

Observation aedc824d-b6f6-442a-a7bd-ae989798ff91 · outbound

This paper cites Composer: Creative and Controllable Image Synthesis with Composable Conditions.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Composer: Creative and Controllable Image Synthesis with Composable Conditions

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T13:58:05.049275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:58:05.049275Z digest=sha256:247c2c37490a45d0a62d72964438983a78269fb002aa9514783dc81c5fd1f85e

Observation 0a970c1c-d1f8-4434-9689-93710fb5c2ff · outbound

This paper cites LLM-grounded Diffusion: Enhancing Prompt Understanding of Text-to-Image Diffusion Models with Large Language Models.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs LLM-grounded Diffusion: Enhancing Prompt Understanding of Text-to-Image Diffusion Models with Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T13:58:05.053158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:58:05.053158Z digest=sha256:0ceadcc1b5dc3c4585f047ba18421481bf04f99a8c866a48a0526947621db7c5

Observation fd0473fa-31b1-4309-9aeb-72ca7fb8cc91 · outbound

This paper cites Adding conditional control to text-to-image diffusion models.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Adding conditional control to text-to-image diffusion models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:06.165248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T13:58:05.057314Z digest=sha256:b4f345f46f1b16c1058a8f2698e61d31f7c4fc5096b7528f6a60c786f4197170

Observation d0195274-c817-4e8b-a3ba-0f901497dd7d · outbound

This paper cites Fourier features let networks learn high frequency functions in low dimen- sional domains.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Fourier features let networks learn high frequency functions in low dimen- sional domains

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:06.152391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T13:58:05.061006Z digest=sha256:a019aaefdf41be5f4125d99cdf626f5538bc7f3982740e40d3f8da9220112e35

Observation 9d934d97-73ad-40df-b623-a83d283a0493 · outbound

This paper cites Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:06.139406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T13:58:05.064778Z digest=sha256:c1d68d4bdfb73e3040e5139b24c012e3fc7360f0ee5aaeaae4b6d52c0f6b43dc

Observation 5dd8cdb2-bf75-499e-b609-e59e05b3f2c8 · outbound

This paper cites Hyperdreambooth: Hypernetworks for fast personalization of text-to-image models.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Hyperdreambooth: Hypernetworks for fast personalization of text-to-image models

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:06.127185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T13:58:05.068589Z digest=sha256:3e7737b2022e20f9a5bbfebcaf4fab97421b9882b01d9cc372428c28ac0e4d21

Observation d3c4b848-ad70-490f-9bbc-fdcc948ddf11 · outbound

This paper cites Spatext: Spatio-textual representation for con- trollable image generation.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Spatext: Spatio-textual representation for con- trollable image generation

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:06.114259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T13:58:05.072153Z digest=sha256:e74d477e4a15772157b5962f09252b2e6965f71b5e6a12eb21fb75decabb7f81

Observation f00b8a3d-e113-493b-ae07-9bb2c96cdf12 · outbound

This paper cites Scaling Rectified Flow Transformers for High-Resolution Image Synthesis.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Scaling Rectified Flow Transformers for High-Resolution Image Synthesis

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T13:58:05.075777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:58:05.075777Z digest=sha256:af8ca613ecb561382e96346686c69334af46d41670b9b53bb54b8ed5907f907c

Observation 5f9b0f5c-a565-44a1-91c0-e3870f289e9c · outbound

This paper cites Diffusion models beat gans on image synthesis.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Diffusion models beat gans on image synthesis

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:06.101424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T13:58:05.080272Z digest=sha256:f3fb18a26adbb2a48190d7301e336b08c4208cc8dda3b44d90682be303c2fcfb

Observation a820c14e-3831-45c8-b988-edf429f9c255 · outbound

This paper cites An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T13:58:05.084309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:58:05.084309Z digest=sha256:b4cfe1e5e9f1486ad7cc4ea8b447c346ba1c5d445fa98bb26b247f36570458af

Observation e18275d2-9c57-48f6-a3e7-aca56999271f · outbound

This paper cites High-resolution image syn- thesis with latent diffusion models.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs High-resolution image syn- thesis with latent diffusion models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:06.089138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T13:58:05.088465Z digest=sha256:ba40f2acaaefd774c2a5e9d05ded84df90019ba4a8479e3ccea066b6819637e8

Observation 1590a588-bddb-4a99-9aa1-282788ff9afe · outbound

This paper cites Generative ad- versarial text to image synthesis.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Generative ad- versarial text to image synthesis

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:06.076676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T13:58:05.091929Z digest=sha256:fe61f83d4c9fb4211acb7b72848e63ea10b8c0759ef6ca4a93dd16efeb9a22e4

Observation 2882007e-6483-443f-b211-b9d621daf465 · outbound

This paper cites Younes Mirinezhad.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Younes Mirinezhad

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:06.063281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T13:58:05.095580Z digest=sha256:73a50bfa1d3b11ce512fc05ab663c36f1c025397f9c60d0dffc874eb18c4ed73

Observation aa69289a-f977-41b8-8bce-40f4f1818b70 · outbound

This paper cites an unresolved cited work.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:58:06.050460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T13:58:05.099809Z digest=sha256:5abf835ed9efbde1403c0ec78c7b2b79fad2d769af5d28569a35ba918b754d91

Observation 8b607d78-345b-402f-9019-477285db138b · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T13:58:05.103489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:58:05.103489Z digest=sha256:d022d889d79596ac5128cc4b1bb555bc9f9cbd56c10cdf25869862ca6feecd9d

Observation dff75b55-9d13-4164-81f5-79445e019411 · outbound

This paper cites Freecontrol: Training-free spatial control of any text-to-image diffusion model with any condition.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Freecontrol: Training-free spatial control of any text-to-image diffusion model with any condition

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:06.037096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T13:58:05.107430Z digest=sha256:12e6e51cf2f712b230b31fdd7e375bced030cbb735172f642569bb08965a7c11

Observation 9f202afb-5277-4256-9229-d79520cce8a1 · outbound

This paper cites Deep unsupervised learning using nonequilibrium thermodynamics.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Deep unsupervised learning using nonequilibrium thermodynamics

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:06.024283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T13:58:05.111372Z digest=sha256:2cbcb06a8e00c66683782d68d91a76766177e199721f82257e3bfd6340329b7c

Observation 365b3635-be53-442c-bcc9-5976687a29ed · outbound

This paper cites Score-based 9 generative modeling through stochastic differential equa- tions.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Score-based 9 generative modeling through stochastic differential equa- tions

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:06.011333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T13:58:05.115135Z digest=sha256:55601631bd85488703d3e2338971c8d1985f6b71268879013f6ce85d48acc9e1

Observation f4a4d41f-fe1c-4a7e-9c56-368377856526 · outbound

This paper cites It’s all about your sketch: Democratising sketch control in diffusion models.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs It’s all about your sketch: Democratising sketch control in diffusion models

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:05.997752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T13:58:05.119314Z digest=sha256:e18ec57f2ff2435d635bbe9b5787a6d203855c8a7aeb536a981021dcbfb97c96

Observation ad81617a-d9db-4dc4-867e-6a734b62bd54 · outbound

This paper cites Attngan: Fine- grained text to image generation with attentional generative adversarial networks.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Attngan: Fine- grained text to image generation with attentional generative adversarial networks

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:05.984472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T13:58:05.123447Z digest=sha256:39a59deca82467a6c3adb09efbfb01325f65b24d8f969be1b19f47a0432f00c1

Observation 042213c5-e384-4891-81bd-13bd131496f3 · outbound

This paper cites Plug-and-play diffusion features for text-driven image-to-image translation.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Plug-and-play diffusion features for text-driven image-to-image translation

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:05.971428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T13:58:05.127180Z digest=sha256:b39e61133f7594ad44aaeb3034de513e75dd504ada3c6c3b34eea126f3f0c55f

Observation 40174456-e7e8-4dc2-aaae-aec564f43fd0 · outbound

This paper cites Compositional Text-to-Image Generation with Dense Blob Representations.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Compositional Text-to-Image Generation with Dense Blob Representations

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T13:58:05.130972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:58:05.130972Z digest=sha256:f6924d2fc9ae6f3968ba0e414b9fd824661ff322cc5e27dfcad2be5ce6cd30e4

Observation 85990893-f0a0-45e3-b0ac-bfd9e6f7ab99 · outbound

This paper cites Layoutgpt: Compositional visual plan- ning and generation with large language models.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Layoutgpt: Compositional visual plan- ning and generation with large language models

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:05.957221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T13:58:05.135424Z digest=sha256:901dbd9b23a5d578a5bbef1a097b891e3b38df33ec4b0b6f91372bd432890178

Observation 6625bbde-6465-4073-9aa9-78dc70909b7a · outbound

This paper cites ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T13:58:05.138991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:58:05.138991Z digest=sha256:f0f110cf28a84fa928ac1e54c3b16b045b14102ce9929e00f25b430bfa273386

Observation dac461c6-ad6a-48a2-86ce-d515f6dd62ed · outbound

This paper cites Humansd: A native skeleton-guided diffusion model for human image generation.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Humansd: A native skeleton-guided diffusion model for human image generation

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:05.943248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T13:58:05.142871Z digest=sha256:6e42f4e2385e4ee0c9ff29cf99f3c2736f16af8f49d301605ca94762d9233c4b

Observation 2be2634a-9867-456f-97c1-6ceb237ba7e2 · outbound

This paper cites eDiff-I: Text-to-Image Diffusion Models with an Ensemble of Expert Denoisers.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs eDiff-I: Text-to-Image Diffusion Models with an Ensemble of Expert Denoisers

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T13:58:05.146539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:58:05.146539Z digest=sha256:56b75b2a065f86e2221e2b2fb314625ed316d98361e633bb3c6537cec4e48500

Observation 8b6efc69-2ffd-4ae7-bccb-0251a4f3104e · outbound

This paper cites Lafite2: Few-shot Text-to-Image Generation.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Lafite2: Few-shot Text-to-Image Generation

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-08-06T13:58:05.211626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T13:58:05.150622Z digest=sha256:30c8e60ae6991e63fdfee00db13067f8ec666e56d24044c1b5b005358f1686f4

Observation b1d736c3-2c9e-4aed-b529-7dc8e21c151e · outbound

This paper cites Gligen: Open-set grounded text-to-image generation.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Gligen: Open-set grounded text-to-image generation

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:05.929027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T13:58:05.154643Z digest=sha256:4951aff23ba2d01ec705b17ecfd4fc6690742a434996d5277dee85d58825c734

Observation 64e4b6a5-2c47-43dc-998c-bf7fb7cd7e00 · outbound

This paper cites Ranni: Taming text-to-image diffusion for accurate instruction following.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Ranni: Taming text-to-image diffusion for accurate instruction following

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:05.915734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T13:58:05.158515Z digest=sha256:8f6b700eae120c94d759a6d93fc15a8dc5e420e7a6a2dd2da130fcf508b56dca

Observation c05889d0-fa51-4631-a53c-603f36e5e8fb · outbound

This paper cites Smartedit: Ex- ploring complex instruction-based image editing with multi- modal large language models.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Smartedit: Ex- ploring complex instruction-based image editing with multi- modal large language models

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:05.902359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T13:58:05.162408Z digest=sha256:a4cba4478fcd97ec2fa8a6f10b71d0e8ca388c89b9f4a2e9478a082345f651b1

Observation 0fbbccd9-48a4-463e-90e8-709a42ec2a94 · outbound

This paper cites Layoutdiffusion: Controllable diffu- sion model for layout-to-image generation.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Layoutdiffusion: Controllable diffu- sion model for layout-to-image generation

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:05.888090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T13:58:05.166417Z digest=sha256:20d9288eb621ab9500e552e4965ec639b20400cb4cc2ceca0d9ec534d7b91f99

Observation aa617503-f3f4-4e57-a2f1-b75cbf59137a · outbound

This paper cites Reco: Region- controlled text-to-image generation.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Reco: Region- controlled text-to-image generation

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:05.875093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T13:58:05.169969Z digest=sha256:99f2b7f5932aa4d44f588b167c5c1d39fbbb6ccb89e5ee239026ceb11516b6ea

Pith citing papers

Observation 136e6b49-1b99-449a-b459-2878b5d978c5 · inbound

EventOD: Event-Aware OD Flow Generation via LLM-Guided Semantic Modulation cites this paper.

EventOD: Event-Aware OD Flow Generation via LLM-Guided Semantic Modulation LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T09:52:01.768852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:52:01.768852Z digest=sha256:ef3c653b0461c576bbe598730b48890f17e68db98cca275a5c67b24f00afb940