Pith. sign in

Paper Citation Record · LEDGER

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation

As of 13 August 2026, this Paper Citation Record lists 87 of 87 outbound references and 1 inbound Pith citation observation for arXiv:2412.05818.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.05818 v2

Coverage vector

measured 87 of 87 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T20:25:15.110746Z

measured 88 of 88 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:30:27.103280Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T05:30:27.496049Z

Reference resolution

87 of 87 outbound references displayed

  • verified exact0
  • verified fuzzy42
  • unresolved45
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cea0948c-e204-4389-98f3-0ab474309b98 · outbound

This paper cites A general theoretical paradigm to un- derstand learning from human preferences.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation A general theoretical paradigm to un- derstand learning from human preferences

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:14.425670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:14.425670Z digest=sha256:0d74d11c34664cb5d8dcea38ffbb5e663cb6307d971055c7686852f56acbb748

Observation 87972546-4366-430f-99ce-a603d9a9cab7 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:14.432116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:14.432116Z digest=sha256:eac4e986620aca2241b445ecdd9909f7f9649c530d882e086d699f90950a0c9c

Observation 3e2308b4-26cd-4692-a7b1-0ab937bf4918 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Constitutional AI: Harmlessness from AI Feedback

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:14.438022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:14.438022Z digest=sha256:411735c8d3d835154931466cc1baacaa1eb9acce6586141e773427634d222191

Observation 0c51e86d-9037-43a5-b399-b44baeeca17f · outbound

This paper cites Improving image generation with better captions.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Improving image generation with better captions

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:14.444537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:14.444537Z digest=sha256:19bdb21aa7cb536dfb2ae1bc64482ec81250c658757451be04689665865f2527

Observation 9df390e6-4d00-4133-b67d-81db4ee16d9b · outbound

This paper cites Training diffusion models with reinforce- ment learning.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Training diffusion models with reinforce- ment learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:14.450209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:14.450209Z digest=sha256:468902b63bab076074207ef34f74b817fea006848e52397773ef97e4cd96e2fe

Observation e9e42c48-eed0-4469-9c66-9a18b80a00d9 · outbound

This paper cites Attend-and-excite: Attention-based se- mantic guidance for text-to-image diffusion models.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Attend-and-excite: Attention-based se- mantic guidance for text-to-image diffusion models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:14.456093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:14.456093Z digest=sha256:cdc5dd8ef39c54d32db14c4970d53bbd550792bd0cf48b98e76b8d5c98d790d7

Observation b83a1fc6-8c7b-4ed9-bc02-a5b7bcb64cef · outbound

This paper cites Spatialvlm: Endow- ing vision-language models with spatial reasoning capabili- ties.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Spatialvlm: Endow- ing vision-language models with spatial reasoning capabili- ties

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:14.461932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:14.461932Z digest=sha256:d4ddfff5f0f32cdf5eec7f4548d4b9b17af6ffcf73ed185e42f035774d4701ed

Observation e8d922ed-9047-4648-8215-7797aa739e7b · outbound

This paper cites PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:14.467979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:14.467979Z digest=sha256:f5a81e9f0f32a693708ad969dab9a8b58d9e8e73d014b6136baa5ead636328d9

Observation b939c50c-136c-4707-bb3e-2aceab3c850a · outbound

This paper cites Multimodal dialog systems with dual knowledge-enhanced generative pretrained language model.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Multimodal dialog systems with dual knowledge-enhanced generative pretrained language model

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:14.474582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:14.474582Z digest=sha256:5b9944c65d040d0c76df0f53aa0b23937c815176f7e38e640e4e6c9ec667b165

Observation 5c756234-837b-4fc9-ae05-f335eea6c755 · outbound

This paper cites Ultrafeedback: Boosting language mod- els with scaled ai feedback.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Ultrafeedback: Boosting language mod- els with scaled ai feedback

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:14.479995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:14.479995Z digest=sha256:91b562f62ad8c5f519af4b28e358ea9c451058d18e377a7caa3467c0c4d47d5f

Observation e2b9dcb9-7d2e-47b7-aea4-1fbbb6f6809e · outbound

This paper cites How Abilities in Large Language Models are Affected by Supervised Fine-tuning Data Composition.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation How Abilities in Large Language Models are Affected by Supervised Fine-tuning Data Composition

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:14.485181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:14.485181Z digest=sha256:8d04519bc4a71c026b9274f12723b054dc32bf79020767cdcf59a21c7540ad40

Observation 570ae0a7-16f7-49e8-a123-ccc5b30618cc · outbound

This paper cites A Survey on In-context Learning.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation A Survey on In-context Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:14.490817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:14.490817Z digest=sha256:a99a02eb963b38caf82e6aac3f36aa7445c6492d1e15f417ca3629f76a9fb998

Observation 86bce015-8c03-4ab8-bc54-ea9998cb64d2 · outbound

This paper cites Dreamllm: Synergistic multimodal com- prehension and creation.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Dreamllm: Synergistic multimodal com- prehension and creation

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:25:17.320857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T20:25:14.496137Z digest=sha256:be36486412206769a344cdcbd3bdd5724d29ccec0f53459e02ee167cec957726

Observation 0e54758c-4964-4bd5-a531-493f70c7c7a9 · outbound

This paper cites The Llama 3 Herd of Models.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation The Llama 3 Herd of Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:14.501287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:14.501287Z digest=sha256:36ef9adb2e2c3058752c5199529f2bdef3756563a80e5acd2a3c8913d92f2b40

Observation 9b6f0ada-8bab-4cb2-a5b4-672a6c967600 · outbound

This paper cites Re- inforcement learning for fine-tuning text-to-image diffusion models.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Re- inforcement learning for fine-tuning text-to-image diffusion models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:25:17.301343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T20:25:14.506597Z digest=sha256:9a79502c0c8863bed1a6221bdba820da253390e5eb81f33e227c627fc1eb99b0

Observation 8ac07f17-bf7b-44c3-b003-1f728254dd55 · outbound

This paper cites Training- free structured diffusion guidance for compositional text-to- image synthesis.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Training- free structured diffusion guidance for compositional text-to- image synthesis

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:25:17.281347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T20:25:14.511744Z digest=sha256:cf79ff1e98a733637307040e999b6b5e210f72b1d262c7cfcbbaaf9b4de9f7dd

Observation 32ed574c-a15e-4127-b44e-fb49d7ed90a8 · outbound

This paper cites LayoutGPT: Compositional Visual Planning and Generation with Large Language Models.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation LayoutGPT: Compositional Visual Planning and Generation with Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:14.516822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:14.516822Z digest=sha256:29d8d8d04a39a3fcd03ba5b2071e918a7186bad28f4f7af7abc58a703986a2f9

Observation c1380e6d-aa3c-4550-8c78-1ec7694fda77 · outbound

This paper cites Dropout as a bayesian approximation: Representing model uncertainty in deep learning.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Dropout as a bayesian approximation: Representing model uncertainty in deep learning

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:25:17.262716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T20:25:14.522862Z digest=sha256:948589f0d22e5dd55db88ff18bf325c8ebbadddeabad55585de8d15763e2d7d1

Observation 76b2b88d-6473-4fb6-9d22-063fcbf954f3 · outbound

This paper cites Making llama see and draw with seed tokenizer.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Making llama see and draw with seed tokenizer

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:25:17.242557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T20:25:14.529239Z digest=sha256:6a16be18b963929777be61eea31b659cc17ca7344358fae9ab3999188c1b20fe

Observation 43f55e14-52cb-43d9-87de-c429d821c81b · outbound

This paper cites SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:14.534962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:14.534962Z digest=sha256:d9a7aa813d293937d63d84cf4646a3883db31585e60003bfa654b56112fbef16

Observation 02dd1084-32f7-4518-8fcb-7253bb27af22 · outbound

This paper cites Pela: Learning parameter-efficient models with low-rank ap- proximation.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Pela: Learning parameter-efficient models with low-rank ap- proximation

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:25:17.222788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T20:25:14.541064Z digest=sha256:cfaa8e26c6d2b1fc076bfe38a54cbe4786fe4c1c4ca26d4ca8940580632c70de

Observation 37d22593-872a-4da7-a0e7-fa0b0031ee45 · outbound

This paper cites Hearst, Susan T Dumais, Edgar Osuna, John Platt, and Bernhard Scholkopf.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Hearst, Susan T Dumais, Edgar Osuna, John Platt, and Bernhard Scholkopf

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:25:17.202807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T20:25:14.546799Z digest=sha256:0b8031fb654c7f0b35d2a2e902fd8bd3f834e1774a03c24f27822a66567bd76e

Observation b6104bc7-3929-4b4f-b642-bbec0342d082 · outbound

This paper cites Clipscore: A reference-free evaluation met- ric for image captioning.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Clipscore: A reference-free evaluation met- ric for image captioning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:14.552087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:14.552087Z digest=sha256:7e73392b1d68541c3e4906ed43b3b11389746e6fd38594469797049569415ddd

Observation ef7b8ac8-23ab-45b7-8521-232b896f4edf · outbound

This paper cites spaCy: Industrial-strength Natural Lan- guage Processing in Python.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation spaCy: Industrial-strength Natural Lan- guage Processing in Python

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:25:17.162098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T20:25:14.557758Z digest=sha256:974d7c6ec7aef77e4a813d996687c815c6f3908abaf9e86728f2dd4d3087b45e

Observation ca859cb2-f157-4de5-a242-4b22b7b76119 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation LoRA: Low-Rank Adaptation of Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:14.563093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:14.563093Z digest=sha256:1720589916df1c8e0796133276a22f3f23d0ba74ab88bbecd7d20898c746b5a2

Observation 8ad3bfda-995a-42fd-8007-f4b48db7f005 · outbound

This paper cites ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:14.569982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:14.569982Z digest=sha256:456353b17b54abb32483cacf68d69b54235ddd03afb9e02703fa8b94036d7258

Observation 8c13b01b-5833-48ac-aca2-b64f90424c1c · outbound

This paper cites Tifa: Accu- rate and interpretable text-to-image faithfulness evaluation with question answering.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Tifa: Accu- rate and interpretable text-to-image faithfulness evaluation with question answering

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:25:17.140222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T20:25:14.575294Z digest=sha256:0d459418b062e0159d2dc1afd3e9bd4f2ddef5031aef8451d3a877e34013434f

Observation ba8b844a-14c1-4197-9215-addeb8761065 · outbound

This paper cites T2i-compbench: A comprehensive bench- mark for open-world compositional text-to-image genera- tion.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation T2i-compbench: A comprehensive bench- mark for open-world compositional text-to-image genera- tion

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:25:17.109749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T20:25:14.580288Z digest=sha256:9371af66eba9f74e6ee3855632ebe7c70b83fee623eea0666009f3a1bd8e8f25

Observation 89755587-7531-4663-8ead-3d3171a34ae1 · outbound

This paper cites Sequence tutor: Conservative fine-tuning of sequence generation models with kl-control.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Sequence tutor: Conservative fine-tuning of sequence generation models with kl-control

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:25:17.085843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T20:25:14.584846Z digest=sha256:e844dd4c68da860a7bc3299bbd5a740b76421ad92bae05a902e774a5279f70c7

Observation be89228f-fb0b-45fc-b0d1-a2e725cbf520 · outbound

This paper cites Human-centric Dialog Training via Offline Reinforcement Learning.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Human-centric Dialog Training via Offline Reinforcement Learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:14.589831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:14.589831Z digest=sha256:ffcfcc13c4eb884a06b16476dfcc7c2c65be516acd6d3da0bd6599ba59fcf82a

Observation c8dac8bf-9d68-40f7-9015-351ea247033e · outbound

This paper cites Pick-a-pic: An open dataset of user preferences for text-to-image generation.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Pick-a-pic: An open dataset of user preferences for text-to-image generation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:25:17.059195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T20:25:14.596060Z digest=sha256:d8753104d58c12c558a4139fdfe5018c7e11d92e65b71448b665cabc1126ed0e

Observation 408d493a-9ba9-43af-bed1-d49fb29ef99a · outbound

This paper cites Rlaif vs.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Rlaif vs

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:25:17.039698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T20:25:14.601314Z digest=sha256:da6a285a04f34bba9a8f12fefd21a7e399c15b4f2f6153319ee90ab2c34eedcd

Observation 2bfce24b-305f-4c3b-914c-5e927ee7b661 · outbound

This paper cites Aligning Text-to-Image Models using Human Feedback.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Aligning Text-to-Image Models using Human Feedback

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:14.606432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:14.606432Z digest=sha256:2357e0829ccedf73017c0e8d1f93fb6374e3cad03b2213620e52a6a2c3981ec4

Observation 35d94c09-4032-4761-942c-a42bd99ed973 · outbound

This paper cites Invariant grounding for video question answering.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Invariant grounding for video question answering

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:25:17.017086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T20:25:14.760767Z digest=sha256:609570949c5257f764a40884c6b4b9396189bfbf8f1730037135ed6590a35e71

Observation 5228c811-4ab1-4013-a631-82f86b8ad6aa · outbound

This paper cites Transformer-empowered invariant grounding for video question answering.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Transformer-empowered invariant grounding for video question answering

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:25:16.997434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T20:25:14.766878Z digest=sha256:83ade2d464605c62763743b1eaaccb0c71627103a9e1404637bc9de497dd9476

Observation 70999c1c-a7c6-4b16-b553-49a1a8147b13 · outbound

This paper cites Attribute-driven disentangled representation learning for multimodal recommendation.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Attribute-driven disentangled representation learning for multimodal recommendation

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:25:16.975428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T20:25:14.773886Z digest=sha256:5d073abaea386a500f3a549b5a606bcf2c7b7efe31e41c58f65a42b6759dbe3a

Observation e06eca76-f3a3-42bc-ad0c-6911f2a7fef2 · outbound

This paper cites LLM-grounded Diffusion: Enhancing Prompt Understanding of Text-to-Image Diffusion Models with Large Language Models.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation LLM-grounded Diffusion: Enhancing Prompt Understanding of Text-to-Image Diffusion Models with Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:14.780369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:14.780369Z digest=sha256:0b8d2c0eba5dad24c760fae18437aca3acce7ad30720ea50c73793506f0bff4f

Observation 320e6226-04a0-46c8-b2f5-c38b42ce22c9 · outbound

This paper cites Llm-grounded video diffusion models.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Llm-grounded video diffusion models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:25:16.952412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T20:25:14.786456Z digest=sha256:6fd343c36631f9f88a2051fca0e452d01d65e3f7fbf048d3b5f685c9c126c9a3

Observation 0e70dd17-b9c3-4039-93f3-576100664f06 · outbound

This paper cites Visual Instruction Tuning.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Visual Instruction Tuning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:14.791940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:14.791940Z digest=sha256:1df231877b905eeab4c3095f028163862a59f1f63cdc8e343f8068a02c637cdc

Observation dd9190e3-a68e-4306-8cb9-74b5bb6d8b48 · outbound

This paper cites Ipo: Interior-point policy optimization under constraints.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Ipo: Interior-point policy optimization under constraints

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:25:16.931447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T20:25:14.798321Z digest=sha256:aaa5e55ac144461e1a29c61a253b91c228bb11117239c945795753ab988a1e0d

Observation ae637458-13c3-4aad-809b-7d8353dda7f6 · outbound

This paper cites Cheap and quick: Efficient vision- language instruction tuning for large language models.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Cheap and quick: Efficient vision- language instruction tuning for large language models

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:25:16.906814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T20:25:14.805016Z digest=sha256:d07cce66fcf18d0e97385dcd21252f39957b0eccc3fa0367ac74f8807e437591

Observation de3e33b1-a0d0-4d4b-a91b-a0e450e21476 · outbound

This paper cites Compositional chain-of-thought prompting for large multimodal models.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Compositional chain-of-thought prompting for large multimodal models

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:25:16.880847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T20:25:14.810931Z digest=sha256:835ebeaff744aa435ec5fe11f0e8aa2d9ff1a229f2bb23a21529db73636c9599

Observation 7c854fc5-b474-4022-902e-04a7e97c823c · outbound

This paper cites Training language models to follow instructions with human feedback.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Training language models to follow instructions with human feedback

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:25:16.857710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T20:25:14.816283Z digest=sha256:9f2be26fd68c7c925df0287026aa3df6a336eccdd1d7ee91e1dbf7a1be1cf12b

Observation c90f62e0-ffb5-49f2-ad00-1f67506be9ac · outbound

This paper cites Lan- guage model self-improvement by reinforcement learning contemplation.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Lan- guage model self-improvement by reinforcement learning contemplation

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:25:16.830206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T20:25:14.822176Z digest=sha256:ea0d3be299f8367b98a14f2f9a4a94ebca42b4da142eed238625e1a9f26c0f2f

Observation a10cdbb2-3f62-43ad-b1a6-358bb7a8d34e · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:14.828251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:14.828251Z digest=sha256:e7be60af6a0a7538fa389d6f85d4af9500852c1c4d7bf5c0ae72e3dff61a992d

Observation 8b8f96f5-f778-4f94-a4aa-efc8f256dfda · outbound

This paper cites Diffusiongpt: Llm-driven text-to-image generation system.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Diffusiongpt: Llm-driven text-to-image generation system

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:14.835904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:14.835904Z digest=sha256:282aeddaf327ba8503d855334f851d8f9fce1604508e878c323affbac2f3dd21

Observation 1c91c364-838c-42e7-9248-30a6077d8814 · outbound

This paper cites 3d-immc: Incomplete multi-modal 3d shape clustering via cross mapping and dual adaptive fu- sion.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation 3d-immc: Incomplete multi-modal 3d shape clustering via cross mapping and dual adaptive fu- sion

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:25:16.801841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T20:25:14.843001Z digest=sha256:08a3fdf8652133f8cc8d4d11e343487fab8ccc5a97e207a22fd4bbdbfd83e993

Observation cc049737-d86c-4483-beab-17e02817bdc0 · outbound

This paper cites Dynamic modality interaction modeling for image-text retrieval.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Dynamic modality interaction modeling for image-text retrieval

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:25:16.780318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T20:25:14.848502Z digest=sha256:490ce908ae024ef8fd355f3bd9f15461c6b0104a6613ad7f2186b4921725a67f

Observation d33920c6-87e8-4169-a282-5973f0c44eae · outbound

This paper cites Layoutllm-t2i: Eliciting layout guidance from llm for text-to-image generation.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Layoutllm-t2i: Eliciting layout guidance from llm for text-to-image generation

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:25:16.755881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T20:25:14.854363Z digest=sha256:0238c19b215a9555a01d5d734418f29e428f6b2b760d592edf6ae2492f1917a0

Observation 887a7e9c-0f78-47e2-85d2-337131800f72 · outbound

This paper cites TIGeR: Unifying Text-to-Image Generation and Retrieval with Large Multimodal Models.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation TIGeR: Unifying Text-to-Image Generation and Retrieval with Large Multimodal Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:14.862417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:14.862417Z digest=sha256:e0b054204a24303e3cdacca25aa07bed81232c309a65cac58c1299ab7cbeb3d8

Observation 84421468-b731-4c91-a58d-75318e869631 · outbound

This paper cites Discriminative probing and tuning for text-to-image generation.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Discriminative probing and tuning for text-to-image generation

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:25:16.716643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T20:25:14.869638Z digest=sha256:baced394e4c0165b0eddc8d3d5c0baa09e0bffa8813c5528bbfe83f3c084a859

Observation 9eecad53-b581-4051-a8a3-f31dffde9fcd · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Learning transferable visual models from natural language supervi- sion

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:14.876127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:14.876127Z digest=sha256:3763505d5d31f1435ad538ffc9dd0fff26bd6b5ab0ecfd6859aab1d8806b77b1

Observation c18a1184-669e-414b-a192-a2b12e67b21d · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Direct preference optimization: Your language model is secretly a reward model

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:25:16.676752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T20:25:14.881529Z digest=sha256:a35c3975c082a237bc14db731e26c412b6532204f226e016acf0cec7a870050a

Observation 35a3e8a0-81de-42ef-8dbe-c9c8a2588edd · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:14.888536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:14.888536Z digest=sha256:e09d355966fcde51821282d6ae04674e4e5e8bc7f34aa43df3976469cca9e6c6

Observation 0f48b242-ce8c-4d77-b59a-fc39a8394717 · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:14.895419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:14.895419Z digest=sha256:050007cfc7d8e380275db6c728d1854e056fe2d6555600895b479c0b8b064d75

Observation a9bda355-e586-4b4a-81a4-b96b7f080455 · outbound

This paper cites Can generative multimodal models count to ten? In Proceedings of the Annual Meeting of the Cognitive Science Society, 2024.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Can generative multimodal models count to ten? In Proceedings of the Annual Meeting of the Cognitive Science Society, 2024

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:25:16.641581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T20:25:14.901138Z digest=sha256:bf2698603d8272d570fd29c2ad3f9ccc8d0bd6fc334f52210852a3f4789ba20b

Observation 539b5c37-2b5b-44a7-ae54-0c49d9ba829e · outbound

This paper cites Photorealistic text-to-image diffusion models with deep language understanding.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Photorealistic text-to-image diffusion models with deep language understanding

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:14.908071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:14.908071Z digest=sha256:96b86241ce4be432613664718a08972baef97937060078717b2f89999774a42d

Observation f7c48f5c-b5ff-4968-b58b-94b54d96ec4b · outbound

This paper cites Proximal Policy Optimization Algorithms.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Proximal Policy Optimization Algorithms

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:14.916590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:14.916590Z digest=sha256:5e1d2f558ec57a2638404c984d4e0a5e2997fb253d8b68de0d6c3d8806202e52

Observation a4fc7d2c-1932-49f0-a8fc-d4745571505c · outbound

This paper cites Llava-prumerge: Adaptive token reduction for efficient large multimodal models.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Llava-prumerge: Adaptive token reduction for efficient large multimodal models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:14.925715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:14.925715Z digest=sha256:faede739fb4915a997705fbdb039530607cd7ac3272322028f7a5ed696290bc1

Observation 2e42345c-659a-491c-86d0-5f8356bd7620 · outbound

This paper cites Kernel methods for pattern analysis.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Kernel methods for pattern analysis

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:25:16.602990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T20:25:14.936104Z digest=sha256:093f941b30c5915dec6dab28d5a9daa4499c833eb0ffdf5109bc4dec7eb62f35

Observation fb1c32aa-617f-4a1b-bdb7-e33bea58efda · outbound

This paper cites Learning to summarize with human feed- back.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Learning to summarize with human feed- back

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:25:16.574025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T20:25:14.942597Z digest=sha256:d2b06663db9a6dff9c3174c401e5e1393e95c5031d92d1a052ca826390bd4217

Observation 6db52ec4-4bc8-4b8d-9d1b-80ca89fab095 · outbound

This paper cites Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:14.948321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:14.948321Z digest=sha256:15508acff31800a4ad5d6106711a76de6adaf49e301a916d165a5017d624ce20

Observation 4a2f98ab-2a47-461a-af6e-df5fc9dad34f · outbound

This paper cites Emu: Generative pretraining in multimodality.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Emu: Generative pretraining in multimodality

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:25:16.550116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T20:25:14.954167Z digest=sha256:695c51d56f9513d35255008f180edc878a20a2e42cabae081791cb27b749ced6

Observation 4a06d198-7078-4c46-90fa-f3e6db573399 · outbound

This paper cites Generative multimodal mod- els are in-context learners.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Generative multimodal mod- els are in-context learners

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:25:16.525803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T20:25:14.960207Z digest=sha256:3c3de95d493805e93b7178795ee62b43472aea1314659a4b6b49bd8c9920e043

Observation da1636db-9755-48d1-b737-1fc8490a3eba · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation LLaMA: Open and Efficient Foundation Language Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:14.965603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:14.965603Z digest=sha256:e081d8f7166f202b3393e96aa65055c9e2cd1bc310e0870af6c6f4b754603668

Observation 27cde872-daf7-4a47-8a95-a8ecee165677 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:14.972615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:14.972615Z digest=sha256:50c9a9c6a4984b589190924802e20870f3c516c676b13a7eb103cda91c89c571

Observation e784b1d0-b4e5-4819-a765-59301aadd65e · outbound

This paper cites Diffusion model align- ment using direct preference optimization.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Diffusion model align- ment using direct preference optimization

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:25:16.476027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T20:25:14.977961Z digest=sha256:f8f10c8d9be2503a09064f106200741d33454625ed508258304481e425467cec

Observation 7f86175a-2f06-488b-af03-f0d5a78f6480 · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Emu3: Next-Token Prediction is All You Need

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:14.984003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:14.984003Z digest=sha256:4334c633ddafd4921ec31c4b8976497a7e440a1c3d8d2d9a970320a5aa9ce5d3

Observation 386c3901-3ef0-45c6-8a8e-f23fb05974a4 · outbound

This paper cites RL-VLM-F: Reinforcement Learning from Vision Language Foundation Model Feedback.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation RL-VLM-F: Reinforcement Learning from Vision Language Foundation Model Feedback

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:14.989517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:14.989517Z digest=sha256:4b4a1a328869c802e15316095e04e054e48bda58ff095ae71b8c0f81891e6871

Observation 23e8be3f-bce4-4c19-8a68-135c0d0f84dd · outbound

This paper cites Divide and Conquer: Language Models can Plan and Self-Correct for Compositional Text-to-Image Generation.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Divide and Conquer: Language Models can Plan and Self-Correct for Compositional Text-to-Image Generation

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:14.995830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:14.995830Z digest=sha256:f2d3f380d72a3fb845b8c34e944161094df1e61eeeac86ae4776103958511733

Observation f4379731-fbbd-4494-8e0f-4dcbafa00c41 · outbound

This paper cites Comprehensive linguistic-visual composition network for image retrieval.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Comprehensive linguistic-visual composition network for image retrieval

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:25:16.454831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T20:25:15.001428Z digest=sha256:4823582b7452e0ab52daf9cfe53960b6ede0e70e475d27a331c4838e4a5b5689

Observation 8156953c-c6eb-45bc-a587-202660fd0e40 · outbound

This paper cites Next-gpt: Any-to-any multimodal llm.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Next-gpt: Any-to-any multimodal llm

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:25:16.430022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T20:25:15.007045Z digest=sha256:3c041ce6a1c8059b24d6bd64b2923ad9dfbc4f815909218188a064e923bc297b

Observation b6b00d45-594d-4e44-aca3-e58920d56492 · outbound

This paper cites Human Preference Score: Better Aligning Text-to-Image Models with Human Preference.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Human Preference Score: Better Aligning Text-to-Image Models with Human Preference

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:15.012326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:15.012326Z digest=sha256:53843d1ee411e845492774898d316a63e7ca44387c2ad56807fd386b53c45f09

Observation 899eeccf-28a4-4e52-9f08-01e9cb958c8c · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:15.018504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:15.018504Z digest=sha256:c958c308652f6ad3ac7041db15bdacc70cd52470c46deec8f048b5589af93cfe

Observation 18dcb2bf-e914-4e33-a6a3-57e8434978b2 · outbound

This paper cites Mastering text-to-image dif- fusion: Recaptioning, planning, and generating with multi- modal llms.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Mastering text-to-image dif- fusion: Recaptioning, planning, and generating with multi- modal llms

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:25:16.408048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T20:25:15.025317Z digest=sha256:c5e60ba218cc2b5d52b0673186901872286f6af3e045ab71de0a22c7feaf11c3

Observation 187bde08-e143-4317-ac54-438a1d50ef05 · outbound

This paper cites Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:15.031298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:15.031298Z digest=sha256:ce24676ccc4c7505750acaa1dc19be14fabf05753a05002b0cd98ae3c101ea99

Observation 81c91a9c-c72f-4739-b8da-d163f2aba32e · outbound

This paper cites Rlaif-v: Aligning mllms through open-source ai feedback for super gpt-4v trustworthiness.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Rlaif-v: Aligning mllms through open-source ai feedback for super gpt-4v trustworthiness

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:15.036973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:15.036973Z digest=sha256:eadc78d94dfe8f466ff69b395b133ce6c5ef699c59bbdcc53297a724f396341b

Observation 01a12e6a-dc3f-46e9-aa50-b0f1a6c0d1ca · outbound

This paper cites Self-Rewarding Language Models.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Self-Rewarding Language Models

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:15.043167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:15.043167Z digest=sha256:45600d9fd77883cf8d7c7607a8bde8cb12bd546167f65225b89d54f12028ad8f

Observation ce8d0df0-8e3a-43b9-ace0-8fc071093344 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:15.050001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:15.050001Z digest=sha256:88f84a5cab1271c36ebbc302cdd2bae270232b80cbed9e4dde34064413227f92

Observation 32ef74f3-9fd0-437e-b0d4-9ec92225f51a · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Fine-Tuning Language Models from Human Preferences

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:15.058528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:15.058528Z digest=sha256:89b422bcfc50e496f8f183c93c64df67f160e107f79412494e5d9fb78cf59a32

Observation 2ffb9301-f40b-4731-beeb-4e9a3fc017c8 · outbound

This paper cites Attributes such as color, shape, texture, and 2D/3D spatial relations are also incorpo- rated.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Attributes such as color, shape, texture, and 2D/3D spatial relations are also incorpo- rated

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:25:16.374352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T20:25:15.064281Z digest=sha256:2c34e4a349662ab53365043fa91236f690f9b16cf3fa13dbad9651f98637629b

Observation be913525-b847-4757-8445-144e73a0bfb1 · outbound

This paper cites These atomic concepts are then transformed into simple yes-or-no questions.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation These atomic concepts are then transformed into simple yes-or-no questions

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:25:16.347768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T20:25:15.072767Z digest=sha256:c8af3b1973b1307fd8cdefb10f3b48b91bf23dd70d14d78c5ef6c448c3d74716

Observation b76f1ddb-96aa-4c1f-9a24-66df37816450 · outbound

This paper cites log σ − ¯σβ 2 LX i=1 ∥hw i − µi∥2 2 − βC + β 2¯σ LX i=1 ∥hw i − µref i ∥2 2 + βC − β 2¯σ LX i=1 ∥hl i − µi∥2 2 − βC + β 2¯σ LX i=1 ∥hl i − µref i ∥2 2 + βC ! # = −E(x,zw ,zl)∼D.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation log σ − ¯σβ 2 LX i=1 ∥hw i − µi∥2 2 − βC + β 2¯σ LX i=1 ∥hw i − µref i ∥2 2 + βC − β 2¯σ LX i=1 ∥hl i − µi∥2 2 − βC + β 2¯σ LX i=1 ∥hl i − µref i ∥2 2 + βC ! # = −E(x,zw ,zl)∼D

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:25:16.323140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T20:25:15.080018Z digest=sha256:2726d6935a4e1feea06b1da9c2d6992202da6c85b4d8e5ba9c0146b5abda7a0a

Observation ef46e847-313a-4933-8d7b-ac3157509c1f · outbound

This paper cites For SEED-LLaMA, the LLM backbone of DreamLLM is optimized for 1k steps, with a learning rate of 5 × 10−5, 100 warm-up steps, and a cosine learning rate scheduler.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation For SEED-LLaMA, the LLM backbone of DreamLLM is optimized for 1k steps, with a learning rate of 5 × 10−5, 100 warm-up steps, and a cosine learning rate scheduler

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:25:16.294856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T20:25:15.086924Z digest=sha256:05d58bab38846bb37ba7528eb0e07e1ea5a834d773b15605be6828fbb3852747

Observation b059f2be-e759-4159-b1db-cc5fd264105a · outbound

This paper cites an unresolved cited work.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Unresolved cited work

Reference 85

Resolution
unresolved
raw_fallback, observed 2026-08-11T20:25:16.274373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T20:25:15.092981Z digest=sha256:49295daa5a34c5d20d0663070ad0ae536d2b6089825113ff2fc4b581d5e43ad7

Observation a0e4efb8-59b7-445c-b30c-c3f8d2abb4f5 · outbound

This paper cites First Half.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation First Half

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:25:16.254438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T20:25:15.100341Z digest=sha256:ef541ab24284e81587d496d1e8118cf51d898aa694642a05eedc5b82119c1cac

Observation 625205fb-3c24-4363-8277-4bc54bb41e6f · outbound

This paper cites 14 - 24” means the rejected data points are sampled from rank-14 to rank-24 which is a hard range, while “20 - 30.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation 14 - 24” means the rejected data points are sampled from rank-14 to rank-24 which is a hard range, while “20 - 30

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:25:16.235360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T20:25:15.110746Z digest=sha256:c98536a501c45e23fb0d6cbef7a653dfaca523526b9c3752ae0055b56281f15b

Pith citing papers

Observation a6bdf839-3534-404d-a26d-b0076853d7a6 · inbound

SUDER: Self-Improving Unified Large Multimodal Models for Understanding and Generation with Dual Self-Rewards cites this paper.

SUDER: Self-Improving Unified Large Multimodal Models for Understanding and Generation with Dual Self-Rewards SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:30:27.503745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T05:30:27.103280Z digest=sha256:8c47ebd48f1ea744891d9ee7b9f450640cfb6d36dfa4323c0edc8b4647ff4000