Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T20:25:15.110746Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 87 of 87 outbound references and 1 inbound Pith citation observation for arXiv:2412.05818.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T20:25:15.110746Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:30:27.103280Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-07T05:30:27.496049Z
87 of 87 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation cea0948c-e204-4389-98f3-0ab474309b98 · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation A general theoretical paradigm to un- derstand learning from human preferences
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87972546-4366-430f-99ce-a603d9a9cab7 · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e2308b4-26cd-4692-a7b1-0ab937bf4918 · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Constitutional AI: Harmlessness from AI Feedback
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c51e86d-9037-43a5-b399-b44baeeca17f · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Improving image generation with better captions
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9df390e6-4d00-4133-b67d-81db4ee16d9b · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Training diffusion models with reinforce- ment learning
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9e42c48-eed0-4469-9c66-9a18b80a00d9 · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Attend-and-excite: Attention-based se- mantic guidance for text-to-image diffusion models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b83a1fc6-8c7b-4ed9-bc02-a5b7bcb64cef · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Spatialvlm: Endow- ing vision-language models with spatial reasoning capabili- ties
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8d922ed-9047-4648-8215-7797aa739e7b · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b939c50c-136c-4707-bb3e-2aceab3c850a · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Multimodal dialog systems with dual knowledge-enhanced generative pretrained language model
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c756234-837b-4fc9-ae05-f335eea6c755 · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Ultrafeedback: Boosting language mod- els with scaled ai feedback
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2b9dcb9-7d2e-47b7-aea4-1fbbb6f6809e · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation How Abilities in Large Language Models are Affected by Supervised Fine-tuning Data Composition
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 570ae0a7-16f7-49e8-a123-ccc5b30618cc · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation A Survey on In-context Learning
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86bce015-8c03-4ab8-bc54-ea9998cb64d2 · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Dreamllm: Synergistic multimodal com- prehension and creation
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 0e54758c-4964-4bd5-a531-493f70c7c7a9 · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation The Llama 3 Herd of Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b6f0ada-8bab-4cb2-a5b4-672a6c967600 · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Re- inforcement learning for fine-tuning text-to-image diffusion models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 8ac07f17-bf7b-44c3-b003-1f728254dd55 · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Training- free structured diffusion guidance for compositional text-to- image synthesis
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 32ed574c-a15e-4127-b44e-fb49d7ed90a8 · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation LayoutGPT: Compositional Visual Planning and Generation with Large Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1380e6d-aa3c-4550-8c78-1ec7694fda77 · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Dropout as a bayesian approximation: Representing model uncertainty in deep learning
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 76b2b88d-6473-4fb6-9d22-063fcbf954f3 · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Making llama see and draw with seed tokenizer
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 43f55e14-52cb-43d9-87de-c429d821c81b · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02dd1084-32f7-4518-8fcb-7253bb27af22 · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Pela: Learning parameter-efficient models with low-rank ap- proximation
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 37d22593-872a-4da7-a0e7-fa0b0031ee45 · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Hearst, Susan T Dumais, Edgar Osuna, John Platt, and Bernhard Scholkopf
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation b6104bc7-3929-4b4f-b642-bbec0342d082 · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Clipscore: A reference-free evaluation met- ric for image captioning
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef7b8ac8-23ab-45b7-8521-232b896f4edf · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation spaCy: Industrial-strength Natural Lan- guage Processing in Python
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ca859cb2-f157-4de5-a242-4b22b7b76119 · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation LoRA: Low-Rank Adaptation of Large Language Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ad3bfda-995a-42fd-8007-f4b48db7f005 · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c13b01b-5833-48ac-aca2-b64f90424c1c · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Tifa: Accu- rate and interpretable text-to-image faithfulness evaluation with question answering
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ba8b844a-14c1-4197-9215-addeb8761065 · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation T2i-compbench: A comprehensive bench- mark for open-world compositional text-to-image genera- tion
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 89755587-7531-4663-8ead-3d3171a34ae1 · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Sequence tutor: Conservative fine-tuning of sequence generation models with kl-control
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation be89228f-fb0b-45fc-b0d1-a2e725cbf520 · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Human-centric Dialog Training via Offline Reinforcement Learning
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8dac8bf-9d68-40f7-9015-351ea247033e · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Pick-a-pic: An open dataset of user preferences for text-to-image generation
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 408d493a-9ba9-43af-bed1-d49fb29ef99a · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Rlaif vs
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 2bfce24b-305f-4c3b-914c-5e927ee7b661 · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Aligning Text-to-Image Models using Human Feedback
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35d94c09-4032-4761-942c-a42bd99ed973 · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Invariant grounding for video question answering
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 5228c811-4ab1-4013-a631-82f86b8ad6aa · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Transformer-empowered invariant grounding for video question answering
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 70999c1c-a7c6-4b16-b553-49a1a8147b13 · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Attribute-driven disentangled representation learning for multimodal recommendation
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation e06eca76-f3a3-42bc-ad0c-6911f2a7fef2 · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation LLM-grounded Diffusion: Enhancing Prompt Understanding of Text-to-Image Diffusion Models with Large Language Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 320e6226-04a0-46c8-b2f5-c38b42ce22c9 · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Llm-grounded video diffusion models
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 0e70dd17-b9c3-4039-93f3-576100664f06 · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Visual Instruction Tuning
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd9190e3-a68e-4306-8cb9-74b5bb6d8b48 · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Ipo: Interior-point policy optimization under constraints
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ae637458-13c3-4aad-809b-7d8353dda7f6 · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Cheap and quick: Efficient vision- language instruction tuning for large language models
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation de3e33b1-a0d0-4d4b-a91b-a0e450e21476 · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Compositional chain-of-thought prompting for large multimodal models
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 7c854fc5-b474-4022-902e-04a7e97c823c · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Training language models to follow instructions with human feedback
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation c90f62e0-ffb5-49f2-ad00-1f67506be9ac · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Lan- guage model self-improvement by reinforcement learning contemplation
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a10cdbb2-3f62-43ad-b1a6-358bb7a8d34e · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b8f96f5-f778-4f94-a4aa-efc8f256dfda · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Diffusiongpt: Llm-driven text-to-image generation system
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c91c364-838c-42e7-9248-30a6077d8814 · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation 3d-immc: Incomplete multi-modal 3d shape clustering via cross mapping and dual adaptive fu- sion
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation cc049737-d86c-4483-beab-17e02817bdc0 · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Dynamic modality interaction modeling for image-text retrieval
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation d33920c6-87e8-4169-a282-5973f0c44eae · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Layoutllm-t2i: Eliciting layout guidance from llm for text-to-image generation
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 887a7e9c-0f78-47e2-85d2-337131800f72 · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation TIGeR: Unifying Text-to-Image Generation and Retrieval with Large Multimodal Models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84421468-b731-4c91-a58d-75318e869631 · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Discriminative probing and tuning for text-to-image generation
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 9eecad53-b581-4051-a8a3-f31dffde9fcd · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Learning transferable visual models from natural language supervi- sion
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c18a1184-669e-414b-a192-a2b12e67b21d · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Direct preference optimization: Your language model is secretly a reward model
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 35a3e8a0-81de-42ef-8dbe-c9c8a2588edd · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Exploring the limits of transfer learning with a unified text-to-text transformer
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f48b242-ce8c-4d77-b59a-fc39a8394717 · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Hierarchical Text-Conditional Image Generation with CLIP Latents
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9bda355-e586-4b4a-81a4-b96b7f080455 · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Can generative multimodal models count to ten? In Proceedings of the Annual Meeting of the Cognitive Science Society, 2024
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 539b5c37-2b5b-44a7-ae54-0c49d9ba829e · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Photorealistic text-to-image diffusion models with deep language understanding
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7c48f5c-b5ff-4968-b58b-94b54d96ec4b · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Proximal Policy Optimization Algorithms
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4fc7d2c-1932-49f0-a8fc-d4745571505c · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Llava-prumerge: Adaptive token reduction for efficient large multimodal models
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e42345c-659a-491c-86d0-5f8356bd7620 · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Kernel methods for pattern analysis
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation fb1c32aa-617f-4a1b-bdb7-e33bea58efda · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Learning to summarize with human feed- back
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 6db52ec4-4bc8-4b8d-9d1b-80ca89fab095 · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a2f98ab-2a47-461a-af6e-df5fc9dad34f · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Emu: Generative pretraining in multimodality
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 4a06d198-7078-4c46-90fa-f3e6db573399 · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Generative multimodal mod- els are in-context learners
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation da1636db-9755-48d1-b737-1fc8490a3eba · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation LLaMA: Open and Efficient Foundation Language Models
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27cde872-daf7-4a47-8a95-a8ecee165677 · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e784b1d0-b4e5-4819-a765-59301aadd65e · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Diffusion model align- ment using direct preference optimization
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 7f86175a-2f06-488b-af03-f0d5a78f6480 · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Emu3: Next-Token Prediction is All You Need
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 386c3901-3ef0-45c6-8a8e-f23fb05974a4 · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation RL-VLM-F: Reinforcement Learning from Vision Language Foundation Model Feedback
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23e8be3f-bce4-4c19-8a68-135c0d0f84dd · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Divide and Conquer: Language Models can Plan and Self-Correct for Compositional Text-to-Image Generation
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4379731-fbbd-4494-8e0f-4dcbafa00c41 · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Comprehensive linguistic-visual composition network for image retrieval
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 8156953c-c6eb-45bc-a587-202660fd0e40 · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Next-gpt: Any-to-any multimodal llm
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation b6b00d45-594d-4e44-aca3-e58920d56492 · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Human Preference Score: Better Aligning Text-to-Image Models with Human Preference
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 899eeccf-28a4-4e52-9f08-01e9cb958c8c · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Show-o: One Single Transformer to Unify Multimodal Understanding and Generation
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18dcb2bf-e914-4e33-a6a3-57e8434978b2 · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Mastering text-to-image dif- fusion: Recaptioning, planning, and generating with multi- modal llms
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 187bde08-e143-4317-ac54-438a1d50ef05 · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81c91a9c-c72f-4739-b8da-d163f2aba32e · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Rlaif-v: Aligning mllms through open-source ai feedback for super gpt-4v trustworthiness
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01a12e6a-dc3f-46e9-aa50-b0f1a6c0d1ca · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Self-Rewarding Language Models
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce8d0df0-8e3a-43b9-ace0-8fc071093344 · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32ef74f3-9fd0-437e-b0d4-9ec92225f51a · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Fine-Tuning Language Models from Human Preferences
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ffb9301-f40b-4731-beeb-4e9a3fc017c8 · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Attributes such as color, shape, texture, and 2D/3D spatial relations are also incorpo- rated
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation be913525-b847-4757-8445-144e73a0bfb1 · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation These atomic concepts are then transformed into simple yes-or-no questions
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation b76f1ddb-96aa-4c1f-9a24-66df37816450 · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation log σ − ¯σβ 2 LX i=1 ∥hw i − µi∥2 2 − βC + β 2¯σ LX i=1 ∥hw i − µref i ∥2 2 + βC − β 2¯σ LX i=1 ∥hl i − µi∥2 2 − βC + β 2¯σ LX i=1 ∥hl i − µref i ∥2 2 + βC ! # = −E(x,zw ,zl)∼D
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ef46e847-313a-4933-8d7b-ac3157509c1f · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation For SEED-LLaMA, the LLM backbone of DreamLLM is optimized for 1k steps, with a learning rate of 5 × 10−5, 100 warm-up steps, and a cosine learning rate scheduler
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation b059f2be-e759-4159-b1db-cc5fd264105a · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Unresolved cited work
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a0e4efb8-59b7-445c-b30c-c3f8d2abb4f5 · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation First Half
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 625205fb-3c24-4363-8277-4bc54bb41e6f · outbound
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation 14 - 24” means the rejected data points are sampled from rank-14 to rank-24 which is a hard range, while “20 - 30
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a6bdf839-3534-404d-a26d-b0076853d7a6 · inbound
SUDER: Self-Improving Unified Large Multimodal Models for Understanding and Generation with Dual Self-Rewards SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.