Pith. sign in

Paper Citation Record · LEDGER

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation

As of 9 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 0 inbound Pith citation observations for arXiv:2507.04151.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.04151 v1

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:01:50.856287Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

32 of 32 outbound references displayed

  • verified exact1
  • verified fuzzy12
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 36ea33ba-720d-4b0d-a522-6a82941ee6aa · outbound

This paper cites Draw ALL Your Imagine: A Holistic Benchmark and Agent Framework for Complex Instruction-based Image Generation.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation Draw ALL Your Imagine: A Holistic Benchmark and Agent Framework for Complex Instruction-based Image Generation

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:48.603688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:01:48.603688Z digest=sha256:bb09622936bf2418a38b209ad27ab6c3113c8726827230f0c50ea1ef1608c5eb

Observation c0237661-a292-42e9-b8b0-f7c521d28f7a · outbound

This paper cites Score: Story coherence and retrieval enhancement for ai narratives,.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation Score: Story coherence and retrieval enhancement for ai narratives,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:48.681445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:01:48.681445Z digest=sha256:769b5da8c0b6100d7ec55eb216f2bd2def525a6e79ad5a67d16bc93588cfd4b6

Observation 712b153c-1902-48dc-9661-2acf37a83993 · outbound

This paper cites Enhancing Code LLMs with Reinforcement Learning in Code Generation: A Survey.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation Enhancing Code LLMs with Reinforcement Learning in Code Generation: A Survey

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:48.736481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:01:48.736481Z digest=sha256:c167ae51427ce6ae22bed56c3d21ea9c146688cf79a4559ac3c862c139db9c40

Observation 471daa47-9b5e-439c-9265-31438a124170 · outbound

This paper cites Diffusion models beat gans on image synthesis,.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation Diffusion models beat gans on image synthesis,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:01:52.022091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:01:48.799947Z digest=sha256:532cd03fa16830ad1f0b929e4d5e787d3dfba17e3600c0f82c5234791bbce29f

Observation c1544d88-8c49-4e86-89c5-bce1d492e2cb · outbound

This paper cites Triple sequence generative adversarial nets for unsupervised image captioning,.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation Triple sequence generative adversarial nets for unsupervised image captioning,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:01:52.012625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:01:48.852793Z digest=sha256:954759a40641ed6b343e3f726884776972a9f09b619114b446786bfbde0063a8

Observation 0380e3cc-e4dd-4c90-8ed4-29e8044e2766 · outbound

This paper cites Visual in-context learning for large vision-language models,.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation Visual in-context learning for large vision-language models,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:48.915151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:01:48.915151Z digest=sha256:9b997fa21b138a2a7562e064da803ae2a1954b3fd3ce0d2888e02ff66cbd2f31

Observation 5f51352c-1325-4fa7-8205-16168498195e · outbound

This paper cites Weak to strong generalization for large language models with multi-capabilities,.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation Weak to strong generalization for large language models with multi-capabilities,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:48.968028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:01:48.968028Z digest=sha256:84d3107117550e74773abd747f43e256fe5e8dd8ca30604c03e3d86445e137fc

Observation 7c2e7e97-51ae-48d3-89ca-cb45602dea6d · outbound

This paper cites Improving Medical Large Vision-Language Models with Abnormal-Aware Feedback.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation Improving Medical Large Vision-Language Models with Abnormal-Aware Feedback

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:49.027115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:01:49.027115Z digest=sha256:23192cd37c7ae042a69adab43caf7f40a5b626c6f1fabe91de6215d25a14487a

Observation e769d58f-c0b2-436b-b2a5-4884d2b33751 · outbound

This paper cites LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:49.072704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:01:49.072704Z digest=sha256:0defe22f81dddae7e7bd1ea90ccd1f52331f31d81129db927634d2d8f6209611

Observation a64bd966-0714-446e-baf7-df3e2de89e81 · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image cap- tioning,.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image cap- tioning,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:49.124936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:01:49.124936Z digest=sha256:80fb5220da2f260e18fa2fa6cbc701b56407653be25e66113b8147e75cdd7c14

Observation 3e983ef3-f2c4-4479-84e6-ed3e42306a8d · outbound

This paper cites Microsoft COCO: common objects in context,.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation Microsoft COCO: common objects in context,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:49.182156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:01:49.182156Z digest=sha256:680aee100ef9148d560aaf670598100e7192d90ab68e29f9f4e9b98a64ec6f30

Observation b51647f4-05da-472d-8aa9-a3e55bbafaea · outbound

This paper cites Beyond empathy: Integrating diagnostic and therapeutic reasoning with large language models for mental health counseling,.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation Beyond empathy: Integrating diagnostic and therapeutic reasoning with large language models for mental health counseling,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:49.238458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:01:49.238458Z digest=sha256:f1193683d4a4cab557d997ae9701fc581dbbf8111f702d2aa902ce63f6de36c1

Observation c4bb82b5-5d46-4de5-a61e-31fa8521f3f0 · outbound

This paper cites Multimodal event transformer for image-guided story ending generation,.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation Multimodal event transformer for image-guided story ending generation,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:01:51.987595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:01:49.291914Z digest=sha256:94a3291e6c57e99b0d415896dbf85c0dad0fbacde76f9087f2c63e6786c4eb00

Observation 52b8e76b-5685-4a07-9ed3-d7699c55c7d8 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale,.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation An image is worth 16x16 words: Transformers for image recognition at scale,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:01:51.979779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:01:49.346212Z digest=sha256:1a9f21d66cf389833dc570c4990fc53a1f62f57dfb028021e9a7c43591bff600

Observation 303dce78-eb93-4e1d-af3f-46553e5732e9 · outbound

This paper cites Language models with image descriptors are strong few- shot video-language learners,.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation Language models with image descriptors are strong few- shot video-language learners,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:49.415102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:01:49.415102Z digest=sha256:dcb75e8493fa87ddf2d3be890853cebfe01e062b1f84c8ccf325c350ab84ae6b

Observation e2b9ee77-6052-4fa3-a5ca-aa9cdc7d36a8 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation Learning transferable visual models from natural language supervision,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:49.477464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:01:49.477464Z digest=sha256:bf6e5d832c2c83e2d2041087194db3ebdee7c0b3df2366357016c49801877bfb

Observation f6b2fe0d-0c95-403f-b6f8-29463105be68 · outbound

This paper cites BLIP: bootstrapping language-image pre-training for unified vision-language understanding and generation,.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation BLIP: bootstrapping language-image pre-training for unified vision-language understanding and generation,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:49.537759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:01:49.537759Z digest=sha256:5b15d939979e4474f5d9febab473cb899a35ceb312489e873b1adb899c3b9ccd

Observation 11c69b06-b6ac-46c8-b741-6f2949ee603e · outbound

This paper cites Flamingo: a visual language model for few-shot learning,.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation Flamingo: a visual language model for few-shot learning,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:01:51.950631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:01:49.665137Z digest=sha256:f562855ddb637abff93a90450db80c721345e264743c89b17dd27e4ae75d5ea1

Observation c8d1758d-90c1-4e67-b954-f884343bde56 · outbound

This paper cites Cross-lingual transfer of large language model by visually- derived supervision toward low-resource languages,.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation Cross-lingual transfer of large language model by visually- derived supervision toward low-resource languages,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:01:51.933684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:01:49.777370Z digest=sha256:5e1a9fd1f07c38787f3d9e5f49c86819d4b514b360d887c14961b413507877c0

Observation bbb1fff5-f82a-4315-a287-ce1fe51623f6 · outbound

This paper cites Coca: Contrastive captioners are image-text foundation models,.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation Coca: Contrastive captioners are image-text foundation models,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:49.857465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:01:49.857465Z digest=sha256:e11e9175f1984e9a13338f451bb3a99e70f77ecc76bdd3c820a6b53b12d0500b

Observation 8a8f4c49-5753-4a90-8d2b-349c24d96839 · outbound

This paper cites Tx-llava: Large language and vision assistant for temporal changes in chest x-rays,.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation Tx-llava: Large language and vision assistant for temporal changes in chest x-rays,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:49.932790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:01:49.932790Z digest=sha256:d899c066edce8812bc0cb25b860c5d07d66eb9cb85b36dd6ca07cf3daf35f9b3

Observation c0f02167-b11d-450b-84ee-efd249df8742 · outbound

This paper cites MAM: Modular Multi-Agent Framework for Multi-Modal Medical Diagnosis via Role-Specialized Collaboration.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation MAM: Modular Multi-Agent Framework for Multi-Modal Medical Diagnosis via Role-Specialized Collaboration

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:49.990147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:01:49.990147Z digest=sha256:126988d2a01eefc9f7caa3bc889dc0c9b4d20a7d1b0c1790b7e850a902b86bb4

Observation 862c957f-faa2-484a-8d51-e71cc023b4db · outbound

This paper cites GLIDE: towards photorealistic image generation and editing with text-guided diffusion models,.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation GLIDE: towards photorealistic image generation and editing with text-guided diffusion models,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:01:51.919344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:01:50.130327Z digest=sha256:55bdc9d1cf2bee7f9a5a341b4c1e0cfbdf204aa5c590608bed958c1066f39d9e

Observation 297a636c-188f-4afa-b695-eec47deb6d44 · outbound

This paper cites Taming transformers for high-resolution image synthesis,.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation Taming transformers for high-resolution image synthesis,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:01:51.910644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:01:50.236512Z digest=sha256:11f5abc08b6cd9ba8d769d3cc16cb284b183a9a62e447c942fd0564a7de8591f

Observation 58fbf2b1-988a-45fb-aa34-34272b0e1b4e · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:50.429219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:01:50.429219Z digest=sha256:b3b830ff739cdc74f2947eecbb929abb505dc27678a8d8beaca6d49d75a6e874

Observation c9bfb984-33e0-4f54-a912-407aa9600c49 · outbound

This paper cites BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models,.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:01:51.891777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:01:50.532097Z digest=sha256:6c956d217bf06f94dbf750396ddef1f0a9f71031ff91c60704c0b5cdd4b9f653

Observation 23f68944-0b40-4f1c-8bc8-6956db206c4c · outbound

This paper cites Self-rewarding large vision- language models for optimizing prompts in text-to-image generation,.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation Self-rewarding large vision- language models for optimizing prompts in text-to-image generation,

Reference 27

Resolution
verified exact
raw_fallback, observed 2026-08-06T20:01:51.120901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:01:50.623402Z digest=sha256:5599823dfa1ce590d2ea679e037eab5c92569e41f8e97cc1310afb567f75fe68

Observation a03b9272-b4da-457a-93c2-b024dea37138 · outbound

This paper cites Evolvedirector: Approaching advanced text-to-image generation with large vision- language models,.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation Evolvedirector: Approaching advanced text-to-image generation with large vision- language models,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:01:51.618270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:01:50.782997Z digest=sha256:a7379eb8ee13788d0b2ec961ef099f2c7efc1a4d0a47ee7db12815b97d8c1c43

Observation d773c6ae-7ce0-4acc-90e3-fc520c889a19 · outbound

This paper cites Improving Compositional Text-to-image Generation with Large Vision-Language Models.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation Improving Compositional Text-to-image Generation with Large Vision-Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:50.856287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:01:50.856287Z digest=sha256:14b77a15b11e732674cf72d6e7764bd4aef97c5c742c545247847d7d59ba1bbd

Observation ce90422c-0cd8-4b94-aee9-40bbe469c126 · outbound

This paper cites 12 888–12 900.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation 12 888–12 900

Reference 162

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:49.603500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:01:49.603500Z digest=sha256:261be0c0786614e1bbbec07d0b7f17d0f49c93a76b0e592a66e6f47c50f23bfe

Observation 36bc1f5d-8b4d-4268-81a7-3680fabc8de5 · outbound

This paper cites 12 873–12 883.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation 12 873–12 883

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:01:51.901437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:01:50.329073Z digest=sha256:629bbf58a37f079076956abcc930d100737a1539e3c4a68528042378917558fa

Observation e42b8339-0647-416f-ab6f-96e8422d85ea · outbound

This paper cites Available: http://papers.nips.cc/paper\ files/paper/2022/ hash/960a172bc7fbf0177ccccbb411a7d800-Abstract-Conference.html 11.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation Available: http://papers.nips.cc/paper\ files/paper/2022/ hash/960a172bc7fbf0177ccccbb411a7d800-Abstract-Conference.html 11

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:01:51.942532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:01:49.714980Z digest=sha256:e06dd58afcd6b74e34eb6e5e7f0345013e13d818b164ce397d5e2b5d2d7edad0

Pith citing papers

No inbound Pith citation observations are available.