Pith. sign in

Paper Citation Record · LEDGER

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation

As of 10 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 0 inbound Pith citation observations for arXiv:2507.04151.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.04151 v1

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:01:50.856287Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

32 of 32 outbound references displayed

  • verified exact1
  • verified fuzzy12
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 36ea33ba-720d-4b0d-a522-6a82941ee6aa · outbound

This paper cites Draw ALL Your Imagine: A Holistic Benchmark and Agent Framework for Complex Instruction-based Image Generation.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation Draw ALL Your Imagine: A Holistic Benchmark and Agent Framework for Complex Instruction-based Image Generation

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:48.603688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:01:48.603688Z digest=sha256:a3cbf8cce6787889e371298acd1c433bfb4e1fa524ec8e12db2700ad857fca78

Observation c0237661-a292-42e9-b8b0-f7c521d28f7a · outbound

This paper cites Score: Story coherence and retrieval enhancement for ai narratives,.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation Score: Story coherence and retrieval enhancement for ai narratives,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:48.681445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:01:48.681445Z digest=sha256:84224da1984f9f99e38b6eab3fd0996976124e085e0cccfe3b64799dfc69b476

Observation 712b153c-1902-48dc-9661-2acf37a83993 · outbound

This paper cites Enhancing Code LLMs with Reinforcement Learning in Code Generation: A Survey.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation Enhancing Code LLMs with Reinforcement Learning in Code Generation: A Survey

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:48.736481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:01:48.736481Z digest=sha256:72d2ccb80040c952218db722099794a0ca5c488acc8e4d5c6f7256ee980b7095

Observation 471daa47-9b5e-439c-9265-31438a124170 · outbound

This paper cites Diffusion models beat gans on image synthesis,.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation Diffusion models beat gans on image synthesis,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:01:52.022091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:01:48.799947Z digest=sha256:a1527eb21e38a8cd53cd937ecd1f7d65f8145b2c2d51bf8a18e244aea75c9ea8

Observation c1544d88-8c49-4e86-89c5-bce1d492e2cb · outbound

This paper cites Triple sequence generative adversarial nets for unsupervised image captioning,.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation Triple sequence generative adversarial nets for unsupervised image captioning,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:01:52.012625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:01:48.852793Z digest=sha256:4a4b427065c1815bfca98bfd7afbe723cab220e2cba382c160488122869659e6

Observation 0380e3cc-e4dd-4c90-8ed4-29e8044e2766 · outbound

This paper cites Visual in-context learning for large vision-language models,.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation Visual in-context learning for large vision-language models,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:48.915151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:01:48.915151Z digest=sha256:b67c9c039bf928e2e6253d8e756b2f4c1f92d21b630e7e5a2b947f43c1007c5f

Observation 5f51352c-1325-4fa7-8205-16168498195e · outbound

This paper cites Weak to strong generalization for large language models with multi-capabilities,.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation Weak to strong generalization for large language models with multi-capabilities,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:48.968028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:01:48.968028Z digest=sha256:bb2025d0e33af9992421e804f8fa47ea813c3bc0b9ff7bdfb87f20b7983161b2

Observation 7c2e7e97-51ae-48d3-89ca-cb45602dea6d · outbound

This paper cites Improving Medical Large Vision-Language Models with Abnormal-Aware Feedback.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation Improving Medical Large Vision-Language Models with Abnormal-Aware Feedback

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:49.027115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:01:49.027115Z digest=sha256:e0e0022679afd0480844f7b2f8de6473259c3945e7f90d04dddc425ed98ce1d9

Observation e769d58f-c0b2-436b-b2a5-4884d2b33751 · outbound

This paper cites LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:49.072704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:01:49.072704Z digest=sha256:9356f9507b6e5ae125943eb69fe91a3b2df029cecb6eff8d198055d61561551a

Observation a64bd966-0714-446e-baf7-df3e2de89e81 · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image cap- tioning,.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image cap- tioning,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:49.124936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:01:49.124936Z digest=sha256:7c53cb73b074f19cb3deef191087df0c64193af24868091fa23e0b4bddfa037e

Observation 3e983ef3-f2c4-4479-84e6-ed3e42306a8d · outbound

This paper cites Microsoft COCO: common objects in context,.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation Microsoft COCO: common objects in context,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:49.182156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:01:49.182156Z digest=sha256:219ce9c074e19f61999de19b15f6cf1e0699b6f2c251e6d041d4520c18acb336

Observation b51647f4-05da-472d-8aa9-a3e55bbafaea · outbound

This paper cites Beyond empathy: Integrating diagnostic and therapeutic reasoning with large language models for mental health counseling,.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation Beyond empathy: Integrating diagnostic and therapeutic reasoning with large language models for mental health counseling,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:49.238458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:01:49.238458Z digest=sha256:eafa12fddcb92e024b2873d788efa6cbe9c855dfebf5f06a26ca41fce701a2ca

Observation c4bb82b5-5d46-4de5-a61e-31fa8521f3f0 · outbound

This paper cites Multimodal event transformer for image-guided story ending generation,.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation Multimodal event transformer for image-guided story ending generation,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:01:51.987595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:01:49.291914Z digest=sha256:b13ca58c41e52558e5afc46a75d19b4d8dcf75bf2ae2798b35128945d11b42cb

Observation 52b8e76b-5685-4a07-9ed3-d7699c55c7d8 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale,.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation An image is worth 16x16 words: Transformers for image recognition at scale,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:01:51.979779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:01:49.346212Z digest=sha256:0eb57428069e15ceed0f6cff008f13d2c5a18a3091bd1027018b88c56c154894

Observation 303dce78-eb93-4e1d-af3f-46553e5732e9 · outbound

This paper cites Language models with image descriptors are strong few- shot video-language learners,.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation Language models with image descriptors are strong few- shot video-language learners,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:49.415102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:01:49.415102Z digest=sha256:63ca88dc5df96e7b6c5dc02c430f6b41318f54d1a8b555542b96b7106cbce226

Observation e2b9ee77-6052-4fa3-a5ca-aa9cdc7d36a8 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation Learning transferable visual models from natural language supervision,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:49.477464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:01:49.477464Z digest=sha256:e3636daffe14f05928211c7f8aa3cfb3abbb4b9ca082e1f6f922e3e84bccfee1

Observation f6b2fe0d-0c95-403f-b6f8-29463105be68 · outbound

This paper cites BLIP: bootstrapping language-image pre-training for unified vision-language understanding and generation,.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation BLIP: bootstrapping language-image pre-training for unified vision-language understanding and generation,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:49.537759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:01:49.537759Z digest=sha256:6dfea3aa11de0d543d9c6a271798cfc633cafb523497460ff2db81b4747d555b

Observation 11c69b06-b6ac-46c8-b741-6f2949ee603e · outbound

This paper cites Flamingo: a visual language model for few-shot learning,.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation Flamingo: a visual language model for few-shot learning,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:01:51.950631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:01:49.665137Z digest=sha256:541d5baa819caffc0e7091f49589dfef36ad8392fbdde13bd8465060add86144

Observation c8d1758d-90c1-4e67-b954-f884343bde56 · outbound

This paper cites Cross-lingual transfer of large language model by visually- derived supervision toward low-resource languages,.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation Cross-lingual transfer of large language model by visually- derived supervision toward low-resource languages,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:01:51.933684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:01:49.777370Z digest=sha256:09c2de68a7a8779ff8e3546237c0075fcef5658d56db02beefd00f6bcf734dbf

Observation bbb1fff5-f82a-4315-a287-ce1fe51623f6 · outbound

This paper cites Coca: Contrastive captioners are image-text foundation models,.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation Coca: Contrastive captioners are image-text foundation models,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:49.857465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:01:49.857465Z digest=sha256:96805839b42de41069b4ba8759d716e815538687103e8a999c45cb33a54e1af2

Observation 8a8f4c49-5753-4a90-8d2b-349c24d96839 · outbound

This paper cites Tx-llava: Large language and vision assistant for temporal changes in chest x-rays,.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation Tx-llava: Large language and vision assistant for temporal changes in chest x-rays,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:49.932790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:01:49.932790Z digest=sha256:85a24754d803ee06389ff13b7a8953016aa1a0f96d770bac5299b081c275dd1c

Observation c0f02167-b11d-450b-84ee-efd249df8742 · outbound

This paper cites MAM: Modular Multi-Agent Framework for Multi-Modal Medical Diagnosis via Role-Specialized Collaboration.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation MAM: Modular Multi-Agent Framework for Multi-Modal Medical Diagnosis via Role-Specialized Collaboration

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:49.990147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:01:49.990147Z digest=sha256:52133c625acd54a38c7c7ddb55bdcbbae299238658b18da8342291e41621a20e

Observation 862c957f-faa2-484a-8d51-e71cc023b4db · outbound

This paper cites GLIDE: towards photorealistic image generation and editing with text-guided diffusion models,.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation GLIDE: towards photorealistic image generation and editing with text-guided diffusion models,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:01:51.919344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:01:50.130327Z digest=sha256:f89ed5e39f9317d1d5ed0de657a51e07672c4d732eadf2ba9723a5732e421454

Observation 297a636c-188f-4afa-b695-eec47deb6d44 · outbound

This paper cites Taming transformers for high-resolution image synthesis,.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation Taming transformers for high-resolution image synthesis,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:01:51.910644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:01:50.236512Z digest=sha256:657e92a53d9691d7837372268261d17b1ea688b946b5f3edaf2a4ec24a1490e1

Observation 58fbf2b1-988a-45fb-aa34-34272b0e1b4e · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:50.429219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:01:50.429219Z digest=sha256:8f6a88df29aa0be6f2750fa8d1c2b561cee09974b2866079a0df8fe4f58da8e5

Observation c9bfb984-33e0-4f54-a912-407aa9600c49 · outbound

This paper cites BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models,.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:01:51.891777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:01:50.532097Z digest=sha256:5bae2a40c971264e3b8bd19e6dd76983ad9ea60c4d7cae3ccf85008cd16bd0e7

Observation 23f68944-0b40-4f1c-8bc8-6956db206c4c · outbound

This paper cites Self-rewarding large vision- language models for optimizing prompts in text-to-image generation,.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation Self-rewarding large vision- language models for optimizing prompts in text-to-image generation,

Reference 27

Resolution
verified exact
raw_fallback, observed 2026-08-06T20:01:51.120901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:01:50.623402Z digest=sha256:3827e9a8e18bfe8b8dbafaaff6c0ed0e3e42a477b05ce90781bdd128242d5063

Observation a03b9272-b4da-457a-93c2-b024dea37138 · outbound

This paper cites Evolvedirector: Approaching advanced text-to-image generation with large vision- language models,.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation Evolvedirector: Approaching advanced text-to-image generation with large vision- language models,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:01:51.618270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:01:50.782997Z digest=sha256:e816c6609c5b6d6cf6b723955fdbfc66be48de3b6d8b054438153acb28851a41

Observation d773c6ae-7ce0-4acc-90e3-fc520c889a19 · outbound

This paper cites Improving Compositional Text-to-image Generation with Large Vision-Language Models.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation Improving Compositional Text-to-image Generation with Large Vision-Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:50.856287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:01:50.856287Z digest=sha256:7825ad4a7a3b87912539bd4b4279279f16d2e90183ed79e5587c032378df4f71

Observation ce90422c-0cd8-4b94-aee9-40bbe469c126 · outbound

This paper cites 12 888–12 900.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation 12 888–12 900

Reference 162

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:49.603500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:01:49.603500Z digest=sha256:fc60c8895bb23c9e51e8b757de28bf84118eb52d73c56e20a21aa53d99e29f41

Observation 36bc1f5d-8b4d-4268-81a7-3680fabc8de5 · outbound

This paper cites 12 873–12 883.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation 12 873–12 883

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:01:51.901437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:01:50.329073Z digest=sha256:3659ebf2cd8803d19719abd38227e60d9d2c6b10405e6db35a6c45d155287028

Observation e42b8339-0647-416f-ab6f-96e8422d85ea · outbound

This paper cites Available: http://papers.nips.cc/paper\ files/paper/2022/ hash/960a172bc7fbf0177ccccbb411a7d800-Abstract-Conference.html 11.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation Available: http://papers.nips.cc/paper\ files/paper/2022/ hash/960a172bc7fbf0177ccccbb411a7d800-Abstract-Conference.html 11

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:01:51.942532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:01:49.714980Z digest=sha256:329faf717568c2438f83f94e99b65b9a6ae0f07f8cdb6e58bacdfe8fbec689b4

Pith citing papers

No inbound Pith citation observations are available.