Pith. sign in

Paper Citation Record · LEDGER

LVLM-Composer's Explicit Planning for Image Generation

As of 8 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2507.04152.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.04152 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:58:05.710909Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

31 of 31 outbound references displayed

  • verified exact3
  • verified fuzzy4
  • unresolved22
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 24e861b9-5e24-4d5c-b14f-c57bdddf08a3 · outbound

This paper cites Less is more: Vision representation compression for efficient video generation with large language models,.

LVLM-Composer's Explicit Planning for Image Generation Less is more: Vision representation compression for efficient video generation with large language models,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T19:58:01.645874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:58:01.645874Z digest=sha256:cde1168ae5e5ba5dafe3d09d9c1d3e8a7455461bc337b60e2789c1433cc4e265

Observation bd2cd77b-15f5-4468-8e93-83283b1ebdde · outbound

This paper cites Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling.

LVLM-Composer's Explicit Planning for Image Generation Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T19:58:01.720824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:58:01.720824Z digest=sha256:16e9f1f00ab78c4ecf488603583f4e41ee9669ea3a3d8b921debbff84468b321

Observation 21eb0ce8-b794-4f5c-af95-7c5121abd867 · outbound

This paper cites Draw ALL Your Imagine: A Holistic Benchmark and Agent Framework for Complex Instruction-based Image Generation.

LVLM-Composer's Explicit Planning for Image Generation Draw ALL Your Imagine: A Holistic Benchmark and Agent Framework for Complex Instruction-based Image Generation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T19:58:01.777284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:58:01.777284Z digest=sha256:166dda54c02e6c67d227c088a34e7a6a0a30b30e017cb71a52f39bb189eced49

Observation 054b8275-6e6b-4a44-b369-6d7e14b4fa48 · outbound

This paper cites ComplexBench-Edit: Benchmarking Complex Instruction-Driven Image Editing via Compositional Dependencies.

LVLM-Composer's Explicit Planning for Image Generation ComplexBench-Edit: Benchmarking Complex Instruction-Driven Image Editing via Compositional Dependencies

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T19:58:01.897972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:58:01.897972Z digest=sha256:9da1f56338e565b375ae1ac4056663de7f3b06edee462f1632797a68f34bc37f

Observation 32798ebe-4040-4612-baa4-705948485c0f · outbound

This paper cites Weak to strong generalization for large language models with multi-capabilities,.

LVLM-Composer's Explicit Planning for Image Generation Weak to strong generalization for large language models with multi-capabilities,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T19:58:02.079218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:58:02.079218Z digest=sha256:91cf66d13babe9eb409a6e939c26947c21585577b9acc95e6183144d643f150e

Observation 41fc1007-da5b-4686-a218-a3d82559baac · outbound

This paper cites Divide-Then-Aggregate: An Efficient Tool Learning Method via Parallel Tool Invocation.

LVLM-Composer's Explicit Planning for Image Generation Divide-Then-Aggregate: An Efficient Tool Learning Method via Parallel Tool Invocation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T19:58:02.300391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:58:02.300391Z digest=sha256:28477e6f45d21bb3fd714541eaa3a79d12b7350f4d752b8b293de1bb1ba5c05a

Observation 1365ae65-0aae-44aa-870f-f3af36727cbd · outbound

This paper cites Thread of Thought Unraveling Chaotic Contexts.

LVLM-Composer's Explicit Planning for Image Generation Thread of Thought Unraveling Chaotic Contexts

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T19:58:02.468597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:58:02.468597Z digest=sha256:7f84c32fe6b7dee1c5883acf1ec9900cb96ffc472311cfacbf92e451c9313da2

Observation 0b17ec1a-4b49-4687-8b59-22e9adaa2801 · outbound

This paper cites Hierarchical reinforcement learning for handling sparse rewards in multi-goal navigation,.

LVLM-Composer's Explicit Planning for Image Generation Hierarchical reinforcement learning for handling sparse rewards in multi-goal navigation,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:58:08.041512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:58:02.608469Z digest=sha256:d499529f9dba1828abcc7d586a49473e0d6b907a9320a947e0ee5923ae69e228

Observation a915607f-1aca-4343-b396-04f7df7932c0 · outbound

This paper cites Training Latent Variable Models with Auto-encoding Variational Bayes: A Tutorial.

LVLM-Composer's Explicit Planning for Image Generation Training Latent Variable Models with Auto-encoding Variational Bayes: A Tutorial

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:58:06.318947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:58:02.725421Z digest=sha256:55a69af754df7e6ba61dbf0a02e0d8e459a5690d8971abbdb217d5871e214af3

Observation b8063857-6afa-433f-a5f5-9f827a0a859d · outbound

This paper cites AT-GAN: An Adversarial Generator Model for Non-constrained Adversarial Examples.

LVLM-Composer's Explicit Planning for Image Generation AT-GAN: An Adversarial Generator Model for Non-constrained Adversarial Examples

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:58:07.424502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:58:02.846209Z digest=sha256:26674a4f351787359b7c426b1af95ec7004b1ae32f49db312c455199ce024524

Observation bbccaa72-c661-4b23-8a95-7325dcaa7de9 · outbound

This paper cites Progressive growing of gans for improved quality, stability, and variation,.

LVLM-Composer's Explicit Planning for Image Generation Progressive growing of gans for improved quality, stability, and variation,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:58:07.885775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:58:02.990587Z digest=sha256:38f80bc287e940271c04c0c65ce900853ce2b4e3e5dca6fed55b5a7e8af53173

Observation f15cd496-2cbc-41f8-b3a0-33f4a0c5b42d · outbound

This paper cites Unpaired image-to-image translation using cycle-consistent adversarial networks,.

LVLM-Composer's Explicit Planning for Image Generation Unpaired image-to-image translation using cycle-consistent adversarial networks,

Reference 12

Resolution
malformed identifier
no resolver link, observed 2026-08-06T19:58:03.146784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:58:03.146784Z digest=sha256:0aca589a459b1903acd477317446b54d0dea764f002a2f85a7d7e048da5047f7

Observation 0a9ed52f-6c7c-4c57-b43b-03abfe0f0bf6 · outbound

This paper cites Wrapped phase denoising using denoising diffusion probabilistic models,.

LVLM-Composer's Explicit Planning for Image Generation Wrapped phase denoising using denoising diffusion probabilistic models,

Reference 13

Resolution
metadata mismatch
raw_fallback, observed 2026-08-06T19:58:07.171960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:58:03.245760Z digest=sha256:fff400372b7a227ccc6b1d3c6f8369b5ac711ca779f37fdd52b7bdd26cb4dc52

Observation 4a3db813-1076-4c35-9802-5e7b85d283d0 · outbound

This paper cites Diffusion-4k: Ultra-high-resolution image synthesis with latent diffusion models,.

LVLM-Composer's Explicit Planning for Image Generation Diffusion-4k: Ultra-high-resolution image synthesis with latent diffusion models,

Reference 14

Resolution
verified exact
doi, observed 2026-08-06T19:58:06.084601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:58:03.368276Z digest=sha256:6ff83cdc777c4bf0a7943964cda6f49a8c8965234ed8112824b3a24be0e5c2ea

Observation 84253ed0-da30-4925-a238-d56d8b0b0521 · outbound

This paper cites Photorealistic text-to-image diffusion models with deep language understanding,.

LVLM-Composer's Explicit Planning for Image Generation Photorealistic text-to-image diffusion models with deep language understanding,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T19:58:03.546204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:58:03.546204Z digest=sha256:ee0fbd62917ecab61ec3b405b1d43adaeb137d5886e46ae16400ecc0b3a2631c

Observation 978225b2-86ff-4309-a5d0-29f1faac460e · outbound

This paper cites Masked autoencoders are scalable vision learners,.

LVLM-Composer's Explicit Planning for Image Generation Masked autoencoders are scalable vision learners,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T19:58:03.698375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:58:03.698375Z digest=sha256:e8fa1d6a9d38113a13c616bf904edc96ed9b5a996e1faef099e39ad52b94cc33

Observation 5e74baa5-ad98-4b2e-ab48-5e18f78c046c · outbound

This paper cites Improving cross-modal alignment for text- guided image inpainting,.

LVLM-Composer's Explicit Planning for Image Generation Improving cross-modal alignment for text- guided image inpainting,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T19:58:03.794578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:58:03.794578Z digest=sha256:f0157b02c35ccf5ab24a745e75bbd57219d30b82c5b29ea805c796a13ae9d0d5

Observation 1ed7a406-62f0-4de9-8b3a-a1ee0a56717b · outbound

This paper cites Improving Medical Large Vision-Language Models with Abnormal-Aware Feedback.

LVLM-Composer's Explicit Planning for Image Generation Improving Medical Large Vision-Language Models with Abnormal-Aware Feedback

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T19:58:03.938240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:58:03.938240Z digest=sha256:002a2c4a9dd2acf07c5ee866b8a29c61eae6b4a7232a319bc65e101ba5d8f629

Observation 499d3422-1ebd-486c-bcdf-e5d7068012e7 · outbound

This paper cites Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks,.

LVLM-Composer's Explicit Planning for Image Generation Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T19:58:04.039751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:58:04.039751Z digest=sha256:acb9add1b663a8e7f8cd30313b9a8f4ad9b642ffa7b9d705e2bc70edbec67f4a

Observation 7e7b9a1f-e213-48e4-9b90-8a219cd8f120 · outbound

This paper cites VisualBERT: A Simple and Performant Baseline for Vision and Language.

LVLM-Composer's Explicit Planning for Image Generation VisualBERT: A Simple and Performant Baseline for Vision and Language

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T19:58:04.104937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:58:04.104937Z digest=sha256:050973505dfcfa9bb4726a066bb0001d8928288f36007a88744cfaae1766d4ea

Observation 5e8eeb73-a962-4392-b228-2bb6c6c2f836 · outbound

This paper cites LXMERT: learning cross-modality encoder representations from transformers,.

LVLM-Composer's Explicit Planning for Image Generation LXMERT: learning cross-modality encoder representations from transformers,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T19:58:04.146702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:58:04.146702Z digest=sha256:f2cc5138231ce64a899b0459673a3d4bfb47289240ea3e9485901c5d14396c0f

Observation e7c31312-2bd0-45e9-9f18-a46deaa6b514 · outbound

This paper cites UNITER: universal image-text representation learning,.

LVLM-Composer's Explicit Planning for Image Generation UNITER: universal image-text representation learning,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T19:58:04.269413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:58:04.269413Z digest=sha256:361399d76930e9827e0bb04d7469a7049ba8ab3280fbd326494d66f6e03e2e6b

Observation c15234cd-319a-4739-9e93-7817ad3283c1 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

LVLM-Composer's Explicit Planning for Image Generation Learning transferable visual models from natural language supervision,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:58:07.749125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:58:04.467187Z digest=sha256:7fc2cc97c78406e2046237dd9e541bb304858bba3e88cab8d07de49011c65190

Observation 99bd6c51-ee5f-4d94-bb95-730fa9ffb1ea · outbound

This paper cites Interleave-vla: Enhancing robot manipulation with interleaved image-text instructions,.

LVLM-Composer's Explicit Planning for Image Generation Interleave-vla: Enhancing robot manipulation with interleaved image-text instructions,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T19:58:04.673194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:58:04.673194Z digest=sha256:fdaa26b390411555d17dfd27433a6d6505bc09b80e98810aa4c9bdba1163b4e2

Observation a18d5673-de44-4be3-89d3-fd2fed68f536 · outbound

This paper cites SDA: Simple Discrete Augmentation for Contrastive Sentence Representation Learning.

LVLM-Composer's Explicit Planning for Image Generation SDA: Simple Discrete Augmentation for Contrastive Sentence Representation Learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T19:58:04.807213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:58:04.807213Z digest=sha256:d8ece4cce48ad6b013ceb177dd5e4830608815e043f7fbd8faaf4b41d13123de

Observation 6c414907-a55c-4a80-ac8b-b0d78fc63b9c · outbound

This paper cites Florence: A New Foundation Model for Computer Vision.

LVLM-Composer's Explicit Planning for Image Generation Florence: A New Foundation Model for Computer Vision

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T19:58:04.995920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:58:04.995920Z digest=sha256:228fb561972cf06551eea18dfb397b44b75f577bc05ab7eac1cbf27726b10fb8

Observation 1d874a5a-7bf3-4e04-870c-719f23f40135 · outbound

This paper cites Coca: Contrastive captioners are image-text foundation models,.

LVLM-Composer's Explicit Planning for Image Generation Coca: Contrastive captioners are image-text foundation models,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:58:07.569594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:58:05.142667Z digest=sha256:022500c3c6ab5c395238fc287f992345329a143e1a692a1d40388301f3c4d016

Observation f190b315-f376-478c-b7ee-7d3a5aa52a98 · outbound

This paper cites Tx-llava: Large language and vision assistant for temporal changes in chest x-rays,.

LVLM-Composer's Explicit Planning for Image Generation Tx-llava: Large language and vision assistant for temporal changes in chest x-rays,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T19:58:05.327150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:58:05.327150Z digest=sha256:8b3fdc361b4e3311a1b9152936cb26418a43dd8b7119bf34a214b377f0c8ea39

Observation 607f8bea-edfd-4106-a2e5-7259d5e0d65f · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

LVLM-Composer's Explicit Planning for Image Generation MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T19:58:05.420033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:58:05.420033Z digest=sha256:37b58b39135b003a4467d15ad48013f1c276a79181be75dc9c693779d6d9964e

Observation db4e1c6d-90c6-4b13-a17c-7e5d6f8bdaad · outbound

This paper cites VisLingInstruct: Elevating Zero-Shot Learning in Multi-Modal Language Models with Autonomous Instruction Optimization.

LVLM-Composer's Explicit Planning for Image Generation VisLingInstruct: Elevating Zero-Shot Learning in Multi-Modal Language Models with Autonomous Instruction Optimization

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T19:58:05.562611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:58:05.562611Z digest=sha256:197a0f83b39584c7128f49fa66ef5d3535f4bf0a4a630883a4ace60880cb33c4

Observation 682ff5f4-01e0-4636-a69f-8cffd40bb5fe · outbound

This paper cites MAM: Modular Multi-Agent Framework for Multi-Modal Medical Diagnosis via Role-Specialized Collaboration.

LVLM-Composer's Explicit Planning for Image Generation MAM: Modular Multi-Agent Framework for Multi-Modal Medical Diagnosis via Role-Specialized Collaboration

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T19:58:05.710909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:58:05.710909Z digest=sha256:849c085738a720c6c3dc61b8b032ea531434e067cf6b03949a787e4fc5a33493

Pith citing papers

No inbound Pith citation observations are available.