Pith. sign in

Paper Citation Record · LEDGER

Learning to Generate Multiple Objects from Dense and Occluded Layouts

As of 10 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2607.03488.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.03488 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-12T02:09:14.783011Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ad8d662a-ab82-4948-86cb-b979f6a9e1a8 · outbound

This paper cites Seethrough3d: Occlusion aware 3d control in text-to-image generation.

Learning to Generate Multiple Objects from Dense and Occluded Layouts Seethrough3d: Occlusion aware 3d control in text-to-image generation

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-12T02:09:14.783011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T02:09:14.783011Z digest=sha256:3c981fce8204a6d6dd88e3c9afcc349ef2a9af0e3c333fa1b81b3cbf6dce6b24

Observation 9c1b1e73-3785-4846-af01-11dbe46dfd98 · outbound

This paper cites Improving image generation with better captions.

Learning to Generate Multiple Objects from Dense and Occluded Layouts Improving image generation with better captions

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-12T02:09:14.783011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T02:09:14.783011Z digest=sha256:b3cc43c241800deada0376175b6f609185f4932e4518fdbcf439ece7fd9c03a9

Observation 609b6dc2-fa71-458b-8de3-600d9bae1806 · outbound

This paper cites Make it count: Text-to-image generation with an accurate number of objects.

Learning to Generate Multiple Objects from Dense and Occluded Layouts Make it count: Text-to-image generation with an accurate number of objects

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-12T02:09:14.783011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T02:09:14.783011Z digest=sha256:d510a97ea33f8c18312325ca9344a7b44290b57380c821c6c1201fd35dba12ee

Observation 546bdf69-a1d4-43c7-90cf-8492b2781205 · outbound

This paper cites Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer.

Learning to Generate Multiple Objects from Dense and Occluded Layouts Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-12T02:09:14.783011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T02:09:14.783011Z digest=sha256:332e8adb9d80acee48b36e93eb0dfa6e57527d8ba0e623fb973e660cd681759f

Observation b3a504ab-495b-4838-a2ed-dbe46c2affe8 · outbound

This paper cites SAM 3: Segment Anything with Concepts.

Learning to Generate Multiple Objects from Dense and Occluded Layouts SAM 3: Segment Anything with Concepts

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-12T02:09:14.783011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T02:09:14.783011Z digest=sha256:a195e2015366c21eec04292a7e6e67b863023f727d09c08b698da53bf2c9c828

Observation 1ad75d48-e6cb-4eeb-a44f-01b76f5a3e8a · outbound

This paper cites Attend-and-excite: Attention-based semantic guidance for text-to-image diffusion models.ACM Trans.

Learning to Generate Multiple Objects from Dense and Occluded Layouts Attend-and-excite: Attention-based semantic guidance for text-to-image diffusion models.ACM Trans

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-12T02:09:14.783011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T02:09:14.783011Z digest=sha256:06acf8e7b0fe774b96138f3d850c41b5c97412cc8163fe8f80423fdc889e5eef

Observation 2b72de42-dbe0-4d19-afca-bbe551e4b2cf · outbound

This paper cites Training-free layout control with cross-attention guidance.

Learning to Generate Multiple Objects from Dense and Occluded Layouts Training-free layout control with cross-attention guidance

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-12T02:09:14.783011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T02:09:14.783011Z digest=sha256:dce576fea3efd54274893ea929a3332a97ac0fda1c7f92aea815e58c52c5b3e8

Observation bac52182-86d1-4124-a3d4-19a7fbdc2304 · outbound

This paper cites Reason out Your Layout: Evoking the Layout Master from Large Language Models for Text-to-Image Synthesis.

Learning to Generate Multiple Objects from Dense and Occluded Layouts Reason out Your Layout: Evoking the Layout Master from Large Language Models for Text-to-Image Synthesis

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-12T02:09:14.783011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T02:09:14.783011Z digest=sha256:e6f119910ae8821be45c86dd5bd1f97b7c3498a7028420988447c40ebfce03bf

Observation 1e2c5967-567b-4291-b48c-eb80a79963db · outbound

This paper cites Prompt-to-Prompt Image Editing with Cross Attention Control.

Learning to Generate Multiple Objects from Dense and Occluded Layouts Prompt-to-Prompt Image Editing with Cross Attention Control

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-12T02:09:14.783011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T02:09:14.783011Z digest=sha256:d353d139530c3a99fe5f3080862afe5163fac0e92ce9985b3c7a158729b39ea9

Observation 2891cf76-d7f7-49dc-b61c-588227f78569 · outbound

This paper cites T2i-compbench: A compre- hensive benchmark for open-world compositional text-to-image generation.Advances in Neural Information Processing Systems, 36:78723–78747, 2023.

Learning to Generate Multiple Objects from Dense and Occluded Layouts T2i-compbench: A compre- hensive benchmark for open-world compositional text-to-image generation.Advances in Neural Information Processing Systems, 36:78723–78747, 2023

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-12T02:09:14.783011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T02:09:14.783011Z digest=sha256:6be054b0a3b1b067f06fd4749fb16b8e58eeff31b0912b86d889b6442e0943d6

Observation f30a5ed3-2d5d-44e2-96ad-173f6abda7f6 · outbound

This paper cites Counting guidance for high fidelity text-to-image synthesis.

Learning to Generate Multiple Objects from Dense and Occluded Layouts Counting guidance for high fidelity text-to-image synthesis

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-12T02:09:14.783011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T02:09:14.783011Z digest=sha256:7d6f970de7dc3f79dc6f4cd5fdafe95599536b14549a356b620a09cbf0e817df

Observation ee752f6c-2d7f-4828-8124-88fb8d5e856d · outbound

This paper cites Counting Guidance for High Fidelity Text-to-Image Synthesis.

Learning to Generate Multiple Objects from Dense and Occluded Layouts Counting Guidance for High Fidelity Text-to-Image Synthesis

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-12T02:09:14.783011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T02:09:14.783011Z digest=sha256:58345d610679deaa0a497d45580648cc9159bb16f2f1995bc4bca0e3642ee6a7

Observation 8e3aea05-25e2-4707-bea1-0525d029be58 · outbound

This paper cites Omg: Occlusion-friendly personalized multi-concept generation in diffusion models.

Learning to Generate Multiple Objects from Dense and Occluded Layouts Omg: Occlusion-friendly personalized multi-concept generation in diffusion models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-12T02:09:14.783011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T02:09:14.783011Z digest=sha256:7d32703ec6f5a70494db09621f3e49f06ddca1ed7bea332c2831aeb6ca78bbe0

Observation 5d43416f-8a12-4c91-aac8-8a17abad44e5 · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annotations.International Journal of Computer Vision, 123(1):32–73, 2017.

Learning to Generate Multiple Objects from Dense and Occluded Layouts Visual genome: Connecting language and vision using crowdsourced dense image annotations.International Journal of Computer Vision, 123(1):32–73, 2017

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-12T02:09:14.783011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T02:09:14.783011Z digest=sha256:fa07c6196395f8af1d2a5e7e4254bb9c488345e7ba7f1df02cc799aa9c7a0e8e

Observation f73cb970-9b98-4996-ae31-80138e8bc470 · outbound

This paper cites Flux.https://github.com/black-forest-labs/flux, 2024.

Learning to Generate Multiple Objects from Dense and Occluded Layouts Flux.https://github.com/black-forest-labs/flux, 2024

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-12T02:09:14.783011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T02:09:14.783011Z digest=sha256:4e2d43f285bca0e191dc95a828bee290680467a3433e6ccdab1fdaba6c7499be

Observation 69c3e2cb-3e67-411e-aeca-6af2e6782f90 · outbound

This paper cites Armand, Divyansh Srivastava, Xiaojun Shan, Zeyuan Chen, Jianwen Xie, and Zhuowen Tu.

Learning to Generate Multiple Objects from Dense and Occluded Layouts Armand, Divyansh Srivastava, Xiaojun Shan, Zeyuan Chen, Jianwen Xie, and Zhuowen Tu

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-12T02:09:14.783011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T02:09:14.783011Z digest=sha256:fd77889d3ce661a2fa88bf94b282982d8641c1de68614a5e8677ad4158faccd1

Observation 1c02271e-0704-453b-9280-97ed2c0d62d9 · outbound

This paper cites GLIGEN: Open-set grounded text-to-image generation.

Learning to Generate Multiple Objects from Dense and Occluded Layouts GLIGEN: Open-set grounded text-to-image generation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-12T02:09:14.783011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T02:09:14.783011Z digest=sha256:ba34daf1401d1b317c7aae4e5983b3d0485f79517abe0cbb1132e3a628c5ea4e

Observation 0a561043-dbf4-4c0e-82b9-b709d647e2b8 · outbound

This paper cites Microsoft COCO: Common objects in context.

Learning to Generate Multiple Objects from Dense and Occluded Layouts Microsoft COCO: Common objects in context

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-12T02:09:14.783011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T02:09:14.783011Z digest=sha256:87d72da6ac8fabea117f19b40c9adeeeabdd67f7411d1d368c90ba2fd92dadae

Observation 1c496f91-fb53-45fa-a68c-32640b822c81 · outbound

This paper cites an unresolved cited work.

Learning to Generate Multiple Objects from Dense and Occluded Layouts Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-12T02:09:14.783011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T02:09:14.783011Z digest=sha256:0e6fc561c7bcddcb91bc392b57a071f36eb6b58d4c0c343fbc14c9cc9edd4814

Observation fb917b53-28ad-42c0-95a9-177487774d36 · outbound

This paper cites T2i-adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models.

Learning to Generate Multiple Objects from Dense and Occluded Layouts T2i-adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-12T02:09:14.783011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T02:09:14.783011Z digest=sha256:cc3a9f89bc21238e6b54080a7c1694f06c10550ca24368f109b56ca2320c5ba9

Observation 68610aff-99af-4919-99a0-790253431a02 · outbound

This paper cites Grounded text-to-image synthesis with attention refocusing.2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 7932–7942, 2024.

Learning to Generate Multiple Objects from Dense and Occluded Layouts Grounded text-to-image synthesis with attention refocusing.2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 7932–7942, 2024

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-12T02:09:14.783011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T02:09:14.783011Z digest=sha256:7c1e93944a9d072db1f121f2237b5bcefe522b4e8913f5f2bda5f29fc06de419

Observation c8be5934-5eb3-420a-b95c-8687020a0404 · outbound

This paper cites SDXL: Improving latent diffusion models for high-resolution image synthesis.

Learning to Generate Multiple Objects from Dense and Occluded Layouts SDXL: Improving latent diffusion models for high-resolution image synthesis

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-12T02:09:14.783011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T02:09:14.783011Z digest=sha256:b24920e4fb26f9b981ae308a19c3a87861ad9672755d721ba2077f64630c5eb6

Observation c910f002-b4a6-4931-8618-4922b9f6f2fe · outbound

This paper cites In- stancediffusion: Instance-level control for image generation.2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 6232–6242, 2024.

Learning to Generate Multiple Objects from Dense and Occluded Layouts In- stancediffusion: Instance-level control for image generation.2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 6232–6242, 2024

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-12T02:09:14.783011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T02:09:14.783011Z digest=sha256:bcad51ddfd38124021274a8d021e9fc982f050db7c820ddf04169c7fb72254d3

Observation f0f2bd11-c132-492b-b0ab-4665d3e5c6ea · outbound

This paper cites Adding conditional control to text-to-image diffusion models.

Learning to Generate Multiple Objects from Dense and Occluded Layouts Adding conditional control to text-to-image diffusion models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-12T02:09:14.783011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T02:09:14.783011Z digest=sha256:190dd913f8e7c1845cf61f8064e456e565f150fd25f3fd9fc4ba8f1d87e84378

Observation c449535a-7ada-459c-a6f9-eae17f653397 · outbound

This paper cites Lay- outdiffusion: Controllable diffusion model for layout-to-image generation.

Learning to Generate Multiple Objects from Dense and Occluded Layouts Lay- outdiffusion: Controllable diffusion model for layout-to-image generation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-12T02:09:14.783011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T02:09:14.783011Z digest=sha256:92cc189d30f3fcc1b6b7174e1617e1f8157186f1bc8ae41bed2f537fdb077405

Observation 80495e0e-ac8d-4d1a-adfe-6b280c9b5856 · outbound

This paper cites Migc: Multi-instance generation controller for text-to-image synthesis.

Learning to Generate Multiple Objects from Dense and Occluded Layouts Migc: Multi-instance generation controller for text-to-image synthesis

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-12T02:09:14.783011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T02:09:14.783011Z digest=sha256:fbd6ac35105d8224ba3e320a055dfef87d3d8d40ab50cfef2b5fcaaab8d9c01f

Observation 2be1af61-4229-48f0-a74c-0e077f803af8 · outbound

This paper cites 3DIS: Depth-driven decoupled image synthesis for universal multi-instance generation.

Learning to Generate Multiple Objects from Dense and Occluded Layouts 3DIS: Depth-driven decoupled image synthesis for universal multi-instance generation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-12T02:09:14.783011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T02:09:14.783011Z digest=sha256:cae9349ec85a1ee3508a7a6040cf0173716fb303e7a26819cc15b3d580c3d433

Pith citing papers

No inbound Pith citation observations are available.