Pith. sign in

Paper Citation Record · LEDGER

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions

As of 20 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 0 inbound Pith citation observations for arXiv:2501.12173.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.12173 v1

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T17:30:48.007551Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

54 of 54 outbound references displayed

  • verified exact1
  • verified fuzzy27
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d490b31e-3d0d-4c43-b68c-25bcb4026d10 · outbound

This paper cites Spatext: Spatio-textual representation for con- trollable image generation.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions Spatext: Spatio-textual representation for con- trollable image generation

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:30:48.686884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T17:30:47.768865Z digest=sha256:7c950110a681a1f3fa4ddf70b6765929a77d4e54fb98e066b0d95ed36128aa2f

Observation df8c0798-d017-47ee-be16-10ae77b31b18 · outbound

This paper cites MultiDiffusion: Fusing Diffusion Paths for Controlled Image Generation.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions MultiDiffusion: Fusing Diffusion Paths for Controlled Image Generation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T17:30:47.773975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:30:47.773975Z digest=sha256:44854f091fc297d8ce0a2af8e82d61ffca984a23756b26a281666f311599ec27

Observation ba481bdf-bd9b-4bab-8d27-4b19380350ed · outbound

This paper cites Sutherland, Michael Arbel, and Arthur Gretton.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions Sutherland, Michael Arbel, and Arthur Gretton

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:30:48.673792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T17:30:47.778737Z digest=sha256:70fde3a429c89c99dc4589c53c861576f13fc4214f87929edaaa347adac41a7b

Observation e15659c6-0fe5-466e-a797-2181377d1e4c · outbound

This paper cites InstructPix2Pix: Learning to Follow Image Editing Instructions.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions InstructPix2Pix: Learning to Follow Image Editing Instructions

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T17:30:47.782994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:30:47.782994Z digest=sha256:154746ecc5f1bbf145195a1544316d9ad93ad9d5e633691c50957202b4b0829f

Observation a7e13581-a7e8-4a81-a5d6-46a2e4255ddd · outbound

This paper cites PhotoVerse: Tuning-Free Image Customization with Text-to-Image Diffusion Models.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions PhotoVerse: Tuning-Free Image Customization with Text-to-Image Diffusion Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T17:30:47.787809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:30:47.787809Z digest=sha256:581dc0d337b062e8cce4fafa6c97f2a2f52d2886f618dda3dedec20f5cf9d065

Observation b51779b0-dd86-4f76-8c11-dd7cd6fe8f80 · outbound

This paper cites Training-free layout control with cross-attention guidance.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions Training-free layout control with cross-attention guidance

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T17:30:47.792688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:30:47.792688Z digest=sha256:7c37a5f13892ffc6802ad85abdd4af0c8949575d8b31cbc54868e836fcfe2183

Observation 64544040-90d9-42c6-b024-18658ba570df · outbound

This paper cites AnyDoor: Zero-shot Object-level Image Customization.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions AnyDoor: Zero-shot Object-level Image Customization

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T17:30:47.798320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:30:47.798320Z digest=sha256:417318a982b2c1695fa358fcbd6cad9a93d701ee74892153ef04c4d65845c38e

Observation 84e5c1ab-d65e-4699-993f-ed758ebd7b9d · outbound

This paper cites LayoutDiffuse: Adapting Foundational Diffusion Models for Layout-to-Image Generation.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions LayoutDiffuse: Adapting Foundational Diffusion Models for Layout-to-Image Generation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T17:30:47.802873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:30:47.802873Z digest=sha256:00ec22e0877dd2a1f14d0ad0b0ee67486b8e050743e9492486f28f1aca5a1a51

Observation 74322a34-e077-4956-a9bb-5c5d8e95b8ae · outbound

This paper cites KPE: Keypoint Pose Encoding for Transformer-based Image Generation.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions KPE: Keypoint Pose Encoding for Transformer-based Image Generation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T17:30:47.807837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:30:47.807837Z digest=sha256:a4280491143507b3d93a7639b8dce9d105d5926b90a773da2a0ff83b1b4aecec

Observation 7463dbd3-605e-4864-920a-c7f8f865d4b9 · outbound

This paper cites Viton-hd: High-resolution virtual try-on via misalignment-aware normalization.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions Viton-hd: High-resolution virtual try-on via misalignment-aware normalization

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:30:48.653415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T17:30:47.812542Z digest=sha256:848a01dbf9259f223ac3149ca7ae7effbd85ce16b81a372bc7335149d3ad6e2a

Observation a75b5cc1-77df-4446-a201-24c77098d514 · outbound

This paper cites Measures of the amount of ecologic association between species.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions Measures of the amount of ecologic association between species

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:30:48.641094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T17:30:47.816586Z digest=sha256:6b280b1008b73a66baf8d3ee70b0bcb54a8b3a8bba9dc0731bab54db054f8058

Observation f4ce4cb4-362e-4645-b07b-e210759a69ec · outbound

This paper cites Stylegan-human: A data-centric odyssey of human genera- tion.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions Stylegan-human: A data-centric odyssey of human genera- tion

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:30:48.628108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T17:30:47.820786Z digest=sha256:e6fdad0acf609e9b2f4c2dc6f43fd02a10549435a1b5a3d51f93213c8773e48e

Observation c72b606b-1c4d-4089-bd16-dd1af06003e9 · outbound

This paper cites An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T17:30:47.825923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:30:47.825923Z digest=sha256:032a0e617add097819dbb9035b4f99e3e0e2cc63446106661b635137c283066d

Observation ba458290-f62a-4d56-9c3a-441c41d42976 · outbound

This paper cites A versatile benchmark for de- tection, pose estimation, segmentation and re-identification of clothing images.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions A versatile benchmark for de- tection, pose estimation, segmentation and re-identification of clothing images

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:30:48.615396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T17:30:47.830402Z digest=sha256:7d60911bf4d9b6fbe273522b22bd8b408b49ea839efba02637683cb529654d71

Observation 6536aca5-b052-495a-9982-daf7022e9f41 · outbound

This paper cites Context- aware layout to image generation with enhanced object ap- pearance.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions Context- aware layout to image generation with enhanced object ap- pearance

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:30:48.602017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T17:30:47.835299Z digest=sha256:f5e7370583e3a3346bd022bb9475f6f0c40db4453ca56d8852f86ea136a944bf

Observation 75ae43e7-c555-42b4-af0a-fe35d4fb7642 · outbound

This paper cites Denoising dif- fusion probabilistic models.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions Denoising dif- fusion probabilistic models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T17:30:47.839765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:30:47.839765Z digest=sha256:7eb4de52bdc67943c63fdefb72b2111d697257e6f78aa23b9bd26d9f54f390d3

Observation e0be945d-b4ad-488d-954a-f4da9bd4ba9b · outbound

This paper cites CogVLM2: Visual Language Models for Image and Video Understanding.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions CogVLM2: Visual Language Models for Image and Video Understanding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T17:30:47.844735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:30:47.844735Z digest=sha256:3cc3189f9a59fb59b41dac6e8efd0d61a5432621f77fcd0d2d56ea793bb7e357

Observation 66cc599d-d0ce-48dd-bff7-158d4bc5a32d · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions LoRA: Low-Rank Adaptation of Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T17:30:47.849204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:30:47.849204Z digest=sha256:fe718b24c45b479f6ec3dfc16aa8b01d75c12509a5cf6e89d192d4b0c955bc14

Observation d177a0fb-e45b-4063-aacc-13e6ea135bf0 · outbound

This paper cites ´Etude comparative de la distribution florale dans une portion des alpes et des jura.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions ´Etude comparative de la distribution florale dans une portion des alpes et des jura

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:30:48.580058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T17:30:47.854040Z digest=sha256:59d353257d830c0cd0814f4d5db65d016772c98aa52faf8d62508fdee0b2f50e

Observation ee5ccada-a3ec-471c-9e24-3d1f12f62f34 · outbound

This paper cites Text2human: Text-driven controllable human image generation.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions Text2human: Text-driven controllable human image generation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:30:48.567316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T17:30:47.858407Z digest=sha256:37163436595b7e0e2b7fe9bf5593b6156b5f3af2f2348633ba36ea28ef5db9ee

Observation fd99474c-d18a-49c5-8832-fbd29d4872e3 · outbound

This paper cites Dense text-to-image generation with attention modulation.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions Dense text-to-image generation with attention modulation

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:30:48.554819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T17:30:47.862430Z digest=sha256:5fb0fe8b9112afde41dece859d599f105a4d5b5ffdc6dd2d1119c3322ad192f3

Observation e86d74a4-fa59-48f2-adf9-23d908bed295 · outbound

This paper cites Auto-Encoding Variational Bayes.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions Auto-Encoding Variational Bayes

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T17:30:47.866740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:30:47.866740Z digest=sha256:5b64558608b9edae4ff32aa860e35dfbe5f69775be2b1235b2e8329f722a75f3

Observation c50343be-f86d-46d3-9e95-fb61a07c58c5 · outbound

This paper cites Segment Anything.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions Segment Anything

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T17:30:47.871094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:30:47.871094Z digest=sha256:c84e6cda4cb9657f39230cd56e69b2adb731c4feee109d581feca21ea53b7b46

Observation 9a33e0e4-b2cd-445b-ad39-14aa23fdc43c · outbound

This paper cites Multi-concept customization of text-to-image diffusion.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions Multi-concept customization of text-to-image diffusion

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T17:30:47.875119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:30:47.875119Z digest=sha256:bd29833c65dd8dc47a87ec5f5074e37ed7cd651069be6de3f0284d7478669a26

Observation b658aa63-f2b8-4934-8adc-0ad7dbfcf72c · outbound

This paper cites Self- correction for human parsing.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions Self- correction for human parsing

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T17:30:47.879911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:30:47.879911Z digest=sha256:acfb547eee6d2ed83c0818ce81ff3c94e50d19f933c3a6224088ec238062a447

Observation 94595758-8b25-4f86-8660-f8c8f87ae733 · outbound

This paper cites Gligen: Open-set grounded text-to-image generation.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions Gligen: Open-set grounded text-to-image generation

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:30:48.527010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T17:30:47.883877Z digest=sha256:b3487b7a3809272a5cd81a35486f50657f51210d4ef813aa68a06016226e28be

Observation 71814155-4120-4460-bfe9-62f5873f8a01 · outbound

This paper cites Image synthesis from layout with locality- aware mask adaption.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions Image synthesis from layout with locality- aware mask adaption

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:30:48.514255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T17:30:47.889188Z digest=sha256:53d409816c3e1e2dca73da3a69c3c2ff3f4abd519664b29c7b8b3c77924ab579

Observation d292a7d4-5b9c-433c-ab03-2970a6ca3eb8 · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T17:30:47.894131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:30:47.894131Z digest=sha256:8e82d86bc4955d14aa537c1e05203da742be2959ba4aabb5427a7068e9ac1b93

Observation 1438a259-7aa5-4b42-8dcb-6d800e6bf84b · outbound

This paper cites Dress code: High- resolution multi-category virtual try-on, 2022.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions Dress code: High- resolution multi-category virtual try-on, 2022

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:30:48.501571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T17:30:47.901085Z digest=sha256:d5f9a905a1a2bdffd0adf52f583d4aedddec76285ecde64fd57afa1101186afc

Observation f07dde4c-a53b-47c4-832f-dc8412757e0f · outbound

This paper cites $\lambda$-ECLIPSE: Multi-Concept Personalized Text-to-Image Diffusion Models by Leveraging CLIP Latent Space.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions $\lambda$-ECLIPSE: Multi-Concept Personalized Text-to-Image Diffusion Models by Leveraging CLIP Latent Space

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T17:30:47.906551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:30:47.906551Z digest=sha256:78b8bfd281c44285d65809e232c8002f6c552572ec60267167d28ce454e9f921

Observation 6c705960-f287-4452-90e0-032760beeb40 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions Learning transferable visual models from natural language supervi- sion

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T17:30:47.911049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:30:47.911049Z digest=sha256:c855d3aa423a1c81bfeceffc2249a9bfd604e77c2f2b0b724dc8957633b2ed9d

Observation 50d1761e-a282-46e0-bc46-e26e004913e6 · outbound

This paper cites High-resolution image syn- thesis with latent diffusion models, 2021.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions High-resolution image syn- thesis with latent diffusion models, 2021

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:30:48.479983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T17:30:47.915124Z digest=sha256:1b898a5cc702d870d7080d56293aef069a3c77a8fa2858c60b190b633d31fca7

Observation f680884f-6307-44e3-80dc-cb74615c614b · outbound

This paper cites Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T17:30:47.918826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:30:47.918826Z digest=sha256:61af05b1396cab397a77d74e32e4955bcdb369256527bf86d1fff92fe56f4369

Observation 30885056-b095-4de1-8e94-778a40c7062f · outbound

This paper cites Humangan: A generative model of hu- man images.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions Humangan: A generative model of hu- man images

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:30:48.459736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T17:30:47.924213Z digest=sha256:c82f5a059c40d60300cfed7ac7ba87f99a60d746c56fdcbae1a186aea3f64751

Observation f95a8c05-5756-442a-8ad7-55e739b573ce · outbound

This paper cites pytorch-fid: FID Score for PyTorch.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions pytorch-fid: FID Score for PyTorch

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T17:30:47.928078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:30:47.928078Z digest=sha256:e000c624d033d8729b815f21252bd46049b83bb821be71e14e281f9ab7c95ac0

Observation 1f523453-f07d-4eea-bfd7-c5e5ae280c68 · outbound

This paper cites In- stantbooth: Personalized text-to-image generation without test-time finetuning.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions In- stantbooth: Personalized text-to-image generation without test-time finetuning

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:30:48.438335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T17:30:47.932515Z digest=sha256:41dd0ab058449d7e5a56d7060005e34b6ba822d5557ab3039ca607bff07d4fa8

Observation cc25dca7-73fc-4be4-82c6-f8d7fcdf2ad4 · outbound

This paper cites Image synthesis from reconfig- urable layout and style.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions Image synthesis from reconfig- urable layout and style

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:30:48.425427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T17:30:47.936451Z digest=sha256:9315be324655ccdffc731aed8824923652ee35a990a147bb039fafc8e5fda56e

Observation df2be0fc-bf8b-4631-9efb-cdf5184e4130 · outbound

This paper cites Object-centric image genera- tion from layouts.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions Object-centric image genera- tion from layouts

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:30:48.412497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T17:30:47.940545Z digest=sha256:7c2f5b17f58feffad6a4936913da0cf33af43b9e91b59e45ce8c4d02de6d51d6

Observation 236f6967-d908-4754-a255-3195ece747eb · outbound

This paper cites Interactive image synthesis with panoptic layout generation.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions Interactive image synthesis with panoptic layout generation

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:30:48.398448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T17:30:47.944723Z digest=sha256:617afba4e6c41670d1c5d1ec7bc29a739a3a2a2c3265d2f786705df6817bec45

Observation 99752318-c0cb-4658-902b-360198564b6e · outbound

This paper cites InstantID: Zero-shot Identity-Preserving Generation in Seconds.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions InstantID: Zero-shot Identity-Preserving Generation in Seconds

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T17:30:47.949730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:30:47.949730Z digest=sha256:250b5675b0411d74c847c8c88390f96b3b074e5bc5bb13fb590971984e7fdc60

Observation 485acd33-d9bb-4334-87ce-05158caaa609 · outbound

This paper cites Instancediffusion: Instance-level control for image generation, 2024.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions Instancediffusion: Instance-level control for image generation, 2024

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:30:48.385731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T17:30:47.954162Z digest=sha256:7593ca12ce529989d5b9cfc9019b6846e24aefd0e8193fc559ae6becade7cca1

Observation a2f20f46-d721-4072-aa23-3776efa3ea22 · outbound

This paper cites Image quality assessment: from error visibility to structural similarity.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions Image quality assessment: from error visibility to structural similarity

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:30:48.372765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T17:30:47.958179Z digest=sha256:6004226f7a72152b9f643664382552a613cb8c1099e6d83f97b887b68fbd3424

Observation 43066a6f-3058-4fa9-8fab-dc8d4c3bc853 · outbound

This paper cites ELITE: Encoding Visual Concepts into Textual Embeddings for Customized Text-to-Image Generation.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions ELITE: Encoding Visual Concepts into Textual Embeddings for Customized Text-to-Image Generation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T17:30:47.961855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:30:47.961855Z digest=sha256:4998ea00901bf3357d208bbdba5d13e13272adb22ea8cd1db3247ef49e1bd01d

Observation 414dc079-dcff-402a-b380-7eb015485c60 · outbound

This paper cites Fastcomposer: Tuning-free multi- subject image generation with localized attention.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions Fastcomposer: Tuning-free multi- subject image generation with localized attention

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:30:48.359668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T17:30:47.966274Z digest=sha256:966e11e87fcfedfb59717b9db4a8f5aab6edee8b483cf6d3445ae9e06ab2561e

Observation 04342a8f-1d53-4922-b7e2-4683bc520249 · outbound

This paper cites R&B: Region and Boundary Aware Zero-shot Grounded Text-to-image Generation.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions R&B: Region and Boundary Aware Zero-shot Grounded Text-to-image Generation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T17:30:47.970015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:30:47.970015Z digest=sha256:83335ba0cc2870b731b41bd63bc8e2e9f7e48ae7b6137e1f78b973ea9e738c19

Observation 46f84a73-4883-4025-8cd0-92a06168b2fc · outbound

This paper cites Boxdiff: Text-to-image synthesis with training-free box-constrained diffusion.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions Boxdiff: Text-to-image synthesis with training-free box-constrained diffusion

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:30:48.346658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T17:30:47.974418Z digest=sha256:c6604673ccbc7996780dc74cd5a87d23ba13699ca96b5c084cb82d0d5db30435

Observation 78a9c7b3-b863-49bf-9550-172e70599253 · outbound

This paper cites Reco: Region-controlled text-to-image genera- tion.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions Reco: Region-controlled text-to-image genera- tion

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T17:30:47.978379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:30:47.978379Z digest=sha256:dc60a290dd45829cb0dae4323a5fc37e95f3d5b5c525ce9e1b257fc477e22ce1

Observation 038fce44-f971-4687-a31e-755c0091c412 · outbound

This paper cites IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T17:30:47.982276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:30:47.982276Z digest=sha256:4cc8ffecec70334b17114033881a63618cb3f56d179245e873da046128c60ded

Observation e5f886bb-68c3-462f-b0e4-f26e54b30344 · outbound

This paper cites Customnet: Zero-shot object customization with variable-viewpoints in text-to-image dif- fusion models, 2023.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions Customnet: Zero-shot object customization with variable-viewpoints in text-to-image dif- fusion models, 2023

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:30:48.326113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T17:30:47.986643Z digest=sha256:1074f44d4e0d960254f9a5b6d7f9cb610b9d14a6ec4e7c16162899d914f2df1a

Observation ba63e5ea-e2e2-4041-96e5-31f3b0664d9b · outbound

This paper cites HumanDiffusion: a Coarse-to-Fine Alignment Diffusion Framework for Controllable Text-Driven Person Image Generation.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions HumanDiffusion: a Coarse-to-Fine Alignment Diffusion Framework for Controllable Text-Driven Person Image Generation

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-08-10T17:30:48.049415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T17:30:47.991487Z digest=sha256:60f000e5491a87836b36ab21631094ae7a6e3f76371fbd046abd3a78c995bf55

Observation 784dc3f6-544e-41cb-874d-bae8c2f5aef7 · outbound

This paper cites The unreasonable effectiveness of deep features as a perceptual metric.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions The unreasonable effectiveness of deep features as a perceptual metric

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T17:30:47.995603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:30:47.995603Z digest=sha256:fb6fd27892740fee898aa62cdb5e78a21203110429c42679e585eedb6b65ca3d

Observation 544697bd-b9d8-430b-a76c-8fa90abb9735 · outbound

This paper cites Layoutdiffusion: Controllable diffu- sion model for layout-to-image generation.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions Layoutdiffusion: Controllable diffu- sion model for layout-to-image generation

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:30:48.305842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T17:30:47.999705Z digest=sha256:b400f85e3ad2107f86b289547c00fb1560b005df2de225375d3cdf6b33887b4f

Observation 70dcb08b-9447-4c06-9930-be3b2f0b5f9b · outbound

This paper cites clip-score: CLIP Score for Py- Torch.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions clip-score: CLIP Score for Py- Torch

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:30:48.292780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T17:30:48.003474Z digest=sha256:53e93c8a01065674c081f8a7ff4c422f94b01c2a8382e6a15144ad6ba4475981

Observation f573d113-9743-4574-a9d8-2b343db30785 · outbound

This paper cites Migc: Multi-instance generation controller for text-to-image synthesis, 2024.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions Migc: Multi-instance generation controller for text-to-image synthesis, 2024

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:30:48.280060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T17:30:48.007551Z digest=sha256:f06f77c731130823e5eaa85358882e4ebce365ad5903386b21ec5c1f875b8317

Pith citing papers

No inbound Pith citation observations are available.