Pith. sign in

Paper Citation Record · LEDGER

Multitwine: Multi-Object Compositing with Text and Layout Control

As of 10 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 1 inbound Pith citation observation for arXiv:2502.05165.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.05165 v1

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T20:05:16.256044Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:35:15.341406Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T23:35:15.923842Z

Reference resolution

58 of 58 outbound references displayed

  • verified exact0
  • verified fuzzy27
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a09b5255-306e-4e4f-874b-fa53aa34b126 · outbound

This paper cites Cross-image attention for zero- shot appearance transfer.

Multitwine: Multi-Object Compositing with Text and Layout Control Cross-image attention for zero- shot appearance transfer

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:05:17.161126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T20:05:16.003397Z digest=sha256:8e6f729a491819a58b1ff4c9f4e3142020f83010475876e1478a0a5292e2502b

Observation 9b0ac045-08d8-484c-90f3-58f2e18265de · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Multitwine: Multi-Object Compositing with Text and Layout Control Flamingo: a visual language model for few-shot learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.009424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.009424Z digest=sha256:15be98e872bc759f4a3c7bed6dd8fd0ef4587f207bb3dae24db1e1795cda46ae

Observation e1305d9a-3c8c-4819-b371-1232a3b83acc · outbound

This paper cites Vip- llava: Making large multimodal models understand arbitrary visual prompts.

Multitwine: Multi-Object Compositing with Text and Layout Control Vip- llava: Making large multimodal models understand arbitrary visual prompts

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:05:17.134681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T20:05:16.014215Z digest=sha256:47c636b1669597fcf638be3387f2256055ae5e784a880cd230b6aa53da729645

Observation e06f5dcb-2bfb-4d6e-95c5-b1b6f20e0033 · outbound

This paper cites Masactrl: Tuning-free mu- tual self-attention control for consistent image synthesis and editing.

Multitwine: Multi-Object Compositing with Text and Layout Control Masactrl: Tuning-free mu- tual self-attention control for consistent image synthesis and editing

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:05:17.119564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T20:05:16.018799Z digest=sha256:870db8cd158432c20ea193a66296d4605b830110adea77c82e94d9dbfb668204

Observation 0b29a90d-0dba-41e3-8e3e-9d8987dc54ce · outbound

This paper cites Re-Imagen: Retrieval-Augmented Text-to-Image Generator.

Multitwine: Multi-Object Compositing with Text and Layout Control Re-Imagen: Retrieval-Augmented Text-to-Image Generator

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.023288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.023288Z digest=sha256:6f77730e4ffc5243b52acd65e41c41a2590ea997b3fd0509565dd582b61648fa

Observation fc45db59-27d1-4316-8531-c3c37be8189a · outbound

This paper cites Subject-driven text-to-image generation via apprenticeship learning.

Multitwine: Multi-Object Compositing with Text and Layout Control Subject-driven text-to-image generation via apprenticeship learning

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:05:17.103540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T20:05:16.028321Z digest=sha256:39476c69568ccbc2e970b5501750e86fe81e3daebed3b1bd303eb1892d065ba5

Observation e9684e9a-4df2-466f-bb84-2292e57918ea · outbound

This paper cites AnyDoor: Zero-shot Object-level Image Customization.

Multitwine: Multi-Object Compositing with Text and Layout Control AnyDoor: Zero-shot Object-level Image Customization

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.032810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.032810Z digest=sha256:6d9f3feac01ee9cd81051efcc36c2d492f6b4836a45698448f0b6f06e4d03a28

Observation c81b1664-3105-45b1-a436-1affe30cc0f3 · outbound

This paper cites Be Yourself: Bounded Attention for Multi-Subject Text-to-Image Generation.

Multitwine: Multi-Object Compositing with Text and Layout Control Be Yourself: Bounded Attention for Multi-Subject Text-to-Image Generation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.038213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.038213Z digest=sha256:f7496fbe08d9d4e563c3aaba9c54786a60af9e40c7899073a40f92188677b47d

Observation 1d1a345a-e7fb-47e4-a468-6d2ca1e9b7ac · outbound

This paper cites DreamLLM: Synergistic Multimodal Comprehension and Creation.

Multitwine: Multi-Object Compositing with Text and Layout Control DreamLLM: Synergistic Multimodal Comprehension and Creation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.042916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.042916Z digest=sha256:6a705501573123793b59f829eafd732318ff0e15639de245695d04b790d072bb

Observation dd0abeb5-e9a7-4178-94c4-1f5c602025b3 · outbound

This paper cites PaLM-E: An Embodied Multimodal Language Model.

Multitwine: Multi-Object Compositing with Text and Layout Control PaLM-E: An Embodied Multimodal Language Model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.047459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.047459Z digest=sha256:b9a7bff752ac93a3df650986cfc8466479472ab6c32099d6c4d48bda062a2b66

Observation 837e58cb-59c3-4b1b-889b-a222bf3bd5a2 · outbound

This paper cites Scaling recti- fied flow transformers for high-resolution image synthesis.

Multitwine: Multi-Object Compositing with Text and Layout Control Scaling recti- fied flow transformers for high-resolution image synthesis

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:05:17.087545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T20:05:16.051903Z digest=sha256:17c7ad4595ca9fc3a1c9d4fd99638b7f73e21041d250cb7f985628f031eb8d73

Observation 341e031a-cf5f-497a-9792-6bcec4aacf70 · outbound

This paper cites An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion.

Multitwine: Multi-Object Compositing with Text and Layout Control An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.056096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.056096Z digest=sha256:e0bb5a6cc74d30792813f4f3f9aeabab4515a9ee5f882aa42c95d5b7c44e6da2

Observation 2aa2c571-9ff0-4ef6-92f2-19c13569d0d1 · outbound

This paper cites Prompt-to-Prompt Image Editing with Cross Attention Control.

Multitwine: Multi-Object Compositing with Text and Layout Control Prompt-to-Prompt Image Editing with Cross Attention Control

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.060539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.060539Z digest=sha256:f02c324dd9caf969bc741bd607e9a1c4799c0ce7836df323639386c23fee43c4

Observation 27d371a3-5e0a-4a91-93c2-0ed07b72e59f · outbound

This paper cites Clipscore: A reference-free evaluation met- ric for image captioning.

Multitwine: Multi-Object Compositing with Text and Layout Control Clipscore: A reference-free evaluation met- ric for image captioning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:05:17.071654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T20:05:16.064651Z digest=sha256:fea2df982d8361967adf518938b6b920bc1d26d181c7fe60c36a48c9fe430cce

Observation 05e0f83e-8f4a-4001-8fd9-3a8c887f304c · outbound

This paper cites Taming Encoder for Zero Fine-tuning Image Customization with Text-to-Image Diffusion Models.

Multitwine: Multi-Object Compositing with Text and Layout Control Taming Encoder for Zero Fine-tuning Image Customization with Text-to-Image Diffusion Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.068679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.068679Z digest=sha256:0db4bcbea216be7aaacf27a289d3f9b23fce1e97f06aa8613871af4db6e3ada8

Observation 5c2e602d-e6af-4798-b16f-8c8a4187f114 · outbound

This paper cites Rendering synthetic objects into legacy photographs.

Multitwine: Multi-Object Compositing with Text and Layout Control Rendering synthetic objects into legacy photographs

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:05:17.055767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T20:05:16.073123Z digest=sha256:e8048f6dea4ea881e907854aa8ef0b17c6812c3f4e5134b5de94750a759b7539

Observation a365751e-6298-45bf-abae-01df37cb318c · outbound

This paper cites 3d object manipulation in a single photograph using stock 3d models.

Multitwine: Multi-Object Compositing with Text and Layout Control 3d object manipulation in a single photograph using stock 3d models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:05:17.040988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T20:05:16.077249Z digest=sha256:de5865055d73aab413f6c19804918f97f76782c626854464a207d3c955de9d01

Observation fc442340-45ae-4eaf-aa15-629daae51b1b · outbound

This paper cites Gen- erating images with multimodal language models.

Multitwine: Multi-Object Compositing with Text and Layout Control Gen- erating images with multimodal language models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.081647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.081647Z digest=sha256:104270cf9bbd79b3af8fa5b0f85b97a72f8092789bfe53795e780fefe3c43500

Observation 67e4255b-c2fd-400f-a0da-e82010b9a3b4 · outbound

This paper cites Open images v5 text annotation and yet another mask text spotter.

Multitwine: Multi-Object Compositing with Text and Layout Control Open images v5 text annotation and yet another mask text spotter

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:05:17.016080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T20:05:16.086200Z digest=sha256:24bcf481c7c5d2a124e72c7179d02ef9752a437be4e23593497c88c0d129dd6e

Observation bb5a39c4-c7a1-4822-925e-b6c12b0d0de9 · outbound

This paper cites Putting people in their place: Affordance-aware hu- man insertion into scenes.

Multitwine: Multi-Object Compositing with Text and Layout Control Putting people in their place: Affordance-aware hu- man insertion into scenes

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:05:17.000101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T20:05:16.090712Z digest=sha256:b239f7cc649327bee02d0dcbb81eecc020615e32bc29416b72edeb5b8e8c78a8

Observation 5c546ea4-3ffb-4918-a06e-95704f0b3072 · outbound

This paper cites Photo clip art.

Multitwine: Multi-Object Compositing with Text and Layout Control Photo clip art

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:05:16.984562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T20:05:16.095184Z digest=sha256:89efce7db81df4b51acc3d0d1729e98127bda252a619285f40e9e79ef0f7b4cd

Observation fb1d63c8-216a-4531-bc3c-b1b7af3a797f · outbound

This paper cites Blip-diffusion: Pre- trained subject representation for controllable text-to-image generation and editing.

Multitwine: Multi-Object Compositing with Text and Layout Control Blip-diffusion: Pre- trained subject representation for controllable text-to-image generation and editing

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:05:16.970266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T20:05:16.099723Z digest=sha256:5da1fa9aa6dcc827dbdb7e327e342ccbfe58781b611b0aca835afe96ec380233

Observation ba07cc92-e700-4c65-b17c-d117e8ad268a · outbound

This paper cites UNIMO-G: Unified Image Generation through Multimodal Conditional Diffusion.

Multitwine: Multi-Object Compositing with Text and Layout Control UNIMO-G: Unified Image Generation through Multimodal Conditional Diffusion

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.103942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.103942Z digest=sha256:3634432df16cbc3cdd386f2c5bbb50b60ea7c38df848e8102a85e4ca0998cf6a

Observation 44f52a1d-a541-41f1-b616-4545299ffd2e · outbound

This paper cites Improved baselines with visual instruction tuning, 2023.

Multitwine: Multi-Object Compositing with Text and Layout Control Improved baselines with visual instruction tuning, 2023

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:05:16.955457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T20:05:16.108408Z digest=sha256:b5162986540596ef29544a9d4196b814340be6a9ffc27ca08ce6bd1ae9c6ca2c

Observation 27ca87c1-5ba6-48a3-a799-d75cc0711370 · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

Multitwine: Multi-Object Compositing with Text and Layout Control Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.112663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.112663Z digest=sha256:c048064eb670b25dae9482bd6504ab00a286ca5d22833a3486867b314e6fd406

Observation 304baa49-21d6-44ff-9d49-97752c29c1cd · outbound

This paper cites Tf-icon: Diffusion-based training-free cross-domain image composi- tion.

Multitwine: Multi-Object Compositing with Text and Layout Control Tf-icon: Diffusion-based training-free cross-domain image composi- tion

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:05:16.940940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T20:05:16.117452Z digest=sha256:fa08cd4ddac907b6ab6b125e6406c33a77c7fb2baaf1c0c4421fda862ba088a1

Observation 55d4a265-1e4a-4841-a59c-7c2027df956a · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Multitwine: Multi-Object Compositing with Text and Layout Control DINOv2: Learning Robust Visual Features without Supervision

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.121656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.121656Z digest=sha256:fdae7f225d30f5e816b90c73d63fcfa077f62813958c857ddbe5e511c00dbd35

Observation 7ee35b5d-4d11-4b72-8216-e62c5eb15e9c · outbound

This paper cites Kosmos-G: Generating Images in Context with Multimodal Large Language Models.

Multitwine: Multi-Object Compositing with Text and Layout Control Kosmos-G: Generating Images in Context with Multimodal Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.125929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.125929Z digest=sha256:8ec37991bc38fd7960c6f88345e055f7bb79112dabaf4f72d7482b64933e0f6f

Observation 05aa7bc1-d977-40cc-874c-3723d47fdd57 · outbound

This paper cites https://pixabay.com/, 2024.

Multitwine: Multi-Object Compositing with Text and Layout Control https://pixabay.com/, 2024

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:05:16.925460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T20:05:16.130241Z digest=sha256:52c309a134723d91ff9bf6f680a54ba459a49bb7180f1c8282e8a072e512974a

Observation 2f083b3d-2002-4b00-b2e3-be4473e9468f · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Multitwine: Multi-Object Compositing with Text and Layout Control SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.134464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.134464Z digest=sha256:c13697db7b70b367ff22e10407002eecf2701985c8a059cf90b1c3a13c67e7e6

Observation 04b7e7bd-be0f-4737-a9bc-320e07c1cee5 · outbound

This paper cites High-Quality Entity Segmentation.

Multitwine: Multi-Object Compositing with Text and Layout Control High-Quality Entity Segmentation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.138710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.138710Z digest=sha256:0fe1d8e5c1760697b25c7644dec0007448ec0cc144a6be3d3012221554a6de43

Observation ddc5b5da-451c-4048-8daa-bb80f0dc28b5 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Multitwine: Multi-Object Compositing with Text and Layout Control Learning transferable visual models from natural language supervi- sion

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.143110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.143110Z digest=sha256:dd5dac4783ec2a0d2093f5670d936ac3ca52ddf2f5816c8f82c26a1b989ee0b9

Observation fbf8423c-7629-4a99-b439-08159bd32e96 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Multitwine: Multi-Object Compositing with Text and Layout Control High-resolution image synthesis with latent diffusion models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.147319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.147319Z digest=sha256:64738e0fb4e85c67966400efe9d786cca0774b3d5385dbdb37598653b423a974

Observation bf7d524c-438c-4870-87c3-c5bf63b1fdfb · outbound

This paper cites Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation.

Multitwine: Multi-Object Compositing with Text and Layout Control Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:05:16.888041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T20:05:16.151779Z digest=sha256:b15288cfdbb8eb35953eb5cfbaa418081eea39366fe16aff9ab75219fe07bd3d

Observation d800c370-6ce3-41b4-8b68-a767ed7aaa08 · outbound

This paper cites Diffuse to Choose: Enriching Image Conditioned Inpainting in Latent Diffusion Models for Virtual Try-All.

Multitwine: Multi-Object Compositing with Text and Layout Control Diffuse to Choose: Enriching Image Conditioned Inpainting in Latent Diffusion Models for Virtual Try-All

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.156007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.156007Z digest=sha256:1d64bb43b2ab55a879197734644d9aef7bb9645bc23b4c049b850c008d40d24e

Observation 355e135b-1342-4b10-aede-3384f133622d · outbound

This paper cites Video visual relation detection.

Multitwine: Multi-Object Compositing with Text and Layout Control Video visual relation detection

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:05:16.871595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T20:05:16.160475Z digest=sha256:fa135c214c91941544291ae026172f8137b2ba071e91f497eeae1bb0ad5378a1

Observation 28cf2d92-51e0-480e-8eff-dd3d8276fdd4 · outbound

This paper cites Annotating objects and relations in user- generated videos.

Multitwine: Multi-Object Compositing with Text and Layout Control Annotating objects and relations in user- generated videos

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:05:16.856146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T20:05:16.164892Z digest=sha256:4b4839731cb532178d3a544eaef8c70988ca0eb72b54766b590aa8bed7b4318e

Observation 3304e5c2-fee3-44c9-bbde-0da93e303fdb · outbound

This paper cites InstantBooth: Personalized Text-to-Image Generation without Test-Time Finetuning.

Multitwine: Multi-Object Compositing with Text and Layout Control InstantBooth: Personalized Text-to-Image Generation without Test-Time Finetuning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.169140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.169140Z digest=sha256:b3d4fb0cb912849f4cdbc0a1f685a371032c0ed5b01de08972416151882a8d30

Observation 69b55bb5-e0e3-454d-85c4-d8487227bf77 · outbound

This paper cites ObjectStitch: Generative Object Compositing.

Multitwine: Multi-Object Compositing with Text and Layout Control ObjectStitch: Generative Object Compositing

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.173596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.173596Z digest=sha256:1b84cc97f4270f9c5bab135789ebd3b0ffc4408a4d360458d483aa0a3fae9005

Observation 607a3426-bca4-4f7d-ad04-7fcc67d41958 · outbound

This paper cites Imprint: Generative object compositing by learning identity-preserving representation.

Multitwine: Multi-Object Compositing with Text and Layout Control Imprint: Generative object compositing by learning identity-preserving representation

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:05:16.840319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T20:05:16.177752Z digest=sha256:f72c3578d10481974d7857c9cf6b9f449f26469a88e3c1478880034261436701

Observation 483e5ba2-0012-4a11-8aa4-8364c98eb43b · outbound

This paper cites Emu: Generative Pretraining in Multimodality.

Multitwine: Multi-Object Compositing with Text and Layout Control Emu: Generative Pretraining in Multimodality

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.182015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.182015Z digest=sha256:c43f8829b347dd9634b13d106cee582c6fc96277c650ba3cf5c4a6bffd116527

Observation d5d21661-b75f-462b-bbde-c39f0e36d0be · outbound

This paper cites Generative multimodal mod- els are in-context learners.

Multitwine: Multi-Object Compositing with Text and Layout Control Generative multimodal mod- els are in-context learners

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:05:16.825417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T20:05:16.186695Z digest=sha256:e8d4f8defc9c98a352d9eed54ed1a9aab67f13905ba65cb2736330f5a4f1b682

Observation 58db0e90-c271-4b15-aec1-08024664cc3f · outbound

This paper cites Thinking Outside the BBox: Unconstrained Generative Object Compositing.

Multitwine: Multi-Object Compositing with Text and Layout Control Thinking Outside the BBox: Unconstrained Generative Object Compositing

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.190538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.190538Z digest=sha256:173e854fb6df9a4c1ed21c1329b8476210a389e80cae920faeb1e3c5adaa58ae

Observation 762dcc6a-bb1d-4db7-9f9b-e93c91d3ba83 · outbound

This paper cites CogVLM: Visual Expert for Pretrained Language Models.

Multitwine: Multi-Object Compositing with Text and Layout Control CogVLM: Visual Expert for Pretrained Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.194518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.194518Z digest=sha256:b894c4523a1b514e884ea35fbd346d60140d32dcc4c8313a7fa0a8c40feb05fa

Observation c1e141ef-52ce-4b42-be37-33b9a6a7d42b · outbound

This paper cites Elite: Encoding visual con- cepts into textual embeddings for customized text-to-image generation.

Multitwine: Multi-Object Compositing with Text and Layout Control Elite: Encoding visual con- cepts into textual embeddings for customized text-to-image generation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.198714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.198714Z digest=sha256:ffa79d45a706c33b051bd0faa63f2ab1423105797d0a8951d7affb861fd2eeb2

Observation dbae3f95-8114-431e-bcab-e2b2cf83b116 · outbound

This paper cites Fastcomposer: Tuning-free multi- subject image generation with localized attention.

Multitwine: Multi-Object Compositing with Text and Layout Control Fastcomposer: Tuning-free multi- subject image generation with localized attention

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:05:16.799738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T20:05:16.202729Z digest=sha256:a1820a5cd2e430b9e0b40a07e1d9d4a62602ccdb3822b79e6e4ff03f672e3a54

Observation 2f9f0806-d127-4b70-bb66-5b3301e8a885 · outbound

This paper cites GroundingBooth: Grounding Text-to-Image Customization.

Multitwine: Multi-Object Compositing with Text and Layout Control GroundingBooth: Grounding Text-to-Image Customization

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.206770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.206770Z digest=sha256:fc001f7c3c8e48e6cceadd032ad5aad01527dadd8a9b55b86fd7286c9a782c60

Observation b051d4b7-f7c2-40d4-9b8d-aab3aaf5e34e · outbound

This paper cites Paint by example: Exemplar-based image editing with diffusion mod- els.

Multitwine: Multi-Object Compositing with Text and Layout Control Paint by example: Exemplar-based image editing with diffusion mod- els

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.211720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.211720Z digest=sha256:8a4ce97ee8cd7906066c47fbcc5962af7c6b9de7ab352746441468d9f6ed5cfe

Observation f61748aa-6e5d-4c4e-b0bb-5bedaee888c9 · outbound

This paper cites CustomNet: Zero-shot Object Customization with Variable-Viewpoints in Text-to-Image Diffusion Models.

Multitwine: Multi-Object Compositing with Text and Layout Control CustomNet: Zero-shot Object Customization with Variable-Viewpoints in Text-to-Image Diffusion Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.216485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.216485Z digest=sha256:1d06b2024991a484acf5e0c175f44d74730e0efd44681ba51ef2d55b0ab1f96e

Observation f381e241-b1f2-44ae-bfa9-92faa49e1e97 · outbound

This paper cites ControlCom: Controllable Image Composition using Diffusion Model.

Multitwine: Multi-Object Compositing with Text and Layout Control ControlCom: Controllable Image Composition using Diffusion Model

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.220898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.220898Z digest=sha256:e4479e63fabdad19d383585837c3a0596536a3ba615a176929136e53c4fa23b6

Observation caa0283f-05fd-4101-ad25-87b5b5c818df · outbound

This paper cites LoCo: Locally Constrained Training-Free Layout-to-Image Synthesis.

Multitwine: Multi-Object Compositing with Text and Layout Control LoCo: Locally Constrained Training-Free Layout-to-Image Synthesis

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.225176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.225176Z digest=sha256:74c5a69f762fb2bf9de55594df7084b4c20ec130475061c2a783042183c33418

Observation 1a560b2d-248f-499f-a8f0-9d9bf0d2d100 · outbound

This paper cites Can you provide a grammatically correct one-line caption for the relation <object A> <relation> <object B> in the image?.

Multitwine: Multi-Object Compositing with Text and Layout Control Can you provide a grammatically correct one-line caption for the relation <object A> <relation> <object B> in the image?

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:05:16.774028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T20:05:16.229501Z digest=sha256:f68b68daf68058c57eab403f3b6bab218fd85fc4eca6ac50ced4ca797151229a

Observation 8a454c94-da25-417d-9c2f-aa9ef7a0911a · outbound

This paper cites Figure 4.

Multitwine: Multi-Object Compositing with Text and Layout Control Figure 4

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:05:16.758388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T20:05:16.234124Z digest=sha256:747cdfaef4fb322fc7e869214b9c8e1586b388e678d1b37841173d08bda9eb78

Observation 6594bd52-d1fc-4891-8733-bd6287635ded · outbound

This paper cites used to extract them.

Multitwine: Multi-Object Compositing with Text and Layout Control used to extract them

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:05:16.740942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T20:05:16.238462Z digest=sha256:d66be1efc40e21b8288fde9a2da859f3bb4a1b6752f89adfe97755e923696a0a

Observation e4a496a8-e693-4bfe-919b-41542389de35 · outbound

This paper cites Background images are sourced from Pixabay [29], while objects are from Pixabay [29], MultiBench [23], and DreamBooth [34].

Multitwine: Multi-Object Compositing with Text and Layout Control Background images are sourced from Pixabay [29], while objects are from Pixabay [29], MultiBench [23], and DreamBooth [34]

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:05:16.725567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T20:05:16.242546Z digest=sha256:7139875bea234e4f7c35541585cea563dc5615c0b4ac412e8831f444cdd89ffe

Observation 6a58adb2-d0ad-42c8-ad3e-6fb483ece5b6 · outbound

This paper cites Further details on user studies can be found in Section 3.3.

Multitwine: Multi-Object Compositing with Text and Layout Control Further details on user studies can be found in Section 3.3

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:05:16.710737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T20:05:16.246690Z digest=sha256:b456d63ca16027150811a8b2429f41ddf6b1ade65259cc59617ee142b67367b9

Observation 8cbbc66a-c750-4349-91f2-6ecbaa65528c · outbound

This paper cites Without multi-view data ( i.e., video data, manually collected data), the model struggles to prop- erly repose and combine objects to align with the textual description.

Multitwine: Multi-Object Compositing with Text and Layout Control Without multi-view data ( i.e., video data, manually collected data), the model struggles to prop- erly repose and combine objects to align with the textual description

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:05:16.694958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T20:05:16.251155Z digest=sha256:207f2c135746a043ed1ae57355123a86e0ce6d4a8d26495793c5bb75258fc18b

Observation 6d49043d-648e-4ee1-a8ae-04f25ddf9c8c · outbound

This paper cites an unresolved cited work.

Multitwine: Multi-Object Compositing with Text and Layout Control Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-08T20:05:16.677728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T20:05:16.256044Z digest=sha256:a3a9ca459494c10fea277d78072ac6d2aa235216519a5020b72dcf8c25877481

Pith citing papers

Observation 58b4d835-8f33-4c81-bfa6-ae638a5cc0b7 · inbound

BlenderFusion: 3D-Grounded Visual Editing and Generative Compositing cites this paper.

BlenderFusion: 3D-Grounded Visual Editing and Generative Compositing Multitwine: Multi-Object Compositing with Text and Layout Control

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-08-06T23:35:15.948136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T23:35:15.341406Z digest=sha256:d5c0ee0beb4f122a470be073076e37e751f607a8573ed88031d2b384ecb90a7f