Pith. sign in

Paper Citation Record · LEDGER

Multitwine: Multi-Object Compositing with Text and Layout Control

As of 9 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 1 inbound Pith citation observation for arXiv:2502.05165.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.05165 v1

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T20:05:16.256044Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:35:15.341406Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T23:35:15.923842Z

Reference resolution

58 of 58 outbound references displayed

  • verified exact0
  • verified fuzzy27
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a09b5255-306e-4e4f-874b-fa53aa34b126 · outbound

This paper cites Cross-image attention for zero- shot appearance transfer.

Multitwine: Multi-Object Compositing with Text and Layout Control Cross-image attention for zero- shot appearance transfer

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:05:17.161126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:05:16.003397Z digest=sha256:07f6796b5867cc97a3eb304ec18b4cec3378e0efb2ad54d537856271fd09aea0

Observation 9b0ac045-08d8-484c-90f3-58f2e18265de · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Multitwine: Multi-Object Compositing with Text and Layout Control Flamingo: a visual language model for few-shot learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.009424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.009424Z digest=sha256:002dcaf4fef1f5a7ac9a7be0b4ff49aff1dcec3968b45e7baef7ae5f0e0ea67e

Observation e1305d9a-3c8c-4819-b371-1232a3b83acc · outbound

This paper cites Vip- llava: Making large multimodal models understand arbitrary visual prompts.

Multitwine: Multi-Object Compositing with Text and Layout Control Vip- llava: Making large multimodal models understand arbitrary visual prompts

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:05:17.134681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:05:16.014215Z digest=sha256:5b4a49500610b2a57ccfc4f3782be910a55e3c67dadc1eabc8b3d5905f54299b

Observation e06f5dcb-2bfb-4d6e-95c5-b1b6f20e0033 · outbound

This paper cites Masactrl: Tuning-free mu- tual self-attention control for consistent image synthesis and editing.

Multitwine: Multi-Object Compositing with Text and Layout Control Masactrl: Tuning-free mu- tual self-attention control for consistent image synthesis and editing

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:05:17.119564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:05:16.018799Z digest=sha256:d8fdc99ef9b49098b83ba21a5192dd39bdbfafa914c84acde4416d52136843c8

Observation 0b29a90d-0dba-41e3-8e3e-9d8987dc54ce · outbound

This paper cites Re-Imagen: Retrieval-Augmented Text-to-Image Generator.

Multitwine: Multi-Object Compositing with Text and Layout Control Re-Imagen: Retrieval-Augmented Text-to-Image Generator

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.023288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.023288Z digest=sha256:e5233bc24db35c602195c24fe59962355a5e9a83f123a6ea1c708fb51bd132fa

Observation fc45db59-27d1-4316-8531-c3c37be8189a · outbound

This paper cites Subject-driven text-to-image generation via apprenticeship learning.

Multitwine: Multi-Object Compositing with Text and Layout Control Subject-driven text-to-image generation via apprenticeship learning

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:05:17.103540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:05:16.028321Z digest=sha256:eb18d1dae66d3fa0cc867b93c69fa62292752d5dec53b4a7e8e013c990715251

Observation e9684e9a-4df2-466f-bb84-2292e57918ea · outbound

This paper cites AnyDoor: Zero-shot Object-level Image Customization.

Multitwine: Multi-Object Compositing with Text and Layout Control AnyDoor: Zero-shot Object-level Image Customization

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.032810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.032810Z digest=sha256:ce4c8183b2d3be9fc14ab63944b4daf7a7f898e018ee1172297834eec2addad5

Observation c81b1664-3105-45b1-a436-1affe30cc0f3 · outbound

This paper cites Be Yourself: Bounded Attention for Multi-Subject Text-to-Image Generation.

Multitwine: Multi-Object Compositing with Text and Layout Control Be Yourself: Bounded Attention for Multi-Subject Text-to-Image Generation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.038213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.038213Z digest=sha256:c62c7e012f5422550aa3d0b452a9b3b6a9b427ef731960dce3c887fdc572b2af

Observation 1d1a345a-e7fb-47e4-a468-6d2ca1e9b7ac · outbound

This paper cites DreamLLM: Synergistic Multimodal Comprehension and Creation.

Multitwine: Multi-Object Compositing with Text and Layout Control DreamLLM: Synergistic Multimodal Comprehension and Creation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.042916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.042916Z digest=sha256:4d5162162b02ed113830e5e97cc9d8e60ed830f30ca72a1dfc4308fc4546c402

Observation dd0abeb5-e9a7-4178-94c4-1f5c602025b3 · outbound

This paper cites PaLM-E: An Embodied Multimodal Language Model.

Multitwine: Multi-Object Compositing with Text and Layout Control PaLM-E: An Embodied Multimodal Language Model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.047459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.047459Z digest=sha256:ece6afb6854b0a979d451344ed62ff52857f075e088347f4b99ada6f5cc33f8b

Observation 837e58cb-59c3-4b1b-889b-a222bf3bd5a2 · outbound

This paper cites Scaling recti- fied flow transformers for high-resolution image synthesis.

Multitwine: Multi-Object Compositing with Text and Layout Control Scaling recti- fied flow transformers for high-resolution image synthesis

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:05:17.087545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:05:16.051903Z digest=sha256:9f40e5a5eb67de6395efe43b598ab5a5ab1e93c6e7375b2fa0c9aef297d9ee66

Observation 341e031a-cf5f-497a-9792-6bcec4aacf70 · outbound

This paper cites An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion.

Multitwine: Multi-Object Compositing with Text and Layout Control An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.056096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.056096Z digest=sha256:0cc6406d44bb205d5a68caf3280195f9715edb3742b5e18193a3c96fffebdda0

Observation 2aa2c571-9ff0-4ef6-92f2-19c13569d0d1 · outbound

This paper cites Prompt-to-Prompt Image Editing with Cross Attention Control.

Multitwine: Multi-Object Compositing with Text and Layout Control Prompt-to-Prompt Image Editing with Cross Attention Control

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.060539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.060539Z digest=sha256:5468d60631333de8a61565b4f68e82f531c83fca80be31f4a2223aa638f91ded

Observation 27d371a3-5e0a-4a91-93c2-0ed07b72e59f · outbound

This paper cites Clipscore: A reference-free evaluation met- ric for image captioning.

Multitwine: Multi-Object Compositing with Text and Layout Control Clipscore: A reference-free evaluation met- ric for image captioning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:05:17.071654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:05:16.064651Z digest=sha256:69a39178b52682f04cd7795a688ad70c478cfe0373e8daabccddee6328c92993

Observation 05e0f83e-8f4a-4001-8fd9-3a8c887f304c · outbound

This paper cites Taming Encoder for Zero Fine-tuning Image Customization with Text-to-Image Diffusion Models.

Multitwine: Multi-Object Compositing with Text and Layout Control Taming Encoder for Zero Fine-tuning Image Customization with Text-to-Image Diffusion Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.068679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.068679Z digest=sha256:a39cbfff4a975a39c8e542301d3e0ee7e1450ab5fad15a400108b85bc85fe4f6

Observation 5c2e602d-e6af-4798-b16f-8c8a4187f114 · outbound

This paper cites Rendering synthetic objects into legacy photographs.

Multitwine: Multi-Object Compositing with Text and Layout Control Rendering synthetic objects into legacy photographs

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:05:17.055767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:05:16.073123Z digest=sha256:1b8b6fe85c23898cd05a7f4bd52ba463953ffb9631833ffbef6be598432567d7

Observation a365751e-6298-45bf-abae-01df37cb318c · outbound

This paper cites 3d object manipulation in a single photograph using stock 3d models.

Multitwine: Multi-Object Compositing with Text and Layout Control 3d object manipulation in a single photograph using stock 3d models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:05:17.040988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:05:16.077249Z digest=sha256:2fc2750351c8318b9bf2b193c156f2fb64471a62fe877568b3034b65aab14b00

Observation fc442340-45ae-4eaf-aa15-629daae51b1b · outbound

This paper cites Gen- erating images with multimodal language models.

Multitwine: Multi-Object Compositing with Text and Layout Control Gen- erating images with multimodal language models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.081647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.081647Z digest=sha256:8f981ef337076713df36220e6ae0538a13fb83b416700ed84d6416284865d442

Observation 67e4255b-c2fd-400f-a0da-e82010b9a3b4 · outbound

This paper cites Open images v5 text annotation and yet another mask text spotter.

Multitwine: Multi-Object Compositing with Text and Layout Control Open images v5 text annotation and yet another mask text spotter

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:05:17.016080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:05:16.086200Z digest=sha256:c003278c3d1730958ef6c7368a159b22ca8794900fc5dcab540ef0d2b8c4620d

Observation bb5a39c4-c7a1-4822-925e-b6c12b0d0de9 · outbound

This paper cites Putting people in their place: Affordance-aware hu- man insertion into scenes.

Multitwine: Multi-Object Compositing with Text and Layout Control Putting people in their place: Affordance-aware hu- man insertion into scenes

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:05:17.000101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:05:16.090712Z digest=sha256:4877aeeee69eb30ac9df87451eb2288b5906f5cb5ea25029c2327377047b15af

Observation 5c546ea4-3ffb-4918-a06e-95704f0b3072 · outbound

This paper cites Photo clip art.

Multitwine: Multi-Object Compositing with Text and Layout Control Photo clip art

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:05:16.984562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:05:16.095184Z digest=sha256:670adf9f6733b27d6648a835bcecd4880ab60d0d9a8a811e548b198a9f4d2d61

Observation fb1d63c8-216a-4531-bc3c-b1b7af3a797f · outbound

This paper cites Blip-diffusion: Pre- trained subject representation for controllable text-to-image generation and editing.

Multitwine: Multi-Object Compositing with Text and Layout Control Blip-diffusion: Pre- trained subject representation for controllable text-to-image generation and editing

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:05:16.970266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:05:16.099723Z digest=sha256:bf64048f831ee38654a06196c5ff49ad41ef6c3c3024efbeb0a84ce1f51270c8

Observation ba07cc92-e700-4c65-b17c-d117e8ad268a · outbound

This paper cites UNIMO-G: Unified Image Generation through Multimodal Conditional Diffusion.

Multitwine: Multi-Object Compositing with Text and Layout Control UNIMO-G: Unified Image Generation through Multimodal Conditional Diffusion

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.103942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.103942Z digest=sha256:22ea5e9806d35b642ecb20946c0f39dc9a9bf92c35343a1aa534c0c7aad2e676

Observation 44f52a1d-a541-41f1-b616-4545299ffd2e · outbound

This paper cites Improved baselines with visual instruction tuning, 2023.

Multitwine: Multi-Object Compositing with Text and Layout Control Improved baselines with visual instruction tuning, 2023

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:05:16.955457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:05:16.108408Z digest=sha256:625dc7dc3339faf7ed53cf0a4e6483fe0c1a3c295f0a5189c5b648535d948398

Observation 27ca87c1-5ba6-48a3-a799-d75cc0711370 · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

Multitwine: Multi-Object Compositing with Text and Layout Control Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.112663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.112663Z digest=sha256:b3da0bf0e5b7e44c777a01347d5fdbcebb458f20c770d4f14303e192bd534a59

Observation 304baa49-21d6-44ff-9d49-97752c29c1cd · outbound

This paper cites Tf-icon: Diffusion-based training-free cross-domain image composi- tion.

Multitwine: Multi-Object Compositing with Text and Layout Control Tf-icon: Diffusion-based training-free cross-domain image composi- tion

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:05:16.940940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:05:16.117452Z digest=sha256:78afb51a4f3f639a5fabd865a488fe4ac6d2fc12461ff8af461df6db17e27142

Observation 55d4a265-1e4a-4841-a59c-7c2027df956a · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Multitwine: Multi-Object Compositing with Text and Layout Control DINOv2: Learning Robust Visual Features without Supervision

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.121656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.121656Z digest=sha256:23515eca3af1cb1549248f062d36ec5c1080174ec2d4255349cb7b1ff7344339

Observation 7ee35b5d-4d11-4b72-8216-e62c5eb15e9c · outbound

This paper cites Kosmos-G: Generating Images in Context with Multimodal Large Language Models.

Multitwine: Multi-Object Compositing with Text and Layout Control Kosmos-G: Generating Images in Context with Multimodal Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.125929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.125929Z digest=sha256:b9c5b428a373a1986807be8b3376300241389feb1cc13167eb8c3c3ebd591f57

Observation 05aa7bc1-d977-40cc-874c-3723d47fdd57 · outbound

This paper cites https://pixabay.com/, 2024.

Multitwine: Multi-Object Compositing with Text and Layout Control https://pixabay.com/, 2024

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:05:16.925460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:05:16.130241Z digest=sha256:640f300a1e65b8493ff7b21e5705b2a1fe37cae56a6f45577f079407028c868c

Observation 2f083b3d-2002-4b00-b2e3-be4473e9468f · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Multitwine: Multi-Object Compositing with Text and Layout Control SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.134464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.134464Z digest=sha256:594ab558363779bc1a7048a51e243a3d00d56b5e9f0c99a0596a0189c7e6afe5

Observation 04b7e7bd-be0f-4737-a9bc-320e07c1cee5 · outbound

This paper cites High-Quality Entity Segmentation.

Multitwine: Multi-Object Compositing with Text and Layout Control High-Quality Entity Segmentation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.138710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.138710Z digest=sha256:4003a62766c52fc8f4d631d715a1972a4442fa60dcd93c76b6f544ccb1be272a

Observation ddc5b5da-451c-4048-8daa-bb80f0dc28b5 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Multitwine: Multi-Object Compositing with Text and Layout Control Learning transferable visual models from natural language supervi- sion

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.143110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.143110Z digest=sha256:6cef054ecf734d8f8f4fd688ac5ec97e161b6d20620938068785bd9e0d91142e

Observation fbf8423c-7629-4a99-b439-08159bd32e96 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Multitwine: Multi-Object Compositing with Text and Layout Control High-resolution image synthesis with latent diffusion models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.147319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.147319Z digest=sha256:ab15bd2a0e504ef6054f51bd9964fa74bec2b7d97a9b8dbbd3af71a3d1a3c46c

Observation bf7d524c-438c-4870-87c3-c5bf63b1fdfb · outbound

This paper cites Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation.

Multitwine: Multi-Object Compositing with Text and Layout Control Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:05:16.888041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:05:16.151779Z digest=sha256:79848742776743d82aa8df91e130d9c7b56953627d11561b3e78e3857e58867b

Observation d800c370-6ce3-41b4-8b68-a767ed7aaa08 · outbound

This paper cites Diffuse to Choose: Enriching Image Conditioned Inpainting in Latent Diffusion Models for Virtual Try-All.

Multitwine: Multi-Object Compositing with Text and Layout Control Diffuse to Choose: Enriching Image Conditioned Inpainting in Latent Diffusion Models for Virtual Try-All

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.156007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.156007Z digest=sha256:2665d622446cd2a5c5c24f50ba8e17b65319202a70df6811113d369e95851150

Observation 355e135b-1342-4b10-aede-3384f133622d · outbound

This paper cites Video visual relation detection.

Multitwine: Multi-Object Compositing with Text and Layout Control Video visual relation detection

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:05:16.871595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:05:16.160475Z digest=sha256:34eea9dfdbb37e491a5072b62cda4002f1a7d0e20998eeb62e456c680db24332

Observation 28cf2d92-51e0-480e-8eff-dd3d8276fdd4 · outbound

This paper cites Annotating objects and relations in user- generated videos.

Multitwine: Multi-Object Compositing with Text and Layout Control Annotating objects and relations in user- generated videos

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:05:16.856146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:05:16.164892Z digest=sha256:89b673c4d02d86c6b93a9bd28d12a37374799afcbdc1eb5d964d58b64380c741

Observation 3304e5c2-fee3-44c9-bbde-0da93e303fdb · outbound

This paper cites InstantBooth: Personalized Text-to-Image Generation without Test-Time Finetuning.

Multitwine: Multi-Object Compositing with Text and Layout Control InstantBooth: Personalized Text-to-Image Generation without Test-Time Finetuning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.169140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.169140Z digest=sha256:f4950194ca65a7d0fed030ae933879b816f26983382bc0cd183220a6f35307e1

Observation 69b55bb5-e0e3-454d-85c4-d8487227bf77 · outbound

This paper cites ObjectStitch: Generative Object Compositing.

Multitwine: Multi-Object Compositing with Text and Layout Control ObjectStitch: Generative Object Compositing

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.173596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.173596Z digest=sha256:d4ca4d177ef0d6ababdad12a2c5ba11db0572a8b0da838cc5a739ecd81fd2e4d

Observation 607a3426-bca4-4f7d-ad04-7fcc67d41958 · outbound

This paper cites Imprint: Generative object compositing by learning identity-preserving representation.

Multitwine: Multi-Object Compositing with Text and Layout Control Imprint: Generative object compositing by learning identity-preserving representation

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:05:16.840319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:05:16.177752Z digest=sha256:0a661747adc989a278abb96bae6dda69bd47ff00fef3767ac26c72a1623722ac

Observation 483e5ba2-0012-4a11-8aa4-8364c98eb43b · outbound

This paper cites Emu: Generative Pretraining in Multimodality.

Multitwine: Multi-Object Compositing with Text and Layout Control Emu: Generative Pretraining in Multimodality

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.182015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.182015Z digest=sha256:95df77423e2f032bd4b5339c9645baeb14019189d75b022d816205dac952b8ed

Observation d5d21661-b75f-462b-bbde-c39f0e36d0be · outbound

This paper cites Generative multimodal mod- els are in-context learners.

Multitwine: Multi-Object Compositing with Text and Layout Control Generative multimodal mod- els are in-context learners

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:05:16.825417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:05:16.186695Z digest=sha256:b22b685e39dffa71432c2f0cb31c75fd523d0068a72c6c909b3591b7b3d8787a

Observation 58db0e90-c271-4b15-aec1-08024664cc3f · outbound

This paper cites Thinking Outside the BBox: Unconstrained Generative Object Compositing.

Multitwine: Multi-Object Compositing with Text and Layout Control Thinking Outside the BBox: Unconstrained Generative Object Compositing

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.190538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.190538Z digest=sha256:cbb7e39d980fbe71cacaa01e1d1e7401b7d7154cba4577e60310ebc6c60546af

Observation 762dcc6a-bb1d-4db7-9f9b-e93c91d3ba83 · outbound

This paper cites CogVLM: Visual Expert for Pretrained Language Models.

Multitwine: Multi-Object Compositing with Text and Layout Control CogVLM: Visual Expert for Pretrained Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.194518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.194518Z digest=sha256:105e875983fa1f6aec53d859b51e8ed9f19f88fba7a559f4a6a1aefdbefd0420

Observation c1e141ef-52ce-4b42-be37-33b9a6a7d42b · outbound

This paper cites Elite: Encoding visual con- cepts into textual embeddings for customized text-to-image generation.

Multitwine: Multi-Object Compositing with Text and Layout Control Elite: Encoding visual con- cepts into textual embeddings for customized text-to-image generation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.198714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.198714Z digest=sha256:91fa8277eabfec68e3d2d776b95ef15b94200cdf647ed4c2d958686b394b6dbc

Observation dbae3f95-8114-431e-bcab-e2b2cf83b116 · outbound

This paper cites Fastcomposer: Tuning-free multi- subject image generation with localized attention.

Multitwine: Multi-Object Compositing with Text and Layout Control Fastcomposer: Tuning-free multi- subject image generation with localized attention

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:05:16.799738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:05:16.202729Z digest=sha256:19d9b1e3a0ed5d22f8e5e789b64f6d9612fc3d11d853554cba29c434f149244e

Observation 2f9f0806-d127-4b70-bb66-5b3301e8a885 · outbound

This paper cites GroundingBooth: Grounding Text-to-Image Customization.

Multitwine: Multi-Object Compositing with Text and Layout Control GroundingBooth: Grounding Text-to-Image Customization

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.206770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.206770Z digest=sha256:d8101e09f3fc3cc6519bdd4af6454de160e96d1c1215d843a5c91855d330c6f7

Observation b051d4b7-f7c2-40d4-9b8d-aab3aaf5e34e · outbound

This paper cites Paint by example: Exemplar-based image editing with diffusion mod- els.

Multitwine: Multi-Object Compositing with Text and Layout Control Paint by example: Exemplar-based image editing with diffusion mod- els

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.211720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.211720Z digest=sha256:6cef0f303f95a4b4cbf71fd7e8ad59e04aa3e32bd5db1c375c0084e0cafc14f7

Observation f61748aa-6e5d-4c4e-b0bb-5bedaee888c9 · outbound

This paper cites CustomNet: Zero-shot Object Customization with Variable-Viewpoints in Text-to-Image Diffusion Models.

Multitwine: Multi-Object Compositing with Text and Layout Control CustomNet: Zero-shot Object Customization with Variable-Viewpoints in Text-to-Image Diffusion Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.216485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.216485Z digest=sha256:ea1877e1f72a231d35d40a632cec2be712a99ea3bf5599bcafe2145949369d35

Observation f381e241-b1f2-44ae-bfa9-92faa49e1e97 · outbound

This paper cites ControlCom: Controllable Image Composition using Diffusion Model.

Multitwine: Multi-Object Compositing with Text and Layout Control ControlCom: Controllable Image Composition using Diffusion Model

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.220898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.220898Z digest=sha256:759b3f30ddebf80b25e9805c24b26a88e78fdbf04245abf6020b3fca8a90ef15

Observation caa0283f-05fd-4101-ad25-87b5b5c818df · outbound

This paper cites LoCo: Locally Constrained Training-Free Layout-to-Image Synthesis.

Multitwine: Multi-Object Compositing with Text and Layout Control LoCo: Locally Constrained Training-Free Layout-to-Image Synthesis

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.225176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.225176Z digest=sha256:3dfeebb78c653840b22db23aec3cbdb5edea3d2e9eab4ad1aa4127f9ab3eeb64

Observation 1a560b2d-248f-499f-a8f0-9d9bf0d2d100 · outbound

This paper cites Can you provide a grammatically correct one-line caption for the relation <object A> <relation> <object B> in the image?.

Multitwine: Multi-Object Compositing with Text and Layout Control Can you provide a grammatically correct one-line caption for the relation <object A> <relation> <object B> in the image?

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:05:16.774028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:05:16.229501Z digest=sha256:e91bc2baa68811bf99bf30dc53e04b9e584e8431162c70ff46e7cfaa8a57ccaa

Observation 8a454c94-da25-417d-9c2f-aa9ef7a0911a · outbound

This paper cites Figure 4.

Multitwine: Multi-Object Compositing with Text and Layout Control Figure 4

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:05:16.758388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:05:16.234124Z digest=sha256:c7521ece8708c2a09ae95c351ed1e02e332e5378564a21f8a304d7e93172fe5c

Observation 6594bd52-d1fc-4891-8733-bd6287635ded · outbound

This paper cites used to extract them.

Multitwine: Multi-Object Compositing with Text and Layout Control used to extract them

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:05:16.740942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:05:16.238462Z digest=sha256:fb6973f0e5491c83419179f4aded1490d6f8dd85bed3930632ae7874333f3884

Observation e4a496a8-e693-4bfe-919b-41542389de35 · outbound

This paper cites Background images are sourced from Pixabay [29], while objects are from Pixabay [29], MultiBench [23], and DreamBooth [34].

Multitwine: Multi-Object Compositing with Text and Layout Control Background images are sourced from Pixabay [29], while objects are from Pixabay [29], MultiBench [23], and DreamBooth [34]

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:05:16.725567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:05:16.242546Z digest=sha256:109b11e21bd67988ded5c9f8979500b46650a51607e41426a54e5a3b1cd34f11

Observation 6a58adb2-d0ad-42c8-ad3e-6fb483ece5b6 · outbound

This paper cites Further details on user studies can be found in Section 3.3.

Multitwine: Multi-Object Compositing with Text and Layout Control Further details on user studies can be found in Section 3.3

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:05:16.710737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:05:16.246690Z digest=sha256:4eb3d5d51bdde9a729cecd795cb54c83b909ffa437b8a59da6a63775d1b7d627

Observation 8cbbc66a-c750-4349-91f2-6ecbaa65528c · outbound

This paper cites Without multi-view data ( i.e., video data, manually collected data), the model struggles to prop- erly repose and combine objects to align with the textual description.

Multitwine: Multi-Object Compositing with Text and Layout Control Without multi-view data ( i.e., video data, manually collected data), the model struggles to prop- erly repose and combine objects to align with the textual description

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:05:16.694958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:05:16.251155Z digest=sha256:fedc0367f9a92a6acc2511c6808010a4afbe7f0aa60cb4b30834be5eb0f09d6a

Observation 6d49043d-648e-4ee1-a8ae-04f25ddf9c8c · outbound

This paper cites an unresolved cited work.

Multitwine: Multi-Object Compositing with Text and Layout Control Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-08T20:05:16.677728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:05:16.256044Z digest=sha256:67623368f064c8688db4da7d849c79f6af438400d561a688b28fa848832bcdf8

Pith citing papers

Observation 58b4d835-8f33-4c81-bfa6-ae638a5cc0b7 · inbound

BlenderFusion: 3D-Grounded Visual Editing and Generative Compositing cites this paper.

BlenderFusion: 3D-Grounded Visual Editing and Generative Compositing Multitwine: Multi-Object Compositing with Text and Layout Control

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-08-06T23:35:15.948136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:35:15.341406Z digest=sha256:3057253bc3972b11f5275517a24ac3de671089c4177c5268e56b4b2d863986f0