Pith. sign in

Paper Citation Record · LEDGER

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation

As of 13 August 2026, this Paper Citation Record lists 77 of 77 outbound references and 2 inbound Pith citation observations for arXiv:2511.20635.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2511.20635 v3

Coverage vector

measured 77 of 77 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T06:47:12.160381Z

measured 79 of 79 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T19:08:49.594504Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T03:29:31.039583Z

Reference resolution

77 of 77 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved77
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation dbb1bb02-68da-4bb9-bb5b-81dc0779e639 · outbound

This paper cites Qwen2.5-VL Technical Report.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:02.521174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:02.521174Z digest=sha256:394225e360f1e722066aaf320e7be7b759ca676187def450735d7fcdb891a2f6

Observation 544374ec-8a64-4add-937d-82fbd5e53e5a · outbound

This paper cites Lumiere: A space-time diffu- sion model for video generation.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation Lumiere: A space-time diffu- sion model for video generation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:02.678126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:02.678126Z digest=sha256:a8aa1f6416d0d7103d10a7b94045d6e8fe6b39c560a4fb1ced47bbc6a0f47c0a

Observation b7242dd7-d25f-4d0a-b8aa-827df50abd9d · outbound

This paper cites Seedream4.0, 2025.https : / / seed.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation Seedream4.0, 2025.https : / / seed

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:02.818809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:02.818809Z digest=sha256:8ebfd600a3a87534c6120adad03cb98e95782f13a18247112c247afd6905fac3

Observation ea4bcddd-a7e2-4cf2-a067-1ae53cde2aba · outbound

This paper cites HiDream-I1: A High-Efficient Image Generative Foundation Model with Sparse Diffusion Transformer.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation HiDream-I1: A High-Efficient Image Generative Foundation Model with Sparse Diffusion Transformer

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:02.976356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:02.976356Z digest=sha256:c2b5812a067fa3ee9ba2380c3af37ea7ee5fa1dc183fecce5b01c89256330462

Observation ae6c611a-2f56-484a-845c-83de3644cc8c · outbound

This paper cites Openpose: Realtime multi-person 2d pose estimation using part affinity fields.IEEE transactions on pattern analysis and machine intelligence, 43(1):172–186,.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation Openpose: Realtime multi-person 2d pose estimation using part affinity fields.IEEE transactions on pattern analysis and machine intelligence, 43(1):172–186,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:03.114478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:03.114478Z digest=sha256:b509cea85408bf801cd0fad95e6a6c7a1a66b4902d3d457dd23414f20c2cbdd1

Observation 834da371-fb5d-422e-a978-70b69116f43e · outbound

This paper cites Emerg- ing properties in self-supervised vision transformers.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation Emerg- ing properties in self-supervised vision transformers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:03.233924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:03.233924Z digest=sha256:16b6375eb247d9fc370cceeab511b86f02c75bc04109a60d8380868cb66b64e8

Observation c0df67fd-3efc-4aad-b2af-120ea6c208c9 · outbound

This paper cites PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:03.411590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:03.411590Z digest=sha256:56f03b89ebe90010c7a5299cee7a7815ada87f31a1de3af2ed10077e2ed2e0b0

Observation 233482c9-4fb8-49e6-96f0-7c51ef73cc7e · outbound

This paper cites DeepVerse: 4D Autoregressive Video Generation as a World Model.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation DeepVerse: 4D Autoregressive Video Generation as a World Model

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:03.534918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:03.534918Z digest=sha256:4d6a90ef28dc2a4931e2c4fccd3dec87ed3d771eda1fdf524c537d3398638275

Observation cfca67ec-e377-438a-b390-a5741ff25a3e · outbound

This paper cites Univid: Unifying vi- sion tasks with pre-trained video generation models.arXiv preprint arXiv:2509.21760, 2025.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation Univid: Unifying vi- sion tasks with pre-trained video generation models.arXiv preprint arXiv:2509.21760, 2025

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:03.642270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:03.642270Z digest=sha256:777e3bb822183491508ac6ed6a3fce06b010632efe62b826c903ca661430ae62

Observation 5e545dbb-c9a7-4593-90bb-e0a3eb1220c8 · outbound

This paper cites Unireal: Universal image generation and editing via learning real-world dynamics.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation Unireal: Universal image generation and editing via learning real-world dynamics

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:03.758730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:03.758730Z digest=sha256:cdede532909f94fc7597cfe18246528263dfd63f83b181fd7930337021233270

Observation c491ff43-79cd-40bf-9599-455153a9b709 · outbound

This paper cites UMO: Scaling Multi-Identity Consistency for Image Customization via Matching Reward.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation UMO: Scaling Multi-Identity Consistency for Image Customization via Matching Reward

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:03.874746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:03.874746Z digest=sha256:ed37717465700d35fc12997f0f24afbe75c455792fddafead79f23f4aab36445

Observation 89ebe157-906e-4531-b344-3e903c4f4837 · outbound

This paper cites Gemini2.5, 2025.https : / / deepmind.google/models/gemini/pro/.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation Gemini2.5, 2025.https : / / deepmind.google/models/gemini/pro/

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:04.050275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:04.050275Z digest=sha256:b5f752b63bbb172765d2a5f68b7ae053cb221c58e205bb1c5939dc2ffe1f0e88

Observation df3f48df-8c46-4420-98f2-1f9dad40af31 · outbound

This paper cites Veo3, 2025.https://deepmind.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation Veo3, 2025.https://deepmind

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:04.195649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:04.195649Z digest=sha256:91b20962b5c0513e59d0db1e6d2c536f38f7673b56475e3da9b4fc4a253dd6a6

Observation 7fd60b4b-f7a4-4183-8bde-3ab79de1b739 · outbound

This paper cites Emerging Properties in Unified Multimodal Pretraining.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation Emerging Properties in Unified Multimodal Pretraining

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:04.340696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:04.340696Z digest=sha256:a8c5aaa96e5e3e1b47768c9e1fd725b4f193faea076dbe7a297a59992f96a338

Observation 5634c617-34c3-446f-8757-adcaa3ce3a45 · outbound

This paper cites Yolov8-face-detection, 2024.https : //huggingface.co/arnabdhar/YOLOv8- Face- Detection.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation Yolov8-face-detection, 2024.https : //huggingface.co/arnabdhar/YOLOv8- Face- Detection

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:04.500990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:04.500990Z digest=sha256:72a5a8fd9a01c94cc23a08f8946f7981f55b0f3b68d1747f500ca8134cc6f57d

Observation 153b526b-2c38-4feb-ae71-57b7ed1e5f13 · outbound

This paper cites Scaling recti- fied flow transformers for high-resolution image synthesis.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation Scaling recti- fied flow transformers for high-resolution image synthesis

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:04.674385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:04.674385Z digest=sha256:886afcd677f196e047855396b991467dcb384b8261650ce8b87cef27396a3a5f

Observation dfb39ff3-0d2e-4d03-bc8a-581c4b679924 · outbound

This paper cites UniVG: A Generalist Diffusion Model for Unified Image Generation and Editing.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation UniVG: A Generalist Diffusion Model for Unified Image Generation and Editing

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:04.847612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:04.847612Z digest=sha256:c6b4047341345bb2ed87bbfd2f0331a17fb66259d86672c9e6c7527c05d43b02

Observation e22bf4ad-90c0-485e-9b86-15ba0e62355b · outbound

This paper cites Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:04.960072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:04.960072Z digest=sha256:85784f756ddbe34acdc9f0521b6d6e02861ff4d97a9f6209df4db6586b043d26

Observation db6a5d71-0ece-4bd4-b225-0255de6e02e2 · outbound

This paper cites MVImgNet2.0: A Larger-scale Dataset of Multi-view Images.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation MVImgNet2.0: A Larger-scale Dataset of Multi-view Images

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:05.055172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:05.055172Z digest=sha256:9a9577f40aa2acbb79b037dfcaef46c7af4506499bbb9791ba02bba7063fad20

Observation 6d27e716-eb2b-4890-9a26-5a76ad0398d6 · outbound

This paper cites Hidream-e1-1, 2025.https : / / huggingface.co/HiDream-ai/HiDream-E1-1.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation Hidream-e1-1, 2025.https : / / huggingface.co/HiDream-ai/HiDream-E1-1

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:05.221223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:05.221223Z digest=sha256:553d4381186580bd120b9ad75221db5067710891365baf9e3571fa0cea7fef22

Observation ea1c90d0-21d7-435b-a5b1-5b0a465c8a75 · outbound

This paper cites Vbench: Comprehensive bench- mark suite for video generative models.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation Vbench: Comprehensive bench- mark suite for video generative models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:05.359216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:05.359216Z digest=sha256:6aa51564e817de173e8035b4da4253890a1357fe3bfd307e016b1fd44538a159

Observation a3a24969-3acf-4f36-a1bb-19263558f5e7 · outbound

This paper cites Controlnet auxiliary models.https : //github.com/huggingface/controlnet_aux? tab=readme-ov-file, 2023.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation Controlnet auxiliary models.https : //github.com/huggingface/controlnet_aux? tab=readme-ov-file, 2023

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:05.516200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:05.516200Z digest=sha256:13bf903956ce4d6aac8f93c509fba2662b476df3a84a37b82f285b70fe0e4ca7

Observation 86bf3562-4b8f-4605-a262-6c5acb8e9d04 · outbound

This paper cites InfiniteYou: Flexible Photo Recrafting While Preserving Your Identity.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation InfiniteYou: Flexible Photo Recrafting While Preserving Your Identity

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:05.660498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:05.660498Z digest=sha256:0a7fab77834c6b2f001d33acb3f3175ccbd8b8a07d79386ef93c84c23b2e929c

Observation bef57812-93d2-4e07-a8cf-e4998f49a69b · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:05.784126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:05.784126Z digest=sha256:949da79ccd55ad1b77144236509a75a376e108d039733b8cf1f775041ca91881

Observation 8a4458db-a56d-4c30-b58b-9acbacd29f39 · outbound

This paper cites Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:05.973725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:05.973725Z digest=sha256:2cd61004dba0a8f7da7d3402383acd9ee28e649e7451d4615fe2005184b3b0e9

Observation f1a681b0-f1e2-47e3-b7b9-5cbdfc014ca0 · outbound

This paper cites Flux.1 [dev], 2024.https : / / huggingface.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation Flux.1 [dev], 2024.https : / / huggingface

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:06.138680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:06.138680Z digest=sha256:4b1389622ee3889e537071b037bbff4918157178b58aeb5c6ed7d0f0eef24b29

Observation 939b3a07-c8b7-4aae-8688-1cee7bee5837 · outbound

This paper cites Clip-based nsfw detector, 2021.https : //github.com/LAION- AI/CLIP- based- NSFW- Detector.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation Clip-based nsfw detector, 2021.https : //github.com/LAION- AI/CLIP- based- NSFW- Detector

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:06.255830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:06.255830Z digest=sha256:df318a656656fc78303d8134163b8f9e84c59494c138b15a884383b313abbf1a

Observation 8f7a22fb-330f-43cb-8124-2416e21e5e00 · outbound

This paper cites UNIMO-G: Unified Image Generation through Multimodal Conditional Diffusion.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation UNIMO-G: Unified Image Generation through Multimodal Conditional Diffusion

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:06.384931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:06.384931Z digest=sha256:5c747f26434836256d33ef14590ba3a1644f9371957d350a64806808366d0704

Observation cdf09944-1623-4c3a-8eec-a6edbe2ab54f · outbound

This paper cites UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:06.552623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:06.552623Z digest=sha256:28236fa60855b890ca9a7f9589ab739c31ff08e9011af8dc006535b590409cdb

Observation f62f80e7-74cf-4f58-b552-a6f1ad172d51 · outbound

This paper cites RealGeneral: Unifying Visual Generation via Temporal In-Context Learning with Video Models.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation RealGeneral: Unifying Visual Generation via Temporal In-Context Learning with Video Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:06.699287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:06.699287Z digest=sha256:902824c52be4c0d479a4ba86e695514e5a0d5dec9fe996b03f9daa244bb4df5c

Observation 579d9b89-804f-43d1-9098-27a306c19693 · outbound

This paper cites Flow Matching for Generative Modeling.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation Flow Matching for Generative Modeling

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:06.821286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:06.821286Z digest=sha256:5628b5cbe74c94546c0c6475db8f1a54b571ae006b9f2b28b37af8dfb870d7c1

Observation 54d1c0b1-7c25-4dab-bbe8-624cb88cfcf7 · outbound

This paper cites Audioldm 2: Learning holistic audio gen- eration with self-supervised pretraining.IEEE/ACM Trans- actions on Audio, Speech, and Language Processing, 32: 2871–2883, 2024.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation Audioldm 2: Learning holistic audio gen- eration with self-supervised pretraining.IEEE/ACM Trans- actions on Audio, Speech, and Language Processing, 32: 2871–2883, 2024

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:06.925885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:06.925885Z digest=sha256:b84844720810498b4177cba745715b6fbbcba27bb5e235c3c881c3f784648591

Observation f5fc3315-0aef-4440-8444-b48a322f6ca7 · outbound

This paper cites Step1X-Edit: A Practical Framework for General Image Editing.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation Step1X-Edit: A Practical Framework for General Image Editing

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:07.058673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:07.058673Z digest=sha256:957e0b65a778651ab72d20f6dbfb527901b4794d940e52f5c40e9e73fd59f773

Observation a3ef9fcc-9480-4d6b-be50-8857081a835a · outbound

This paper cites Univid: The open-source unified video model.arXiv preprint arXiv:2509.24200, 2025.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation Univid: The open-source unified video model.arXiv preprint arXiv:2509.24200, 2025

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:07.182018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:07.182018Z digest=sha256:f5f42be055a94513311f66f669272c3c10703748b8c19cf71fbf458bbb154482

Observation f1b6c4e2-b526-4fb1-9e77-3f7c237e4e89 · outbound

This paper cites ACE++: Instruction-Based Image Creation and Editing via Context-Aware Content Filling.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation ACE++: Instruction-Based Image Creation and Editing via Context-Aware Content Filling

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:07.337393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:07.337393Z digest=sha256:4997738a75b372cf91d68cc3fdeb168d588b6f9864227b7ecb906a2aec9fd6d3

Observation 4341b93e-1ebc-4738-a333-ef7647fefc4c · outbound

This paper cites Gpt4o, 2024.https://www.openai.com/.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation Gpt4o, 2024.https://www.openai.com/

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:07.445105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:07.445105Z digest=sha256:b44eb3f02239fe5a9ef9cb62a7766236031872e9c3156d928f040f128b1ec274

Observation b6e49269-494e-484e-bba9-2dae4769ec21 · outbound

This paper cites Sora2, 2025.https://openai.com/index/ sora-2/.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation Sora2, 2025.https://openai.com/index/ sora-2/

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:07.542754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:07.542754Z digest=sha256:1a7c1f99fedcf9a7b60357423871e150376cdb442e67dd2fd06cd4a0fc3ada79

Observation 2c0b917d-e397-4cc1-94d9-9a5cfe995558 · outbound

This paper cites Scalable diffusion models with transformers.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation Scalable diffusion models with transformers

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:07.647770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:07.647770Z digest=sha256:f799243939b5a3c2272d2875f9c2f673692e2e4eb7d90225af19d6b70a77dc22

Observation b05af249-c56d-429f-816d-d6f449bae16d · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:07.767318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:07.767318Z digest=sha256:344a370954cd5b721f93a506982697630d4b5a1eede4fb9d14b469742ec89194

Observation edb03314-eab6-4b4d-bcdb-be41978b3855 · outbound

This paper cites Lumina-Image 2.0: A Unified and Efficient Image Generative Framework.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation Lumina-Image 2.0: A Unified and Efficient Image Generative Framework

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:07.873473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:07.873473Z digest=sha256:58c3c024740f4e875817fa4a9f8a5bb3dfb5a4ea60d4ef15d970c4cdbc0bdf80

Observation 03b10944-c7ca-4d04-8723-9cd6e51ff104 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation Learning transferable visual models from natural language supervi- sion

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:08.020620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:08.020620Z digest=sha256:9813ce372a91c0d536fb64e110a09308d0941050c042ed549e3a0d3fb4069850

Observation 82d39e5b-8256-4159-9f93-d2df42ef8617 · outbound

This paper cites Make-a-story: Visual memory conditioned consistent story generation.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation Make-a-story: Visual memory conditioned consistent story generation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:08.151908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:08.151908Z digest=sha256:ff4403bcdef6f57b49143bc1032c0f9a70abecbf4679ca56e1db254b2f3de482

Observation 855b24d5-0b9e-4d05-91fe-08f815e53b20 · outbound

This paper cites Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:08.251845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:08.251845Z digest=sha256:449e88f8c1f8934a7478f451b7b8be1f994d3cb9aeb40347e202b33ea7eabd83

Observation 9f78cc0b-b93a-4168-b640-1ab405708b96 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation High-resolution image synthesis with latent diffusion models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:08.346628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:08.346628Z digest=sha256:c51719e4e7aed9f84c3bc41361dac4238584242a1b35903f7d17d2a59bc9e265

Observation aabc1dd7-94e0-4512-a15f-e3f5e0fb6295 · outbound

This paper cites Pathways on the image manifold: Image editing via video generation.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation Pathways on the image manifold: Image editing via video generation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:08.450256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:08.450256Z digest=sha256:207c0b1fd657902b3b3031231dd13daca634def7aa3debd228a4f242792fdc5c

Observation 11b55f3f-1f22-4257-9fcc-85987052d015 · outbound

This paper cites Photorealistic text-to-image diffusion models with deep language understanding.Advances in neural information processing systems, 35:36479–36494, 2022.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation Photorealistic text-to-image diffusion models with deep language understanding.Advances in neural information processing systems, 35:36479–36494, 2022

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:08.559024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:08.559024Z digest=sha256:60280e72c62360bc47e21f2875b136736e522fcc1122b52f2b8e74bce69bedea

Observation 120f5d7d-03a0-429f-8f58-2eeb5103adaf · outbound

This paper cites HunyuanVideo-Foley: Multimodal Diffusion with Representation Alignment for High-Fidelity Foley Audio Generation.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation HunyuanVideo-Foley: Multimodal Diffusion with Representation Alignment for High-Fidelity Foley Audio Generation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:08.625247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:08.625247Z digest=sha256:2191dfdaf6752417208dd5825c9ff43873a903257924cb9b3f23a9935233f656

Observation 27d13411-0f0b-4ff9-ba90-53307b885476 · outbound

This paper cites Aether: Geometric-Aware Unified World Modeling.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation Aether: Geometric-Aware Unified World Modeling

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:08.711461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:08.711461Z digest=sha256:40225d8f2237e9f7039158437e246e94176eda66d43852a3e48df5090eb8c9ec

Observation 620b5149-c115-419a-9a7d-099a1b7c3c37 · outbound

This paper cites Raft: Recurrent all-pairs field transforms for optical flow.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation Raft: Recurrent all-pairs field transforms for optical flow

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:08.797000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:08.797000Z digest=sha256:7cfc4101956afed8f8aeb5b13b9566dde434ddf03e842da39c6366405b3a9ed5

Observation abe87faf-64e1-4cfe-9174-d11e04825029 · outbound

This paper cites Audiobox: Unified Audio Generation with Natural Language Prompts.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:08.922024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:08.922024Z digest=sha256:474bc108e154ee2c90b409b1bcba3cfb61ec2753127e54771dbe64fcbb62bc33

Observation 98cd8732-bba7-4c19-a051-0aa3d3eae5c3 · outbound

This paper cites InstantID: Zero-shot Identity-Preserving Generation in Seconds.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation InstantID: Zero-shot Identity-Preserving Generation in Seconds

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:09.006094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:09.006094Z digest=sha256:de7c1baf0eb0a8e2ade26fe0ef3a2e281f087ac0b52df5af3e810e8e03929b2b

Observation b22e0ecf-34f4-4222-a246-8505d0bfe88a · outbound

This paper cites Autostory: Generating diverse storytelling images with minimal human efforts.Interna- tional Journal of Computer Vision, pages 1–22, 2024.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation Autostory: Generating diverse storytelling images with minimal human efforts.Interna- tional Journal of Computer Vision, pages 1–22, 2024

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:09.122053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:09.122053Z digest=sha256:c70ebdda8a7d0e6262688d91b37f284da0d2eb977a8c2f64c909b8a1f47abd8a

Observation 5b95dbdd-b7d7-40cd-8b07-15eb457e9171 · outbound

This paper cites UniVideo: Unified Understanding, Generation, and Editing for Videos.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation UniVideo: Unified Understanding, Generation, and Editing for Videos

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:09.213780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:09.213780Z digest=sha256:fb27f5f4826f0caa2711ce342deedae2d6d7b8043a13c4d28cf694a9d3f47597

Observation fe4599e1-f29a-4df5-9c6c-ec9a0030f03f · outbound

This paper cites OmniGen2: Towards Instruction-Aligned Multimodal Generation.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:09.354756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:09.354756Z digest=sha256:5db656154bd845f519eacf29d778d9649076955e4f1c42e501b792562e3a794a

Observation 090f917d-1ab2-4756-9670-8ceaf496dc68 · outbound

This paper cites Chronoedit: Towards temporal reason- ing for image editing and world simulation.arXiv preprint arXiv:2510.04290, 2025.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation Chronoedit: Towards temporal reason- ing for image editing and world simulation.arXiv preprint arXiv:2510.04290, 2025

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:09.491668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:09.491668Z digest=sha256:b5547942778bc8ae4197c696a8f4e9a094edfeaaa30abd261c42377c7ed13e9c

Observation b453d790-b836-49af-bfa6-d49d5f7e63d5 · outbound

This paper cites USO: Unified Style and Subject-Driven Generation via Disentangled and Reward Learning.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation USO: Unified Style and Subject-Driven Generation via Disentangled and Reward Learning

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:09.609041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:09.609041Z digest=sha256:3d533359da219bbf600d94a506e469b21817ce3b751f21c9ec4061046131a9c3

Observation 9ec24dd0-c196-4990-81ed-38cfda623f4c · outbound

This paper cites Less-to-More Generalization: Unlocking More Controllability by In-Context Generation.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation Less-to-More Generalization: Unlocking More Controllability by In-Context Generation

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:09.770582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:09.770582Z digest=sha256:b044874603eb2f65c862139faaccec4129d3775ae94c7702b8ecd5e1e36e18dc

Observation 27c86e6a-e60e-4684-9736-24f0fb670d28 · outbound

This paper cites Dreamomni: Unified image generation and editing.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation Dreamomni: Unified image generation and editing

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:09.894429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:09.894429Z digest=sha256:2caecb21a95b8b6deefc80c4ee7bf5df33240e4eb6c8f3914b18f72823a426aa

Observation 3b7a6045-e1a2-43d3-9564-57f4e992f09f · outbound

This paper cites Omnigen: Unified image genera- tion.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation Omnigen: Unified image genera- tion

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:10.046196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:10.046196Z digest=sha256:5a4f8d849d54b250f2b3aeddbdafc5ef37140ad1390891cc5d7c1f93d707e9a2

Observation 81d87c44-f2d2-41c6-89e0-4d9a13bd4347 · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:10.169710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:10.169710Z digest=sha256:a934b36e4797e73e7f4b0f2ad56df5019357af44e66ee0df13a2e94064504026

Observation 0dfa7745-8b8d-4803-b8f7-4b5edcad2b7c · outbound

This paper cites CSGO: Content-Style Composition in Text-to-Image Generation.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation CSGO: Content-Style Composition in Text-to-Image Generation

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:10.328950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:10.328950Z digest=sha256:d196bcc13e4d43cb8ffed0a27a74463120f53c5b5507606d3ec9eac26ab7dbbe

Observation c16ad3f9-1e63-47d0-a882-c7d3c5f377f3 · outbound

This paper cites Withanyone: Towards control- lable and id consistent image generation.arXiv preprint arXiv:2510.14975, 2025.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation Withanyone: Towards control- lable and id consistent image generation.arXiv preprint arXiv:2510.14975, 2025

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:10.390873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:10.390873Z digest=sha256:f00cc3c7f1df501e38ceb123459152b1dabcd5f388dd0a67bd59c4ee184a4b3c

Observation 3fb4263e-3399-4e4f-b0ee-b933c3288000 · outbound

This paper cites UniAudio: An Audio Foundation Model Toward Universal Audio Generation.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:10.473180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:10.473180Z digest=sha256:3840c63ff90142090d9d3bf605ef29cc47feab43115e6e0f8505479a5adec964

Observation ed23f208-f311-426c-b74c-e7ebf9059ae4 · outbound

This paper cites Depth any- thing v2.Advances in Neural Information Processing Sys- tems, 37:21875–21911, 2024.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation Depth any- thing v2.Advances in Neural Information Processing Sys- tems, 37:21875–21911, 2024

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:10.539305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:10.539305Z digest=sha256:00d7a3f2ef9c260e08bd1b81f3350a8a0fb995af81ed2437f7e6d0960d57e5df

Observation 2b5ee6e0-242a-49d6-9209-000c2d6bb892 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:10.554757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:10.554757Z digest=sha256:3e50ee4290dc37a29c4c625233279e08c65a6114ddcb5dda4b3f172d6912e8b3

Observation 34440ad9-f74a-43e0-816f-c8d31e660894 · outbound

This paper cites IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:10.638722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:10.638722Z digest=sha256:cde817772787c2062145ea827e953de62f34dafe579e639cf5c18cc4990c29c4

Observation 930cac71-df1c-4ade-85a8-a73f533250b2 · outbound

This paper cites Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:10.728416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:10.728416Z digest=sha256:7b2009c4afe2ca2a0b3851f7e1d288dfdd01b2840f1ac3ca1caef257261dee9c

Observation 7a6b1941-4420-4b9d-b3fa-ea4cc57e6778 · outbound

This paper cites ImgEdit: A Unified Image Editing Dataset and Benchmark.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation ImgEdit: A Unified Image Editing Dataset and Benchmark

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:10.830159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:10.830159Z digest=sha256:c0c8e73ce8e51f1dbcc091efc42dea767bfa10c329ba248f49e33e7c54b1d70e

Observation b1f5b6d1-ed67-41a6-89ee-3564bae7afc7 · outbound

This paper cites Adding conditional control to text-to-image diffusion models.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation Adding conditional control to text-to-image diffusion models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:10.929454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:10.929454Z digest=sha256:beee27c1d088ee8360efc47a95064f8d00aab788fa25c71a465a3daa67da6d4e

Observation 4e43feea-a87f-439a-9f7a-1c4a7f78eee5 · outbound

This paper cites In-Context Edit: Enabling Instructional Image Editing with In-Context Generation in Large Scale Diffusion Transformer.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation In-Context Edit: Enabling Instructional Image Editing with In-Context Generation in Large Scale Diffusion Transformer

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:11.109575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:11.109575Z digest=sha256:26c0e1e868455e14639603e7b7c11984d208a38a301bb3b8213fe49dbc98e89c

Observation 6a715269-7317-4d48-bd85-48f9045caa8f · outbound

This paper cites VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:11.222875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:11.222875Z digest=sha256:6aecb0618756d3bde37a00464871f49e73346917ee98ecbf9fe9e208d785d2fd

Observation e17bd8a3-7b86-4592-ab72-e872ba646a3d · outbound

This paper cites Open-Sora: Democratizing Efficient Video Production for All.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation Open-Sora: Democratizing Efficient Video Production for All

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:11.394627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:11.394627Z digest=sha256:136aa05a5d684c36ab342a4eba1b790912f720144b8262314010ca5efd268808

Observation 07bda8c2-5a14-440b-92ee-f4e36c3c7e56 · outbound

This paper cites Storydiffusion: Consistent self- attention for long-range image and video generation.Ad- vances in Neural Information Processing Systems, 37: 110315–110340, 2024.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation Storydiffusion: Consistent self- attention for long-range image and video generation.Ad- vances in Neural Information Processing Systems, 37: 110315–110340, 2024

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:11.584253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:11.584253Z digest=sha256:447857845c243439fb6782498efeac72f142e86e5aa7bf59e660187fddea589d

Observation ca84c11a-e5f1-4be1-b78e-5c0caa3d8f09 · outbound

This paper cites In pre- training stage, we start with more video clip data and less image editing data, then gradually counting more image editing data for better instruction following capability.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation In pre- training stage, we start with more video clip data and less image editing data, then gradually counting more image editing data for better instruction following capability

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:11.754908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:11.754908Z digest=sha256:624c8442ee3bb701bf12c4dd1dd4b3e43b8c2c7b6b9c65c1d844218c523358ee

Observation 8b79f121-43a9-444a-84bd-23a83d6a3c5c · outbound

This paper cites Please find our image editing results in Fig.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation Please find our image editing results in Fig

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:11.930219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:11.930219Z digest=sha256:0addbcfdbfe2a51e989bf5c72c391b457d825aa7ee3f909abd5f0a13b2480f47

Observation 0133ca88-3140-498a-b29e-eade5444e621 · outbound

This paper cites Storyboard Generation Evaluation For a comprehensive evaluation on our many-to-many set- ting, we choose storyboard generation to report numerical metrics.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation Storyboard Generation Evaluation For a comprehensive evaluation on our many-to-many set- ting, we choose storyboard generation to report numerical metrics

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:12.052977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:12.052977Z digest=sha256:a559730eff62efa8c850b31cc2352ec2f7911656f99111729f3b0d2515e68342

Observation 72300f1b-3abf-465a-9cf1-e0b7f802ec3d · outbound

This paper cites vi- sual sentences,.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation vi- sual sentences,

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:12.160381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:12.160381Z digest=sha256:ddba2bb15a3d60ea222634f24b54b1750508f81f65689b891f8c760711c6580a

Pith citing papers

Observation 4777e389-dbf3-463b-8163-6f5f9b762b5d · inbound

FreeStyle: Free Control of Style-Content Dual-Reference Generation from Community LoRA Mining cites this paper.

FreeStyle: Free Control of Style-Content Dual-Reference Generation from Community LoRA Mining iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation

Reference 99

Resolution
verified exact
arxiv_id, observed 2026-08-04T02:24:10.739854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-26T17:59:00.685869Z digest=sha256:87956879340ec95f3ee4a49ff2d71e705fdbbbeeae48824b0ac3ef1362a5403c

Observation ab42dc16-8c92-4fa1-8467-e2109f11d4cf · inbound

ControlRef: Efficient Layout-Guided Multi-Instance Generation via Anchored 4D-RoPE cites this paper.

ControlRef: Efficient Layout-Guided Multi-Instance Generation via Anchored 4D-RoPE iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T19:08:49.594504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:08:49.594504Z digest=sha256:d1fe4ecfee0235c3e0097eefe8ff1a8bc43e8323b63d79d3ac4df37cb03cd46c