Pith. sign in

Paper Citation Record · LEDGER

DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation

As of 13 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 1 inbound Pith citation observation for arXiv:2412.07589.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.07589 v2

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T18:45:41.356187Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-18T19:02:01.962726Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T19:02:48.647832Z

Reference resolution

52 of 52 outbound references displayed

  • verified exact1
  • verified fuzzy28
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 773722b5-c8ca-4726-ba22-04aad92fefdd · outbound

This paper cites manga109.

DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation manga109

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:45:42.156162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T18:45:41.105552Z digest=sha256:b98482a7fa0846fb5ff84b5cba3072fd823af8e46113b5875dfc1ba584248732

Observation 09dc404a-c473-4ed9-8251-311940ecfcf0 · outbound

This paper cites Meissonic: Revitalizing Masked Generative Transformers for Efficient High-Resolution Text-to-Image Synthesis.

DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation Meissonic: Revitalizing Masked Generative Transformers for Efficient High-Resolution Text-to-Image Synthesis

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T18:45:41.110816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:45:41.110816Z digest=sha256:e5dfc0fdb08d15a6ebcff211807b6c41d5b4b10c5bc2164d68f995ae7214c23f

Observation 0ecc44a4-dd73-40a4-ad60-580dee0df27f · outbound

This paper cites Anydoor: Zero-shot object-level im- age customization.

DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation Anydoor: Zero-shot object-level im- age customization

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T18:45:41.120749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:45:41.120749Z digest=sha256:4b05f7808bafe51f0be3e5b59f7d4b605930c09062b7f064c507241c24782b2e

Observation 8953e0ff-0aab-419f-81af-be3b77de731a · outbound

This paper cites AutoStudio: Crafting Consistent Subjects in Multi-turn Interactive Image Generation.

DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation AutoStudio: Crafting Consistent Subjects in Multi-turn Interactive Image Generation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T18:45:41.125747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:45:41.125747Z digest=sha256:9e821f089a01f19774cd588078f8e95379fb8ed495636f4998018fed59eb8591

Observation 042f6b93-1a6c-4514-9dd0-6cd8b9e32425 · outbound

This paper cites Guiding instruction-based im- age editing via multimodal large language models.

DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation Guiding instruction-based im- age editing via multimodal large language models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T18:45:41.130770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:45:41.130770Z digest=sha256:b30c6d5123e1b5141ea9c213ae21bb341db7b7e2095291a4de395692cc35f68f

Observation b0559dba-793b-46ec-a997-0335231e905b · outbound

This paper cites An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion.

DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T18:45:41.135530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:45:41.135530Z digest=sha256:4d77163d44075c49478116912337ad5975d9ffefcf2732bda348eceff5236d9f

Observation 38aecb32-e63b-4caa-b7f0-98e31cb6a19e · outbound

This paper cites SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation.

DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T18:45:41.140601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:45:41.140601Z digest=sha256:e2f06a0271d6a28a135157242da30e99bbc6bdce15a66f14d0418e027094668c

Observation cf510d77-bf79-4660-93eb-1e9ae3fb968c · outbound

This paper cites Mix-of-show: Decentralized low- rank adaptation for multi-concept customization of diffusion models.

DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation Mix-of-show: Decentralized low- rank adaptation for multi-concept customization of diffusion models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:45:42.124010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T18:45:41.145801Z digest=sha256:19ac2020d92eabcd9a55a7b7660b08a3fde4eaab5bfe52ffa079c1e5c1199b58

Observation c7a1a330-130f-4241-9a1b-c1dababf9763 · outbound

This paper cites Imagine this! scripts to composi- tions to videos.

DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation Imagine this! scripts to composi- tions to videos

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:45:42.109739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T18:45:41.150811Z digest=sha256:802bbf7f7e21311881f5ae41ea806eb50d2b4c2a18f0ceb999f56eb55dee2d50

Observation 72a46c8e-03a5-4cd9-8422-c2d781dc0039 · outbound

This paper cites A Generalist FaceX via Learning Unified Facial Representation.

DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation A Generalist FaceX via Learning Unified Facial Representation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T18:45:41.155966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:45:41.155966Z digest=sha256:d71c55937fab7290581673b243255643b9d84724005c9aea4f775273de153057

Observation feb4980f-38fe-4a9b-ac37-eb21b2ccd753 · outbound

This paper cites Face adapter for pre-trained diffusion models with fine- grained id and attribute control.

DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation Face adapter for pre-trained diffusion models with fine- grained id and attribute control

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:45:42.095272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T18:45:41.160934Z digest=sha256:7bee6ba4352704902f9bbe4d38cc0937e2572044bc9d9a6599a4c640754634cf

Observation f94fc723-4a74-4a31-9bc2-ac07888183fa · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilib- rium.

DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation Gans trained by a two time-scale update rule converge to a local nash equilib- rium

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:45:42.080150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T18:45:41.165972Z digest=sha256:61fbad8236933a2b95c76a6abe76b4d8fc4e5a0e1fcfc3d92a73e95813304e5a

Observation d10fdce0-62ed-4181-8e76-f5e41b69593d · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation LoRA: Low-Rank Adaptation of Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T18:45:41.170571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:45:41.170571Z digest=sha256:d8572684090be8f9839a3771ada070836608ed93b9aae1339259076c4855f81a

Observation 0f0143e9-d1b0-4d58-a821-9066f587e2af · outbound

This paper cites Visual storytelling.

DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation Visual storytelling

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:45:42.065506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T18:45:41.175482Z digest=sha256:e80af4ab80a78cf22843470678cb79a692016200fd19ee4e33086c4dec33c390

Observation 7d4b624a-9cfe-408b-a36d-770db086e425 · outbound

This paper cites Smartedit: Exploring com- plex instruction-based image editing with multimodal large language models.

DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation Smartedit: Exploring com- plex instruction-based image editing with multimodal large language models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:45:42.050567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T18:45:41.180002Z digest=sha256:2eef6935bcf7ea9ad91766ba7595978e5b53b8265536a69e3c88cb153e870ca1

Observation 9a93f684-3463-4ac4-95ca-4bd167c28f57 · outbound

This paper cites Announcing black forest labs, 2024.

DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation Announcing black forest labs, 2024

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:45:42.035948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T18:45:41.184441Z digest=sha256:9e910c524431b022a7a39ea77043037519e1ec7f4a1e8045266d63db55db581d

Observation 39662e98-5410-4bf7-b631-0ab6ebcc86b0 · outbound

This paper cites Storygan: A sequential conditional gan for story visu- alization.

DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation Storygan: A sequential conditional gan for story visu- alization

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:45:42.019867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T18:45:41.188857Z digest=sha256:4a196768d4e3299a6413210d0a33bafccc05c840c8bacaf638e575d0359be3b2

Observation c8efdead-3656-42f0-9ac9-2cc6803d7672 · outbound

This paper cites Photomaker: Customizing re- alistic human photos via stacked id embedding.

DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation Photomaker: Customizing re- alistic human photos via stacked id embedding

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:45:42.003796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T18:45:41.193235Z digest=sha256:ca4532c28b708c35e70a7bc369580414fc07fbbb1fd5445be5e4d07d71440610

Observation 7066cf9e-6c22-441e-92b3-b37107bb6332 · outbound

This paper cites Sketch2manga: Shaded manga screening from sketch with diffusion models.

DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation Sketch2manga: Shaded manga screening from sketch with diffusion models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:45:41.987661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T18:45:41.198232Z digest=sha256:715b7636226be52436320747ec24517d13653c881b7c6da72fe0196f7f7535ba

Observation 15a55421-9903-453f-bb84-5e59c5d9ed69 · outbound

This paper cites Intelligent grimm-open-ended vi- sual storytelling via latent diffusion models.

DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation Intelligent grimm-open-ended vi- sual storytelling via latent diffusion models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:45:41.973143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T18:45:41.202694Z digest=sha256:a86a5a783021b1fd00ce016a3645df24ee23b65c8a80484a9edaf265b22ef7ce

Observation de84f89a-25ed-420e-81e1-07403cf16c70 · outbound

This paper cites Visual Instruction Tuning.

DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation Visual Instruction Tuning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T18:45:41.207732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:45:41.207732Z digest=sha256:63cf46faf3b7c65ebb35332b532172a29d684551fbe3f28728a4ce23801dcde9

Observation 2eaab4a8-7e68-48d6-b17c-2dad4fcebee8 · outbound

This paper cites Decoupled Weight Decay Regularization.

DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation Decoupled Weight Decay Regularization

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T18:45:41.213276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:45:41.213276Z digest=sha256:85b973fcc411d7d7c3ff1cdf625ea66b3ac2fc1a3dc3249c8c8cdb3396e10fc0

Observation 35842c52-c4c3-432c-80dc-2d57d42fc901 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation DINOv2: Learning Robust Visual Features without Supervision

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T18:45:41.218197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:45:41.218197Z digest=sha256:b9042b169c1440932574f59c0bb6728d9cd5d3c4ee3d9e15ddb6365086a4e730

Observation f33af489-9f6f-42dd-bc1f-65dc11ff6177 · outbound

This paper cites Synthesizing coherent story with auto-regressive la- tent diffusion models.

DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation Synthesizing coherent story with auto-regressive la- tent diffusion models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:45:41.957447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T18:45:41.222919Z digest=sha256:e41aacf152e466421862a89ac8d02226ae4c9b3f6da11d11c9e1cb8fee6792bd

Observation 1b265ddb-ea1b-460c-b55b-5a034c51b40b · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T18:45:41.227442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:45:41.227442Z digest=sha256:b72d0e5d376ed7b02630382aafbfe1c5eb91d19266c8bf16d4a58b2f45fdf9d1

Observation 1785eb3b-9cd8-42e8-a303-d67ae6bbcd3e · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation Learn- ing transferable visual models from natural language super- vision

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:45:41.942102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T18:45:41.232390Z digest=sha256:6dc3bece808b7c87b1ccc495a8341c0cab1a3a1430d4484541d867f9992f57d8

Observation 1d37f950-01e3-4ac5-9791-5e375e91f9ef · outbound

This paper cites High-resolution image syn- thesis with latent diffusion models.

DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation High-resolution image syn- thesis with latent diffusion models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T18:45:41.237857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:45:41.237857Z digest=sha256:fccf174e8ff781b7394ed59e16dddc673a0045fd4401c5735b6cf171805d7632

Observation f6f91206-ea74-47db-b11a-34011e0303a9 · outbound

This paper cites Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation.

DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T18:45:41.242483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:45:41.242483Z digest=sha256:9aac599c3339f856db12a4de9e3ae12d18f6968dd89c1e2c4e38a661985a1bb0

Observation 05520532-f666-4d4d-be16-9cfc745ae795 · outbound

This paper cites The manga whis- perer: Automatically generating transcriptions for comics.

DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation The manga whis- perer: Automatically generating transcriptions for comics

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:45:41.907651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T18:45:41.247185Z digest=sha256:410a26d46b7528eb245ec63495721572aecc4274a7629550e79cb6d2c2eb0df6

Observation 15b4a452-d8c3-4081-87aa-316d192f256d · outbound

This paper cites Tails Tell Tales: Chapter-Wide Manga Transcriptions with Character Names.

DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation Tails Tell Tales: Chapter-Wide Manga Transcriptions with Character Names

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-11T18:45:41.527477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T18:45:41.251697Z digest=sha256:8eea3391dfc9cb14660775f986c64f427eb23ea00c55f92b43abe3dfc410929f

Observation a3503d35-712f-441d-a6c9-951752f1bc66 · outbound

This paper cites DreamRelation: Bridging Customization and Relation Generation.

DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation DreamRelation: Bridging Customization and Relation Generation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T18:45:41.256153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:45:41.256153Z digest=sha256:84002c75afc1148127c1675ecb149bcf76f8b2fd449ea1794252c2604b8495a6

Observation c43dd1c0-6ff5-482c-9357-deaf35d1d481 · outbound

This paper cites Mangagan: Unpaired photo-to-manga transla- tion based on the methodology of manga drawing.

DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation Mangagan: Unpaired photo-to-manga transla- tion based on the methodology of manga drawing

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:45:41.892212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T18:45:41.260946Z digest=sha256:5db633e2d430836d0a38f37dd3023afb000791b128ece8c0eff95c62c66eb496

Observation 893ebc0a-2a88-43d1-a486-d9272179f7bc · outbound

This paper cites Generative multimodal models are in-context learners.

DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation Generative multimodal models are in-context learners

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:45:41.874894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T18:45:41.266358Z digest=sha256:2151dbb4adc5b4eb38e73098bc22e6ac7af13bc2dc9af80c25770bd094efb511

Observation da4e0ec0-b27f-46d1-bdb1-3655d2984f80 · outbound

This paper cites One missing piece in Vision and Language: A Survey on Comics Understanding.

DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation One missing piece in Vision and Language: A Survey on Comics Understanding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T18:45:41.271087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:45:41.271087Z digest=sha256:0e944e5a734c2067605a206820ac62ff662d6a4b69166e36c3e8bec03dcd89db

Observation 10d5035c-fe53-4954-89b2-29dd1b12630d · outbound

This paper cites InstantID: Zero-shot Identity-Preserving Generation in Seconds.

DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation InstantID: Zero-shot Identity-Preserving Generation in Seconds

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T18:45:41.276182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:45:41.276182Z digest=sha256:13bbeb05fde31d3ccc2b2fb6e51e74eeb84baca67daa8a25b362ee549857c61d

Observation 5167e3a6-5333-48d8-8d9d-aead243697cf · outbound

This paper cites AutoStory: Generating Diverse Storytelling Images with Minimal Human Effort.

DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation AutoStory: Generating Diverse Storytelling Images with Minimal Human Effort

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T18:45:41.281026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:45:41.281026Z digest=sha256:94f5299ac822e2ae1846f6cb8674223b7a4a014c598ca08b4fbc23b4f17bac1d

Observation 2412e76d-bd1c-4386-aa83-b8d0a1456698 · outbound

This paper cites Instancediffusion: Instance-level control for image generation.

DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation Instancediffusion: Instance-level control for image generation

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:45:41.859815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T18:45:41.285595Z digest=sha256:a6f9ea652e958b146dc8300108109a1597185f123a1e6106ad7615fc666e61fb

Observation 5c9e2ff2-4851-4628-98ba-d9049b3e2f36 · outbound

This paper cites MS-Diffusion: Multi-subject Zero-shot Image Personalization with Layout Guidance.

DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation MS-Diffusion: Multi-subject Zero-shot Image Personalization with Layout Guidance

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T18:45:41.290187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:45:41.290187Z digest=sha256:6348c898dbbbfaf0950f59adf32189cde3579519b3365e3b6874cdfa52eb3530

Observation e2195d38-288a-4ffb-b1a5-44f7985bcd0d · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation Emu3: Next-Token Prediction is All You Need

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T18:45:41.294864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:45:41.294864Z digest=sha256:984f539582acc5817b713b40c7e69e027ec1654628ae854524aeb2a7d8620c06

Observation 3f451203-056f-48d8-bef4-90808f168a5c · outbound

This paper cites Genartist: Multimodal llm as an agent for unified image gen- eration and editing.

DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation Genartist: Multimodal llm as an agent for unified image gen- eration and editing

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:45:41.843972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T18:45:41.299529Z digest=sha256:c2a05062b01a23a05c2ad3e496254cc9f6385c3c81bde3a2607eadb9a48e2bfe

Observation 0e21964f-f30e-4a25-863e-e03236d8a4b3 · outbound

This paper cites Towards language-driven video inpainting via multimodal large language models.

DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation Towards language-driven video inpainting via multimodal large language models

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:45:41.828426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T18:45:41.304207Z digest=sha256:b4dc2797bf9dbe8c65a3d65cac3372451641bc7c553c45eac510cb131c30cd00

Observation a8541bb7-9eed-493e-8155-c920cba572a3 · outbound

This paper cites Towards open vocabulary learning: A survey.

DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation Towards open vocabulary learning: A survey

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:45:41.813274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T18:45:41.308953Z digest=sha256:8fdac7a24f54a67476b6511ab0daca48701ca9d2463ef1fc08119f9a8f98c7eb

Observation d7b7c7f3-b8c7-43a0-b7ca-6dcce347f284 · outbound

This paper cites Mo- tionbooth: Motion-aware customized text-to-video genera- tion.

DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation Mo- tionbooth: Motion-aware customized text-to-video genera- tion

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:45:41.797095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T18:45:41.314153Z digest=sha256:9ef2bb1ba6b7a6d29450590eeb2a1fccea640ee9a4e64cf49df6d3cb5a3798c1

Observation 1ac71443-3422-45bb-a9c5-afb609291f86 · outbound

This paper cites Boxdiff: Text-to-image synthesis with training-free box-constrained diffusion.

DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation Boxdiff: Text-to-image synthesis with training-free box-constrained diffusion

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:45:41.782276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T18:45:41.318761Z digest=sha256:2d2db9f3a06334824deadb7c7064de6b455dc9c1b847d8e29ed8b0dbfb06b2dc

Observation 5c8e3fd8-4ab9-4867-a02e-6da472943178 · outbound

This paper cites SEED-Story: Multimodal Long Story Generation with Large Language Model.

DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation SEED-Story: Multimodal Long Story Generation with Large Language Model

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T18:45:41.323022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:45:41.323022Z digest=sha256:f17910845e29f52a85ffc3acd9d41a741db2c12c952a319faa0352846da08606

Observation 772b6d5b-2d2e-4c8c-8ea1-2fe7bc3032b1 · outbound

This paper cites IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models.

DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T18:45:41.327685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:45:41.327685Z digest=sha256:c62eca51d3abeaa1f9a2318c242eaa6c4f678c91234baf070746493ddddae08f

Observation 441f541f-b9f5-411c-9415-c93aa10e5888 · outbound

This paper cites Ai-driven background generation for manga illustrations: A deep generative model approach.

DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation Ai-driven background generation for manga illustrations: A deep generative model approach

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:45:41.766703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T18:45:41.332599Z digest=sha256:7b1dcfdf61cca86a0067f8252b1899505995e3dd72dfed9666499fd51fbffd29

Observation 5900f635-7cb0-4b9c-aa86-02c66ccd6e0a · outbound

This paper cites Magicbrush: A manually annotated dataset for instruction- guided image editing.

DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation Magicbrush: A manually annotated dataset for instruction- guided image editing

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:45:41.751701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T18:45:41.337047Z digest=sha256:72a530a0a6c1fee2ca14b0935bf20c739980246f181aba05a4f529ca1ce276a7

Observation 2738b748-fe71-429f-9f5a-1418b1e90656 · outbound

This paper cites Generating manga from illustrations via mimicking manga creation workflow.

DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation Generating manga from illustrations via mimicking manga creation workflow

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:45:41.736194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T18:45:41.341911Z digest=sha256:3696ff059c4947b01273b79d3bcc77496e62046d05207aebc8cfc8eac8b643f9

Observation 307b0686-c2b0-469e-9d30-85228e859f9c · outbound

This paper cites Adding conditional control to text-to-image diffusion models.

DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation Adding conditional control to text-to-image diffusion models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T18:45:41.346403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:45:41.346403Z digest=sha256:bdcf74a6efa393ce0423a1b539757432a672ed2c0297321228efca23aaac74ba

Observation 43d6116f-c0e3-4478-8e46-862bc698980c · outbound

This paper cites Cus- tomization assistant for text-to-image generation.

DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation Cus- tomization assistant for text-to-image generation

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:45:41.710555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T18:45:41.351553Z digest=sha256:663c5fd85544bba4175431b303685c3710381ac96c9c39ce50301d3688d60f75

Observation 76b7bd0c-20ef-416e-bf54-2321f1e90161 · outbound

This paper cites page results.pdf.

DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation page results.pdf

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:45:41.694419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T18:45:41.356187Z digest=sha256:11e42550d55dcf110d8e6435590a2513efbcdaefca6f2182d77d7de3002de630

Pith citing papers

Observation 314501bf-3cad-4343-9640-d7315b222c17 · inbound

TaleDiffusion: Multi-Character Story Generation with Dialogue Rendering cites this paper.

TaleDiffusion: Multi-Character Story Generation with Dialogue Rendering DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:02:48.651255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T19:02:01.962726Z digest=sha256:40438d65126d50987dd4a2bfc1077f898b2780ecfb55a529e2fd7d2c5957d71c