Pith. sign in

Paper Citation Record · LEDGER

Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling

As of 17 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 0 inbound Pith citation observations for arXiv:2507.16240.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.16240 v1

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:19:24.257877Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

57 of 57 outbound references displayed

  • verified exact0
  • verified fuzzy33
  • unresolved23
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 82e22da6-7af7-4216-b8cb-68f35c4886d9 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T15:19:23.901137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:19:23.901137Z digest=sha256:a9386edd567d96858c6679f17c51953ace80a9323fe6a2bc04fd76520979a2b5

Observation 7d1a0870-e191-4c3b-8b64-e96107903fa2 · outbound

This paper cites an unresolved cited work.

Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:19:25.741836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:19:23.907126Z digest=sha256:f8ab129bfffcaf4bf2215fd22390c359e45444d8c867091ebde08f760eff3899

Observation fe6fee38-1fc1-4cc1-b5e5-f64c74966f7d · outbound

This paper cites Masactrl: Tuning-free mu- tual self-attention control for consistent image synthesis and editing.

Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling Masactrl: Tuning-free mu- tual self-attention control for consistent image synthesis and editing

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T15:19:23.913095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:19:23.913095Z digest=sha256:1c0c973309dd15e43f0b497d037f4d137b4f6146afafa247b067628177dbe0d4

Observation 3b863ed6-9f6e-4267-9dee-a1bea7b3ac69 · outbound

This paper cites an unresolved cited work.

Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:19:25.697120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:19:23.921432Z digest=sha256:f118c665c590036e6f1992242e7b50a70227939eeca8ec7f3826d3a741d09f88

Observation 97b3ee4c-60db-4403-850d-3c4d59aa7f3b · outbound

This paper cites Attend-and-excite: Attention-based se- mantic guidance for text-to-image diffusion models.

Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling Attend-and-excite: Attention-based se- mantic guidance for text-to-image diffusion models

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:19:25.659950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:19:23.927577Z digest=sha256:bb039fd10951d1cede2eb4d5892cfb782cfbe0cc452797c66f55d3340ea217e6

Observation 58746ee0-3aa5-44d6-aba6-22f43eac1092 · outbound

This paper cites How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites.

Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T15:19:23.936196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:19:23.936196Z digest=sha256:0131e9780f6431a6b6771d8d8068c4b78472edc35b4a1814984bc6f109fd3e9f

Observation 96a8ed07-f74d-44b4-b552-fc49584b0716 · outbound

This paper cites Intern vl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling Intern vl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:19:25.618397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:19:23.942285Z digest=sha256:b4e1748b6eeaa686bb0fd4bf4c7c41fdfc87be946e82b4ed3c9729ddee840af0

Observation bd4a5c15-87db-4a8c-aef9-6860fbff8ea5 · outbound

This paper cites Tam- ing Transformers for High-Resolution Image Synthesis.

Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling Tam- ing Transformers for High-Resolution Image Synthesis

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:19:25.592095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:19:23.948584Z digest=sha256:1dd8bad6277af986efe83a29f5cca471cf0b274b0fe1834a0980a71c8e5b2087

Observation f16ef858-c072-4bf0-a1f3-15e9f967ddb5 · outbound

This paper cites Scaling rec- tified flow transformers for high-resolution image synthesis.

Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling Scaling rec- tified flow transformers for high-resolution image synthesis

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:19:25.562607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:19:23.955342Z digest=sha256:5be74f62150ae6ebfa1bf75a7ce0ca727bdc6d8b1cca8bc31b015765bb5ac979

Observation d79e8615-81de-4136-a2e1-453e8157184d · outbound

This paper cites Focus on your instruction: Fine-grained and multi-instruction image editing by atten- tion modulation.

Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling Focus on your instruction: Fine-grained and multi-instruction image editing by atten- tion modulation

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:19:25.537902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:19:23.962728Z digest=sha256:2528583605b72515e496075343e3294930b614e025ac8bb473aa071cb5b652ef

Observation 5f1fdf2c-5dc9-49a2-81ea-6a41ba5f6d95 · outbound

This paper cites Prompt-to-prompt image editing with cross-attention control.

Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling Prompt-to-prompt image editing with cross-attention control

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:19:25.505743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:19:23.967636Z digest=sha256:3e02816e5584a22fcd31c88b588ff7b2880444973d703e51dd44c48f5dfadd49

Observation 4066999a-c684-4a0f-b306-d7449d211313 · outbound

This paper cites Classifier-Free Diffusion Guidance.

Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling Classifier-Free Diffusion Guidance

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T15:19:23.975314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:19:23.975314Z digest=sha256:01545f87d24d5fffa40d745fc2f05b3ec3e8f0e29821c17752a0330b16e7b65e

Observation ada86dd1-736d-4b06-95f1-1fcb31f4c771 · outbound

This paper cites An edit friendly ddpm noise space: Inversion and manipulations.

Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling An edit friendly ddpm noise space: Inversion and manipulations

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:19:25.477070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:19:23.981305Z digest=sha256:1abe6b6add562724f1776f3c430fe5fb9e1f8f621c99f78a07cad4b22f0a9420

Observation d86c799f-3a2d-49f0-ad32-2eae8ef5e2d4 · outbound

This paper cites Pnp inversion: Boosting diffusion-based editing with 3 lines of code.

Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling Pnp inversion: Boosting diffusion-based editing with 3 lines of code

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:19:25.449443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:19:23.987784Z digest=sha256:0f48140d90654a42b0f8d60fb91282e523986e8cd5f8799d93b9a29c53a1a8e3

Observation 42272f09-424f-4567-976e-8fde399d364d · outbound

This paper cites What's in the Image? A Deep-Dive into the Vision of Vision Language Models.

Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling What's in the Image? A Deep-Dive into the Vision of Vision Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T15:19:23.993702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:19:23.993702Z digest=sha256:957c367f1693fbf0476d1ebdcf702933dba3d357fc0a8f9af50f6f0b21a69557

Observation 097f2fe7-c164-4d62-83b6-09d13aae35f2 · outbound

This paper cites Diffusion models for open-vocabulary segmenta- tion.

Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling Diffusion models for open-vocabulary segmenta- tion

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:19:25.419848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:19:23.999720Z digest=sha256:5822eac441fb9de25a35581c2fdc8a173455c1234136f6b6aa7b694cece10952

Observation f6d84c08-5432-4388-b903-c5698cb01ac4 · outbound

This paper cites Pick-a-pic: An open dataset of user preferences for text-to-image generation.

Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling Pick-a-pic: An open dataset of user preferences for text-to-image generation

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:19:25.392231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:19:24.006128Z digest=sha256:dcc6bfae0b17c677de97007f084176f32d3baa8d499d230b2422dfb6634b9173

Observation a5524cf3-931c-4922-804a-d7a65949e9e3 · outbound

This paper cites Open-vocabulary object segmenta- tion with diffusion models.

Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling Open-vocabulary object segmenta- tion with diffusion models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:19:25.370876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:19:24.012496Z digest=sha256:66d5a686f17143709538d4c652e0e49e8a13366ac5ae01afece32c614b346f06

Observation 1d1f865b-3ce7-4e67-b5b3-9c9fa565ce0e · outbound

This paper cites an unresolved cited work.

Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:19:25.341640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:19:24.017126Z digest=sha256:e37006042c4ef1c0ec2379327d7a42b5bfa7270f63ae5b4a7e3f736b30b69392

Observation 3701e39a-af9f-43a6-9586-38f2c5d8f78c · outbound

This paper cites ACE++: Instruction-Based Image Creation and Editing via Context-Aware Content Filling.

Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling ACE++: Instruction-Based Image Creation and Editing via Context-Aware Content Filling

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T15:19:24.022111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:19:24.022111Z digest=sha256:e41d43838466b100d4ca6fa08b72a841ee1589884f855b18a99c520891d904af

Observation cd0aa095-1217-4e07-8acc-546c199af7f4 · outbound

This paper cites Null-text inversion for editing real images using guided diffusion models.

Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling Null-text inversion for editing real images using guided diffusion models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T15:19:24.029689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:19:24.029689Z digest=sha256:a71b77cb61e7a8b79f7aaa4042cbb694dd10449827372fd05ff6a5bd3dbb58ff

Observation f040496e-dbc1-4d1a-a091-e996500aee35 · outbound

This paper cites T2i-adapter: learning adapters to dig out more controllable ability for text-to-image diffusion models.

Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling T2i-adapter: learning adapters to dig out more controllable ability for text-to-image diffusion models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:19:25.297465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:19:24.035152Z digest=sha256:3b6b649aff472a8a6b11c4ffcba367c81dca011a92c9fbacfcafcfe6c129f89f

Observation 8500d275-4dda-4a51-a83b-76ac18f3892f · outbound

This paper cites an unresolved cited work.

Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:19:25.273496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:19:24.043386Z digest=sha256:7617aa90a7fabb1d6f7a27e2f44c586e8677c62d4738032a7e8294deddd1d647

Observation 0970514e-b682-4cfc-b6dc-5f74142b3d3f · outbound

This paper cites DiffuseV AE: Efficient, controllable and high- fidelity generation from low-dimensional latents.

Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling DiffuseV AE: Efficient, controllable and high- fidelity generation from low-dimensional latents

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:19:25.249347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:19:24.050306Z digest=sha256:7bce02d487e9e3842e49dfc32b7d9afafc1d34d553c748bd982e362b0ce05169

Observation 21907f3e-f9aa-4ba9-a4bc-8be1e5df7a37 · outbound

This paper cites Peebles and Saining Xie.

Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling Peebles and Saining Xie

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:19:25.219844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:19:24.058975Z digest=sha256:00be5afd2a937267c9690776c38616b55f6136934e0571f70c53ab37b9978ab1

Observation b1bfc2cc-a40a-4866-bd03-75deb27861a3 · outbound

This paper cites SDXL: Improving latent diffusion models for high-resolution image synthesis.

Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling SDXL: Improving latent diffusion models for high-resolution image synthesis

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:19:25.186321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:19:24.067815Z digest=sha256:5c3c6a0afb43d3946ce2f82fb333b37dd8bb06596147926970ca5e3d2c707228

Observation 27ae3419-6c5a-42ff-b383-2305643ffa07 · outbound

This paper cites Unicontrol: A unified diffusion model for controllable visual generation in the wild.

Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling Unicontrol: A unified diffusion model for controllable visual generation in the wild

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:19:25.137318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:19:24.073913Z digest=sha256:e71818fc1572d44490a63395bd0af1b1db59f660dc702569a37c1efd856862bc

Observation 469fd176-17d5-4835-9962-6f84226b7639 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling Learning transferable visual models from natural language supervision

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:19:25.098284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:19:24.079415Z digest=sha256:a3b5b6fca03e0d9e98ca5ebbfe37c09fa2a37255a1ac653a2b410e7e708728d1

Observation 6886e00c-110a-4705-b517-6afb3e335e1c · outbound

This paper cites an unresolved cited work.

Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:19:25.042318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:19:24.084085Z digest=sha256:332c40c4857f249a2dede29bcb4b2bc8c1de807ae8fd9cfad59e3e51f7f87197

Observation 421ebc2c-84f8-4970-9ca1-4e4c396c1f21 · outbound

This paper cites High-Resolution Image Synthesis with Latent Diffusion Models.

Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling High-Resolution Image Synthesis with Latent Diffusion Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T15:19:24.090828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:19:24.090828Z digest=sha256:d39af838e580c1c122ae37411117a518cd30ad8f76181792e23bec2a3eefdf87

Observation 73eb8019-e0dc-4a69-a8c6-898afb802334 · outbound

This paper cites Semantic Image Inversion and Editing using Rectified Stochastic Differential Equations.

Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling Semantic Image Inversion and Editing using Rectified Stochastic Differential Equations

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T15:19:24.106867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:19:24.106867Z digest=sha256:84b914d5677beef2be767f89bbe854b8cbf121fec5e9054f8510357377252ea7

Observation f8f50047-6e17-4832-97e4-c01b7d63384b · outbound

This paper cites Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation.

Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:19:24.973539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:19:24.114803Z digest=sha256:1d29d84299ad16eacf5317288c444a6f16e12b551c3cef409fdd4ff1a7afe91c

Observation 63c812f5-c7a9-4169-ad91-63d03f382504 · outbound

This paper cites Emu edit: Precise image editing via recognition and gener- ation tasks.

Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling Emu edit: Precise image editing via recognition and gener- ation tasks

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:19:24.943142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:19:24.122529Z digest=sha256:37192bf7b740885b7539bfd970defdd0560173fc0a6fb634c4dd25c47edb9cf8

Observation 0dc44844-861b-479d-9e21-61efb79c9bb3 · outbound

This paper cites Generative Multimodal Mod- els are In-Context Learners.

Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling Generative Multimodal Mod- els are In-Context Learners

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:19:24.916902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:19:24.127028Z digest=sha256:4ff84d3aa47a0dbcf45876e9ec340be15c4b0aa5c3baac72a6f553bc3574ea32

Observation 7830b0de-7e90-4287-a571-4de59968d0a3 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T15:19:24.133948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:19:24.133948Z digest=sha256:b9858c5475ccb18854453302263ec5f2731850e0d5027b1acdce95c4c339d20a

Observation defb99e8-7f5c-4709-9320-4922ace41136 · outbound

This paper cites Qwen2.5-vl, 2025.

Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling Qwen2.5-vl, 2025

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:19:24.889930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:19:24.141983Z digest=sha256:2448caabbe10325a8be9cfff8d15e12a76461384f71b227c7286a1b71bf90172

Observation 66d5f01d-cf16-449a-84dc-015b32c9b60c · outbound

This paper cites MIGE: Mutually Enhanced Multimodal Instruction-Based Image Generation and Editing.

Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling MIGE: Mutually Enhanced Multimodal Instruction-Based Image Generation and Editing

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T15:19:24.148030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:19:24.148030Z digest=sha256:322bbfca608289a35603ee91a89a7a578adbc695ac3c5c17a13857506a922a30

Observation 7effc298-eb7c-4254-9f1e-e2bfdf993674 · outbound

This paper cites Plug-and-play diffusion features for text-driven image-to-image translation.

Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling Plug-and-play diffusion features for text-driven image-to-image translation

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:19:24.862758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:19:24.154545Z digest=sha256:60bd23fcf79ce6e474c26cc0ded5e75f11bad1c4f847c0aa9215fac4a96ce6d9

Observation 0b2db0eb-c4f8-4d16-9084-39298d544190 · outbound

This paper cites Neural discrete representation learning.

Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling Neural discrete representation learning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:19:24.840158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:19:24.160838Z digest=sha256:b4f61dbde7c032b3e3a1f36f689088cde5866b4d0e5e1c598182e2ab34555275

Observation 78592b5c-cc42-430e-b650-b8b4c82c822f · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T15:19:24.166327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:19:24.166327Z digest=sha256:3ea3ecb873c98ab5b4ad4d7e415d765c26719f2cb2011191e2738590e2d92fbf

Observation af51db51-5bb3-4eaa-9322-4323b5386b56 · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling Emu3: Next-Token Prediction is All You Need

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T15:19:24.172573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:19:24.172573Z digest=sha256:a700f7993b69f683af260c1ac1e0f2c261a8e15ea987b2df63cdce3ddc99357e

Observation db83409d-11ae-4438-ad2f-ce4e2896f416 · outbound

This paper cites Hairclipv2: Unifying hair editing via proxy feature blending.

Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling Hairclipv2: Unifying hair editing via proxy feature blending

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:19:24.814086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:19:24.177280Z digest=sha256:f1c37cb321547b19d677f089adf9b2df7bfbeff48efe98b945439a3d05ab13ac

Observation 4de35d14-f731-4746-ba70-481e00658dbf · outbound

This paper cites Enhancing MMDiT-Based Text-to-Image Models for Similar Subject Generation.

Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling Enhancing MMDiT-Based Text-to-Image Models for Similar Subject Generation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T15:19:24.182520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:19:24.182520Z digest=sha256:1eb98c6b7cbe4d7fd2b030cdada5b50d3407f4e9262505579e226c54c815820f

Observation a6d14d49-943c-4a8f-9043-a4b9bb487e02 · outbound

This paper cites FreeFlux: Understanding and Exploiting Layer-Specific Roles in RoPE-Based MMDiT for Versatile Image Editing.

Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling FreeFlux: Understanding and Exploiting Layer-Specific Roles in RoPE-Based MMDiT for Versatile Image Editing

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T15:19:24.190577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:19:24.190577Z digest=sha256:0ab873f38800b9f38f6137fbbadab2b9401f1ddf930743f8336a637d2ba31be8

Observation 57d084f3-01eb-4f6e-acee-5fd3e93b8c5f · outbound

This paper cites Diffumask: Synthesizing images with pixel-level annotations for semantic segmentation using dif- fusion models.

Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling Diffumask: Synthesizing images with pixel-level annotations for semantic segmentation using dif- fusion models

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:19:24.790384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:19:24.196217Z digest=sha256:cd21426de1cf59647b42b6be116a87e37f936787eacce4edd02564cac07194c4

Observation 6ea76899-4891-4a1f-b045-4b18f4c644e7 · outbound

This paper cites OmniGen: Unified Image Generation.

Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling OmniGen: Unified Image Generation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T15:19:24.201599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:19:24.201599Z digest=sha256:d932c9ab32156604c3395726365d33fe85809c6718ecfdc66f1a1f909bbfee56

Observation 06795431-7097-48ee-9292-6642fcfa5a40 · outbound

This paper cites Show-o: One single transformer to unify multimodal understanding and generation.

Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling Show-o: One single transformer to unify multimodal understanding and generation

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:19:24.772357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:19:24.208916Z digest=sha256:021f0a7a155f89e1671b81d6db405a124a24131cdfaf3ed621a268a6969a0bd2

Observation 41f2d3de-e58c-4b5b-93c1-a8b8110a45f7 · outbound

This paper cites Characteristic analysis of otsu threshold and its ap- plications.

Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling Characteristic analysis of otsu threshold and its ap- plications

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:19:24.753503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:19:24.215130Z digest=sha256:43f7eb711a2a1889210bab3fb6d08971b818e953045f8327791b1eea5bbc385e

Observation ba3da3bd-99a7-4c67-87fa-f21e8ebb26f3 · outbound

This paper cites Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy.

Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T15:19:24.219539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:19:24.219539Z digest=sha256:5388ffae0d3cbe895d62e5d14dcf25c5da8005c80c76286145d514a68ad75c97

Observation bdf60ab8-15ef-4015-98c2-5c0f297e0f2b · outbound

This paper cites RoCC: Robust Covert Communication Based on Cross-Modal Information Retrieval.Journal of Im- age and Graphics, 29(2):369–381, 2024.

Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling RoCC: Robust Covert Communication Based on Cross-Modal Information Retrieval.Journal of Im- age and Graphics, 29(2):369–381, 2024

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:19:24.727379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:19:24.225189Z digest=sha256:5c3e08514b5e07310e61dd1825dfebc73be625c09e73dd0bedc3fd8583494b63

Observation 3a5bce58-4e60-4502-9821-184934fb97cc · outbound

This paper cites Peeling back the layers: Interpreting the story- telling of vit.

Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling Peeling back the layers: Interpreting the story- telling of vit

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:19:24.703322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:19:24.230896Z digest=sha256:e723ca6b71d8f1b67d9dbd665a0cf64b11e1564d5177d76775d177c1c6bdd995

Observation 10092bef-6215-4931-971b-fe381d049155 · outbound

This paper cites Magicbrush: A manually annotated dataset for instruction- guided image editing.

Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling Magicbrush: A manually annotated dataset for instruction- guided image editing

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:19:24.682925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:19:24.236139Z digest=sha256:a161e880e5ece269021aa7d080f4487845e99ca8439521cd2aa4c6adfdf3b8aa

Observation 59d5aea2-faed-47da-b4b3-0cef14a2af6a · outbound

This paper cites Adding conditional control to text-to-image diffusion models.

Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling Adding conditional control to text-to-image diffusion models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T15:19:24.241243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:19:24.241243Z digest=sha256:ec7c83b5f1155ef6d28669b4fb19661d0d4d9de61ab90a160e695d82e0aabd2e

Observation 849468f0-4926-4eb2-868a-6bee860082a8 · outbound

This paper cites Ultraedit: Instruction-based fine-grained image editing at scale.

Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling Ultraedit: Instruction-based fine-grained image editing at scale

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:19:24.641989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:19:24.246620Z digest=sha256:d4fd19884b8c1b491d5f6e4972bfad33995a37717120908a69ecf79074c32ab4

Observation 1a218d12-f8db-4506-8b20-9a1ce3bc723f · outbound

This paper cites Semantic under- standing of scenes through the ade20k dataset.

Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling Semantic under- standing of scenes through the ade20k dataset

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:19:24.624514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:19:24.251392Z digest=sha256:97822a5722b8d5c06f37a70dc9f04899d96900216e35add5505bee1fade4b6ed

Observation c0c91d74-648f-497e-9ead-10c5fe6c2257 · outbound

This paper cites graffiti.

Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling graffiti

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:19:24.607230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:19:24.257877Z digest=sha256:2b4c8b62d54daab0433f48024fd7e7b1b244cb43ceaf22f6c9ae81d7964856f4

Observation 828e88a1-7c49-43a7-a490-377b6eabdc43 · outbound

This paper cites an unresolved cited work.

Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling Unresolved cited work

Reference 2022

Resolution
parse uncertain
raw_fallback, observed 2026-08-06T15:19:24.997503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:19:24.100705Z digest=sha256:7473785558948d9baf21dd9ff376872c01ce27c61f917e43ff55c5d39919aab4

Pith citing papers

No inbound Pith citation observations are available.