Pith. sign in

Paper Citation Record · LEDGER

Meta-CoT: Enhancing Granularity and Generalization in Image Editing

As of 11 August 2026, this Paper Citation Record lists 87 of 87 outbound references and 3 inbound Pith citation observations for arXiv:2604.24625.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.24625 v1

Coverage vector

measured 87 of 87 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-08T04:30:28.636915Z

measured 90 of 90 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-03T21:23:41.271521Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-03T21:28:58.249219Z

Reference resolution

87 of 87 outbound references displayed

  • verified exact51
  • verified fuzzy32
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2874846a-c909-4187-bb09-5518e054697c · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:41:19.360977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:b7875c257b5845bff0e8e70776a2aa8b3b3eb621d1672abaa2c106a14c4fc46c

Observation a78a23f0-fd8a-41dd-a087-df2d0e9e740c · outbound

This paper cites Qwen2.5-VL Technical Report.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Qwen2.5-VL Technical Report

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:41:19.383070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:68cbc2f846c335d74e56985d3c55f0fc14a8a31d762d876a70dcc1bfb6131606

Observation 9451513a-ac71-49f7-83b3-8e578c0c2cba · outbound

This paper cites In- structpix2pix: Learning to follow image editing instructions.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing In- structpix2pix: Learning to follow image editing instructions

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:53:02.747717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:c156a621a064de6c522545c7d9aa03f5d156358421d2c3e40c5b5ea33f24223a

Observation c64ff326-3b9b-4328-a666-044ec560aa0d · outbound

This paper cites BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:19.216281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:f1f13ce13e9b81708fec80bd03540c6cb4fe3fb81af05d4f1ef8545da4f76041

Observation abe6e895-6ab0-4119-bfb8-9dfd70fff197 · outbound

This paper cites Blip3o-next: Next frontier of native image generation.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Blip3o-next: Next frontier of native image generation

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:19.222593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:2631ab332d0e11a8219205a6888b0029cd3ef6e52ab84b1154d13b4a232b6902

Observation e8ea0ca8-cd9e-45c2-9baf-38e0b3f0acbf · outbound

This paper cites Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:46:14.055858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:e8b3460926ecc3882a7aa282eef53131f6214f83af2bd67ab8347d26a31310e5

Observation 38fdb81e-6299-4b00-bdb2-00d92e0fe034 · outbound

This paper cites ChatUMM: Robust Context Tracking for Conversational Interleaved Generation.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing ChatUMM: Robust Context Tracking for Conversational Interleaved Generation

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-06-02T04:04:28.060339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:f7f5003a9fe797f32bafad8bbfa8853e941bbaa181a009bb731b513cdac74f1d

Observation a67b3420-a60c-476f-94ec-7d97852d5a32 · outbound

This paper cites Emerging Properties in Unified Multimodal Pretraining.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Emerging Properties in Unified Multimodal Pretraining

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:46:15.005174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:6dc14cd36b879d04baed3b725d8813e2dcafca3f64348a8144cc2ceee94b85db

Observation e0f0ab6a-cd87-46e2-989e-2f305261a917 · outbound

This paper cites Dreamllm: Synergistic multimodal com- prehension and creation.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Dreamllm: Synergistic multimodal com- prehension and creation

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:53:02.721789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:15a0262dc25c0b79c7149adc291b5da52c7a0ec1e01645b0c382bf66cc1bb15f

Observation e500cf13-84a6-4c94-95e4-5c45e0173aef · outbound

This paper cites GoT-R1: Unleashing Reasoning Capability of MLLM for Visual Generation with Reinforcement Learning.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing GoT-R1: Unleashing Reasoning Capability of MLLM for Visual Generation with Reinforcement Learning

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:46:13.279370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:140f9d394e3ab54ff0a68092989ec16d0e29bfd9da28f145699be20f495ab267

Observation c251ec09-3171-4a66-8787-bfef63b28c40 · outbound

This paper cites GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:19.274947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:b1f8d6fee048dc1f7078cef57e33ac262aecceba47ff6d02cc630ae5a9851dba

Observation f704a458-45c9-43e1-8da5-c1b1fe9776f4 · outbound

This paper cites ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:46:12.382365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:816fb7b2946c47cda07b40655aa399917b762f13b0e98ff7fc03392e71dabb55

Observation 7cd4543e-76bc-4aef-82f1-3d065e399684 · outbound

This paper cites MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:19.313335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:afa4996e621dacfaa7b817e3c288c428385238a53819dd869b3b688476c5aa59

Observation 2e74d709-3e64-4ee9-9ac0-419bbc4f41fb · outbound

This paper cites Visual program- ming: Compositional visual reasoning without training.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Visual program- ming: Compositional visual reasoning without training

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:53:02.731617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:58993c24f32ac0eb3989527f0ae4e562f1dca3c9bf34c4d911c83bbda6607e1a

Observation 99275ec5-d445-46e7-863e-b3cad79f3f72 · outbound

This paper cites Controlthinker: Unveiling latent semantics for controllable image generation through visual reasoning.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Controlthinker: Unveiling latent semantics for controllable image generation through visual reasoning

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:46:13.486416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:513ced80161543fa51abec07e697516815481348b9cdc47f641077f63604c012

Observation a5a5a9e4-34e3-42e7-9d68-fa2c3498179d · outbound

This paper cites Freeedit: Mask-free reference-based image editing with multi-modal instruction.IEEE Transactions on Pattern Anal- ysis and Machine Intelligence.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Freeedit: Mask-free reference-based image editing with multi-modal instruction.IEEE Transactions on Pattern Anal- ysis and Machine Intelligence

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:53:02.738145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:27693a5c8cdcc4d598fc502b33d08f293b4b0bdf7b924635d3428303a790928f

Observation 88dc8d79-2fae-4b63-a436-cc0ab1b8eb68 · outbound

This paper cites Re-align: Structured reasoning-guided alignment for in-context image generation and editing.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Re-align: Structured reasoning-guided alignment for in-context image generation and editing

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:19.157170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:d35c335657975bc2757c83eac847b8c254afe0f14677da09f3e2a5f6383065eb

Observation f312a434-02c4-4b04-845b-00bc95d56784 · outbound

This paper cites Visual sketchpad: Sketching as a visual chain of thought for multimodal language models.NeurIPS, 37:139348–139379.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Visual sketchpad: Sketching as a visual chain of thought for multimodal language models.NeurIPS, 37:139348–139379

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:53:02.753884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:f9bfc5501c25740edf33991bd081332600233478696ed3a8941e5616c5045d46

Observation e3172cd7-0357-4060-9b32-6c9124654e42 · outbound

This paper cites Image Editing As Programs with Diffusion Models.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Image Editing As Programs with Diffusion Models

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:46:12.360685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:8b3155a746332cd35790b1f255a7764befb5d3725953539763ce36c39284d4bb

Observation f4689a43-976b-4aa1-9c08-4105a9fbbe3f · outbound

This paper cites Large Language Models Can Self-Improve.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Large Language Models Can Self-Improve

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-18T17:00:48.316458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:e32fe6efd02419e110a09a677ad607f8b13fa3d04acf225ece4affac5230ee93

Observation 78c1b218-e674-4b05-b64c-8811eeaa877c · outbound

This paper cites Interleaving Reasoning for Better Text-to-Image Generation.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Interleaving Reasoning for Better Text-to-Image Generation

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:19.288085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:aefa497ad07bb35f6f3d41c07d3dc783e7fc376bedaa906cbb2b950fd48056a2

Observation a73b6f31-13d3-4787-a20f-47c5863a97e3 · outbound

This paper cites Ming-univision: Joint image understanding and generation with a unified continuous tokenizer.arXiv preprint arXiv:2510.06590, 2025a.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Ming-univision: Joint image understanding and generation with a unified continuous tokenizer.arXiv preprint arXiv:2510.06590, 2025a

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:46:12.750885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:69c43f9be127db30bb89e384fb71f28c8c796e73c0f537ef9c993954ea7a4279

Observation 660225a3-82dc-469d-bbf7-5c2231f9b071 · outbound

This paper cites T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:19.265561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:2faa22a15b0b9ad9e22a5562d6da58aafdcd479398794b4d30f21910f05487e4

Observation 81b3d0eb-df2e-4e77-ac49-c87c051354fd · outbound

This paper cites an unresolved cited work.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-05-26T20:53:02.724507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:00a0c733bf0b47d15d9bb349c58be333e3d3f17f51ef4f2a0f9708de959799c3

Observation d48f321e-3acf-45a1-943f-faf5e2b1b736 · outbound

This paper cites VIEScore: Towards Explainable Metrics for Conditional Image Synthesis Evaluation.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing VIEScore: Towards Explainable Metrics for Conditional Image Synthesis Evaluation

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:41:19.200450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:3835490018d69847142b998ef1554f4003835ed7a17ef33a411a9d4332cea7f0

Observation d306eb2f-bc74-4aaf-a228-a81bff10a227 · outbound

This paper cites FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:41:19.248810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:7dbf90f322924dfd2d6f5baa7cc03ea642c1fced7073067065eb3814ed0e19b5

Observation 872c16db-089b-4603-a488-9ee45b070041 · outbound

This paper cites Prometheus-vision: Vision-language model as a judge for fine-grained evaluation.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Prometheus-vision: Vision-language model as a judge for fine-grained evaluation

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:53:02.727850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:751685f22734d6df281e02c5e9306c14637e81f07d1476d7bc0f6583655e6561

Observation f929c851-287c-483d-a507-922dd376e1a7 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing LLaVA-OneVision: Easy Visual Task Transfer

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:41:19.168309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:790a6ef1e05cab8fb14485953f91de3193aec4779d2d7fe23a6ecf20e92f00be

Observation 0724f602-58d2-42b9-9adf-cf1c64ca76c1 · outbound

This paper cites Imagine while Reasoning in Space: Multimodal Visualization-of-Thought.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Imagine while Reasoning in Space: Multimodal Visualization-of-Thought

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-16T23:09:34.995208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:2af15f26b45afb0ca150e21ec2163eb9996e002ce01c8f55dbc76ccf03ea418b

Observation 2bdb0531-36c5-4ea3-8f4b-b7963a30f91e · outbound

This paper cites ToonComposer: Streamlining Cartoon Production with Generative Post-Keyframing.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing ToonComposer: Streamlining Cartoon Production with Generative Post-Keyframing

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:46:12.593382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:74605d8bca4ad2087f204cd1938c13867b571989210017e3ec51ccaf5c77113d

Observation e1df8cdf-e095-434f-ae49-2dac8b1e6f08 · outbound

This paper cites Nvcomposer: Boosting generative novel view synthesis with multiple sparse and unposed images.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Nvcomposer: Boosting generative novel view synthesis with multiple sparse and unposed images

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:53:02.734711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:3bf009ce79b6bf8104aca9e72f1a05f4411b26310fd683094524f432bb7f8176

Observation 756f4290-0c5c-4196-bd51-fc05f3e49e5e · outbound

This paper cites Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-18T02:48:45.158778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:30edcdc21e453d40097b95406463affac9fc152ff2a560205742fe896c02cdbe

Observation e7352c77-fad4-4aaf-ad4c-57d751daddd9 · outbound

This paper cites Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-17T07:24:05.047503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:61c6fe569ffc2e8741432ab56aa26a4d2d36ee20d0c63d8cfb44360e94fba6ef

Observation cc8a5292-4147-40e5-866f-cbb95b73c7be · outbound

This paper cites UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-12T17:34:27.304691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:302705b2b65b7c9694420f6f102c575443250ba3973e91869f7a9ad4d85f426b

Observation 7c33a7aa-d548-43d8-bbef-b586cbd9fb8f · outbound

This paper cites JarvisArt: Liberating Human Artistic Creativity via an Intelligent Photo Retouching Agent.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing JarvisArt: Liberating Human Artistic Creativity via an Intelligent Photo Retouching Agent

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:19.240606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:e1419a783737c5819a7ca87e9b38a6e16efeaaaf22016877c3eaa936ea2be0e9

Observation 55bda6c0-36bc-4bf9-9f8e-95778907e6dd · outbound

This paper cites JarvisEvo: Towards a self-evolving photo editing agent with synergistic editor-evaluator optimization.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing JarvisEvo: Towards a self-evolving photo editing agent with synergistic editor-evaluator optimization

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:19.178955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:4b6cbc91f769e44a80f074a46d8b3074501222dae15b0db0cddaf5734ec45890

Observation 10895829-856d-46f2-9f14-1a283cc64d6f · outbound

This paper cites Flow-GRPO: Training Flow Matching Models via Online RL.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Flow-GRPO: Training Flow Matching Models via Online RL

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:46:15.121348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:45ead9ecc4f443e46d8d034bdcbee46f4510e32af5bba59af0e77bf711a2b4ac

Observation 21440f07-a99d-4170-b9e5-072bbb325427 · outbound

This paper cites Step1X-Edit: A Practical Framework for General Image Editing.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Step1X-Edit: A Practical Framework for General Image Editing

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:41:19.319109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:8924c5fb402480f48d0f6fd88b2f03ef5162d3804abca152019582c00452b130

Observation c0f9f15c-61ab-4142-8d71-357f084edb35 · outbound

This paper cites Unified-io 2: Scaling autoregressive multimodal models with vision language audio and action.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Unified-io 2: Scaling autoregressive multimodal models with vision language audio and action

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:53:02.741290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:0bfc87e1a44a169f4e80c13d43cf23e3c44bc51bc94857a81a9298e839a4d36f

Observation f6bb620b-1011-4536-b087-b4b9e68fec27 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.NeurIPS, 35: 2507–2521.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Learn to explain: Multimodal reasoning via thought chains for science question answering.NeurIPS, 35: 2507–2521

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:53:02.756900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:7497d64c522175e1b66260cc5a846ec24acc8bccd944b86ab369a3831a5a6f7e

Observation 5cdc5a57-7228-4b44-8af1-82bd98dad454 · outbound

This paper cites Follow your pose: Pose- guided text-to-video generation using pose-free videos.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Follow your pose: Pose- guided text-to-video generation using pose-free videos

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:53:02.759771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:c89184fdb693a1ed44e376f28df19d37a51cd360b65363c653ee38375b0aa76c

Observation 60008d92-5c65-45e4-8f25-fa10b455564d · outbound

This paper cites JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:46:12.347080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:24aeb177bf10941ddf71172c5a46ecbe18d0886636908ac1ad3478568fe8cb1f

Observation 67ca32ca-fb9e-4c6e-9a85-b183b026f6d1 · outbound

This paper cites Fastvmt: Eliminat- ing redundancy in video motion transfer.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Fastvmt: Eliminat- ing redundancy in video motion transfer

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:46:16.183811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:ab3b2cf5ac99e1420d4a867b968b4841ee0190d17c8420cb6599757fc95f343d

Observation b716aeb6-631a-453f-a487-15391d81d25b · outbound

This paper cites Gpt-image-1.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Gpt-image-1

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:53:02.744311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:be46f1f2a573d9cbff3f9480a1e634917000b2851ae874e1f2fbd20ef23c6f81

Observation 0594f594-434e-41ea-afa6-31cef20c1094 · outbound

This paper cites Introducing 4o image generation.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Introducing 4o image generation

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:53:02.702266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:ddddd3f8dced6b556566bcc063d54b13659da070c9455386ead12db9b7539b86

Observation 1c5628e3-af67-4ad0-ad82-63288334c44e · outbound

This paper cites Transfer between Modalities with MetaQueries.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Transfer between Modalities with MetaQueries

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:49:23.311693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:63342ad09107b4e10b1ead8fd0036bd633627d483eb98f07353111949ed0842c

Observation be94396a-a4d1-46cd-a52a-af3776768c43 · outbound

This paper cites Scalable diffusion models with transformers.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Scalable diffusion models with transformers

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:53:02.685938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:3a26f320110a7371746a4f4c4fca2ed25ebd8ab94ad417ad546b131766f0fcb0

Observation fec4d6af-7265-4a0e-a833-957939712b01 · outbound

This paper cites TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:19.367534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:eb781a4413387db3639fcc242d72369def6b98c7b38b3c2450e535923202930a

Observation 9b98ae0b-e1a8-43d9-9f6d-162ce98e9c99 · outbound

This paper cites Laion-5b: An open large-scale dataset for train- ing next generation image-text models.NeurIPS, 35:25278– 25294.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Laion-5b: An open large-scale dataset for train- ing next generation image-text models.NeurIPS, 35:25278– 25294

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:53:02.698904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:38620c1c05bf9b5560a31946dd76895e27a592b1ced48874586ac777a8c803fa

Observation e52a145d-34a0-40e6-8743-b0a0781312ac · outbound

This paper cites Visual cot: Advancing multi-modal language models with a comprehen- sive dataset and benchmark for chain-of-thought reasoning.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Visual cot: Advancing multi-modal language models with a comprehen- sive dataset and benchmark for chain-of-thought reasoning

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:53:02.692609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:e379ed1d3f8b4d33c9d3f3957d9f427eb2e5de23cf29515fb65c6d5186c69798

Observation 3d594a3f-26bd-4e08-956d-66a4a06485f8 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:41:19.232172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:54a9d05f84ef85139972bd9d9f01a456dd340e2354f832eb23c8fce061651da3

Observation 66ee8cee-fdc4-41e5-ab3a-66281e56cb1d · outbound

This paper cites LMFusion: Adapting Pretrained Language Models for Multimodal Generation.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 52

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:46:14.691378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:836f302c70b7a04581c95246bf93151b5ab65d1a8d34681fcc60a0b80335666d

Observation 161c3c34-404e-4c82-9490-1f4baa4ec993 · outbound

This paper cites Emu: Generative pretraining in multimodality.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Emu: Generative pretraining in multimodality

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:53:02.695626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:db3a96601b7dd07939f36872a6713841303c4f80b047fe044a110c464248584e

Observation cc5f570f-630f-453d-84a3-f96aa4ebf832 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:41:19.340833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:bc45dc823a26b5632805324750746d80f362839263cbfaf51e7824296d964a3c

Observation 465c08dc-a505-4d6f-9d24-1dc2ffdb1557 · outbound

This paper cites On the estimation of relationships involving qualitative variables.American Journal of Sociology, 76(1): 103–154.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing On the estimation of relationships involving qualitative variables.American Journal of Sociology, 76(1): 103–154

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:53:02.661332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:41930b818cb27ba83af85c61003c225a6e286e2d364eca9a19d54c882920f3a9

Observation eceae1de-462f-4abb-8b01-5d30f5df3546 · outbound

This paper cites MetaMorph: Multimodal Understanding and Generation via Instruction Tuning.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing MetaMorph: Multimodal Understanding and Generation via Instruction Tuning

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-17T07:51:13.882527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:1838ece8fe3646c8cc4f655da5c1be262ba915057095352bea20006040bb2205

Observation 9649be60-f329-4e82-94cc-5e60515e225e · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:46:12.896794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:496a14c3ae376e9af0fe9b84b042dd2ee42dd1a028b9fc61ae6808d0cf72375a

Observation e136563d-1708-4680-9821-b224abf183f7 · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Emu3: Next-Token Prediction is All You Need

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:41:19.293621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:93b18261243436649fbd5a267e39baaae78a584b2265c6f941840f6efeef0603

Observation 5dc32682-7d7e-4624-a17d-5b97596ce755 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large lan- guage models.NeurIPS, 35:24824–24837.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Chain-of-thought prompting elicits reasoning in large lan- guage models.NeurIPS, 35:24824–24837

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:53:02.675239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:a704d1bed33688a4e72543ea15d170d1afb45771cd42fe46549e0673c432ce61

Observation 510a65ff-dae2-48c0-a321-0d6264a02ce1 · outbound

This paper cites Janus: Decoupling visual encod- ing for unified multimodal understanding and generation.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Janus: Decoupling visual encod- ing for unified multimodal understanding and generation

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:53:02.709369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:f8e9eff1de08843103d549a1107ae014ad317439269a9f46a775c3f316e6d886

Observation 3b0951c5-58ba-4e25-a937-a4fe2693acdf · outbound

This paper cites Qwen-Image Technical Report.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Qwen-Image Technical Report

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:46:15.895704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:ff255fe8b64d5f9ccdb7eef4cf622b0e329601f939f13f4e92e9432bff87e6e2

Observation 36dc96e8-3034-440d-9ee5-ea5454ba3ee7 · outbound

This paper cites OmniGen2: Towards Instruction-Aligned Multimodal Generation.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:46:12.247983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:637a634634128987094718dd063d668fee1bca2b339bb0545857c575cd50d55b

Observation 8802737e-8f9e-45ec-9e3b-31b192b74c79 · outbound

This paper cites V?: Guided visual search as a core mechanism in multimodal llms.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing V?: Guided visual search as a core mechanism in multimodal llms

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:53:02.672018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:7c60bd9f57df1eb17cf7fd0a9874d0b126ed85ecaa7811d63a3211b8dfc3c0b4

Observation 1f2d14b7-c754-4f67-b28e-a1cc6d0abef9 · outbound

This paper cites Next-gpt: Any-to-any multimodal llm.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Next-gpt: Any-to-any multimodal llm

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:53:02.682123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:483fe8134b0838128484e9baf3f8ece5d95efa6f9b13d0fc82456943014f4113

Observation 31a92358-c2fc-4b2f-9642-a82f0129ff5f · outbound

This paper cites Omnigen: Unified image genera- tion.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Omnigen: Unified image genera- tion

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:53:02.678741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:6df81e25ca06aeee80787e8445fab2186460bc6ba85dc34bb3aaeeda33866d07

Observation fd93135a-5da6-4330-99b7-9a43a7b57548 · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:41:19.375413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:abeae3d71a4f492e8b12aa347d9747a60a7c4b255c87c46ad20fba7e57567894

Observation 0ee58263-b97d-40d9-8dfc-66c1e6a3e39a · outbound

This paper cites Show-o2: Improved Native Unified Multimodal Models.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Show-o2: Improved Native Unified Multimodal Models

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:51:16.424084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:df1e4299411ea70dcf5146decd777c3a71f4eb897332dad9a1645b9ac39d24aa

Observation 2a62e55b-571b-4b3a-9b9c-189ddd77d5b1 · outbound

This paper cites HeadRouter: A Training-free Image Editing Framework for MM-DiTs by Adaptively Routing Attention Heads.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing HeadRouter: A Training-free Image Editing Framework for MM-DiTs by Adaptively Routing Attention Heads

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:19.256048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:d50d0608e80014c3e09f1c6cb78d3dc8dcf8e4bfc43090a41708af3769801a41

Observation cb10bf85-c4dd-46fb-9be9-edc8b41d8094 · outbound

This paper cites Tag-moe: Task-aware gating for unified generative mixture-of-experts.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Tag-moe: Task-aware gating for unified generative mixture-of-experts

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:46:14.436382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:fa9f8504f882fd67d0d32c9ad4dd5b8409bf4a3e622485282d7d3a7fb7d7c9ab

Observation 46081f00-a955-4a7e-94dd-dff49f1de582 · outbound

This paper cites Qwen2.5 Technical Report.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Qwen2.5 Technical Report

Reference 70

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:46:13.357969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:dc7f766bcc9d57a4d7487a33480dac2300cb18d6b2df2509a1d3ef59f5c28bae

Observation 71e30a2b-0cb7-4087-b6ea-f1a9a9f6e029 · outbound

This paper cites Uni-paint: A unified framework for multimodal image inpainting with pretrained diffusion model.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Uni-paint: A unified framework for multimodal image inpainting with pretrained diffusion model

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:53:02.705890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:3bf46255a20f237035d202e4e6c533240d99971ad62df5b15dfc84afc686f3ee

Observation cc25ebff-3400-45ed-b60b-55baa2014840 · outbound

This paper cites Direct-a-video: Customized video generation with user- directed camera movement and object motion.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Direct-a-video: Customized video generation with user- directed camera movement and object motion

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:53:02.664587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:8e39b1ed2f8442ebc047073eb11c6515fe9b10cd40b651c4fbd09379ea431bf1

Observation efde17d6-ad2d-44dd-aaea-9b9f2aa228af · outbound

This paper cites $\texttt{Complex-Edit}$: CoT-Like Instruction Generation for Complexity-Controllable Image Editing Benchmark.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing $\texttt{Complex-Edit}$: CoT-Like Instruction Generation for Complexity-Controllable Image Editing Benchmark

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:19.354347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:94eeb0d4c33d65aa9bc17e3914361b15772b2f7203c27e1fbf7aff215996a518

Observation 1a19edff-0b81-4eb3-a4aa-fea9d1253aa0 · outbound

This paper cites Multimodal rewardbench: Holistic evalu- ation of reward models for vision language models.URL https://api.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Multimodal rewardbench: Holistic evalu- ation of reward models for vision language models.URL https://api

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:53:02.750831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:0dd720c0acaa2f236819d51a153b921852b08e71b46dc08a81d5b019ecf0c0dc

Observation c0dd0fc4-9bfd-48f1-8795-d705e41db21c · outbound

This paper cites ImgEdit: A Unified Image Editing Dataset and Benchmark.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing ImgEdit: A Unified Image Editing Dataset and Benchmark

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:17:45.693018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:cf7cef82dfd4af459b6fa79f1ae20cb9c1b3104b494022bc8960c1bfd5a3a0cb

Observation 9db8e75f-a664-4a45-8a2f-1bb1886b8f9c · outbound

This paper cites Anyedit: Mastering unified high-quality image editing for any idea.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Anyedit: Mastering unified high-quality image editing for any idea

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:53:02.715699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:9a0bcd84dca536933812f6059cb543d61396b42d9749ad2c9cfd88568be6f7f0

Observation 3fad433e-784f-4789-9415-48455184384a · outbound

This paper cites Magicbrush: A manually annotated dataset for instruction- guided image editing.NeurIPS, 36:31428–31449.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Magicbrush: A manually annotated dataset for instruction- guided image editing.NeurIPS, 36:31428–31449

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:53:02.718964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:a1b70b30c093bfafc9a102a048d518aa18eee469f5b899eb70587e50c175bb5c

Observation 9b5f7015-e7c9-4fe7-8366-46901c619d5a · outbound

This paper cites Logo: A long-form video dataset for group action quality assessment.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Logo: A long-form video dataset for group action quality assessment

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:53:02.712401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:4762b2e667c909c12699c73cf2f8b05235917b8bbd4b28abdfdde1be37e4e5fd

Observation d05b10c7-fc7a-4060-9475-ba70dfdfa0ad · outbound

This paper cites Narrative action evaluation with prompt-guided multimodal interaction.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Narrative action evaluation with prompt-guided multimodal interaction

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:53:02.689075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:636afd684e528ae1cc797028e54e10864d1d9ee459b62450897b0a1e530b2bbd

Observation 83242267-11e7-4be0-9e72-e108910d5a22 · outbound

This paper cites Flexiact: Towards flexible action control in heterogeneous scenarios.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Flexiact: Towards flexible action control in heterogeneous scenarios

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:53:02.668687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:1a16eb624747036f85b660179a4abff781befd27c418b2d6352f05e089b68a7e

Observation b6533477-c6b4-43e5-8708-eb4dc00684ca · outbound

This paper cites Multimodal Chain-of-Thought Reasoning in Language Models.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Multimodal Chain-of-Thought Reasoning in Language Models

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:12:27.664923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:faa3a21eb8b8136f723c41a159b5eeca638ca87bb0b4c95e01f106351e270b6b

Observation 413de24d-53a8-4232-a463-30151c361b93 · outbound

This paper cites In-Context Edit: Enabling Instructional Image Editing with In-Context Generation in Large Scale Diffusion Transformer.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing In-Context Edit: Enabling Instructional Image Editing with In-Context Generation in Large Scale Diffusion Transformer

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-05-16T16:07:53.247084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:0f098c50cb78b909388bd16502680b54b4e1b6d39be86bbfcc9f2665771e7df7

Observation e9a17fd4-de51-4ae6-bd7f-c38be8a334ab · outbound

This paper cites Ultraedit: Instruction-based fine-grained image editing at scale.NeurIPS, 37:3058–3093.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Ultraedit: Instruction-based fine-grained image editing at scale.NeurIPS, 37:3058–3093

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:53:02.654541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:ed63009e2b8f9f1d928fc10e86982db397f903268bdfbb84189b79cb3a035c31

Observation 037c69fb-a8b6-40f6-bb6f-ccdf74121747 · outbound

This paper cites Envisioning Beyond the Pixels: Benchmarking Reasoning-Informed Visual Editing.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Envisioning Beyond the Pixels: Benchmarking Reasoning-Informed Visual Editing

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:46:13.957691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:76f25c2d968dc99af8ea3a9f0a9a37d1c306b2fea2ce43756e8b981953f780f0

Observation ee8cee85-5f76-40f0-abf5-7c297284993e · outbound

This paper cites Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

Reference 85

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:57:27.005672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:2c7674119d579befc3047ad7fb9027907bbef2914789b38d248bdb49a7c01858

Observation 8efda024-697b-4795-8dd5-fec457353b44 · outbound

This paper cites Kv-edit: Training-free image editing for precise background preservation.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Kv-edit: Training-free image editing for precise background preservation

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:53:02.657787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:bd756b9027cc9a91069506b650676c075cc8830d110e38f9aad3d3107fd28d7f

Observation f097a609-7b8b-458a-922a-e6b8fb7312fa · outbound

This paper cites ColorFlow: Retrieval-Augmented Image Sequence Colorization.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing ColorFlow: Retrieval-Augmented Image Sequence Colorization

Reference 87

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:19.148799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:c8e32161a0a4eb63163b5d54702ad1cac4f8df4875b08c2c8fb69f0faaef221a

Pith citing papers

Observation ed7c5895-4410-421c-a7c9-af892091f48f · inbound

Reasoning to Align: Implicit Reasoning in Diffusion Transformers for Video Editing cites this paper.

Reasoning to Align: Implicit Reasoning in Diffusion Transformers for Video Editing Meta-CoT: Enhancing Granularity and Generalization in Image Editing

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-06-30T13:44:41.150231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-30T13:37:09.892339Z digest=sha256:f9f7a29234ec8c3f23abfdf29a68008cb7b2d7f201883c0a8da754f1ffa02a80

Observation 81170c89-5dcd-4806-b952-1138c72f9bc5 · inbound

GMO-E$^2$DIT: Grounded Multi-Operation Editing for E-Commerce Images cites this paper.

GMO-E$^2$DIT: Grounded Multi-Operation Editing for E-Commerce Images Meta-CoT: Enhancing Granularity and Generalization in Image Editing

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-02T13:56:59.012818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-02T13:56:43.671622Z digest=sha256:d11917f0f2dab85ea9c7c4dff4394e62e61df651b27f0b08f1d732609a722521

Observation 2cac8ef2-7c57-4dd2-a656-8f1d4d6fb016 · inbound

GMO-E$^2$DIT: Grounded Multi-Operation Editing for E-Commerce Images cites this paper.

GMO-E$^2$DIT: Grounded Multi-Operation Editing for E-Commerce Images Meta-CoT: Enhancing Granularity and Generalization in Image Editing

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-03T21:28:58.250641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-03T21:23:41.271521Z digest=sha256:c713fd63fe81a741c4977cc2b2da772b4e0123f51074b3435a5e1a6ccc83f5dd