Pith. sign in

Paper Citation Record · LEDGER

GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 40 inbound Pith citation observations for arXiv:2503.10639.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.10639 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 40 of 40 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:02:42.907988Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T21:28:58.240235Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 43cde051-d7e4-4926-8375-5a61830f08ec · inbound

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey cites this paper.

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 255

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T17:18:53.577665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T17:18:52.996467Z digest=sha256:96cd07fa06469f424bbfd22ac83f19233792d264f8aac765eb74f4b54fa5734c

Observation a311df4b-048b-46b2-b3fa-fad1dae13097 · inbound

Generative AI Act II: Test Time Scaling Drives Cognition Engineering cites this paper.

Generative AI Act II: Test Time Scaling Drives Cognition Engineering GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-16T12:02:42.907988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:02:42.907988Z digest=sha256:d8855485529c67ec7abe35f9d4ecec5422179e8831172e79c3c02ac368d8dbc6

Observation 05e16346-56ee-4424-8657-735ad92e1cc8 · inbound

Step1X-Edit: A Practical Framework for General Image Editing cites this paper.

Step1X-Edit: A Practical Framework for General Image Editing GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T14:36:41.949615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-11T14:36:41.467429Z digest=sha256:6a9f4f0d9e9e70cf8bd1996ed3018be18cd876a14a3e7289f8e4421313cff8e7

Observation 2d10fbf6-fa10-4d28-8599-c33d4693be75 · inbound

WorldGenBench: A World-Knowledge-Integrated Benchmark for Reasoning-Driven Text-to-Image Generation cites this paper.

WorldGenBench: A World-Knowledge-Integrated Benchmark for Reasoning-Driven Text-to-Image Generation GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T04:21:03.804768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:21:03.804768Z digest=sha256:3fca53622156773e4beaadae02450bc5ff1e9d8308268450c0235ef1bc6f30c3

Observation 71cab7d0-ea13-4b25-88de-9852fedef3cd · inbound

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO cites this paper.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.139315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.139315Z digest=sha256:75c8f5e5a2861de4a535f2ef6e311a16df063ffbcfc9fdd1aa4be3c634d72a94

Observation bb2098e1-9e53-41ce-a64b-85e74d3182cd · inbound

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation cites this paper.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:52.941000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:52.941000Z digest=sha256:6712974d3170a74433bcef3ea6b544a5f8d0d9f32bcf461686c3bd12f7679677

Observation 2a0928ba-56dc-4f6e-a987-3068b786be3e · inbound

KRIS-Bench: Benchmarking Next-Level Intelligent Image Editing Models cites this paper.

KRIS-Bench: Benchmarking Next-Level Intelligent Image Editing Models GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:16.805570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:16.805570Z digest=sha256:7af40c21c1a2e11b4cf37053ecc56aaea186a25ea996177b99982ff7c4da6382

Observation be32fd37-cc73-4585-8623-2b3b7073afb3 · inbound

R-Genie: Reasoning-Guided Generative Image Editing cites this paper.

R-Genie: Reasoning-Guided Generative Image Editing GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:46:34.827824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:46:34.827824Z digest=sha256:013dee76c04ea2c8d1a2cc96da8fe002af04449bf397b7e2f8b843e2636bcc5d

Observation c9887bfd-010f-452f-aed6-0082d6982958 · inbound

ComfyMind: Toward General-Purpose Generation via Tree-Based Planning and Reactive Feedback cites this paper.

ComfyMind: Toward General-Purpose Generation via Tree-Based Planning and Reactive Feedback GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:12.548057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:12.548057Z digest=sha256:61fa0143e95a3b291e741d065efec23a79ac7418c7887eebc131a5b34df5a9c8

Observation 1614e1a3-3500-4594-a6e6-ef53ad82de7f · inbound

Beyond Editing Pairs: Fine-Grained Instructional Image Editing via Multi-Scale Learnable Regions cites this paper.

Beyond Editing Pairs: Fine-Grained Instructional Image Editing via Multi-Scale Learnable Regions GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:41.061314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:41.061314Z digest=sha256:63614219def48eadb72fe72fdd927564790cfad3553b9ae91e7c8a09d446cfde

Observation db554c65-c3ad-45c8-9371-8fe68adb408c · inbound

ImgEdit: A Unified Image Editing Dataset and Benchmark cites this paper.

ImgEdit: A Unified Image Editing Dataset and Benchmark GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T18:17:45.508850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-12T18:17:45.123690Z digest=sha256:d639974e8e585a6e9bef01049e4f80a4da06e2c38751b6e91f1b852a95703c49

Observation f9e05ca0-be08-4147-969c-b18a3d36e5a5 · inbound

Draw ALL Your Imagine: A Holistic Benchmark and Agent Framework for Complex Instruction-based Image Generation cites this paper.

Draw ALL Your Imagine: A Holistic Benchmark and Agent Framework for Complex Instruction-based Image Generation GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:18:32.501235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:18:32.501235Z digest=sha256:784fb0ab75aa4b1cfdfa69f1e23598d9c4f041d3df8bed7fc400b8752c365cc3

Observation e85fb9d6-0103-4439-b9cc-e33646806b97 · inbound

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics cites this paper.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:24.165875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:24.165875Z digest=sha256:36460306094b6009566f950d7897e9dec117f8ff47585401a3bf564b8465ac56

Observation 74ba5345-9cb4-4895-ac99-6692f515617b · inbound

ComplexBench-Edit: Benchmarking Complex Instruction-Driven Image Editing via Compositional Dependencies cites this paper.

ComplexBench-Edit: Benchmarking Complex Instruction-Driven Image Editing via Compositional Dependencies GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T00:40:19.605957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:40:19.605957Z digest=sha256:4ca8b6e093da1e163334a7a8a8b0028a511956d5ed2b32204606b6751ac85ed9

Observation 0aea6666-678d-4dcf-bd4f-d7c7a8949f0c · inbound

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations cites this paper.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:10.647842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:10.647842Z digest=sha256:2d99cf32c9448fbc764e9dea56892adf6954f87565f685251b7fbd3528f41f4c

Observation 0feb2e14-1f15-4d9d-b0ec-05c9ae66f644 · inbound

Via Score to Performance: Efficient Human-Controllable Long Song Generation with Bar-Level Symbolic Notation cites this paper.

Via Score to Performance: Efficient Human-Controllable Long Song Generation with Bar-Level Symbolic Notation GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T05:40:56.881671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:40:56.881671Z digest=sha256:968e4e8c4ccc7848e15dd638a43e82f8383818fead3487cacc5cc8c07e3481de

Observation e584b0c6-8dab-44f3-8a1c-0762cb7b68e5 · inbound

Visual-CoG: Stage-Aware Reinforcement Learning with Chain of Guidance for Text-to-Image Generation cites this paper.

Visual-CoG: Stage-Aware Reinforcement Learning with Chain of Guidance for Text-to-Image Generation GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-15T17:03:42.934010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:03:42.934010Z digest=sha256:0079f0639839e10e32e7219489acb57f01f228804481d18d977bf1538ced4be9

Observation 5be1d564-b5ec-4ea8-b747-5d355dacab17 · inbound

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey cites this paper.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 256

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:21:48.132193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:aeaeb8d2f1498cffb6f0425b0ffa84c984a3bf70904c8c8c52b8abfc1e5b0e69

Observation 86739b7b-e6d6-4ebd-a99e-dd28b5e1bf17 · inbound

MEPG:Multi-Expert Planning and Generation for Compositionally-Rich Image Generation cites this paper.

MEPG:Multi-Expert Planning and Generation for Compositionally-Rich Image Generation GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T10:24:55.232863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:24:55.232863Z digest=sha256:80833d5a75e4f3a5a404e5ccaef43947f5ee2bfba314370730c2bbaaa9a19cad

Observation 9f56752a-f678-4670-b871-473500b61ff1 · inbound

Interleaving Reasoning for Better Text-to-Image Generation cites this paper.

Interleaving Reasoning for Better Text-to-Image Generation GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:44.877594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:44.877594Z digest=sha256:3d6f0f048dab722f39f9d0feb864f3f7e89f597724860ee18ba36ad409c242dc

Observation 1ad52245-d29a-460c-bd5d-77ca3b5b34fb · inbound

Reconstruction Alignment Improves Unified Multimodal Models cites this paper.

Reconstruction Alignment Improves Unified Multimodal Models GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T22:36:07.936009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T22:36:07.936009Z digest=sha256:843eb0b0aff32a11a63ac5cccec0acc57b629cf8ae14cd1b6fcd0448b0506441

Observation 15feb0bb-9423-4260-a030-ccd50ba0b956 · inbound

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark cites this paper.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.645811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.645811Z digest=sha256:ebd674f0dc3b6b9eab83dd75f87aa2f7d6eb60b3a526f91c549cf0d0f2dde6f3

Observation c317703b-34b2-4365-81ba-33f4da95b3ef · inbound

SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models cites this paper.

SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T09:54:24.239016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:54:24.239016Z digest=sha256:86e1a46dbadee5ab290aa5ebdf876cdadb0bff4608c5de1655b4f46ddaefa7bd

Observation b5450353-038d-4dac-a264-4d7fa7247780 · inbound

IRG-MotionLLM: Interleaving Motion Generation, Assessment and Refinement for Text-to-Motion Generation cites this paper.

IRG-MotionLLM: Interleaving Motion Generation, Assessment and Refinement for Text-to-Motion Generation GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T17:06:54.722954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:06:54.722954Z digest=sha256:d248e7054fc2197c536c985a28266b21fbbcc5235175f5abf5b56d702c1e122d

Observation f8b3094d-71f9-4f13-968f-a51d8674b083 · inbound

Do-Undo Bench: Reversibility for Action Understanding in Image Generation cites this paper.

Do-Undo Bench: Reversibility for Action Understanding in Image Generation GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:11:18.473960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-16T22:10:29.304091Z digest=sha256:48b664805c6e35501785059854e6354433095e2454b7f41f7f23f527004d4842

Observation b4f0b072-f1dc-495e-9848-97a5fb37d484 · inbound

Demystifying Video Reasoning cites this paper.

Demystifying Video Reasoning GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-13T23:27:11.006580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:27:11.006580Z digest=sha256:8cbcbb3ade68b3c2cfdfab9918426410c04a944c6b0794d83f2ceead5de4e3ef

Observation 3766ccda-d2e8-4b89-a5ea-a612eab982a5 · inbound

Demystifying Video Reasoning cites this paper.

Demystifying Video Reasoning GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T02:33:53.930732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:33:53.930732Z digest=sha256:8fd3fafccc58b086ac9c5e6a49f5184b69c7dc30d58ebf7fc7571110d32fa0f9

Observation 14adf7f1-dbb3-474f-be97-4dc5bf312958 · inbound

AIM-Bench: Benchmarking and Improving Affective Image Manipulation via Fine-Grained Hierarchical Control cites this paper.

AIM-Bench: Benchmarking and Improving Affective Image Manipulation via Fine-Grained Hierarchical Control GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T10:11:04.280660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T15:35:49.575589Z digest=sha256:9d58e8d6b301d87a59c8a38a6faf2f529a11449b6f695b4af85c11a46a501e1b

Observation 62b9b7dc-027b-415e-9552-c6643be237f4 · inbound

Edit Where You Mean: Region-Aware Adapter Injection for Mask-Free Local Image Editing cites this paper.

Edit Where You Mean: Region-Aware Adapter Injection for Mask-Free Local Image Editing GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:06:13.722077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-08T06:47:22.451132Z digest=sha256:d7f988def7023755c3959eb0e345d062184a70954394f54463dc468be7dbedc3

Observation c251ec09-3171-4a66-8787-bfef63b28c40 · inbound

Meta-CoT: Enhancing Granularity and Generalization in Image Editing cites this paper.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:19.274947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:3ee39af20c4731ebff8ddea4987e19f3702a5c50ea91cae1cb49ebfdba2cc2fe

Observation 2cbfaf2e-1b78-4469-9421-422f690c10d1 · inbound

RCoT-Seg: Reinforced Chain-of-Thought for Video Reasoning and Segmentation cites this paper.

RCoT-Seg: Reinforced Chain-of-Thought for Video Reasoning and Segmentation GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:31:00.029682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-11T01:16:25.031349Z digest=sha256:863c852c4a6ba7fcf74800b16efec99523c759f9431db91d996399b3110ee1df

Observation af9e71d7-eece-42b0-b3e3-f446fe819732 · inbound

Masked Generative Transformer Is What You Need for Image Editing cites this paper.

Masked Generative Transformer Is What You Need for Image Editing GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:06:26.027511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-12T04:35:52.355925Z digest=sha256:3a2023e816c6281355a87af76398fa7d2b383c9460b6e0c06ff233dce8cce1e6

Observation 2eb83584-ced1-4f53-b90d-a9e56d3d8c6c · inbound

UniPath: Adaptive Coordination of Understanding and Generation for Unified Multimodal Reasoning cites this paper.

UniPath: Adaptive Coordination of Understanding and Generation for Unified Multimodal Reasoning GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:47:04.960167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-13T01:29:23.774196Z digest=sha256:c586cadbefcedc4dc3dd2a39fca740a8716d8fa4362b66a4ed5cf584b94220c2

Observation 65c4a275-426e-46cd-b7ad-8b45e09534c1 · inbound

Editor's Choice: Evaluating Abstract Intent in Image Editing through Atomic Entity Analysis cites this paper.

Editor's Choice: Evaluating Abstract Intent in Image Editing through Atomic Entity Analysis GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:25:46.043811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T21:41:10.852265Z digest=sha256:03ad3c61e8389f8ca28a139ce3b1a7508278de22b3200c1648a08fb4c4858837

Observation a57079f0-abd8-431f-ac79-e66a5a1ac02e · inbound

Are Tools Always Beneficial? Learning to Invoke Tools Adaptively for Dual-Mode Multimodal LLM Reasoning cites this paper.

Are Tools Always Beneficial? Learning to Invoke Tools Adaptively for Dual-Mode Multimodal LLM Reasoning GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:18:05.182007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-20T06:16:47.650748Z digest=sha256:2450ae0616c2731753ad93eb5ede6a3aff74ae218241e67fcc571aaa9b5c7413

Observation 14463bd6-eece-4c48-8105-59f7a7b1fb00 · inbound

Evaluating Reasoning Fidelity in Visual Text Generation cites this paper.

Evaluating Reasoning Fidelity in Visual Text Generation GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T07:16:44.516479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-28T07:03:38.967856Z digest=sha256:c024d2b056906150de7dcb54ddd4e7979a591d991a430135bdf0a329848f2adf

Observation 814bee08-d142-41f8-bbd4-71e59b4c17f1 · inbound

MetaPoint: Unlocking Precise Spatial Control in Agentic Visual Generation cites this paper.

MetaPoint: Unlocking Precise Spatial Control in Agentic Visual Generation GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T08:26:47.752091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-28T06:07:06.056441Z digest=sha256:f12ee6761bb8989ad5c12440e982ae9367e9d84ffeffbda3620d3b537aeeab62

Observation 364406dd-a65b-4b65-8622-0a880d401306 · inbound

GMO-E$^2$DIT: Grounded Multi-Operation Editing for E-Commerce Images cites this paper.

GMO-E$^2$DIT: Grounded Multi-Operation Editing for E-Commerce Images GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:56:58.969221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-02T13:56:43.671622Z digest=sha256:bb4b18057fb9f4b999d3bc4bd7a2b5e281d2bc12bddef09f56d04c17af53bba4

Observation e6a8179a-8d93-479a-90dc-1b99efc7a669 · inbound

GMO-E$^2$DIT: Grounded Multi-Operation Editing for E-Commerce Images cites this paper.

GMO-E$^2$DIT: Grounded Multi-Operation Editing for E-Commerce Images GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-03T21:28:58.241767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-03T21:23:41.271521Z digest=sha256:ffbe728efedf62c7ca401b6b325e7e647c3febc510b49c9eac72e1bddf6d28ed

Observation cec7d7d6-c797-4468-a227-9ac89c904e7a · inbound

Model Guides You How to Draw: Adaptive Visual Gating for Unified Multimodal Reasoning cites this paper.

Model Guides You How to Draw: Adaptive Visual Gating for Unified Multimodal Reasoning GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-14T01:07:06.861298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T01:07:06.861298Z digest=sha256:38ed06481224b728c545616aaa4e3b1207c5d7db9d23a34c3b88e40b5d190bbf