Pith. sign in

Paper Citation Record · LEDGER

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing

As of 10 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 0 inbound Pith citation observations for arXiv:2607.08497.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.08497 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-10T06:30:28.462917Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

45 of 45 outbound references displayed

  • verified exact26
  • verified fuzzy18
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 118692d1-fc3c-44b5-b6d3-f4442849222f · outbound

This paper cites GPT-4 Technical Report.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing GPT-4 Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-10T06:36:52.563355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:ef62b56ec838e26027cdc61519768933eda87e144755e72f32faab998315cff5

Observation 11e1e380-6d78-435e-ab1d-989592fae595 · outbound

This paper cites Qwen Technical Report.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Qwen Technical Report

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-10T06:36:52.572324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:aff7aa0e654ee359470ad8b4bad26b98b894db9b9710a8bf4c574351367004ea

Observation ee0602c2-0354-4ea2-bb65-7a2d4e0d4002 · outbound

This paper cites Qwen-vl: A versatile vision-language model for understand- ing, localization, text reading, and beyond.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Qwen-vl: A versatile vision-language model for understand- ing, localization, text reading, and beyond

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T06:36:52.975236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:7c10f6f0073fda6f67eb17a1bf621134732a8d14ed3bc24fbfcd52aff4144c8e

Observation e567e0b5-1cc4-4e0d-915f-6656802e901d · outbound

This paper cites Qwen3-VL Technical Report.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Qwen3-VL Technical Report

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-10T06:36:52.558779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:89690876b06e3c4e1723f8ff0a5cbae976b508d04272d3285587face824ec7c0

Observation acc2e0eb-e4c1-4c16-a46a-6db0cfd7c988 · outbound

This paper cites In- structpix2pix: Learning to follow image editing instructions.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing In- structpix2pix: Learning to follow image editing instructions

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T06:36:52.973493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:b4691cee2416441220153682c8a8692455a5ad72db2de4de9e5f82bdcc7d6f85

Observation 7a6a227c-33ce-4bf4-9ceb-f9880e313684 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-10T06:36:52.570188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:8d425a1eb4591ee5ba510c15c80095a60291a53531a2a47cc0233244d2d969c4

Observation 238ab7e4-30dd-4ca3-a8c5-82e808772937 · outbound

This paper cites Emerging Properties in Unified Multimodal Pretraining.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Emerging Properties in Unified Multimodal Pretraining

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-07-10T06:36:52.567988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:0d18d87d4c8be88fabdbae42981781ac03ba13b258b4bf4703fc92ffb455f7ce

Observation 2a2f523f-2b61-476e-8eee-b7856980268b · outbound

This paper cites Videoagent: A memory-augmented multi- modal agent for video understanding.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Videoagent: A memory-augmented multi- modal agent for video understanding

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T06:36:52.969766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:24b10b8397dafe1e7a39cdd3696abd1a8e9221b18252b31e5c5f7622787c5bb9

Observation 614e779c-392d-4a80-a56f-5936bfe66faf · outbound

This paper cites Multi-modal Agent Tuning: Building a VLM-Driven Agent for Efficient Tool Usage.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Multi-modal Agent Tuning: Building a VLM-Driven Agent for Efficient Tool Usage

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-10T06:36:52.553796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:74cef498efa4b68e33192db25fb77e2857f3840ea2d5f9cbbe46bad3b978fe8f

Observation 3d65c253-8658-4c0e-a091-d89bfa607dbe · outbound

This paper cites Metagpt: Meta programming for a multi-agent collaborative framework.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Metagpt: Meta programming for a multi-agent collaborative framework

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T06:36:52.968036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:08234847ff2c9b24c66c28490edf918964dacfcd6928fe29810f4a1a55c482bb

Observation 4c0f5849-0085-4fe7-a67c-4d89f32e16da · outbound

This paper cites Talkphoto: A versatile training-free conversa- tional assistant for intelligent image editing.arXiv preprint arXiv:2601.01915, 2026.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Talkphoto: A versatile training-free conversa- tional assistant for intelligent image editing.arXiv preprint arXiv:2601.01915, 2026

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-10T06:36:52.526893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:106bbbfa9778340163db5f111506f1c88be953ac5f3cb6d80a8e7866dcfe70a8

Observation 451c7b3c-4205-45bc-bd1f-1b7927b3efd1 · outbound

This paper cites Wegen: A unified model for interac- tive multimodal generation as we chat.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Wegen: A unified model for interac- tive multimodal generation as we chat

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T06:36:52.971543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:c30eccd7bddc1e5201473322398ad3aa2ea85cb8b0184b4e0c3a96d48d25afa4

Observation 4250f6db-719a-48d3-938d-93dcd311fe5c · outbound

This paper cites SYNAPSE: Synergistic associative processing & semantic encoding.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing SYNAPSE: Synergistic associative processing & semantic encoding

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-07-10T06:36:52.549002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:302326b24976f34094ebd6423ffbba280eb224ae4495048fea2f11446ac0fc41

Observation 868cf897-d97a-45b8-b55e-48b8ee1be68b · outbound

This paper cites Videomem: En- hancing ultra-long video understanding via adaptive memory management.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Videomem: En- hancing ultra-long video understanding via adaptive memory management

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-10T06:36:52.535630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:92cad7e08c325b74d58d42ed6021c809202950083f57735b0f6d978c6348d111

Observation 364476b8-984f-488a-ba55-9972166d8930 · outbound

This paper cites MMCTAgent: Multi-modal Critical Thinking Agent Framework for Complex Visual Reasoning.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing MMCTAgent: Multi-modal Critical Thinking Agent Framework for Complex Visual Reasoning

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-10T06:36:52.551375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:4d618b814b17ea1d43f87540037743f7dcdecd97b5ea91666fbd8b717bdac121

Observation afa2fd87-cf78-499d-a3e0-80358efa4eb8 · outbound

This paper cites Camel: Com- municative agents for” mind” exploration of large language model society.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Camel: Com- municative agents for” mind” exploration of large language model society

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T06:36:52.962455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:6e382500328972b3b8c25eab9410beac6837c4befdab64339d04ce25c829a423

Observation cb6cc774-ac5c-40ea-bb86-b702ebf7bb22 · outbound

This paper cites Blip- 2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Blip- 2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T06:36:52.966212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:53c55e538dfd0384f358a93e507e48b2dd89ad95dccbdb09164f5e08bdbdfbab

Observation bfa69c8c-ae61-41f7-8130-cf3306aba638 · outbound

This paper cites Iterative trajectory exploration for multi- modal agents.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Iterative trajectory exploration for multi- modal agents

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T06:36:52.958907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:aed65bcb7e749bb1a9df22a66ed37e863789046a8d51324e839d69c7a45edb25

Observation 1a600f12-802a-4133-9427-85247b16fb6a · outbound

This paper cites Llava-next: Improved reason- ing, ocr, and world knowledge, 2024.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Llava-next: Improved reason- ing, ocr, and world knowledge, 2024

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T06:36:52.960735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:cb9c61e63b77aadd2a85f7e7e6ff613ddcf7b17e967e8ef88a598737401004eb

Observation 86aecfdd-2dbb-4b59-9ad8-8604735bb7df · outbound

This paper cites Agent0 -vl: Exploring self -evolving agent for tool -integrated vision -language reasoning.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Agent0 -vl: Exploring self -evolving agent for tool -integrated vision -language reasoning

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-10T06:36:52.541423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:ff6d8bd6f99465ff36f8e8f83a80427bc748051ae3a84875a0c593bb0c192d97

Observation 86d93bfe-f407-41d6-a0f6-a399cf026589 · outbound

This paper cites Seeing, listening, remembering, and reasoning: A multi- modal agent with long-term memory.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Seeing, listening, remembering, and reasoning: A multi- modal agent with long-term memory

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-10T06:36:52.538636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:8d673fce3fe0a9b09ea22308b4db97598ee18127509fd2d0f3f8999920994e9d

Observation d28de4e5-65ed-4b91-9306-d7bc82277630 · outbound

This paper cites Janusflow: Harmonizing autoregression and rectified flow for unified multimodal understanding and generation.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Janusflow: Harmonizing autoregression and rectified flow for unified multimodal understanding and generation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T06:36:52.956862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:490aa5883c3dd86fa67ad7fef4f6396e9637eb329d0868a2181f1485c9a96344

Observation 4044ca7d-07cc-42da-bdcc-eed259536201 · outbound

This paper cites ReasoningBank: Scaling Agent Self-Evolving with Reasoning Memory.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing ReasoningBank: Scaling Agent Self-Evolving with Reasoning Memory

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-07-10T06:36:52.546322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:b599a6b7e75d7c3b896360d47ad9cf19baf345a7256af3bf1487636a9fb9a726

Observation 9bd478fc-be42-4874-84ef-3c93137fb3e5 · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-07-10T06:36:52.543951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:799acdf1c1898134e6148185efed0386186683ec776f49df3c920afd2d2db317

Observation f010496c-7246-4375-b471-caec90bc69a6 · outbound

This paper cites Chatdev: Communicative agents for software devel- opment.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Chatdev: Communicative agents for software devel- opment

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T06:36:52.955062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:f032ef4e166b45006b3743123a52e3b04d3f11cfbd865726a92911b1e08b2661

Observation d8ca0203-eac4-4df3-8090-c8ae668ff738 · outbound

This paper cites quiet" and.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing quiet" and

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T06:36:52.953298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:6141581ce8a61bf9ab74219fb4250c49a95802c4c9b968ceb4e99e747a25cb0d

Observation 1632f313-2d01-48c0-a98e-7ad050f361fa · outbound

This paper cites Unilip: Adapting clip for unified multimodal understanding, generation and editing.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Unilip: Adapting clip for unified multimodal understanding, generation and editing

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-10T06:36:52.529778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:1b736453ad63e4cb43367e963efb8e183fb9e8efed1d081a49062ae0a2f6274c

Observation 3ac3107a-04e5-4f82-8ccd-8cb50a706ec4 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Gemini: A Family of Highly Capable Multimodal Models

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-07-10T06:36:52.576992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:0e79b2a6d2bbe61db432625871042b03f7defb9788c02318185a44fc0ce1497c

Observation d770cabd-c3fc-4b06-a29b-5cbfc367fbcf · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing LLaMA: Open and Efficient Foundation Language Models

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-07-10T06:36:52.516126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:e9ce559d44d6a058c9b76f1fafdef2008d4cea5b14ac1b73699ea2c92b53ce64

Observation 659ce5ac-57e5-4001-82c4-708bdb4ff38d · outbound

This paper cites Crea: A collaborative multi-agent framework for creative content generation with diffusion mod- els.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Crea: A collaborative multi-agent framework for creative content generation with diffusion mod- els

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T06:36:52.948389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:ab9a285f16f8222a48b7a74cc4fba65a6f8cacf39d5f837acf3692d76f05c558

Observation 141f7f3d-8d8e-4a1b-9082-f83562ca4b1d · outbound

This paper cites Yanyun-3: Enabling cross-platform strategy game operation with vision-language models.arXiv preprint arXiv:2511.12937, 2025.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Yanyun-3: Enabling cross-platform strategy game operation with vision-language models.arXiv preprint arXiv:2511.12937, 2025

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-07-10T06:36:52.532440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:a872987885e52c531aec98bc99bc1431c3187be26cae53b4eb4897645dff863f

Observation a69fc2c7-550d-4561-9c66-dac08667474c · outbound

This paper cites Multimodal needle in a haystack.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Multimodal needle in a haystack

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T06:36:52.950001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:4045588bf5ecf2afa72cd7e19b4f8fea792f678181f648d6641650c6dedca7b5

Observation c5b8d679-e97e-4e89-8220-6f2fb9b67d5f · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-07-10T06:36:52.556292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:5c265d833892038fb716e4f0a9d546e578953a92c9049bc1b87f82c6f1741240

Observation 1fab9c4f-c2c3-49c5-a1d7-b5ad56bbebd2 · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Emu3: Next-Token Prediction is All You Need

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-07-10T06:36:52.574682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:0c658ee419f51013ee0f92a0572369d0eaeb0fb6a21ca8800b06e4ed15b56269

Observation 39bdbb5b-3444-48b1-a8a7-5554bbce8171 · outbound

This paper cites Videoagent: Long-form video understanding with large language model as agent.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Videoagent: Long-form video understanding with large language model as agent

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T06:36:52.951712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:93e183ef1cce6719dfe724221e77fda597d53ac160449fdebbc26cf86cc89ada

Observation 7ba7a256-fdbf-4092-b012-feb8aa1ee298 · outbound

This paper cites Genartist: Multimodal llm as an agent for unified image gener- ation and editing.Advances in Neural Information Processing Systems, 37:128374–128395, 2024.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Genartist: Multimodal llm as an agent for unified image gener- ation and editing.Advances in Neural Information Processing Systems, 37:128374–128395, 2024

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T06:36:52.964348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:696e79d87be8dd71af97b9c880026e4fdeda643ecdf6741c07b8566a542189f9

Observation 11c699b5-43dc-49f3-b604-801e2d9687d7 · outbound

This paper cites Qwen-image technical report, 2025.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Qwen-image technical report, 2025

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T06:36:52.946383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:4bd4164b381ecff7e51f27adb67584919c15b47177c8f15c11a69b0d8c668bac

Observation 7fdc33c5-da4a-458a-b748-ea7865372227 · outbound

This paper cites Visual Haystacks: A Vision-Centric Needle-In-A-Haystack Benchmark.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Visual Haystacks: A Vision-Centric Needle-In-A-Haystack Benchmark

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-07-10T06:36:52.561097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:44c84e07020dfe0b1e98b5b10922a046d1db907b7ccb13a81a6220790a8af90d

Observation 981b8bcc-a252-40a0-996c-1ef5eb69dac3 · outbound

This paper cites Show-o2: Improved Native Unified Multimodal Models.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Show-o2: Improved Native Unified Multimodal Models

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-07-10T06:36:52.521661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:edac5a34bff109f032d5011c283c2abbfa5740549f079cf01175548cbdffad0b

Observation ebc799a5-f80d-4189-8158-bb72c5c21345 · outbound

This paper cites Qwen3 Technical Report.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Qwen3 Technical Report

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-07-10T06:36:52.510592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:74b3c4bd0ab41598b912500cc1f5f44d19fb7314de28914430ea7590ca50bfa2

Observation 62c9e693-29be-4823-999c-b826b20f61d8 · outbound

This paper cites Agentfold: Long-horizon web agents with proactive context management.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Agentfold: Long-horizon web agents with proactive context management

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-07-10T06:36:52.518800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:4a4a6755c855ed369bcf876adae021484d1dcecccecdf61743d0a69bfc50e677

Observation 7b42fe35-f250-400d-a039-20e52b447e34 · outbound

This paper cites Worldmm: Dynamic multimodal memory agent for long video reasoning.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Worldmm: Dynamic multimodal memory agent for long video reasoning

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-07-10T06:36:52.513341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:4abc08194549c8d6285b335be8ae876b8673639b2ec7b2a09d22943531716798

Observation ed3a82d3-2bbf-40c7-928b-a0d057c187df · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-07-10T06:36:52.565720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:a3e8936ec69b07e309372acca2ff5423db4ad802e61bd3a7f94da11692424d65

Observation e9555c9c-9e65-4e8e-beb9-46a671eaf0d0 · outbound

This paper cites Multi-turn consistent image editing.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Multi-turn consistent image editing

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T06:36:52.944652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:d263e1b0afcba190cf18deec4e218caa092027088aa7b1a12bcc2a62939260b2

Observation 8e0fb181-2552-434d-9b5e-cafdf3777fa3 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-07-10T06:36:52.524292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:90d4c53bd782a4062b934cbaf3b03c5c77872f6655130b93bf685bfdb1dd6e6e

Pith citing papers

No inbound Pith citation observations are available.