Pith. sign in

Paper Citation Record · LEDGER

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization

As of 6 August 2026, this Paper Citation Record lists 82 of 82 outbound references and 1 inbound Pith citation observation for arXiv:2604.06777.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.06777 v1

Coverage vector

measured 82 of 82 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T18:20:02.559108Z

measured 83 of 83 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-30T21:38:08.611674Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

82 of 82 outbound references displayed

  • verified exact45
  • verified fuzzy29
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch4

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0824279e-b1b2-48f0-90fc-9e29c3eb427d · outbound

This paper cites Qwen2.5-VL Technical Report.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Qwen2.5-VL Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-11T00:45:50.219986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:cee823d2854c317b9f381612ba63674d10678a2c755d076c5f20cd4c39ed2954

Observation 51bd245e-7b26-474a-af68-f51e37231b67 · outbound

This paper cites Seed1.5-VL Technical Report.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Seed1.5-VL Technical Report

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:06.820639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:abc9ed52af273a9586792359c2e3ad9404a3edc429e60ecf574a479baad7e888

Observation d4849641-f7b0-4652-aea1-879283ec4a98 · outbound

This paper cites Visual sketchpad: Sketching as a visual chain of thought for multimodal language models.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Visual sketchpad: Sketching as a visual chain of thought for multimodal language models

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T05:14:05.448632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:813f9a6eb3746d4939c7866c70d170bbfef552de3f34f3f7831627671d241113

Observation 57061bdd-2297-40f8-8564-6eed3fef0bba · outbound

This paper cites Thinking with images.https://openai.com/index/thinking-with-images/.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Thinking with images.https://openai.com/index/thinking-with-images/

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T05:14:05.451898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:2abd0fe085085c43b1e1cdc50c1287a66f23662246d56c35747b3eeaac218927

Observation c76b4e67-8fd3-4a5c-a447-1bf63fb47abe · outbound

This paper cites DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:42:57.181531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:31cd1b4e5e76fe730b712fa5ec4cd28c8d533e9608f6f927754d7aa76658afb2

Observation 2734c21c-f751-4c00-b0be-0b6995e451fe · outbound

This paper cites Thyme: Think Beyond Images.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Thyme: Think Beyond Images

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-15T00:33:29.495992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:c2e61d083b62f3402bd25e85f113d06b185f7a7003d4b9dfac165e384d5b26b5

Observation 5faa6fd7-50f9-480d-be80-305e691693eb · outbound

This paper cites Mini-o3: Scaling Up Reasoning Patterns and Interaction Turns for Visual Search.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Mini-o3: Scaling Up Reasoning Patterns and Interaction Turns for Visual Search

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:17:55.768539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:fd3a9f259b2f5895ff00982506ced123c10b0d0e4f08943ee18c7dbb93968b43

Observation 0a36ce3a-39eb-4ed4-a0ae-01d8473eebaf · outbound

This paper cites Deep but reliable: Advancing multi-turn reasoning for thinking with images.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Deep but reliable: Advancing multi-turn reasoning for thinking with images

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:45:50.046372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:c57972aee358997f1682ca44a6065fee30f9da9dd4dc1d9169451272d89a2088

Observation 91322220-7493-4638-8c19-141e56fe2a83 · outbound

This paper cites Llavanext: Improved reasoning, ocr, and world knowledge.https://llava-vl.github.io/blog/ 2024-01-30-llava-next/.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Llavanext: Improved reasoning, ocr, and world knowledge.https://llava-vl.github.io/blog/ 2024-01-30-llava-next/

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T05:14:05.542218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:2d879823c8ebdbdd5dd3a96a30b197c0e268be6b26b0d6f47afd600a0dda9f25

Observation 99136b9d-790d-456a-8fd8-33e6f24f0326 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-11T00:45:50.049454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:17e1468b19eb2bd73a8872cbb1effb0bac18bbf883ecb198c152251b0515a17a

Observation 14d5b53b-7dab-4662-a12b-6e333a1b4f5e · outbound

This paper cites Kwai Keye-VL Technical Report.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Kwai Keye-VL Technical Report

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:45:50.095758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:c107ed95b2b454554207f61dab907ac3f5ccca69d2756d87a9a81bc406c57142

Observation b602e8d3-fff2-446a-bea0-0b11e78b8181 · outbound

This paper cites Ovis2.5 Technical Report.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Ovis2.5 Technical Report

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:30:17.288454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:536184fd7876d142b14d4851b9a78e9c72233f0048a14abafb0dfe0b5daf0777

Observation 9c676ea5-6e55-4fa4-b1ba-e95346b6a391 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Learning transferable visual models from natural language supervision

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T05:14:05.535839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:565099db9a4f82128ba242c0b89bc60793b25bf7bc29c96816078b43520f07f9

Observation 74bfbb62-cd82-4e0c-a785-3622ec31e4ba · outbound

This paper cites BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T05:14:05.539324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:8da63a371afe310f434df5569a27c6937fdbe6d799f8cf21eecc82f2e94017e0

Observation ed505a88-b6fd-424b-ae49-9b8d9cf4d93c · outbound

This paper cites BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T05:14:05.557477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:a7527a1763eed9cf90313074961fb4abc5f17c08951ad140721ee3c93fe614db

Observation aef8e27e-8660-4c24-a180-2b865840a4e7 · outbound

This paper cites Sigmoid loss for language image pre- training.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Sigmoid loss for language image pre- training

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T05:14:05.554326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:e6803100ed7d5fad1e08aa2292316ad79a6945d44497681d058176ca2276b2cc

Observation 7e46cd40-1d3d-4923-ae93-887f65da2b20 · outbound

This paper cites Show and tell: A neural image caption generator.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Show and tell: A neural image caption generator

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T05:14:05.517695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:3a9fd2a1f7329f84de5e2e2a2190ab748d449cea89a26cae5eaf800ee3a6209b

Observation 24f1d174-58b8-41e1-8930-5bd8892a1e65 · outbound

This paper cites Show, attend and tell: Neural image caption generation with visual attention.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Show, attend and tell: Neural image caption generation with visual attention

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T05:14:05.486859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:17e3d0dfbf53f0ad0a2a53d9c9ee56788d7aaa335d7814fb04cb3049df2c0914

Observation 1affeafd-4778-41de-8ea4-492c579c9f8c · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Flamingo: a visual language model for few-shot learning

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T05:14:05.499234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:32f3f73542b133e5a9b763cd31efdbd67f9c6a48bed9a60c025366c0a10d0607

Observation e592f24f-5b3f-44e7-bd48-d6409248ea92 · outbound

This paper cites Visual instruction tuning.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Visual instruction tuning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T05:14:05.512973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:d482433bb9b3a030c1e5b504c71fa4d671440d1fca0a438e97db80fabd9ef211

Observation 3e6fb234-22c5-4528-ab7b-002c0213761a · outbound

This paper cites Language models are few-shot learners.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Language models are few-shot learners

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T05:14:05.551624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:e6eaa99483bc2fd59c7a7aaaa9f55a3bb0b82d1955aa1c6c652c45df0e55234d

Observation 1aad9eaf-692c-4b6c-9d23-69756211b4d7 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization LLaMA: Open and Efficient Foundation Language Models

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-11T00:45:50.208622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:c7a2f73460f5bf989fc9b3b14abc02a2494a07425502997dd8126c374d73b506

Observation e885d802-26b6-456e-bc1a-5c0a1868753f · outbound

This paper cites Training language models to follow instructions with human feedback.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Training language models to follow instructions with human feedback

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T05:14:05.458674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:2701c87006559a1ff032d6c32f4ee42a549cff03d20d6e6316b3c0ac4df96d16

Observation 5a98bd98-33e3-4111-8a37-9cf928554a1f · outbound

This paper cites Improved baselines with visual instruction tuning.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Improved baselines with visual instruction tuning

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T05:14:05.528701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:f879dc31660db1fa561433bbd0e0e9e05cb353104450581a8b5c3fe719835a40

Observation a0cfa46d-cecd-4c59-b462-3ed711aa4c48 · outbound

This paper cites How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites.Science China Information Sciences, page 220101.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites.Science China Information Sciences, page 220101

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T05:14:05.548450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:fb1393fd04b73d2cede16a600f59836a344c4ed600c01fc6c09973b6b67a975d

Observation f859988a-690c-41bb-97c1-c087a1441a09 · outbound

This paper cites Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T05:14:05.545536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:95b8987aab94290eeacc5f430c6943dd847d81b86cf1a29808689e92a29bc1ec

Observation 9c2bf4f9-43de-4769-aff3-12657367f6ae · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-11T00:45:50.212455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:8e8ae94ca6926d47ac9565bea2c6284ffe8c91c0b6f040b2021588989c7729f2

Observation e937d26b-8bb0-4fb0-b115-b42e3001aba0 · outbound

This paper cites Qwen3-VL Technical Report.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Qwen3-VL Technical Report

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-11T00:45:50.200508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:d2040ba0523636bf35ab2e079b9c769ff33a9ae9c031be2b27126145a8fd6785

Observation dc01909e-d858-4ee6-975c-c22717e2873b · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T05:14:05.462568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:c7a767b4f76041e340bc41c382e6a3f11e1253eb7d06c20c52daa877407c6c99

Observation dbb5e3f9-8286-4f06-a8c7-c91cfe99e2a7 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-11T00:45:50.204490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:9dca3392e1b6c3e7ffa34fe3d045eede2119f7c599172d1e40f2f729bf56ff4b

Observation ef8d8350-e08d-4672-ad60-e45d15abe2c8 · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-11T00:45:50.062623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:98b7327e8da278ae1e4557465e39364a647bb87d2c99ff08368179c5d2139cff

Observation c5a2c74d-25d1-474d-85cf-87a7b22718e4 · outbound

This paper cites Ovis: Structural Embedding Alignment for Multimodal Large Language Model.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:45:50.076778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:f315b724f5b0e47b6b16bd9b630b79f1a4a530770a8da0a560bcbcf45f5cea47

Observation 749cd877-3525-4518-b60b-54876c5b3f08 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization LLaVA-OneVision: Easy Visual Task Transfer

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-11T00:45:50.192737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:58a6736a9106614b815c64e9562e655856ebb79ee45c86273de48399c83bc856

Observation d3cff9b5-13fd-4e3f-a775-4f383422ced1 · outbound

This paper cites Multimodal Chain-of-Thought Reasoning in Language Models.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Multimodal Chain-of-Thought Reasoning in Language Models

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:12:27.664923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:ed0cfda2111d1af648d567ab932e287ff86805ba5beb16a03e770762e619164d

Observation eff8e9f7-7997-4900-91e8-1a4f2be97088 · outbound

This paper cites an unresolved cited work.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-05-17T05:14:05.531990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:19b3fb0cb5495088f55f916499d1e4991f68888a65881467d8f80a98a4fe43ca

Observation b6ecd65e-7569-467c-8290-ccb324a1b228 · outbound

This paper cites Satori-r1: Incentivizing multimodal reasoning through explicit visual anchoring.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Satori-r1: Incentivizing multimodal reasoning through explicit visual anchoring

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:45:50.180651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:e3baee05c3f9532bc191e0b5e1e48b1b66004c0606f22c235d76e051993fccfa

Observation fd9edc39-c727-4ea7-9be8-a6c37f628a02 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:30:15.750849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:c553afee35dd9564ce3af06c9307e3ef86cf40bf50fe14fc33edabc0c9dcb7c6

Observation 745a9d40-ea66-4c4d-bbf3-3a16a50535e7 · outbound

This paper cites Toolformer: Language models can teach themselves to use tools.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Toolformer: Language models can teach themselves to use tools

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T05:14:05.476636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:c10b7d0ba5a56f5c15ef6fb94b62eeb800b7e628408b0bc4a6fb1df594cb29cf

Observation 67f11fdf-69db-492a-b679-a7bfa042a3a2 · outbound

This paper cites React: Synergizing reasoning and acting in language models.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization React: Synergizing reasoning and acting in language models

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T05:14:05.473160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:d520f47001d31838ab3b6cf3ef8b69d78ad3f7e3277b93ea1f541865fbd9d7ac

Observation fd2f3094-b3b1-469d-aeba-98b6bcbed5c5 · outbound

This paper cites MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-14T01:17:58.977376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:9b815eb4702fc7c4c4d2dc87ec0b035e8298c85b798398f2a523de6b1b0bd22c

Observation dec36dd2-bf1e-4ce1-9f0c-2b9e015ebdbf · outbound

This paper cites Visual programming: Compositional visual reasoning without train- ing.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Visual programming: Compositional visual reasoning without train- ing

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T05:14:05.483554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:7f0953c7509692190cf31f2ea6b5d4b46d000aeccc4973fec926e4640d26f5df

Observation 3a1d39f3-e894-4a73-ba7c-f4e92d44d3e4 · outbound

This paper cites REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T09:21:45.673967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:8a6c4656d6c93187c12f4b65d5fee002016087cf7479c9e8264b82c4a8a7d44b

Observation 0732976c-d11f-4b3f-933b-3480377d370d · outbound

This paper cites Proximal Policy Optimization Algorithms.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Proximal Policy Optimization Algorithms

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-05-11T00:45:50.162903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:bf6a0ccfc3f19c761e88261b4403e6754417cb2f78a13d0f4186512dfa807f1b

Observation 212f6652-4e15-4cbe-b609-52ebbbbe7b1a · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-11T00:45:50.116532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:562abe688d36765289e8fcce6d4129f35b814ffcfc0c1a8bdc2c874bd4ebf329

Observation a90c4980-cc53-440d-a8c2-8112c29fea9a · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-11T00:45:50.059446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:9ce8713dd5533aafd2fd7088319c0a7c179c81f887da85f2993e0f32dbbae1d6

Observation f020a5f1-104d-4d5c-b581-f4def7661036 · outbound

This paper cites Group Sequence Policy Optimization.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Group Sequence Policy Optimization

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-05-11T00:45:50.184268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:0d2bb02893adc3c4d7069f94cda8fb5dc5411c3a65c9d9bca08c9c4a537b9aa7

Observation 4ee9354d-9266-4470-b354-2f84ff34aca4 · outbound

This paper cites Kimi K2: Open Agentic Intelligence.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Kimi K2: Open Agentic Intelligence

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-05-11T00:45:50.069676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:1bfe78c4007c29a704ced83fd2d4fe7173a075efa7320dc9c04adb9fcb8d4d92

Observation 7b074a6a-42d9-4f4d-b6d9-ebadc9e61e0c · outbound

This paper cites WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:56:24.054213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:41bf8b965e2bb8576ab5cdf2b6dd73bfab3e6086c9a6cc5d0832a5d7fb1ff517

Observation 5cd4db6e-d670-4757-9327-b65a1ed2e3eb · outbound

This paper cites WebSailor: Navigating Super-human Reasoning for Web Agent.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization WebSailor: Navigating Super-human Reasoning for Web Agent

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-17T15:37:09.773663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:3354b30efdc77f3ecc8472b82b3254e428d60a3f5b646b80dde3e61124ebbc3d

Observation 13040d64-8ccd-45cf-b01d-be7d383320dc · outbound

This paper cites WebShaper: Agentically Data Synthesizing via Information-Seeking Formalization.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization WebShaper: Agentically Data Synthesizing via Information-Seeking Formalization

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:45:50.158468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:76c02bd8032bf3daafb9ab1a31c84368dc3a50d2cce18c5ede6f649133804b62

Observation fb5d1019-2f23-4e62-9b69-c82919ee8576 · outbound

This paper cites Agent RL Scaling Law: Agent RL with Spontaneous Code Execution for Mathematical Problem Solving.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Agent RL Scaling Law: Agent RL with Spontaneous Code Execution for Mathematical Problem Solving

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T00:45:50.167268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:026598810d8eb48b869a6ae18ff343aa4af24660f4325b9d6ca33b35d2fb8d2c

Observation 91ff23a9-3d9f-4891-9a74-03c8c64502b9 · outbound

This paper cites SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:45:50.171881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:25c6b0b352f270619dee45908442b5e01d5b58b5ebe2fb585f5108e021d5b51f

Observation 83bf4041-421f-4799-826c-c324198a0115 · outbound

This paper cites Explorer: Scaling exploration-driven web trajectory synthesis for multimodal web agents.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Explorer: Scaling exploration-driven web trajectory synthesis for multimodal web agents

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T05:14:05.502562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:a067bb6e9f22bb62cdc3b3f72f049eb90cce1d6e15ccbd9d5d6d360878d972c5

Observation 943c6789-6844-4526-b83a-c2c61459f484 · outbound

This paper cites RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:13:34.687091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:8dbb9bab16d6f47e6a88cc80a7e4463e99b3c7e37b3d5436a26c5727d5205212

Observation 65d14b92-343e-4f1f-b38a-dfc12d969005 · outbound

This paper cites Thinking with Images for Multimodal Reasoning: Foundations, Methods, and Future Frontiers.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Thinking with Images for Multimodal Reasoning: Foundations, Methods, and Future Frontiers

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-13T08:34:23.540456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:d19494e8c49ea934ca1e6691bac2b2aecff36bae6ca6ce9457b74e4163cefb52

Observation 3a010412-b2be-4ea9-be83-e9400d76d655 · outbound

This paper cites Latent sketchpad: Sketching visual thoughts to elicit multimodal reasoning in mllms.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Latent sketchpad: Sketching visual thoughts to elicit multimodal reasoning in mllms

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:45:50.145785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:6839e5a1a9e05332da9f08f407108f27f07fc99d5bd3dcf33602ebbfb5d1b439

Observation 17de2ce0-8681-41ed-b83c-c66a1fa2614b · outbound

This paper cites Monet: Reasoning in latent visual space beyond images and language.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Monet: Reasoning in latent visual space beyond images and language

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:45:50.149921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:0d4f14e8044ff12de0f10b28fd691ce0bd949ea882e6d4aa6802fd6f858579c9

Observation e1aa2929-42ee-4b61-91c8-3aa65be2d3db · outbound

This paper cites Latent Visual Reasoning.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Latent Visual Reasoning

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:41:30.500607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:8968cfd013b256cc329604572203ede90ba5ca1e2d0eca3473d6985556046aa8

Observation 598ea679-32ab-4f53-878f-369b9ebb2ee2 · outbound

This paper cites Interleaved latent visual reasoning with selective perceptual modeling.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Interleaved latent visual reasoning with selective perceptual modeling

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:45:50.087660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:36fa7583461817914c2eec13e30dfa2743972d83bf4eca97e23f6dd1f508c405

Observation 31f3e513-db5e-4382-af03-15f82102ef10 · outbound

This paper cites DeepEyesV2: Toward Agentic Multimodal Model.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization DeepEyesV2: Toward Agentic Multimodal Model

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:32:29.694372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:104f1726753b0738688d6411c03157a3b57a9209489f697fef6572bb0c75016f

Observation 052aa34e-c3d5-4b46-8620-57bb85519c02 · outbound

This paper cites Adaptive Chain-of-Focus Reasoning via Dynamic Visual Search and Zooming for Efficient VLMs.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Adaptive Chain-of-Focus Reasoning via Dynamic Visual Search and Zooming for Efficient VLMs

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:35:13.359401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:2033a75ef91789cc8a49a85890ba8d35b88c8c2e3763da7234afbf9457e0705b

Observation 4101407c-c167-430d-89bf-2f14a241def1 · outbound

This paper cites Pixel Reasoner: Incentivizing Pixel-Space Reasoning with Curiosity-Driven Reinforcement Learning.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Pixel Reasoner: Incentivizing Pixel-Space Reasoning with Curiosity-Driven Reinforcement Learning

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-14T02:22:27.090901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:7ceefa584d3af4fdf231ff4a7ddd521d2b63b02082bcd64c7a9fcbcdb9cc4d67

Observation f5beb4e0-9243-4940-a28a-9be871b815b3 · outbound

This paper cites High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning

Reference 63

Resolution
verified exact
local_arxiv, observed 2026-05-11T00:45:50.112672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:610c3ec06d9f33fe6ab212772bc02865d9cfb1fc89d44c7303b4bf944ea3ceec

Observation 98f7b2a9-196c-4e0a-bc8e-28aa5b145015 · outbound

This paper cites MMSearch-R1: Incentivizing LMMs to Search.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization MMSearch-R1: Incentivizing LMMs to Search

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-16T15:27:04.573030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:0cd46ac9b5d6cc030220e1b739262c2e1ddafced7af0469aac4798c107f18c2a

Observation f55e675c-9e83-4278-b74b-1bb613346095 · outbound

This paper cites VisionThink: Smart and Efficient Vision Language Model via Reinforcement Learning.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization VisionThink: Smart and Efficient Vision Language Model via Reinforcement Learning

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:45:50.083817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:258f948052c100c880951ec76271e1b7f4c053bd167d3931115107c35ec685c3

Observation a3849417-c099-4ddc-a179-1b1a96fbc4f9 · outbound

This paper cites ACTIVE-o3: Empowering MLLMs with Active Perception via Pure Reinforcement Learning.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization ACTIVE-o3: Empowering MLLMs with Active Perception via Pure Reinforcement Learning

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-06-09T03:07:59.867366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:0eadf08d513b0db16f152f9889853b14b805d4681fb90a92746023ee039bfa86

Observation 68a4a4b5-454f-45a1-b45a-c0602ff4824d · outbound

This paper cites Thinking with programming vision: Towards a unified view for thinking with images.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Thinking with programming vision: Towards a unified view for thinking with images

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:45:50.133336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:8cc23382f7cdc6418e89bf8705af46f8a5e755a08776ad0119dfbdbfe24692a3

Observation 0d0649de-7dc6-4641-9d38-0bd76b0a6402 · outbound

This paper cites Vacot: Rethinking visual data augmentation with vlms.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Vacot: Rethinking visual data augmentation with vlms

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:45:50.052410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:9c7815ef52bbb8ea623899cd931bbc9786ec8967e70feabf995df410dfbb55e5

Observation 260d141c-5131-4670-8b54-78b2e1aa3899 · outbound

This paper cites arXiv preprint arXiv:2602.12916 , year=.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization arXiv preprint arXiv:2602.12916 , year=

Reference 69

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T00:45:50.141554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:0a1705419b3788a9b252fc87b215bec4435822f20e9e85bd0645937f0b948a73

Observation 04c301e5-6b5e-4bec-b822-39af19c9d3ae · outbound

This paper cites Hybridflow: A flexible and efficient rlhf framework.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Hybridflow: A flexible and efficient rlhf framework

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T05:14:05.466157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:f0caa423e2f77ab177579f249a7c5dabe0fba3735b511667035be812346c9c9d

Observation 605af6cb-b296-4c94-a8a3-a7e2f7df4831 · outbound

This paper cites Momentum-based variance reduction in non-convex sgd.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Momentum-based variance reduction in non-convex sgd

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T05:14:05.509515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:35f20376ae43360861a353009195631039ff04e6a1b398fe4d2ffc870c49c308

Observation 6ec6489e-66a8-4588-873a-de8b58b255bf · outbound

This paper cites Gpt-5 system card.https://cdn.openai.com/gpt-5-system-card.pdf.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Gpt-5 system card.https://cdn.openai.com/gpt-5-system-card.pdf

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T05:14:05.455333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:ab98d033ace709ea3d0b23415a9ac28e07cab4c51ca19f5a44b1afe1675573aa

Observation 289660dd-1155-4bb1-9317-45204dd24262 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 73

Resolution
verified exact
local_arxiv, observed 2026-05-11T00:45:50.137506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:85d5898d781745ce9be1680e9d3ff70fc89f5d26c31c117de238ad1c0e2d14e4

Observation 898d3559-3658-4270-b87a-7cb07bf1ace5 · outbound

This paper cites V?: Guided visual search as a core mechanism in multimodal llms.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization V?: Guided visual search as a core mechanism in multimodal llms

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T05:14:05.490077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:0008901d0b610b3434953b5c5dde3791b971cbca4a5a29cb57bc5e1cae403c24

Observation a3d5a24d-02da-4bce-99e1-d61d13540fdb · outbound

This paper cites Divide, conquer and combine: a training-free framework for high-resolution image perception in multimodal large lan- guage models.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Divide, conquer and combine: a training-free framework for high-resolution image perception in multimodal large lan- guage models

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T05:14:05.493060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:821021db62b4f77c2f2ebdcfa0f02cd9129376b7d7f90fa3fc9a0ed94ecfc3fc

Observation c18c734a-b9b2-4098-9036-f621d47f9e38 · outbound

This paper cites MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?

Reference 76

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T07:59:32.958879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:e7251fb1865d26d9e1c7172b922191273e598457e636d48258cdfb20f087efd7

Observation cc9c9c16-8d6b-44c9-b715-aab5dc443c4b · outbound

This paper cites an unresolved cited work.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Unresolved cited work

Reference 77

Resolution
unresolved
raw_fallback, observed 2026-05-17T05:14:05.495829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:195399a7d50f102609e6b1c6bb49186725bc700016aedc8dd3c43c109dd552e6

Observation 548b0822-a4b1-491f-8913-215fcdb13900 · outbound

This paper cites Varr(ˆgout|τ) =∥∇ θ logπ(τ)∥ 2 ·Var(r out) =C τ ·σ 2 out (21) Varr(ˆgsem|τ) =∥∇ θ logπ(τ)∥ 2 ·Var(r sem) =C τ ·σ 2 sem (22) whereC τ =∥∇ θ logπ(τ)∥ 2.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Varr(ˆgout|τ) =∥∇ θ logπ(τ)∥ 2 ·Var(r out) =C τ ·σ 2 out (21) Varr(ˆgsem|τ) =∥∇ θ logπ(τ)∥ 2 ·Var(r sem) =C τ ·σ 2 sem (22) whereC τ =∥∇ θ logπ(τ)∥ 2

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T05:14:05.524735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:9f988656cba5b31d8043e352646a849339457651ffc2ad3b48c68344b1b0f2bc

Observation 1acc7250-554a-45a4-94eb-41be46b8fce1 · outbound

This paper cites an unresolved cited work.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Unresolved cited work

Reference 79

Resolution
unresolved
raw_fallback, observed 2026-05-17T05:14:05.469357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:31b646c5e947ba0a66bfd4ea94f01e5159c6831d40835d8a04444e58cc8c5d98

Observation abf79b2a-3732-49bc-abe1-de71f31464d2 · outbound

This paper cites an unresolved cited work.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Unresolved cited work

Reference 80

Resolution
unresolved
raw_fallback, observed 2026-05-17T05:14:05.521174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:3023102d46a3e24503f5ad4bbb9f842e8614603f65648010112f278526b83ef9

Observation 57af2e20-13f1-428d-8ffc-563dc82ad2a8 · outbound

This paper cites type":"function.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization type":"function

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T05:14:05.480070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:6b32641318658cf779e5167e2a116ab7543d2be05db5b92f823499db5bf2f795

Observation 3cea1cd7-16df-420a-9025-49d5f5ac1de9 · outbound

This paper cites The input image resolution is dynamically handled, with a pixel constraint range of[10 6,2×10 6]pixels to support high-resolution visual reasoning.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization The input image resolution is dynamically handled, with a pixel constraint range of[10 6,2×10 6]pixels to support high-resolution visual reasoning

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T05:14:05.506196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:05e62e9e3cd202fea7eabb2cc2f65d90efc56e35232735ab042e0d28dc0f647d

Pith citing papers

Observation 9b906d49-9f1c-4e96-80b6-e2f12fa6e915 · inbound

See2Think: Do Multimodal Models Really Use Intermediate Visual States? cites this paper.

See2Think: Do Multimodal Models Really Use Intermediate Visual States? Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-30T21:38:08.611674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T21:38:08.611674Z digest=sha256:6ca5b7d4744fdf03f8c391f2697ba0c6442437e08cbdd561efebc338c4a1ca02