Pith. sign in

Paper Citation Record · LEDGER

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification

As of 17 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 8 inbound Pith citation observations for arXiv:2506.07235.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.07235 v1

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:45:30.735139Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T05:31:41.386191Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T13:43:28.625652Z

Reference resolution

38 of 38 outbound references displayed

  • verified exact0
  • verified fuzzy10
  • unresolved27
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 26949dd0-946c-4b20-8aae-73a130af3245 · outbound

This paper cites Blink: Multimodal large language models can see but not perceive.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Blink: Multimodal large language models can see but not perceive

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:45:31.201843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T05:45:30.587797Z digest=sha256:76379d9017507e56251f88bbfa455d50edb3c5519acc9aa66a92cb9f14882e54

Observation 252017fe-5670-4cb6-b23d-e17e2c40e8b6 · outbound

This paper cites V∗: Guided visual search as a core mechanism in multimodal llms.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification V∗: Guided visual search as a core mechanism in multimodal llms

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:45:31.188988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T05:45:30.591770Z digest=sha256:62f45065f8cdbe717a322785fbea8f784553c53d6fddc2070aac031b1320691c

Observation e674edb9-6e54-4a7f-aa85-33d81f32411a · outbound

This paper cites Eyes wide shut? exploring the visual shortcomings of multimodal llms.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Eyes wide shut? exploring the visual shortcomings of multimodal llms

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T05:45:30.595949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:45:30.595949Z digest=sha256:f36e219635ca79075ef3581cdddd581beb0431e999293c5a2207588888444681

Observation e5b51c26-c286-46df-9016-544557ae8b24 · outbound

This paper cites MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:45:30.600959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:45:30.600959Z digest=sha256:a374c0ad647920fc1ad1a3f3824f8f0066a52f7f03dfa271e75ac4312a26e00a

Observation bbb19b7a-cbac-4017-9c6e-430cd949d879 · outbound

This paper cites Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:45:30.605338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:45:30.605338Z digest=sha256:f09d8ba372efacf2dfd36b518cd2ad7e5e75de330e8f71c3500bf808a75d4362

Observation 52e12acf-5b2a-4255-8ca5-94afd5e420d3 · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:45:30.610059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:45:30.610059Z digest=sha256:88a223dd8b6b3391118070db77895837f508015eba8fe11544c038b919effb40

Observation 5d8a4253-37a2-4eec-abcf-5c91400e7576 · outbound

This paper cites Visual program distillation: Distilling tools and programmatic reasoning into vision-language models.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Visual program distillation: Distilling tools and programmatic reasoning into vision-language models

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:45:31.168232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T05:45:30.615181Z digest=sha256:71683f1845abd9502873b04ddb91219446a35e8699ea1b78e0ff453069b9d70b

Observation cbcf874c-d05b-42ec-a1f6-e7368d67ad0b · outbound

This paper cites Visual programming: Compositional visual reasoning without training.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Visual programming: Compositional visual reasoning without training

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:45:30.619311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:45:30.619311Z digest=sha256:7ac576b018c630de51787bcd741e2f6ed96d0d5593e33d74a1f1686a2fa6f5e6

Observation 9b9772bb-683c-4563-8847-1e8b248f6754 · outbound

This paper cites Vipergpt: Visual inference via python execution for reasoning.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Vipergpt: Visual inference via python execution for reasoning

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:45:31.146028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T05:45:30.623036Z digest=sha256:8c43ff65970e3058d07480f245ef79163ecb3803b1e7867ce86e3acf948d118e

Observation 92032de7-d6d7-48c5-9a3f-a5cf9317df0b · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.Advances in neural information processing systems, 35:24824–24837, 2022.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Chain-of-thought prompting elicits reasoning in large language models.Advances in neural information processing systems, 35:24824–24837, 2022

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T05:45:30.627374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:45:30.627374Z digest=sha256:7665bcad00d24a6151f20aa661143b82ef82576169471fa75527f4a7905f36c3

Observation b523a58f-2eef-46ee-b336-3950e403d8e8 · outbound

This paper cites Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T05:45:30.631557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:45:30.631557Z digest=sha256:9d090304d58abdd44f875cce6ffb46f9f03a89515fc8796ac360770f8b54a829

Observation bf3df506-9ebc-4103-a789-956e840cebd0 · outbound

This paper cites Scaffolding Coordinates to Promote Vision-Language Coordination in Large Multi-Modal Models.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Scaffolding Coordinates to Promote Vision-Language Coordination in Large Multi-Modal Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T05:45:30.635922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:45:30.635922Z digest=sha256:a2eb8e6453dc3658db97e989539035ac61623af5f1a435c7f8e97ad29f9587ab

Observation dcf6d6da-1844-4d90-a1f5-fa4898a2e6f5 · outbound

This paper cites MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:45:30.640388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:45:30.640388Z digest=sha256:9a76526feef8dcc9251730b1bccdb7ed80402a5d5a60229bb134ce78a19f6b0f

Observation fc168fb3-3060-4879-b51d-a94c1a8661d8 · outbound

This paper cites Making language models better reasoners with step-aware verifier.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Making language models better reasoners with step-aware verifier

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T05:45:30.644103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:45:30.644103Z digest=sha256:f3a4274fa5b1e8d69e86370c2e8e0ed8daa7357e55ff4d7672909e07e00a96db

Observation cb1f1667-6ae4-482e-a7ad-06dfc7a1262a · outbound

This paper cites Let’s verify step by step.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Let’s verify step by step

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:45:30.648369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:45:30.648369Z digest=sha256:7869d8136173dcd5b3d37be268a16f62af14c69a224a21cc6aff0fd6a936f608

Observation 3e4fe20d-b57c-481a-b5cd-c493fd0e6d42 · outbound

This paper cites Solving math word problems via cooperative reasoning induced language models.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Solving math word problems via cooperative reasoning induced language models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:45:31.109578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T05:45:30.652126Z digest=sha256:de5cad9c7b78c48b4fecb0eb3cfaea32469abdc85bf9713824905482d213df99

Observation fb622090-c761-4f60-8374-09194f456762 · outbound

This paper cites Let's reward step by step: Step-Level reward model as the Navigators for Reasoning.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Let's reward step by step: Step-Level reward model as the Navigators for Reasoning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T05:45:30.655336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:45:30.655336Z digest=sha256:23fc0b42854d8fe98b6bad81a195f1422b3e507c334d00b8ad664c4fe47c8b60

Observation 2e143f65-c3c8-497e-892e-f6efc4614f71 · outbound

This paper cites Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:45:30.658831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:45:30.658831Z digest=sha256:cc42769e04d765b12e1c67dc142d5d85c47afb0db47ed13a9d433617ca09a215

Observation 68ac5975-e95e-4064-b93e-b95e536655c7 · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Fine-Tuning Language Models from Human Preferences

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:45:30.662617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:45:30.662617Z digest=sha256:dffdaacc96b22ffcd5df8baf9a96820761c3a68dbcfefd7e6fa9f78b60985d3e

Observation 184ebd8e-b7e4-47db-80de-756321b7cf7b · outbound

This paper cites Rlcd: Reinforcement learning from contrastive distillation for lm alignment.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Rlcd: Reinforcement learning from contrastive distillation for lm alignment

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:45:31.097704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T05:45:30.666869Z digest=sha256:519eecf5bc3a85ce03d6f48a518f1aae7ccfe5178b8bec0840a39cd85a7f5293

Observation 3ba739d5-16d2-4232-bc94-8f028c22e319 · outbound

This paper cites Pretraining language models with human preferences.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Pretraining language models with human preferences

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:45:31.084696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T05:45:30.671380Z digest=sha256:2d43ff9f49332ba9a7b5569b48545800209c012d1b963b57db3e13fe0037da0b

Observation cdee5b0e-b45c-4eef-b8b3-e02f34ede260 · outbound

This paper cites Improving Large Language Model Fine-tuning for Solving Math Problems.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Improving Large Language Model Fine-tuning for Solving Math Problems

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T05:45:30.675404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:45:30.675404Z digest=sha256:ae0420ab622c4daf1a5becfd36346425868983338db6f4438d34e9d52a1dcdf6

Observation b221ca03-9306-44e2-b182-2b87ba767138 · outbound

This paper cites V-STaR: Training Verifiers for Self-Taught Reasoners.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification V-STaR: Training Verifiers for Self-Taught Reasoners

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T05:45:30.679605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:45:30.679605Z digest=sha256:38b49941df8345c60fc37adff98527205a2e1d38800187f2f387dd43acefaf3a

Observation 0e2ad5c0-c823-4ec9-b278-869e749f1dc5 · outbound

This paper cites Building Math Agents with Multi-Turn Iterative Preference Learning.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Building Math Agents with Multi-Turn Iterative Preference Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T05:45:30.683409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:45:30.683409Z digest=sha256:5c54b6170300053080ec69977e4288366594e3cf5ae36586e4c9742505e39e5f

Observation 4bde0720-d8c2-4b59-923e-cfb12b1e46bc · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification LLaVA-OneVision: Easy Visual Task Transfer

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T05:45:30.687744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:45:30.687744Z digest=sha256:07db7ac70ec5a564b046af7193b749096f7c97f2a35ead26721fe42672bc6624

Observation 7a1641a5-f9cd-41d4-b1f7-03d615ea423b · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:45:30.691615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:45:30.691615Z digest=sha256:2a563327f821df30a574e4cb83cbb1ba926f115960c316f9a47ee95f00ff7723

Observation 5afb27e4-9bc5-454f-a275-f37192943a9f · outbound

This paper cites Qwen2.5-VL Technical Report.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Qwen2.5-VL Technical Report

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T05:45:30.694953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:45:30.694953Z digest=sha256:b86927cb6e94673cd35fcd86aba90b1da073172b7b524875cf236830ddc5def6

Observation f957249a-e9a8-470d-837c-08b495156d34 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T05:45:30.698822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:45:30.698822Z digest=sha256:2dbd289d4a1ff306bf87b94e0ece3c0e46e137c35f398c2d0136a4fb0667961b

Observation 459c0db6-b89b-4b85-948e-7185cb7681f2 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T05:45:30.702326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:45:30.702326Z digest=sha256:f987781c6fc797bedbb244ee55880401dd0bf7b8ca5472b7d97c72fdb4afaab0

Observation 65f1e13d-11f2-4578-9337-f6ffec9cf394 · outbound

This paper cites MMFactory: A Universal Solution Search Engine for Vision-Language Tasks.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification MMFactory: A Universal Solution Search Engine for Vision-Language Tasks

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T05:45:30.706182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:45:30.706182Z digest=sha256:55db94766a1261c6544ec23ea91728c36678688baf85b0a9eba4eb95165c16f9

Observation 0d54322b-c595-4edf-bbea-a1c08399fd84 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Direct preference optimization: Your language model is secretly a reward model

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T05:45:30.709569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:45:30.709569Z digest=sha256:a085d29ae2f2ae8b0ee2417f94088bd73a517f0ea38c737ecc8f68a3f4c675b5

Observation 2c26b258-5e7b-4a1b-8917-bd377116f215 · outbound

This paper cites Mathematical analysis of machine learning algorithms.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Mathematical analysis of machine learning algorithms

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T05:45:30.712751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:45:30.712751Z digest=sha256:10e8e024d73abbb218a21b48231a24c4f15e690a2cc78a8082b3e23e66bddc11

Observation 1fbb9cad-42c1-4b29-9652-ca3470c1e0c0 · outbound

This paper cites Probability: theory and examples, volume 49.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Probability: theory and examples, volume 49

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:45:31.056031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T05:45:30.716671Z digest=sha256:f31a66a62e8aa8af7e91603f17cfbb3ee14ba7a1e314e63767a42e3cf4d77d06

Observation cd360758-56b3-4fd2-a23a-6f26d30c2e42 · outbound

This paper cites an unresolved cited work.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Unresolved cited work

Reference 34

Resolution
parse uncertain
raw_fallback, observed 2026-08-07T05:45:31.043147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T05:45:30.720489Z digest=sha256:b74b61666cd74fcea58b7f3042c5d0e8bdccbe33fc4235f0fe180f1b17de0df3

Observation 44af0ef1-f44d-4789-b130-ec622744a4b5 · outbound

This paper cites If in the last definition,= is replaced by ≤ or ≥, then Xn is said to be a supermartingale or submartingale, respectively.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification If in the last definition,= is replaced by ≤ or ≥, then Xn is said to be a supermartingale or submartingale, respectively

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:45:31.028048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T05:45:30.723796Z digest=sha256:c5d13726c168e4bf2ad38e6279c30b7ae598008c24583e4661e84bebc9e31089

Observation 5c9b9790-6d33-471b-99e3-7bde2da9912d · outbound

This paper cites an unresolved cited work.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:45:31.015628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T05:45:30.727319Z digest=sha256:ec351ddaf1586cbfe9bcb4816536688fcff3bf95ad916bedb4becd7219dca5df

Observation 8af02ebc-195d-427a-b337-ca2f411af61f · outbound

This paper cites an unresolved cited work.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:45:31.004160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T05:45:30.731168Z digest=sha256:60a69b8ba8ea9801cce12dd8d07719a54660bfd7b1a6afb2edf4c3ec9759b329

Observation 5c19b0e8-7570-4206-92b9-a2fd8c9ccfca · outbound

This paper cites log V ˆϕSDPO (th | sh) Vϕ0 (th | sh) +log V ˆϕSDPO (ah | th) Vϕ0 (ah | th) # . Observe that condition (ii) implies that Eth∼R ˆθSFT(·|sh).

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification log V ˆϕSDPO (th | sh) Vϕ0 (th | sh) +log V ˆϕSDPO (ah | th) Vϕ0 (ah | th) # . Observe that condition (ii) implies that Eth∼R ˆθSFT(·|sh)

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:45:30.992308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T05:45:30.735139Z digest=sha256:4604fa28cad50d126f29906d19fd16c97380df13c611764057876cb75cef18f9

Pith citing papers

Observation 39b49e20-785e-4b2e-9bf3-6d436a00e7db · inbound

VABench: A Comprehensive Benchmark for Audio-Video Generation cites this paper.

VABench: A Comprehensive Benchmark for Audio-Video Generation Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:08:43.833402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-17T00:03:45.576961Z digest=sha256:ddff638699ebece9703539a1c50401ea42b9f18e0362962112b6e876aa34c6ce

Observation 12e287ce-b650-4ece-9a3a-d938f677967b · inbound

OpenWorldLib: A Unified Codebase and Definition of Advanced World Models cites this paper.

OpenWorldLib: A Unified Codebase and Definition of Advanced World Models Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:45:49.164337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T19:36:42.100191Z digest=sha256:8d3d80c5319b8dd31b39e40c6560c88bb9f5ceb33296c4de1456c20b59e41703

Observation 979eb267-7af1-4b5b-94d9-ca1c7e9f0ca1 · inbound

OpenWorldLib: A Unified Codebase and Definition of Advanced World Models cites this paper.

OpenWorldLib: A Unified Codebase and Definition of Advanced World Models Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-13T09:42:23.808691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T09:42:23.808691Z digest=sha256:6948f0e8d0c9cc7b963d75ee881ff41ba30efb901bb144c13eae18e643d523a3

Observation d546229b-d332-4831-9195-e3e3a43148fb · inbound

Test-time Scaling over Perception: Resolving the Grounding Paradox in Thinking with Images cites this paper.

Test-time Scaling over Perception: Resolving the Grounding Paradox in Thinking with Images Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:26:01.651378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T16:38:11.785469Z digest=sha256:b287d4a990b9b422a3bb3e272e44c77f3a553fd1ef35cc162833fdde431be9e2

Observation ab9dbc9c-ef0d-4135-bed8-94aa1b062247 · inbound

Test-time Scaling over Perception: Resolving the Grounding Paradox in Thinking with Images cites this paper.

Test-time Scaling over Perception: Resolving the Grounding Paradox in Thinking with Images Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T05:31:41.386191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:31:41.386191Z digest=sha256:f6aa6dfd237656119be788a738f125ad621256ae16cb8f3ea2ca15dcc2fd3287

Observation fee230e5-cb3d-4ff1-923f-61f1c48c44b0 · inbound

Maestro: Reinforcement Learning to Orchestrate Hierarchical Model-Skill Ensembles cites this paper.

Maestro: Reinforcement Learning to Orchestrate Hierarchical Model-Skill Ensembles Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-22T07:51:15.425338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-22T07:51:13.362986Z digest=sha256:b50b17b79436c930670a5a78bc94aca3ba0469f30af63133e0bd08544b140e2a

Observation 37134517-1422-4099-86b4-023ea435cf36 · inbound

ROVER: Routing Object-Centric Visual Evidence for Grounded Multi-Image Reasoning cites this paper.

ROVER: Routing Object-Centric Visual Evidence for Grounded Multi-Image Reasoning Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:43:28.627224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-29T13:41:44.049230Z digest=sha256:fa94f1dcf04dc69180882e7f183205c5b7a35fae4b3383804792e480ea0e623d

Observation e1a2eaa2-1aba-4f20-8b48-d2dcdff3366a · inbound

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception cites this paper.

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification

Reference 164

Resolution
unresolved
no resolver link, observed 2026-07-12T04:17:40.198357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T04:17:40.198357Z digest=sha256:16427ce2e867fdfad565beb2f11ae88be30bdfc6cfc168a9a4ec80829235955b