Pith. sign in

Paper Citation Record · LEDGER

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs

As of 6 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 1 inbound Pith citation observation for arXiv:2605.02735.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.02735 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-09T15:47:49.982564Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-31T16:13:17.573307Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

47 of 47 outbound references displayed

  • verified exact14
  • verified fuzzy31
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3a2c6d5f-8b2d-4508-80e2-23be1da2faab · outbound

This paper cites Qwen2.5-VL Technical Report.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs Qwen2.5-VL Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-11T16:36:09.579251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:f6fbde3b66a4b95b65ee7d0c572b65e6fea1fc00d87802c183c2f8e3b04fbfb3

Observation cc5c100b-b3ca-4721-848f-e2f4db540260 · outbound

This paper cites Qwen2.5-vl technical report.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs Qwen2.5-vl technical report

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:25.991467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:00afda29b3a49a677284c156608d050060bd4f04f05539564bf6427426009d82

Observation 66037d08-075a-43d3-afba-1f975d3c4ec1 · outbound

This paper cites UniVG-R1: Reasoning Guided Universal Visual Grounding with Reinforcement Learning.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs UniVG-R1: Reasoning Guided Universal Visual Grounding with Reinforcement Learning

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:36:09.631193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:1d46791297fe041ea8a201ab3f45148c691172d10bbc164a8f0c2b9230e050bf

Observation 519fb987-1521-4309-885f-e59d3d4a5cf7 · outbound

This paper cites Sft or rl? an early investigation into training r1-like reasoning large vision-language models.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs Sft or rl? an early investigation into training r1-like reasoning large vision-language models

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:25.978226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:a683f2e8a64ced9fd096121389757c57758943cb0b64fdab0343f99c912e1977

Observation ba3b1d82-cfba-413d-aa70-88101c9f74d8 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-12T19:41:44.612219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:a2e3db9ffc8960f7af8f0f52aa6dd372a92e582e472db97c2858e45c6e1a555a

Observation 529ad48e-7849-43da-8eb9-57d01a694020 · outbound

This paper cites Think with 3d: Geometric imagination grounded spatial reasoning from limited views.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs Think with 3d: Geometric imagination grounded spatial reasoning from limited views

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:25.917571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:a324b961f8a078794923945296cddf9e3cc1fe8cb2957610271cdb30d26ba72b

Observation a6291f43-4fd3-49e4-ab53-0d6edfb426e7 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:25.984613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:27eb567f041dfceeacd4a12a1d4916aff7a4124de2db590b2a1c4de267c130a1

Observation ed1cadea-6207-4b5c-8aa3-40ecee404002 · outbound

This paper cites BLINK: Multimodal Large Language Models Can See but Not Perceive.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:18:15.850579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:6b01377d3762abf39520ec8f8449dd370d52139a9bb214b9b6c8bb6ae2557b63

Observation 8f4b7016-b63a-4096-99f4-7ea0ca3e514f · outbound

This paper cites Refocus: Visual editing as a chain of thought for structured image understanding.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs Refocus: Visual editing as a chain of thought for structured image understanding

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:25.998588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:76d243ed47dde3aeeb994358762b4104cf92df96a579e7edce47ecc6bf2883c0

Observation c7ec3e4d-a5c5-44ff-bfb8-fa94413f0e62 · outbound

This paper cites Omni-MATH: A universal olympiad level mathematic benchmark for large language models.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs Omni-MATH: A universal olympiad level mathematic benchmark for large language models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:25.949019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:19dacf305996ad680847a2472f79572da20de9a0c508f8d67cc030151d3c03ed

Observation 8c2ee88f-a43e-4f6a-9657-550f5c015b93 · outbound

This paper cites Interleaved-modal chain-of-thought.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs Interleaved-modal chain-of-thought

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:25.926002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:99baee630e49c259b93bbec694ef63a7e044f53a993697568f2877b3a537104c

Observation 8733b059-e4b1-4135-a7a5-2a44b73fe9cc · outbound

This paper cites Hallusion- bench: An advanced diagnostic suite for entangled language hallucination & visual illusion in large vision-language models.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs Hallusion- bench: An advanced diagnostic suite for entangled language hallucination & visual illusion in large vision-language models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:25.929162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:21787185c1f5fbd7b898cc27583876a9e3fd0f1a2a8da5569c4099929599bd44

Observation 2e6e0cbf-20a7-46d9-8a0a-bb77addb9cae · outbound

This paper cites Training large language models to reason in a continuous latent space.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs Training large language models to reason in a continuous latent space

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:25.988336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:0804985093372caec2e152511f3d30ed28bbf0a1ac5e85009ed426c02af4e55d

Observation 7334ff85-fe9c-46cf-b2cb-208cf32169fb · outbound

This paper cites Visual sketchpad: Sketching as a visual chain of thought for multimodal language models.Advances in Neural Information Processing Systems, 37:139348–139379.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs Visual sketchpad: Sketching as a visual chain of thought for multimodal language models.Advances in Neural Information Processing Systems, 37:139348–139379

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:25.900376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:9ca3e79f7209ab0275829d10875b4c7d9d988270823cf01d78d883e5e3e63ec2

Observation f6492854-552e-441c-886a-1a135ae0c3e4 · outbound

This paper cites Latent visual reasoning.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs Latent visual reasoning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:25.981224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:2e8359b12d0917c315ecadf31d85ba12a91b34d22f959c1b9847174b08a18ff8

Observation 89b6da3e-4c93-40b6-9ea1-d17d776ad77d · outbound

This paper cites Reasoning Within the Mind: Dynamic Multimodal Interleaving in Latent Space.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs Reasoning Within the Mind: Dynamic Multimodal Interleaving in Latent Space

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-11T16:36:09.585493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:ed2de25052491f9089c282d0fc60afefe3f9c8549d686b41a130d3e3ab061727

Observation c32a2eec-c6af-4597-af8e-4d48820cc54a · outbound

This paper cites Deliberation in latent space via differentiable cache augmentation.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs Deliberation in latent space via differentiable cache augmentation

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:25.888577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:0c3b194b8cb534235dd1b178e64fa23054c61aebb291b4bcb0c4841860a56127

Observation 68ca0288-8ca9-4a17-af42-fc0c5bc3dc2d · outbound

This paper cites Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:25.880107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:713e12d1cc2ec839f5862df055e02f957c644b8050d33d06e733af7816559427

Observation b958fd74-c7a4-425d-8d34-807959c82865 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:25.974192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:db9ee9e5d4f637e4505a086d5015320992203e5d7ba635d0fd97c6ee901aa413

Observation 68d1ffcb-7056-4b6f-a2bb-f2ef99f2c862 · outbound

This paper cites A survey on latent reasoning.arxiv.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs A survey on latent reasoning.arxiv

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:25.995586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:e0e103a4e1e24279d2f83cf7e026176a83f6008db8ec194596e9598a78c6648e

Observation 4c330125-d170-4c8c-b72b-e0ff100b5b08 · outbound

This paper cites LaRe: Latent Refocusing for Multimodal Reasoning.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs LaRe: Latent Refocusing for Multimodal Reasoning

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-27T02:05:07.580934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:c497e2bcc76953bd83825c6f827d4f97ecc42aed7a0f9b4e012f15b67d9a4b77

Observation b4a50e53-5de6-4bc5-885f-d95cabce7c8c · outbound

This paper cites Compositional chain-of- thought prompting for large multimodal models.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs Compositional chain-of- thought prompting for large multimodal models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:25.884686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:6c29324bec621dff54d23eeae411d3749dd285d42378aa195776f28b760fecfc

Observation 29a6f4c4-47ec-4893-b174-e1fb73bfccf8 · outbound

This paper cites Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-08-06T02:01:32.731977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:16a7ceb7c0f285c581a6c7058c231f30d04d68bbc5c1520942ae1863ae26d71c

Observation c279abb1-582c-4ea8-a1fa-71ac6c0e3188 · outbound

This paper cites an unresolved cited work.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-05-26T01:06:25.921888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:a3ecb6e6198626105f9532fe1798eb29ac5df60a6cc58ad2b1ee5088e305ffbe

Observation 7f7ffa2f-596d-483c-9e4e-3a37b04ef130 · outbound

This paper cites Codi: Com- pressing chain-of-thought into continuous space via self-distillation.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs Codi: Com- pressing chain-of-thought into continuous space via self-distillation

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:25.913984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:1a27f30cd6342290a21850681a7235649d742d7f4b1ded077be879622cd18f25

Observation f4718f76-c2a1-4d4d-9cbb-71c311ed7cce · outbound

This paper cites Swireasoning: Switch-thinking in latent and explicit for pareto-superior reasoning LLMs.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs Swireasoning: Switch-thinking in latent and explicit for pareto-superior reasoning LLMs

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:25.892022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:b67ec4aea8a5397c372d8e7da0796678ddc3a903ca2ddbba7315c55ef34a7120

Observation 239432c9-d6be-4f55-b0e5-7af43b159614 · outbound

This paper cites Think silently, think fast: Dynamic latent compression of llm reasoning chains.Advances in Neural Information Processing Systems.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs Think silently, think fast: Dynamic latent compression of llm reasoning chains.Advances in Neural Information Processing Systems

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:25.952508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:f5334041fa262048b97e8e4b82c93ed55406713aa9bc3b973462092df394cc3d

Observation df45b34b-7cbc-46f8-9c89-5636f31c0502 · outbound

This paper cites Visual position prompt for mllm based visual grounding.IEEE Transactions on Multimedia.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs Visual position prompt for mllm based visual grounding.IEEE Transactions on Multimedia

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:25.942767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:321a95ede985799282b59f04a55093f6f0641229a2af142cdf119a955d8218ac

Observation eeb952a7-f6e8-4310-a861-0b4fbc97e246 · outbound

This paper cites Eyes wide shut? exploring the visual shortcomings of multimodal llms.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs Eyes wide shut? exploring the visual shortcomings of multimodal llms

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:25.896455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:4e61c313894dbbb55f007298876c8ff9317da41dd14125cd53b080730e8d5ef5

Observation 5283bfd8-3b5b-4718-8c67-41d8414ae3e9 · outbound

This paper cites MLLM can see? dynamic correction decoding for hallucination mitigation.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs MLLM can see? dynamic correction decoding for hallucination mitigation

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:25.933325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:98d1bbcaae17f35a1285d21810ad9e94d0a464f60c683b6d1f903ca16b1ac780

Observation f569092e-9845-4108-8277-7acd9e655a91 · outbound

This paper cites Monet: Reasoning in latent visual space beyond images and language.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs Monet: Reasoning in latent visual space beyond images and language

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:36:09.653976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:1dd0ef82de9cc0fc2466a0c38caa0a1eced94f14a1e6d2e296b008b9cf956934

Observation 12eb500d-62b4-4d9f-aa2d-6fbc85357130 · outbound

This paper cites Image Tokens Matter: Mitigating Hallucination in Discrete Tokenizer-based Large Vision-Language Models via Latent Editing.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs Image Tokens Matter: Mitigating Hallucination in Discrete Tokenizer-based Large Vision-Language Models via Latent Editing

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:36:09.558238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:5e901300a2d7ae5855e0d7b0b0d7c7eb906c1bcf930066facbe0475c39bb10d7

Observation 4cda328c-e917-4ec3-b5fa-2aeecf482a0d · outbound

This paper cites Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:18:53.866178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:ee52c12436c4aaa2c772c02e1a34736cf7921bff9ebb2f2df7ee48c4166952b5

Observation 85060b44-8a35-4beb-a7e4-89efad4d3055 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs Chain-of-thought prompting elicits reasoning in large language models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:25.960259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:552c9e71e734ffd9f79d1ebd319e9979f3a7b9356da9703ed37f32e28ac67cb8

Observation eafd2daa-e09f-407c-8b0e-5ec8fd47dbd8 · outbound

This paper cites Deepscientist: Advancing frontier-pushing scientific findings progressively.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs Deepscientist: Advancing frontier-pushing scientific findings progressively

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:25.963118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:22fff63e1de9c30e499115e9d080c607e41ab94a8e60538acfa997e86c6bd30a

Observation f4a84c9d-0c51-4dc0-baac-ab8e51beee06 · outbound

This paper cites Mind’s eye of llms: visualization-of-thought elicits spatial reasoning in large language models.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs Mind’s eye of llms: visualization-of-thought elicits spatial reasoning in large language models

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:25.936641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:242db02a2406ae59d077d4146e2f31b6c9df6356db99b9e001da7f0ee7b55be1

Observation d98ebe88-1f9a-4510-ba6b-49e297132419 · outbound

This paper cites Mini-omni-reasoner: Token-level thinking-in-speaking in large speech models.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs Mini-omni-reasoner: Token-level thinking-in-speaking in large speech models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:25.906496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:3064111dc87f26d502ebba0c309f261c8540fc393fa55818ee8d156d45b1fdad

Observation 464b87df-fffc-4add-9810-35f6daf5ca5f · outbound

This paper cites Thinking in uncertainty: Mitigating hallucinations in mlrms with latent entropy-aware decoding.arXiv preprint arXiv:2603.13366.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs Thinking in uncertainty: Mitigating hallucinations in mlrms with latent entropy-aware decoding.arXiv preprint arXiv:2603.13366

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:36:09.643792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:ab578e5b4a208ee79be9471986b243b10e47e34c9d63e6869e5e0c4b055e381c

Observation d8d1f104-679d-44e3-a497-29f2b09baa2e · outbound

This paper cites R1-onevision: Advancing generalized multimodal reasoning through cross-modal formalization.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs R1-onevision: Advancing generalized multimodal reasoning through cross-modal formalization

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:25.945927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:1e310761e3b4a309ba124df435636cd0b49cf0a350078eb4a5802c4b91e162ec

Observation 9a002af4-fc4d-42c0-aece-bc49f199e477 · outbound

This paper cites Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:36:09.544273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:67eb2d53ee6bfec2ecb8722c40efa3dd0cdad5a8e2296f083c308915cb862e54

Observation 6aa3f6b0-6b4b-4750-8c29-d17ec79da595 · outbound

This paper cites Diffusion of thought: Chain-of-thought reasoning in diffusion language models.Advances in Neural Information Processing Systems, 37:105345–105374.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs Diffusion of thought: Chain-of-thought reasoning in diffusion language models.Advances in Neural Information Processing Systems, 37:105345–105374

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:25.968119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:2a5686b9e21e5a111c309d26335608f0b99a0313e46366ba17392f77f1c36e2f

Observation a8175318-fcaa-4440-94db-57ad8b7afcd5 · outbound

This paper cites A survey on multimodal large language models.National Science Review, 11(12):nwae403.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs A survey on multimodal large language models.National Science Review, 11(12):nwae403

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:26.002163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:47b440ef8de9dad4ec1eb6013cd5c1e9a1aa1abda00ded6e40bc2456899c4c4d

Observation 15e7bcb6-1472-4bf6-9258-728eb6b02d07 · outbound

This paper cites Mm-cot:a benchmark for probing vi- sual chain-of-thought reasoning in multimodal models.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs Mm-cot:a benchmark for probing vi- sual chain-of-thought reasoning in multimodal models

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:36:09.609725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:146528abd196d40e81736bcc568e16c576951f72129199cf96d4c40bf1ff11e5

Observation 41d65bfa-371a-4fec-a898-e8b6bd0e72bb · outbound

This paper cites Multi- modal chain-of-thought reasoning in language models.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs Multi- modal chain-of-thought reasoning in language models

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:25.910229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:2c150029d3d8ee18b621077e5d6beb48c8217abd648d6d27c87025266d713753

Observation 19905017-6415-4da5-86b3-dd82f0ef7ca5 · outbound

This paper cites Promptcot: Synthesizing olympiad- level problems for mathematical reasoning in large language models.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs Promptcot: Synthesizing olympiad- level problems for mathematical reasoning in large language models

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:25.955941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:29f1552b126e956fa9013c35ad2e19a6245a24320739f6cd3d08b5003d68ae1f

Observation 82f72304-2dd1-4caf-b824-de36c125a1ef · outbound

This paper cites Image-of-Thought Prompting for Visual Reasoning Refinement in Multimodal Large Language Models.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs Image-of-Thought Prompting for Visual Reasoning Refinement in Multimodal Large Language Models

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:36:09.568637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:485536abb2a0c1e26a1d4e9504b448272a8409a7135f4864a48cf3276c75facf

Observation 6c795bb6-3c09-41eb-bc64-b841891ac706 · outbound

This paper cites Intern-s1-pro: Scientific multimodal foundation model at trillion scale.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs Intern-s1-pro: Scientific multimodal foundation model at trillion scale

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:36:09.664141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:02f39c5adb8bd772236d6b3b28acf0b911a95fc4339cf47fa50e3000b3fc97ed

Pith citing papers

Observation 577b9ca7-52c3-4b31-a8f4-fca40918c88f · inbound

OPLD: On-Policy Latent Distillation for Multimodal Reasoning cites this paper.

OPLD: On-Policy Latent Distillation for Multimodal Reasoning Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-31T16:13:17.573307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T16:13:17.573307Z digest=sha256:439b8eaead66a2019c3585ebaf112bdc40f8f7b046c3d7049c042a979cf43f7c