Pith. sign in

Paper Citation Record · LEDGER

ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding

As of 14 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 31 inbound Pith citation observations for arXiv:2501.05452.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.05452 v1

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T21:17:03.098852Z

measured 70 of 70 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 31 of 31 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:12:03.512804Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T19:20:07.268611Z

Reference resolution

39 of 39 outbound references displayed

  • verified exact0
  • verified fuzzy18
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 551b367e-c72b-4d5c-9396-fea8378b5a6a · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T21:17:02.953637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:17:02.953637Z digest=sha256:b49429e93332705430b4b0db7847e1d9be561afa16095163bdfe3afe89092bfb

Observation 890e7e32-0475-4032-b9ab-f5dd97917b89 · outbound

This paper cites Vip- llava: Making large multimodal models understand arbitrary visual prompts.

ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding Vip- llava: Making large multimodal models understand arbitrary visual prompts

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:17:03.502169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:17:02.958604Z digest=sha256:02d7d2d192359831a249857470b3bea681559baa480bab3ba0cb2d1af3c7a9c3

Observation ae235e4b-6838-4c9c-b6bc-a9ca53e6c85f · outbound

This paper cites Bigtable: A distributed storage system for structured data.

ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding Bigtable: A distributed storage system for structured data

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:17:03.492745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:17:02.962854Z digest=sha256:f52849d3b90d806816fb68e2d25436d1378f5ff5b8d3ae517427dfd43180851a

Observation 1bc0f498-95fd-4cfc-9191-e03ed9842b63 · outbound

This paper cites TabFact: A Large-scale Dataset for Table-based Fact Verification.

ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding TabFact: A Large-scale Dataset for Table-based Fact Verification

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T21:17:02.966682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:17:02.966682Z digest=sha256:eb2285131c3469064b5497b1b8c0f9a6be595284a58ab1d604cd93755bb77c4b

Observation f35a8280-5f57-40c7-be88-2dc5ef7ac549 · outbound

This paper cites HybridQA: A Dataset of Multi-Hop Question Answering over Tabular and Textual Data.

ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding HybridQA: A Dataset of Multi-Hop Question Answering over Tabular and Textual Data

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T21:17:02.970923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:17:02.970923Z digest=sha256:e60a486bda37cebf884e1f6a5e1cc78b257448e72726aa006def667b078ff590

Observation 129445eb-5a02-45ef-b9d8-25ca240451ea · outbound

This paper cites Visual chain- of-thought prompting for knowledge-based visual reasoning.

ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding Visual chain- of-thought prompting for knowledge-based visual reasoning

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:17:03.483174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:17:02.975049Z digest=sha256:4e2a8b7eda5dc9aba05e5916540fbc9007ab9213ec5ed0371dc7d8e717039045

Observation 752f8bec-c5f1-4de5-ac7e-268977014af5 · outbound

This paper cites Selective attention and the organization of visual information.

ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding Selective attention and the organization of visual information

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:17:03.473231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:17:02.979467Z digest=sha256:2ad930194ced4e2380118f44f37be14ea86a779d2a9dde7279ff6ef6dc9f148f

Observation fd405925-dcba-4488-a58c-c3c3c3f77105 · outbound

This paper cites BLINK: Multimodal Large Language Models Can See but Not Perceive.

ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T21:17:02.983278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:17:02.983278Z digest=sha256:54634f117b3c1cfdae7eaa2b3c45172049b8df1876f3c5531f4dace3eec8b005

Observation 48de0c73-444f-4ae3-9837-37c80ec70877 · outbound

This paper cites The cambridge structural database.

ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding The cambridge structural database

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:17:03.462479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:17:02.987743Z digest=sha256:cc8110e6b73574f2c3311ff5ae96bd52e70878575cb2a7914a8c4be044170d7d

Observation 12fe13e4-b780-485f-bd09-0cacd42b10c9 · outbound

This paper cites Visual program- ming: Compositional visual reasoning without training.

ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding Visual program- ming: Compositional visual reasoning without training

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:17:03.451440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:17:02.991201Z digest=sha256:04ed48092ee450fb681c3283e8c6be4e852e306b4b9e3a76a8f0c80b7a6e4474

Observation aba33c52-2cba-43ff-8195-ce4c1902d980 · outbound

This paper cites ChartLlama: A Multimodal LLM for Chart Understanding and Generation.

ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding ChartLlama: A Multimodal LLM for Chart Understanding and Generation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T21:17:02.994917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:17:02.994917Z digest=sha256:35e6ae3f85f678aef5123807525a44dc4bce9bf49d4099281be4139c09b5a8f5

Observation c03a0a0e-a2ae-40c3-b22c-1bdda52436ff · outbound

This paper cites LoRA: Low-rank adaptation of large language models.

ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding LoRA: Low-rank adaptation of large language models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:17:03.440185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:17:02.999001Z digest=sha256:93bc6bbfb8ffe926d78aa946c35e9feae25c40189adf33bba842876aa3332be1

Observation 6accc76a-4a80-4525-a1cb-aae5763a6e31 · outbound

This paper cites Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models.

ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T21:17:03.002283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:17:03.002283Z digest=sha256:222fb8c550f4f41a55d580b1b0be072a0eebe65d5c406d8830c9eb7dc04cc803

Observation 2e4d4111-8e2e-4203-b519-db631c3aa707 · outbound

This paper cites Selective attention.

ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding Selective attention

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:17:03.429560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:17:03.006282Z digest=sha256:f174461ada14fdc168fc28def52a0d115d8281ed680a3c39c8d6fc03dd30043c

Observation 5a1c74f1-a1ed-437f-8949-5ad9bc45fd2b · outbound

This paper cites Dvqa: Understanding data visualizations via question answering.

ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding Dvqa: Understanding data visualizations via question answering

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:17:03.419064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:17:03.009484Z digest=sha256:b42c9c9c5482723beb5f2c1a38e6f1e1e7210ce2e820b92bacbfe74f5baa77f8

Observation 10d939ae-e693-44ef-a778-5f7e649cadc1 · outbound

This paper cites TableVQA-Bench: A Visual Question Answering Benchmark on Multiple Table Domains.

ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding TableVQA-Bench: A Visual Question Answering Benchmark on Multiple Table Domains

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T21:17:03.012678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:17:03.012678Z digest=sha256:54a56202aa87884084918b21e8eaf3861baf6e88b7326b270f3f580fab42fed1

Observation fe7bad8c-76f4-4d13-b236-24af8e57085e · outbound

This paper cites Semantic-SAM: Segment and Recognize Anything at Any Granularity.

ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding Semantic-SAM: Segment and Recognize Anything at Any Granularity

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T21:17:03.016060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:17:03.016060Z digest=sha256:4c9f1a6ead4898f8c1b0d0290d7fb3622b0dd4a36f8c3f66f4997873fd0d7425

Observation fd6835a0-9c66-440e-b350-7582cd28631e · outbound

This paper cites MatCha: Enhancing Visual Language Pretraining with Math Reasoning and Chart Derendering.

ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding MatCha: Enhancing Visual Language Pretraining with Math Reasoning and Chart Derendering

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T21:17:03.019352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:17:03.019352Z digest=sha256:7413cf51edcb9b47502a10df9043c8a9b3d905ce6bdff7b50b3881cf9841253f

Observation 71447ef2-9335-42e7-9fa4-9212da114fbc · outbound

This paper cites Deplot: One- shot visual language reasoning by plot-to-table translation.

ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding Deplot: One- shot visual language reasoning by plot-to-table translation

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:17:03.407477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:17:03.023415Z digest=sha256:78af27c48711b2c217ef1cd66a11f3289eb1fe720cff80c090402cc174d3a975

Observation 4a436c11-afc5-4582-88d6-53512bb7049b · outbound

This paper cites Llava-next: Improved reason- ing, ocr, and world knowledge, 2024.

ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding Llava-next: Improved reason- ing, ocr, and world knowledge, 2024

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:17:03.398857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:17:03.026969Z digest=sha256:c13780989d27801de7718e3a0d5a326f8ba0543a44bdb770dfaf64ba4e60167d

Observation 22f212fe-bd87-48cc-bd6a-979246823546 · outbound

This paper cites LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents.

ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T21:17:03.030625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:17:03.030625Z digest=sha256:3ca8d431d11bac0b52bbbe0eaff4bde739628471870e5428ee2dababd7d3e193

Observation 36634454-de3a-4607-9329-e070e89430c8 · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T21:17:03.034408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:17:03.034408Z digest=sha256:76e4d824b2ac5a43a0f964e30e8e77d51f0a788b5e67fe05145a905a97ac6510

Observation 8d041773-43db-41e4-a747-21eae5fc0896 · outbound

This paper cites Ocrbench: On the hidden mystery of ocr in large multimodal models, 2024.

ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding Ocrbench: On the hidden mystery of ocr in large multimodal models, 2024

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:17:03.389605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:17:03.038132Z digest=sha256:c9a2ac83d0610d5601af10750b4191104ab229c9c2bda8f7b3be6c37336a05f5

Observation b31274e7-bc60-47e9-9fb3-865ffedb965a · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T21:17:03.041529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:17:03.041529Z digest=sha256:a0e8f096c51399ea14ac18afa01fe9d1ca3bb4965582a0964a5f5e79653ec2f6

Observation a5664581-51f7-4884-9d9c-5a0593773aab · outbound

This paper cites UniChart: A Universal Vision-language Pretrained Model for Chart Comprehension and Reasoning.

ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding UniChart: A Universal Vision-language Pretrained Model for Chart Comprehension and Reasoning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T21:17:03.045109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:17:03.045109Z digest=sha256:dbf15f8f8fd6620ff70e17d00c572822a56bf64d4692ba8edfece915a89089f3

Observation 434f9a2e-c558-4962-8fc9-049fdb0a98f2 · outbound

This paper cites Gpt-4 technical report, 2023.

ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding Gpt-4 technical report, 2023

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T21:17:03.052101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:17:03.052101Z digest=sha256:654f256300dce6eaa72f9f852a396f6fcad02630563daae17bcb179501c5be04

Observation eaab08c5-a119-467b-be8f-8388fc6fce3a · outbound

This paper cites Compositional Semantic Parsing on Semi-Structured Tables.

ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding Compositional Semantic Parsing on Semi-Structured Tables

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T21:17:03.055546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:17:03.055546Z digest=sha256:7e38e93b3942160297e82f8d5f2b72257ce46688a365c9d69d70bbac0022249d

Observation 228911d1-b60e-4c00-8ea7-d7b70e3c56c7 · outbound

This paper cites Space and selective attention.

ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding Space and selective attention

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:17:03.372547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:17:03.059260Z digest=sha256:fd6a30876dca2fb85d16e3c3abc28d2d08477a76ee08650016e4807dedad5359

Observation 0b96053f-5827-4957-a6b3-d76a212ded52 · outbound

This paper cites Visual cot: Advancing multi-modal language models with a comprehen- sive dataset and benchmark for chain-of-thought reasoning.

ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding Visual cot: Advancing multi-modal language models with a comprehen- sive dataset and benchmark for chain-of-thought reasoning

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:17:03.361969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:17:03.062714Z digest=sha256:98011183797c46b3a7c8bcf04dc498bbb14ae1a6999293951dc53ee810f98357

Observation 2b3046a8-7b7c-4207-a61f-eaf32d4fbaf7 · outbound

This paper cites What does clip know about a red circle? vi- sual prompt engineering for vlms.

ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding What does clip know about a red circle? vi- sual prompt engineering for vlms

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:17:03.351996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:17:03.066204Z digest=sha256:481c3a8c8774ac50bbbae284114f00f5d890f30958bdb99fa353b117a45cd273

Observation 96b638de-b646-417e-a14a-a74937bfeef9 · outbound

This paper cites Vipergpt: Visual inference via python execution for reasoning.

ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding Vipergpt: Visual inference via python execution for reasoning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:17:03.341854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:17:03.069392Z digest=sha256:56463ea361bc44ae90e463a8f30989f11a43f205f185c505f8738b0176f36d2b

Observation c1588615-5afe-4728-8fb3-8b12e9b238c5 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding Gemini: A Family of Highly Capable Multimodal Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T21:17:03.072991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:17:03.072991Z digest=sha256:4d75b35732b9a7bb0603f5afceda03ceb709ed5c051c0551c7c8aae284df8aab

Observation 3db01fd5-aa86-43e9-b497-7b07c4661d31 · outbound

This paper cites CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs.

ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T21:17:03.076817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:17:03.076817Z digest=sha256:d1d23caa67c84265eee079a8023fa703eb080a2f9302a862c7987da87552ff3f

Observation 7019453c-5e3c-4b32-8505-c6115d5c482c · outbound

This paper cites Chain-of- thought prompting elicits reasoning in large language models.

ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding Chain-of- thought prompting elicits reasoning in large language models

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:17:03.331103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:17:03.080711Z digest=sha256:24242624264ab13c5610f5b264d1daed215f0e1d6e3f604c846c854b2a814240

Observation fc8289f2-9294-4a24-b9b1-bbdda74fc2e5 · outbound

This paper cites List Items One by One: A New Data Source and Learning Paradigm for Multimodal LLMs.

ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding List Items One by One: A New Data Source and Learning Paradigm for Multimodal LLMs

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T21:17:03.083937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:17:03.083937Z digest=sha256:fa7beac4c2077a4e50be51d181369266bca343b0f3944e8f93df56667bd20a1e

Observation f8ca0e2e-60f4-44c2-9e1a-ab94a3e89909 · outbound

This paper cites Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V.

ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T21:17:03.087965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:17:03.087965Z digest=sha256:26d7836bddafe03c2f7cc47912568ae0c5a4610c5705d50f2a8a3d72d134bd38

Observation 8aad3ddb-f5d3-4ee6-bb04-eeae1003bac4 · outbound

This paper cites Depth anything: Unleashing the power of large-scale unlabeled data.

ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding Depth anything: Unleashing the power of large-scale unlabeled data

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:17:03.320375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:17:03.091608Z digest=sha256:d34309893d28ac611154902b8ed092875b6046d80033392f3eebb757e3ecdb00

Observation 8e309efd-9dba-4b97-92e4-195337ba7a2c · outbound

This paper cites Multimodal Chain-of-Thought Reasoning in Language Models.

ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding Multimodal Chain-of-Thought Reasoning in Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T21:17:03.094988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:17:03.094988Z digest=sha256:5823f0152aecee5a2fc6b26035edb25806b34b9e32970b4be8fd62533c5f2efd

Observation a46cecd2-4a68-4e7f-983e-da1af320d80a · outbound

This paper cites TAT-QA: A Question Answering Benchmark on a Hybrid of Tabular and Textual Content in Finance.

ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding TAT-QA: A Question Answering Benchmark on a Hybrid of Tabular and Textual Content in Finance

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T21:17:03.098852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:17:03.098852Z digest=sha256:76ac17280e7d3f9f02067a9f242abd66e5d9a7ba550bea1a2edf3e103ef58435

Pith citing papers

Observation e425596d-f5a0-4bab-adea-2f6a8358a6e1 · inbound

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey cites this paper.

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding

Reference 123

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:18:53.182630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T17:18:52.996467Z digest=sha256:0579a9b5fceffb527f8fbb270e966cd90edecd66bde4ac08d756c6d70673c8a8

Observation 77c84b54-142c-4262-9e70-9a820e655541 · inbound

Point-RFT: Improving Multimodal Reasoning with Visually Grounded Reinforcement Finetuning cites this paper.

Point-RFT: Improving Multimodal Reasoning with Visually Grounded Reinforcement Finetuning ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:12:03.512804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:12:03.512804Z digest=sha256:c074ca54dfe85ed76594ec53b4803ce452bd1864d376f7f3e3f6b85d16c3b787

Observation 53ea2b8a-629e-4183-9890-b255564e406e · inbound

Grounded Reinforcement Learning for Visual Reasoning cites this paper.

Grounded Reinforcement Learning for Visual Reasoning ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-22T01:05:52.119661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-22T01:05:18.801388Z digest=sha256:ff91f4cf1b26b0b2f0522175f51b2c0ce5da0689960ae7d3d87dd9a0c05ad6eb

Observation 22bcc49e-abe1-4239-a509-8b6c2363d5ac · inbound

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs cites this paper.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:07.980626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:07.980626Z digest=sha256:ba4228bdbf1a36c4182adca94c9e709b31eb1e72f35ad626a2bb286d19780784

Observation 9bc20143-f1e1-43d1-bd7f-deb87faafb28 · inbound

TableMoE: Neuro-Symbolic Routing for Structured Expert Reasoning in Multimodal Table Understanding cites this paper.

TableMoE: Neuro-Symbolic Routing for Structured Expert Reasoning in Multimodal Table Understanding ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T22:33:12.123031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:33:12.123031Z digest=sha256:88116377211e126284aa373b4c26fa9b69e1a7bd8bda92cf80e6755d8525851c

Observation be5a5374-0176-4028-a3ab-e3510564f0d7 · inbound

Learning Only with Images: Visual Reinforcement Learning with Reasoning, Rendering, and Visual Feedback cites this paper.

Learning Only with Images: Visual Reinforcement Learning with Reasoning, Rendering, and Visual Feedback ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T13:23:30.018055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T13:23:30.018055Z digest=sha256:5c608ddef29e1d734b897d17a43c069b8792d0c972b953ebb3f9c75168e208de

Observation cbc72a80-308f-4e3c-b271-e62f4bef7e53 · inbound

Empowering Nanoscale Connectivity through Molecular Communication: A Case Study of Virus Infection cites this paper.

Empowering Nanoscale Connectivity through Molecular Communication: A Case Study of Virus Infection ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T00:03:27.768763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:03:27.768763Z digest=sha256:54787d1a7f3ba4d2eab1d64a82c9bf1890861c7a246a8417afd97fcb11d73439

Observation ad25cc20-36be-4892-b653-c9ffa711c9c5 · inbound

Empowering Multimodal LLMs with External Tools: A Comprehensive Survey cites this paper.

Empowering Multimodal LLMs with External Tools: A Comprehensive Survey ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding

Reference 130

Resolution
unresolved
no resolver link, observed 2026-08-05T20:28:54.709175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:28:54.709175Z digest=sha256:71da0e046202177072b5720e26f03504507a7a6ec43a251f2c0c5fe0b85d46cb

Observation 9e9a0b5b-0735-4ee7-9491-a579f047b008 · inbound

Visual-TableQA: Open-Domain Benchmark for Reasoning over Table Images cites this paper.

Visual-TableQA: Open-Domain Benchmark for Reasoning over Table Images ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-18T17:42:47.613443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-18T17:37:31.837022Z digest=sha256:c8134c90510278d186aaeef3711ff8ea93a72513d653b68faab70b0ebbe5cf4d

Observation 998508ec-f91f-4376-ac1b-79235175a383 · inbound

DeFacto: Counterfactual Thinking with Images for Enforcing Evidence-Grounded and Faithful Reasoning cites this paper.

DeFacto: Counterfactual Thinking with Images for Enforcing Evidence-Grounded and Faithful Reasoning ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-21T22:35:42.672254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T22:35:36.136639Z digest=sha256:ae53909e970171ab64eb98a2834723670a78435cecd1936b3547612b29d558f8

Observation f37e4270-47d3-4738-ae3b-4ba2172e483d · inbound

DeFacto: Counterfactual Thinking with Images for Enforcing Evidence-Grounded and Faithful Reasoning cites this paper.

DeFacto: Counterfactual Thinking with Images for Enforcing Evidence-Grounded and Faithful Reasoning ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-22T13:36:35.981987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-22T13:36:06.229352Z digest=sha256:489ba4a126e4d728aa729028d7e25a6ee623ddd352a5b6b2e3fa8f212c3b5917

Observation 8a1e917d-b360-4c8a-abcc-d04eaf46d0a8 · inbound

Latent Visual Reasoning cites this paper.

Latent Visual Reasoning ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:41:30.372775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T18:41:30.307521Z digest=sha256:43f09154342b6d7e91a9e576620f9eb3e275e344ddcd2167a426edda3d183fa0

Observation 219fe1a3-7932-4d13-b87d-31de3c9157f6 · inbound

Training Multi-Image Vision Agents via End2End Reinforcement Learning cites this paper.

Training Multi-Image Vision Agents via End2End Reinforcement Learning ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:01:24.319646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-17T00:59:28.618477Z digest=sha256:bc0c6e539515db02378ea051173b956a749cb0572b1f79e084bce03c836eb474

Observation 03055f22-2501-4018-be04-b3f9247b503c · inbound

Reasoning Within the Mind: Dynamic Multimodal Interleaving in Latent Space cites this paper.

Reasoning Within the Mind: Dynamic Multimodal Interleaving in Latent Space ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-16T23:03:38.937182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-16T23:02:28.588225Z digest=sha256:636b0f5f4498f86052ba4d8b474b7f96328f8ed01a63c1155c0ec86d9c345647

Observation b0158267-52e2-4167-bac1-f521860f18d4 · inbound

Fully Spiking Neural Networks with Target Awareness for Energy-Efficient UAV Tracking cites this paper.

Fully Spiking Neural Networks with Target Awareness for Energy-Efficient UAV Tracking ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-13T16:55:20.099628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:55:20.099628Z digest=sha256:2e3eb25164dfcd3f057d1d60bc51796ca28517dd146f9147141b28a241767f9f

Observation 3904687a-9598-41e6-9b8f-70d966161ec2 · inbound

Learning to Focus and Precise Cropping: A Reinforcement Learning Framework with Information Gaps and Grounding Loss for MLLMs cites this paper.

Learning to Focus and Precise Cropping: A Reinforcement Learning Framework with Information Gaps and Grounding Loss for MLLMs ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:38:00.966585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-14T21:35:12.859669Z digest=sha256:4ca7c9bb241848d69249802051c0f378ecbce4acc1979fe26f212acae3e5dc46

Observation 1ee1dbbd-ac54-421e-b3d6-e99745b3c56d · inbound

TableVision: A Large-Scale Benchmark for Spatially Grounded Reasoning over Complex Hierarchical Tables cites this paper.

TableVision: A Large-Scale Benchmark for Spatially Grounded Reasoning over Complex Hierarchical Tables ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-13T17:23:02.410245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-13T17:20:28.036531Z digest=sha256:4218268fad60001fcd01eda28e3b93170ca4d77c8cc11030cb06ba0c48147b9b

Observation 117371bc-87f1-486e-a6d4-b5027cdfda2b · inbound

Decompose, Look, and Reason: Reinforced Latent Reasoning for VLMs cites this paper.

Decompose, Look, and Reason: Reinforced Latent Reasoning for VLMs ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T07:21:00.745050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T17:12:33.231488Z digest=sha256:f4a279c3ccdc2c13ec7f37c56aef19c20c442093a50ef4b7e2ba4d85a8c3ec66

Observation 0ef29106-7f0d-4c59-9e7a-af3a3bb4a19c · inbound

MedLVR: Latent Visual Reasoning for Reliable Medical Visual Question Answering cites this paper.

MedLVR: Latent Visual Reasoning for Reliable Medical Visual Question Answering ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:35:59.528033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:03:14.408704Z digest=sha256:6fef767dbdfe52d6fd0f51ffe4ef06199aca7dc6d3f147da6b273545fe23b8e0

Observation 4c4725e2-eccb-4442-92bb-9d091f470937 · inbound

MedLVR: Latent Visual Reasoning for Reliable Medical Visual Question Answering cites this paper.

MedLVR: Latent Visual Reasoning for Reliable Medical Visual Question Answering ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T16:33:55.457196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:33:55.457196Z digest=sha256:506b1a180e95d9f8ca2d2280a0259b21a98a19c9cea7e99a1e8644ac768f84d9

Observation f704a458-45c9-43e1-8da5-c1b1fe9776f4 · inbound

Meta-CoT: Enhancing Granularity and Generalization in Image Editing cites this paper.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:46:12.382365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:63a9c425ae34b5c48e69add0132a89c7491c1f42b09064d347de60a68ea4eeb8

Observation dff89b1a-891e-48ca-b7e3-46cb9a7a0c52 · inbound

Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation cites this paper.

Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:58:13.403120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T10:58:01.621488Z digest=sha256:c6eab5ed137e13272b1e9a0ac556ff7ab24d195cdaef44086a91bf8010169d8e

Observation f1b918f8-0771-4b8f-aeca-509014fe7a7c · inbound

Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation cites this paper.

Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:35:00.655979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T18:28:21.605646Z digest=sha256:29345ff5598d02f330b0d3c3f80ec02418db1f3ef52a298b952d84e6ff3a51c0

Observation c0f0b526-4d84-449c-935b-dc8265e2006d · inbound

STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models cites this paper.

STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-06-29T23:14:02.204792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-29T23:04:21.463842Z digest=sha256:708198e9863a4e4d54972ee9cd831a85f3150e4bcd44303bd8c14c888bc7ed27

Observation d66e4bf6-435c-4712-bc60-d69a1260106f · inbound

DeepLatent: Think with Images via Parallel Latent Visual Reasoning cites this paper.

DeepLatent: Think with Images via Parallel Latent Visual Reasoning ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:52:36.103666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-28T18:44:39.545911Z digest=sha256:fa6e084227fe5f882f1e4fd9a84c9c5b9b115d119ddddaf451fbf5fd3c2648c7

Observation 34188065-5041-4a10-9a42-0215a8c9592f · inbound

Dive into the Scene: Breaking the Perceptual Bottleneck in Vision-Language Decision Making via Focus Plan Generation cites this paper.

Dive into the Scene: Breaking the Perceptual Bottleneck in Vision-Language Decision Making via Focus Plan Generation ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:56:29.290692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-28T10:28:41.952330Z digest=sha256:95a15ae351e7fbd50438f5cb84d1783d8e39250aabd02d942af9fe3c0a10634e

Observation 39696f71-5678-4cc6-9ead-75adaa80cd11 · inbound

Latent Visual States for Efficient Multimodal Reasoning cites this paper.

Latent Visual States for Efficient Multimodal Reasoning ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:29:57.273422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T00:38:11.619574Z digest=sha256:d0266af9ff2fc1bea2c2823301a6869fbc84e51fd1439415e59c49dc6e0f9baf

Observation ad5ce396-1ef0-4591-b632-ed4e207cee02 · inbound

V-Zero: Answer-Label-Free On-Policy Distillation with Contrastive Evidence Gating for Fine-Grained Visual Reasoning cites this paper.

V-Zero: Answer-Label-Free On-Policy Distillation with Contrastive Evidence Gating for Fine-Grained Visual Reasoning ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-07-04T19:20:07.270404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-25T21:23:10.051805Z digest=sha256:cb4e0222bd7ae970b13ad5a54efb853d58dfd6c416c3d1e49d8dbcadc338b3f5

Observation caddb4f2-0306-4499-9eee-e6e140c10e32 · inbound

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models cites this paper.

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding

Reference 152

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:55.247997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T01:50:54.242508Z digest=sha256:fcfd4148d272233b21afcb7610c5ede5a409aea60cdfb4a36a19f94543d5b50a

Observation 8fa4e668-6590-46fe-9b7c-0c84de0cd499 · inbound

FaithEyes: Towards Faithful Tool Use via Multi-Agent Process-Image Verification cites this paper.

FaithEyes: Towards Faithful Tool Use via Multi-Agent Process-Image Verification ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-31T14:08:11.175717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T14:08:11.175717Z digest=sha256:443f3d6f02e472c710dcc703448af05cd60660b9e1c39de5690f6624e34c45b3

Observation 84d35572-6fcf-4aa5-89c6-796679fd1500 · inbound

VAD: Attributing Visual Evidence for Target Reconstruction in Multimodal On-Policy Distillation cites this paper.

VAD: Attributing Visual Evidence for Target Reconstruction in Multimodal On-Policy Distillation ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding

Reference 59

Resolution
unresolved
no resolver link, observed 2026-07-31T02:56:56.445550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T02:56:56.445550Z digest=sha256:3d686326440cdac7677cf587eb039d9a88f829abd268798326755e1a42d62229