Pith. sign in

Paper Citation Record · LEDGER

Divide, Conquer and Combine: A Training-Free Framework for High-Resolution Image Perception in Multimodal Large Language Models

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 13 inbound Pith citation observations for arXiv:2408.15556.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2408.15556 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:40:40.270764Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T11:46:55.941528Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 498932a9-8030-4f1a-b15d-e3f277ec51ff · inbound

Adaptive Chain-of-Focus Reasoning via Dynamic Visual Search and Zooming for Efficient VLMs cites this paper.

Adaptive Chain-of-Focus Reasoning via Dynamic Visual Search and Zooming for Efficient VLMs Divide, Conquer and Combine: A Training-Free Framework for High-Resolution Image Perception in Multimodal Large Language Models

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:35:13.196401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T05:35:13.118221Z digest=sha256:0f5a47dcbd54ad365f9d25477673ca0d71ccef04b7bcf8adde1306de90fb2166

Observation fd65673c-bd1a-4940-8b85-dd4bc318fbda · inbound

A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images cites this paper.

A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images Divide, Conquer and Combine: A Training-Free Framework for High-Resolution Image Perception in Multimodal Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T17:40:40.270764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:40:40.270764Z digest=sha256:fbe07d0db2e5694733d9509c0dc48b076d1a88bc8cc90599f055866c80cd01ed

Observation 3aeafd0e-95ae-4f16-88e7-3beea45346ff · inbound

SIEVES: Selective Prediction Generalizes through Visual Evidence Scoring cites this paper.

SIEVES: Selective Prediction Generalizes through Visual Evidence Scoring Divide, Conquer and Combine: A Training-Free Framework for High-Resolution Image Perception in Multimodal Large Language Models

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:31:16.703802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-07T16:48:20.100672Z digest=sha256:3a9a8104872be9458162baf2669e4c06ad3abaed8ba7c53ba1c5baa8107daca1

Observation cfccd4ef-e3ed-4276-92c4-9958e26dc9f4 · inbound

SIEVES: Selective Prediction Generalizes through Visual Evidence Scoring cites this paper.

SIEVES: Selective Prediction Generalizes through Visual Evidence Scoring Divide, Conquer and Combine: A Training-Free Framework for High-Resolution Image Perception in Multimodal Large Language Models

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:19:50.048642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T06:15:55.968506Z digest=sha256:294ddf7d8b7a926cea90ee0d2b2ff6327b821bc55bfc241b95ae018132be1a06

Observation a1e7dca6-0b09-4fca-bcf0-1a96bd5b4578 · inbound

MCMit: Mid-Circuit Measurement Error Mitigation cites this paper.

MCMit: Mid-Circuit Measurement Error Mitigation Divide, Conquer and Combine: A Training-Free Framework for High-Resolution Image Perception in Multimodal Large Language Models

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:45:35.071706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T08:36:37.641628Z digest=sha256:9698cf0efae72675d705f761d846bb544de1b2c0f9abd0af43a905048f4db9d4

Observation abc9bb63-8f73-44f8-9c58-219f56106c5c · inbound

Starve to Perceive: Taming Lazy Perception in VLMs with Constrained Visual Bandwidth cites this paper.

Starve to Perceive: Taming Lazy Perception in VLMs with Constrained Visual Bandwidth Divide, Conquer and Combine: A Training-Free Framework for High-Resolution Image Perception in Multimodal Large Language Models

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T10:53:13.083496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T10:52:32.238241Z digest=sha256:c50860aa383f5dfaa7af1df2160451577bfa5513040742e00b31e7c2dce1a6fb

Observation ae552715-b7ea-4c3d-a7e2-005de79615fa · inbound

Starve to Perceive: Taming Lazy Perception in VLMs with Constrained Visual Bandwidth cites this paper.

Starve to Perceive: Taming Lazy Perception in VLMs with Constrained Visual Bandwidth Divide, Conquer and Combine: A Training-Free Framework for High-Resolution Image Perception in Multimodal Large Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-12T16:34:34.452430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T16:34:34.452430Z digest=sha256:e9e7880c7c0360d574f8c2f449c192f8fb41623c5baed27ab225bccf0305929b

Observation 94de064a-f031-4d7e-8ae6-45190a185b19 · inbound

ClaimDiff-RL: Fine-Grained Caption Reinforcement Learning through Visual Claim Comparison cites this paper.

ClaimDiff-RL: Fine-Grained Caption Reinforcement Learning through Visual Claim Comparison Divide, Conquer and Combine: A Training-Free Framework for High-Resolution Image Perception in Multimodal Large Language Models

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:39:53.717075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T08:36:27.676888Z digest=sha256:24aa809aca8caf7fd430703374699ab33999fffe998900409f5916250f7a509c

Observation f5651ffc-427a-40e7-877f-47917c8181a4 · inbound

ClaimDiff-RL: Fine-Grained Caption Reinforcement Learning through Visual Claim Comparison cites this paper.

ClaimDiff-RL: Fine-Grained Caption Reinforcement Learning through Visual Claim Comparison Divide, Conquer and Combine: A Training-Free Framework for High-Resolution Image Perception in Multimodal Large Language Models

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:04:58.341749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T17:57:47.409741Z digest=sha256:64116ac0ab7209537572f9aaebb0196d223344af0cc7f5e88c8710f395040ec4

Observation 2b012226-e0ae-4d93-9408-3da5044fcadc · inbound

UltraVR: A Diagnostic Ultra-Resolution Image-VQA Benchmark for Evidence-Grounded Reasoning cites this paper.

UltraVR: A Diagnostic Ultra-Resolution Image-VQA Benchmark for Evidence-Grounded Reasoning Divide, Conquer and Combine: A Training-Free Framework for High-Resolution Image Perception in Multimodal Large Language Models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-02T11:46:55.943105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T02:58:25.590904Z digest=sha256:44f731a250194540fbdb577edff4f1dd6ab458b3960b0d2c402e6aa56f9881c3

Observation fa4d5bd5-276d-4984-985d-7ee87fed7d0a · inbound

AdaTurn: Budget-Aware Test-Time Scaling for Active Visual Perception Agents cites this paper.

AdaTurn: Budget-Aware Test-Time Scaling for Active Visual Perception Agents Divide, Conquer and Combine: A Training-Free Framework for High-Resolution Image Perception in Multimodal Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T01:50:53.323606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:50:53.323606Z digest=sha256:7000e4bc7657f8d1bd1254863f17a3de59d1da7e7cc4ffc9ef7a6ff32a0d9c82

Observation f14c760b-b865-4160-9267-8a621f6c0ccf · inbound

RP-OPSD: Resolution-Privileged On-Policy Self-Distillation for Multimodal Large Language Models cites this paper.

RP-OPSD: Resolution-Privileged On-Policy Self-Distillation for Multimodal Large Language Models Divide, Conquer and Combine: A Training-Free Framework for High-Resolution Image Perception in Multimodal Large Language Models

Reference 64

Resolution
unresolved
no resolver link, observed 2026-07-31T14:46:30.136196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T14:46:30.136196Z digest=sha256:631da06580249726a8bae3a07ed4990fd7cc51185ff015e39423b13b348abbeb

Observation aa723826-0f6f-48dd-8d98-7c120e7a0deb · inbound

VAD: Attributing Visual Evidence for Target Reconstruction in Multimodal On-Policy Distillation cites this paper.

VAD: Attributing Visual Evidence for Target Reconstruction in Multimodal On-Policy Distillation Divide, Conquer and Combine: A Training-Free Framework for High-Resolution Image Perception in Multimodal Large Language Models

Reference 90

Resolution
unresolved
no resolver link, observed 2026-07-31T02:56:58.284636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T02:56:58.284636Z digest=sha256:6e12e2bbaa207edcb97616fbc9ecd319a79b0b3e53604f9eaa78c5d73ba74834