Pith. sign in

Paper Citation Record · LEDGER

Dragonfly: Multi-Resolution Zoom-In Encoding Enhances Vision-Language Models

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2406.00977.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.00977 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T13:35:05.146089Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 4c1de7a1-5708-4282-846f-1656ee8c09e0 · inbound

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models cites this paper.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Dragonfly: Multi-Resolution Zoom-In Encoding Enhances Vision-Language Models

Reference 105

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:55:12.690324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:e5ad2eccf91c272b24d15d4b03345d357b1d0a08ad34ec5f1294b49b6978f25a

Observation bd7a9c8b-04b3-46e6-acee-abd39ef0c37d · inbound

NanoVLMs: How small can we go and still make coherent Vision Language Models? cites this paper.

NanoVLMs: How small can we go and still make coherent Vision Language Models? Dragonfly: Multi-Resolution Zoom-In Encoding Enhances Vision-Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T13:35:05.146089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:35:05.146089Z digest=sha256:0624d8be6fd7cf76fdd77ebccae1ffe7c7da59e5be7449184a9e1853cf6f9e4d

Observation 54ded20d-c61f-4f0d-afe0-acdd9efd7b58 · inbound

VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model cites this paper.

VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Dragonfly: Multi-Resolution Zoom-In Encoding Enhances Vision-Language Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:13:57.522812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T01:13:57.368874Z digest=sha256:e5f5833694135b5f17905b5e9b4a33810525341c6fd5260b65f15bb1c536540a

Observation c9c35a3a-e7ab-49b8-82e2-6924501282f5 · inbound

Align and Surpass Human Camouflaged Perception: Visual Refocus Reinforcement Fine-Tuning cites this paper.

Align and Surpass Human Camouflaged Perception: Visual Refocus Reinforcement Fine-Tuning Dragonfly: Multi-Resolution Zoom-In Encoding Enhances Vision-Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:53.620718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:53.620718Z digest=sha256:7b39269784c78ec9497db8aaf4222877a478ded992b2ffb146a1b2ad848ecb9b

Observation 95305907-eca2-414a-9d4f-20ead735841f · inbound

HSENet: Hybrid Spatial Encoding Network for 3D Medical Vision-Language Understanding cites this paper.

HSENet: Hybrid Spatial Encoding Network for 3D Medical Vision-Language Understanding Dragonfly: Multi-Resolution Zoom-In Encoding Enhances Vision-Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T04:47:07.644114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:47:07.644114Z digest=sha256:149890d2d040da5d1bbb49a0f07af1f069ff50a2bf5c3f280d64b0fe6964f22c

Observation 4370f78f-ad7e-460a-afac-bbb72edbc717 · inbound

Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints cites this paper.

Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints Dragonfly: Multi-Resolution Zoom-In Encoding Enhances Vision-Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T04:57:50.493602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:57:50.493602Z digest=sha256:7c04042294f710ae32f6f939bc6b12dd23dbcdbec0418c6936d0153dc48efa87

Observation c563a569-7719-4c6a-a3a1-304bf844b4e4 · inbound

LensVLM: Selective Context Expansion for Compressed Visual Representation of Text cites this paper.

LensVLM: Selective Context Expansion for Compressed Visual Representation of Text Dragonfly: Multi-Resolution Zoom-In Encoding Enhances Vision-Language Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:41:00.029294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T01:09:01.612988Z digest=sha256:6af9b41ad8621586638d8a1cf1a923d82a601ed00700d53bbf2c49e2e4bd91fd

Observation efe94f43-145d-4253-9e29-de71d652dc03 · inbound

TextTeacher: What Can Language Teach About Images? cites this paper.

TextTeacher: What Can Language Teach About Images? Dragonfly: Multi-Resolution Zoom-In Encoding Enhances Vision-Language Models

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-22T07:26:12.686085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-22T07:26:03.594414Z digest=sha256:5ab9e6faa139795f72dde784f41782b24831e04f02c36105d12477c3c67bfe60