Pith. sign in

Paper Citation Record · LEDGER

Visual Haystacks: A Vision-Centric Needle-In-A-Haystack Benchmark

As of 19 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2407.13766.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.13766 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:43:17.618225Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T06:36:52.559941Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 56912e52-c4ae-42a5-ba26-ae04ab7d507f · inbound

Document Haystacks: Vision-Language Reasoning Over Piles of 1000+ Documents cites this paper.

Document Haystacks: Vision-Language Reasoning Over Piles of 1000+ Documents Visual Haystacks: A Vision-Centric Needle-In-A-Haystack Benchmark

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T14:11:22.084401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:11:22.084401Z digest=sha256:d56cc0b9f65c62d1744b0328a45ca6c00c62a7a145d321003bd3adef4ab18e87

Observation 5c9abfd0-f032-491f-8137-ab3180bbf295 · inbound

GIRAFFE: Design Choices for Extending the Context Length of Visual Language Models cites this paper.

GIRAFFE: Design Choices for Extending the Context Length of Visual Language Models Visual Haystacks: A Vision-Centric Needle-In-A-Haystack Benchmark

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:57.183419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:57.183419Z digest=sha256:c0ee076f0494cd9dbd6eadd51500a37b7cdd1a414f7211a9aab27f7ba7a5b67e

Observation 354fad58-8c66-4189-8728-22959c02b8fa · inbound

FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding cites this paper.

FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding Visual Haystacks: A Vision-Centric Needle-In-A-Haystack Benchmark

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T10:43:17.618225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:43:17.618225Z digest=sha256:dc5f168d9608b526ff28881c0796b4363fb2e7a1ab94c60b79fecea0b6cb333f

Observation d38f6978-372a-4e46-914f-b76d9b486cff · inbound

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? cites this paper.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Visual Haystacks: A Vision-Centric Needle-In-A-Haystack Benchmark

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-16T10:22:52.641231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:22:52.641231Z digest=sha256:51b71a192399fb85ca93e0e9f3980f2f9bd6f960d7ca347a387fffdcb9af3a89

Observation 03f03589-0a7e-44de-a1b4-023e511b63e0 · inbound

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks cites this paper.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Visual Haystacks: A Vision-Centric Needle-In-A-Haystack Benchmark

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.475228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.475228Z digest=sha256:d700ec3e7aeb9e1a8c3baa1a8cbbc36743e8b7875fbc7b21f3afb16854c71f6a

Observation 1d64bf2a-1c1b-4886-bc56-f5b34b402675 · inbound

From General to Targeted Rewards: Surpassing GPT-4 in Open-Ended Long-Context Generation cites this paper.

From General to Targeted Rewards: Surpassing GPT-4 in Open-Ended Long-Context Generation Visual Haystacks: A Vision-Centric Needle-In-A-Haystack Benchmark

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:48.044179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:48.044179Z digest=sha256:f0375a81acf76728ac4210661fbf85b68955a89e4226f0af8ce8832321304c77

Observation 9d3bcf66-c07c-44de-9637-691bed3f7b76 · inbound

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames cites this paper.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Visual Haystacks: A Vision-Centric Needle-In-A-Haystack Benchmark

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.795465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.795465Z digest=sha256:d0825281feb3bd43d025ba40b9b8421068e39f1441a37e512090a04922609d59

Observation e8e2e635-b3b7-4dc6-b0a6-020f38e29882 · inbound

Decoding the Pulse of Reasoning VLMs in Multi-Image Understanding Tasks cites this paper.

Decoding the Pulse of Reasoning VLMs in Multi-Image Understanding Tasks Visual Haystacks: A Vision-Centric Needle-In-A-Haystack Benchmark

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-15T16:06:14.482917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T16:04:58.851378Z digest=sha256:dcd91b169293381498fe5bae6f827157db96fdfca938c6376dd3cde0ef8901ba

Observation cb694d30-1942-4dcb-969d-6b8ee29c11e1 · inbound

Utility-Oriented Visual Evidence Selection for Multimodal Retrieval-Augmented Generation cites this paper.

Utility-Oriented Visual Evidence Selection for Multimodal Retrieval-Augmented Generation Visual Haystacks: A Vision-Centric Needle-In-A-Haystack Benchmark

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T19:09:22.942744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-14T19:09:18.975682Z digest=sha256:07d0cd6e6566b26c034c128b11c72bacf9a786ab0f67e42dec11dc41a56fe203

Observation 2259e90d-95dd-4db2-83a6-3189f1a20789 · inbound

Training Long-Context Vision-Language Models Effectively with Generalization Beyond 128K Context cites this paper.

Training Long-Context Vision-Language Models Effectively with Generalization Beyond 128K Context Visual Haystacks: A Vision-Centric Needle-In-A-Haystack Benchmark

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:17:50.089420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-14T19:16:07.851098Z digest=sha256:1a26f16d14160d71ec94627544867dc71ec1f25d148bbe8415d13a0193d2fce0

Observation fbce73d7-06dd-444f-a247-b711ec15e776 · inbound

Personal Visual Memory from Explicit and Implicit Evidence cites this paper.

Personal Visual Memory from Explicit and Implicit Evidence Visual Haystacks: A Vision-Centric Needle-In-A-Haystack Benchmark

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-06-29T12:43:25.439201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-29T12:40:34.740804Z digest=sha256:e40952f40fec9e2d05ac6558df2527744cfb8c6c71c3cb42f586e96a22142b88

Observation 71aab449-6c2f-4f8a-aedc-21992cb0b003 · inbound

Invoice Haystack: Benchmarking Document Retrieval and Visual Question Answering Under Strong Visual Homogeneity cites this paper.

Invoice Haystack: Benchmarking Document Retrieval and Visual Question Answering Under Strong Visual Homogeneity Visual Haystacks: A Vision-Centric Needle-In-A-Haystack Benchmark

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-07-04T19:30:07.775820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-25T21:15:12.242955Z digest=sha256:286710db58355940e3725e2faf852ced35f9fe674d71fb37dbad2807438aa1f2

Observation c0c2a662-ca93-4faa-b0f8-66391d037bad · inbound

Invoice Haystack: Benchmarking Document Retrieval and Visual Question Answering Under Strong Visual Homogeneity cites this paper.

Invoice Haystack: Benchmarking Document Retrieval and Visual Question Answering Under Strong Visual Homogeneity Visual Haystacks: A Vision-Centric Needle-In-A-Haystack Benchmark

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:23:51.256632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-29T05:07:28.537679Z digest=sha256:b6dcfb6b8bb027324f1809c9d01a75ee66f0c490f9c690fe0c224ecc96a21e07

Observation 339be329-fadd-4f7f-a564-1a93d51be974 · inbound

C3-Bench: A Context-Aware Change Captioning Benchmark cites this paper.

C3-Bench: A Context-Aware Change Captioning Benchmark Visual Haystacks: A Vision-Centric Needle-In-A-Haystack Benchmark

Reference 93

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T19:50:10.146177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-25T21:02:52.529391Z digest=sha256:7f37850c30ad6865797fabf518e0b409cc11d5c51c459f3d0e2e92af72ff9c9e

Observation 7fdc33c5-da4a-458a-b748-ea7865372227 · inbound

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing cites this paper.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Visual Haystacks: A Vision-Centric Needle-In-A-Haystack Benchmark

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-07-10T06:36:52.561097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:0cc8478f2087e27c5de6789374300e79d935aa9690598eaf2f04ec88be6adf5b

Observation ebfad781-b4d7-45e9-82e1-51941c47bea5 · inbound

ReToken: One Token to Improve Vision-Language Models for Visual Retrieval cites this paper.

ReToken: One Token to Improve Vision-Language Models for Visual Retrieval Visual Haystacks: A Vision-Centric Needle-In-A-Haystack Benchmark

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-31T01:43:24.234557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T01:43:24.234557Z digest=sha256:c2cc31f7127d07aee4a798b3562703996154e87e5af03f7a41bb9f2079c6ea23