Pith. sign in

Paper Citation Record · LEDGER

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding?

As of 19 August 2026, this Paper Citation Record lists 84 of 84 outbound references and 0 inbound Pith citation observations for arXiv:2504.18406.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.18406 v2

Coverage vector

measured 84 of 84 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:22:52.691595Z

measured 84 of 84 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

84 of 84 outbound references displayed

  • verified exact0
  • verified fuzzy51
  • unresolved32
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f9b1899d-2417-4785-8d5f-28b113856ef8 · outbound

This paper cites 12 Kucherlapati Raju 13, Genome data analysis: Baylor College of Medicine Creighton Chad J.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? 12 Kucherlapati Raju 13, Genome data analysis: Baylor College of Medicine Creighton Chad J

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T10:22:52.389235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:22:52.389235Z digest=sha256:734da4b3bff79e3083dcff2a291e08da6064463840821520d81166e219137d8c

Observation 4d0c885f-e035-4a56-9bb4-5492d0030c35 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T10:22:52.393488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:22:52.393488Z digest=sha256:42eef140f57aea601b47b95a145db9ddfc195f19227c04bfc58d81b0dc3f12ee

Observation 1175f303-e6a7-4395-97ef-3c9cd122ee9e · outbound

This paper cites GPT-4 Technical Report.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? GPT-4 Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T10:22:52.397322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:22:52.397322Z digest=sha256:6b5b6c70432c9b916b6a85410b6602f807758b87d76a88526ba0cc57b3bdaab1

Observation cbc5567b-d8f4-429e-9f34-44a9c47784e2 · outbound

This paper cites The claude 3 model family: Opus, sonnet, haiku.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? The claude 3 model family: Opus, sonnet, haiku

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T10:22:52.401320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:22:52.401320Z digest=sha256:84cbba4ab6174906c01da7a12fc07bae939b801f3734ce8559345a55a8787ffb

Observation 15c957c5-8cd7-4cd8-bd1a-b94bde5faa64 · outbound

This paper cites Bach: Grand challenge on breast cancer histology im- ages.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Bach: Grand challenge on breast cancer histology im- ages

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T10:22:52.404867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:22:52.404867Z digest=sha256:b7153b5a7faad81a0ea8837dc43f4bf1d48a0f7dbb60948bc3bbb7fc4a818b93

Observation dd92b8e1-1e15-4332-b442-9ff0bc14226d · outbound

This paper cites Jauregui, and Juan Andr ´es Cardoso.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Jauregui, and Juan Andr ´es Cardoso

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T10:22:52.408735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:22:52.408735Z digest=sha256:8855e1cb503d59eaa3e06e15a0e07c784cf05d5d11f175978d20a170d16e78c3

Observation d3a4f4f7-944b-474c-995c-f7149645fcae · outbound

This paper cites Ef- ficient high-resolution deep learning: A survey.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Ef- ficient high-resolution deep learning: A survey

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.682253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T10:22:52.412513Z digest=sha256:a6084cc6fc7981754fd26cd0d96713ca174bea7460c95b893b7fe3e14739af01

Observation 19bb4625-5def-4fbf-9d05-ce0b6ba78736 · outbound

This paper cites Effi- cient high-resolution deep learning: A survey.ACM Comput.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Effi- cient high-resolution deep learning: A survey.ACM Comput

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.672630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T10:22:52.416326Z digest=sha256:679dc1962b2945f778459d80063e2ffe79dc42f1862f235781ea8ac7917f5cf6

Observation e944eaa2-510a-49db-9955-3c21580f39a6 · outbound

This paper cites Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T10:22:52.419841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:22:52.419841Z digest=sha256:63412a6f6e7b65f317ece799e9dc44887299423a704aa2f0215c13846e4f3db6

Observation 888b5071-0d94-441c-bd07-3b0a4ffc64c3 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T10:22:52.423560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:22:52.423560Z digest=sha256:bfd8c9134f888ddfdcde95f3a1a25042b31dc4d7076abaf6659e854c25037065

Observation e6070e0a-7b27-4f7e-a792-7a5bf964a236 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T10:22:52.427409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:22:52.427409Z digest=sha256:95898c4c5124473a1b036510f1a1d7cc00f82f7f1745c93c4fd3d2a6feb84405

Observation e97359c0-4da3-485b-8341-aa52329fbaa5 · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.663099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T10:22:52.431192Z digest=sha256:96aba8c391c33d64b44cc6db3af3efca4fc48987e0d2dcfb20ace86b8ae96796

Observation 16647775-f89e-440e-b7ef-ec1659f534dd · outbound

This paper cites Functional map of the world.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Functional map of the world

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.652996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T10:22:52.434724Z digest=sha256:0fd3773a5995ffadeb1797596289884725053cd0d80f685587c5dad5a23312d8

Observation c6c8c29c-eb5e-4f8d-89e7-bfb21486adfd · outbound

This paper cites The cityscapes dataset for semantic urban scene understanding.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? The cityscapes dataset for semantic urban scene understanding

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.642930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T10:22:52.438441Z digest=sha256:ccf154fc44b269a879a6c21556efc8c519cb5cca3bd4736cf984b2eee30abd03

Observation 57a333d4-255e-4b52-86fb-ddd5aea2bf8f · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T10:22:52.441852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:22:52.441852Z digest=sha256:6a69ddc855b59a20bb69a635bca19362164fa6ef95ccee2703dd831988c3a64f

Observation e34ee8a0-b7f5-476e-b3ed-e84af2704903 · outbound

This paper cites LungHist700: A dataset of histological images for deep learning in pulmonary pathology.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? LungHist700: A dataset of histological images for deep learning in pulmonary pathology

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.632512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T10:22:52.445363Z digest=sha256:ba6bc3101512d09146b875d3e5c453c47680c285dab2832547997e3e96de425b

Observation b299ab0e-d64d-4bbd-87ba-8cadd9c49f01 · outbound

This paper cites LungHist700: A dataset of histological images for deep learning in pulmonary pathology.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? LungHist700: A dataset of histological images for deep learning in pulmonary pathology

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.621966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T10:22:52.448580Z digest=sha256:68d4260fb3491f997f39ef3d0021237eb4371ec3d198e5e9c65b9ae5d501e0f4

Observation ab042208-0d94-4f79-8ecc-e81ad17c7e34 · outbound

This paper cites Internlm-xcomposer2-4khd: A pioneer- ing large vision-language model handling resolutions from 336 pixels to 4k HD.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Internlm-xcomposer2-4khd: A pioneer- ing large vision-language model handling resolutions from 336 pixels to 4k HD

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.611682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T10:22:52.451866Z digest=sha256:eb37f07022012fdf3219c5ca354ea3ef5d3dcb30dc3e20d897970fe729c509ab

Observation a263f2d1-568b-4dbf-b65f-3d6ffda9c112 · outbound

This paper cites The Llama 3 Herd of Models.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? The Llama 3 Herd of Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T10:22:52.455103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:22:52.455103Z digest=sha256:ac7e37a6f1610cd864d5d0cccefc6516ef777f7c15b8324eb79f344022811fce

Observation 3543c1bb-b91e-41ad-a5a0-6f8843f862db · outbound

This paper cites Describing differences in image sets with natural language.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Describing differences in image sets with natural language

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.601042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T10:22:52.458785Z digest=sha256:423f6082ea45366bfb710ad8a5eeae4a38e6bd45c833347ca71846348524f845

Observation f65c327e-f379-4aad-89a6-efc6604d2669 · outbound

This paper cites Floorplancad: A large-scale cad draw- ing dataset for panoptic symbol spotting.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Floorplancad: A large-scale cad draw- ing dataset for panoptic symbol spotting

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.590413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T10:22:52.462320Z digest=sha256:90e52aa2b33e068dc222bab22317dacd0c0385ffa383806c1114289ac076f0ef

Observation b33ecc19-b636-4868-807b-f9f156c0d2a3 · outbound

This paper cites Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T10:22:52.465714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:22:52.465714Z digest=sha256:98672e8f1752a4b965c3c64d39b4cff9514c3e2c5bdf45c296619c39ef03a1bb

Observation 67467734-d1fc-46cc-8cea-14e87c0e9ea4 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.579762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T10:22:52.469307Z digest=sha256:9aa90a9fc101d9fba8cd17a6579f8e6826250594dcfb4d3e9935987d767801a4

Observation 5dae7f81-ff1f-428d-8965-120282bb8756 · outbound

This paper cites Llava-uhd: An LMM perceiving any aspect ratio and high- resolution images.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Llava-uhd: An LMM perceiving any aspect ratio and high- resolution images

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.569489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T10:22:52.472689Z digest=sha256:c80b0fb535932c82c78b85dd6fb0072babd0573ded3143c19ccc547be171eb32

Observation a7cb86e0-1b89-49ee-8103-6ee52159a44f · outbound

This paper cites Cogagent: A visual lan- guage model for GUI agents.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Cogagent: A visual lan- guage model for GUI agents

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.558847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T10:22:52.476218Z digest=sha256:afc87806ef5119d8f69c254f7df644630b4bed0d2d38285ead37165a6e5905f2

Observation e71e3d52-d4ac-49ac-bb8c-d59361ed0628 · outbound

This paper cites mplug-docowl 1.5: Unified structure learning for ocr-free document understanding.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? mplug-docowl 1.5: Unified structure learning for ocr-free document understanding

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.548052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T10:22:52.479535Z digest=sha256:c025b8f644855e93ea936840e8296d0b7d1cceaf54ccde75c67c9f36c54c9f06

Observation c82a385b-598a-4c8f-87ef-8d3fcbdbcdda · outbound

This paper cites mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T10:22:52.483040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:22:52.483040Z digest=sha256:ccf6e4f28f28c3e279154b8cd497afab01073929d37d56dd5c39bfaca678bdab

Observation 694fe1a0-1e68-4e9d-a306-5ab338b52b13 · outbound

This paper cites Novachart: A large- scale dataset towards chart understanding and generation of multimodal large language models.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Novachart: A large- scale dataset towards chart understanding and generation of multimodal large language models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.538004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T10:22:52.486969Z digest=sha256:2a402654e1b0e082a1667cf80a4ec36198dbc7e4adf01622877f544a1e6595e9

Observation 320bb115-4ebc-421b-a50b-ed91ba62d796 · outbound

This paper cites High reso- lution image quality database.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? High reso- lution image quality database

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.527667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T10:22:52.490267Z digest=sha256:660338793818792e98ab8ed4a5f69eada5146a4663ccdbe6404118557d02debf

Observation dbd2eaae-922f-4a31-b77a-d201cdb248a9 · outbound

This paper cites Mini-monkey: Alleviating the semantic saw- tooth effect for lightweight MLLMs via complementary im- age pyramid.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Mini-monkey: Alleviating the semantic saw- tooth effect for lightweight MLLMs via complementary im- age pyramid

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.517732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T10:22:52.493634Z digest=sha256:97c63edb4975780b8b8ab93d6bf513f96453226c180a02b2bd16580a48898ac8

Observation a3c99a96-fffb-4241-a0cb-e6cceaf6ca5c · outbound

This paper cites The apolloscape dataset for autonomous driving.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? The apolloscape dataset for autonomous driving

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.507439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T10:22:52.496876Z digest=sha256:569337c75cb6c830587c8d349cbed766caf021a22fac120b63fb8c96468af265

Observation f5021136-94e3-4344-99ca-29f1255b301d · outbound

This paper cites GPT-4o System Card.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? GPT-4o System Card

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T10:22:52.500288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:22:52.500288Z digest=sha256:f1e4dfe14a9e47cf355cd5d0b839e702cb6998ca574a8599cba5ea0cbcd0aaf6

Observation 47e467f5-93b5-46f2-8aa4-3c00497d8f1d · outbound

This paper cites Multi-source multi-scale counting in extremely dense crowd images.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Multi-source multi-scale counting in extremely dense crowd images

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.496564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T10:22:52.503899Z digest=sha256:d371c0c63dead2d43d4143aec71378e12bc87abb1c73e81963d3b5b99e2c5117

Observation 7771275f-05a5-40d4-b808-9a8c39b60bd7 · outbound

This paper cites Composition loss for counting, density map estima- tion and localization in dense crowds.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Composition loss for counting, density map estima- tion and localization in dense crowds

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.486389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T10:22:52.507778Z digest=sha256:646081068f4803d182e1620b3543576ee64e48329ad35f164c931d3db9ce287c

Observation bbdc62ea-4066-464f-9852-4d5b711ee143 · outbound

This paper cites CosmoCLIP: Generalizing Large Vision-Language Models for Astronomical Imaging.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? CosmoCLIP: Generalizing Large Vision-Language Models for Astronomical Imaging

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T10:22:52.511131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:22:52.511131Z digest=sha256:b24007338813ee23a867d39b18d723e663d1f97d931b879be44da64cf620ee87

Observation af02aa70-aabd-4317-b138-8e2890980087 · outbound

This paper cites MMAD: A Comprehensive Benchmark for Multimodal Large Language Models in Industrial Anomaly Detection.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? MMAD: A Comprehensive Benchmark for Multimodal Large Language Models in Industrial Anomaly Detection

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T10:22:52.515246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:22:52.515246Z digest=sha256:b76cdc4a3ee9db0fcaadf113fc4e1fde7500a2413c55bd75d044ddc969470c0a

Observation f97d40ff-5c0e-4917-8a81-08ae1b554f3b · outbound

This paper cites A diagram is 10 worth a dozen images.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? A diagram is 10 worth a dozen images

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.475510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T10:22:52.518870Z digest=sha256:708752f40a38a3e7a66a3e746b9560bd23af26379fca7a515cc78d6c41ccf72f

Observation ef01dd84-2979-419b-b91b-f2006978501c · outbound

This paper cites Berg, Wan-Yen Lo, Piotr Doll ´ar, and Ross B.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Berg, Wan-Yen Lo, Piotr Doll ´ar, and Ross B

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.463325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T10:22:52.522188Z digest=sha256:c64a3bd7e016577bcd2d1f985c863109b26176e887c12ed078419719f60aa208

Observation 652cbbda-5d7a-4ae4-b779-259e86ac5003 · outbound

This paper cites an unresolved cited work.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-16T10:22:53.452609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T10:22:52.525350Z digest=sha256:fdd61811398556ef7e7d0eb80391d388c94851ae1ab45d7d730deda348e88e98

Observation ce38ad38-cf11-496e-b34a-23bafbf7087c · outbound

This paper cites A dataset of clinically generated visual questions and answers about radiology images.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? A dataset of clinically generated visual questions and answers about radiology images

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.441274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T10:22:52.528830Z digest=sha256:6fe1699626b27723d6ea65d9c40baa40d9a3153f54c9a0276a6f7c9cb736d0d1

Observation dd87e033-17dd-40d5-a7ac-b5ca145bb28e · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation, 2022.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation, 2022

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.431021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T10:22:52.532223Z digest=sha256:fc10eb18f9228bbe8768db60e9571864bc7c0a4207780773717622f10d2a57cb

Observation 090ce644-3b98-469d-9077-51d19d9bb127 · outbound

This paper cites Hrvqa: A visual question answering benchmark for high-resolution aerial images.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Hrvqa: A visual question answering benchmark for high-resolution aerial images

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.420747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T10:22:52.535599Z digest=sha256:15c39470a6e60d37e83b767e0e649a4cb8601b44ce18367fdfa6c146c0098203

Observation 5ea0ebcb-a443-4f56-bce3-67b7cd2542a7 · outbound

This paper cites To- wards streaming perception.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? To- wards streaming perception

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.410588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T10:22:52.538872Z digest=sha256:568333e75f1d774d4292579452a99c55edcba4096b23666cb064145ea1b2f811

Observation e02da5d8-5072-4c2c-b1f7-49c0d86fa976 · outbound

This paper cites Mon- key: Image resolution and text label are important things for large multi-modal models.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Mon- key: Image resolution and text label are important things for large multi-modal models

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.399929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T10:22:52.542538Z digest=sha256:07ede0dfe0cdc6bbe2abd0dce9c0f6b5f392f661aea97754e79e5d8b642383e7

Observation 8834e4b0-29c1-49bb-ac17-2cf49e132b87 · outbound

This paper cites The ArtBench Dataset: Benchmarking Generative Models with Artworks.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? The ArtBench Dataset: Benchmarking Generative Models with Artworks

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-16T10:22:52.545955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:22:52.545955Z digest=sha256:4a7b6982166f357739b000734e4675afaa8182a54d6412cd120ba9cbf82e4e99

Observation 702d130a-c5fd-47a4-bb23-c923f4826bf1 · outbound

This paper cites Visual instruction tuning.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Visual instruction tuning

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.389280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T10:22:52.549625Z digest=sha256:38b10f3632c197cdb705d3604b98ce14063c9bd2f88b9503c7b8ba376a358d8d

Observation b662ac6b-21e7-47e4-91a6-3b735978b878 · outbound

This paper cites Visual instruction tuning.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Visual instruction tuning

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.378720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T10:22:52.552983Z digest=sha256:402e453ff5c9d1fbc82d472875084658f41603bde40e71eb1d371c692a4b3d50

Observation 5c0f2fe2-da3d-4828-b7db-7357e9005fd2 · outbound

This paper cites Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.368266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T10:22:52.556413Z digest=sha256:08caaeaf7b6195c4cf1825d1937e59be9e5072a421075846c2197153306690b9

Observation f5d8e1cc-4b44-4e31-8afd-14b7e3d78322 · outbound

This paper cites A convnet for the 2020s.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? A convnet for the 2020s

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.357129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T10:22:52.560005Z digest=sha256:38342c55e5fa21ccc51d7cddbb19e86ba7fac0949dc923ed8ad7c8b0c65877d3

Observation f3b305db-fadf-4584-b15e-fa91c6b88242 · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-16T10:22:52.563373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:22:52.563373Z digest=sha256:a9e6d47bdfe004587436aa30a1537b5a2cfb77c0b583df58f61d792ef4341123

Observation 997a93e9-ecbe-44a5-9a1d-bf4f0c1d7d36 · outbound

This paper cites Feast your eyes: Mixture- of-resolution adaptation for multimodal large language mod- els.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Feast your eyes: Mixture- of-resolution adaptation for multimodal large language mod- els

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.344940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T10:22:52.567069Z digest=sha256:78054d3935243e08158e66f846412184ee0f102e074d9494780a7619c48cc7f2

Observation 432380a8-118f-4daf-b63d-c7788b99631b · outbound

This paper cites Infographicvqa.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Infographicvqa

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.333675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T10:22:52.570808Z digest=sha256:301c72fa1469abaef9df5eb345a9e1daca18c3a3ea5f6aae023bde917f72f3eb

Observation 304adae2-21fe-4aed-8c1e-cadeca871c33 · outbound

This paper cites MM1: methods, analysis and insights from multimodal LLM pre-training.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? MM1: methods, analysis and insights from multimodal LLM pre-training

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.322446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T10:22:52.574251Z digest=sha256:2a29e8192400397a8a5a93375906eca8a3005a97c07de6c5e0f579f37bd1b842

Observation 36d208a9-8be8-4387-b4e5-9c87de2704f5 · outbound

This paper cites The mame dataset: on the relevance of high resolution and variable shape image properties.Applied Intelligence, 52(10):11703–11724, 2022.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? The mame dataset: on the relevance of high resolution and variable shape image properties.Applied Intelligence, 52(10):11703–11724, 2022

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.311099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T10:22:52.577807Z digest=sha256:259a276fc77f8220bfc920ba2e6e45f1f2872a7a11062891f969bd4ec85918a0

Observation 6d458e48-6f4a-4836-86b6-9d15bbb56e64 · outbound

This paper cites Learning transferable visual models from natural language supervision.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Learning transferable visual models from natural language supervision

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.300422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T10:22:52.581521Z digest=sha256:0e33dda510adcb88a5514b167690282e921ea270915c18857eeb1740dc17afcb

Observation 418f9234-1f9a-48e7-b5ec-48bed3326402 · outbound

This paper cites The synthia dataset: A large collection of synthetic images for semantic segmentation of urban scenes.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? The synthia dataset: A large collection of synthetic images for semantic segmentation of urban scenes

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-16T10:22:52.584823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:22:52.584823Z digest=sha256:896c16386a3c995e74ff8c64f68fe366b68a48940c0b34ab4ea9bb5b8f036ac1

Observation 6dbd83c3-87e4-46ae-bc16-6a5df342a570 · outbound

This paper cites Jhu-crowd++: Large-scale crowd counting dataset and a benchmark method.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Jhu-crowd++: Large-scale crowd counting dataset and a benchmark method

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.281672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T10:22:52.588328Z digest=sha256:30b8cf6d6f582cafd3fc5aab2055e08cbb651c26a4a9d27138e766f8797f5c5b

Observation 2effe67c-8df2-4ba9-be1d-f9082195663a · outbound

This paper cites MileBench: Benchmarking MLLMs in Long Context.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? MileBench: Benchmarking MLLMs in Long Context

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-16T10:22:52.591664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:22:52.591664Z digest=sha256:d04ac40feb55737e7f33a1780b0e88b467cf8d30ce8f98174222aa449df722ed

Observation 2249752f-28cf-48b5-afd6-ee52df310d5b · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Gemini: A Family of Highly Capable Multimodal Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-16T10:22:52.595291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:22:52.595291Z digest=sha256:c6ef8da95c8b5077a93ea52b286e3ace33742a28000c3f9e8d75aaaf32798c27

Observation 0073b7ca-a1b8-42e1-bd00-913a855c079c · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-16T10:22:52.599306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:22:52.599306Z digest=sha256:6b86f0b1595840d899ae2e7cb9f61e8415568a8e4100d37c43232dfa436e62ef

Observation e8d37346-5381-4628-8b22-f894d82d57e9 · outbound

This paper cites Predicting breast tumor proliferation from whole-slide im- ages: the tupac16 challenge.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Predicting breast tumor proliferation from whole-slide im- ages: the tupac16 challenge

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.270358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T10:22:52.602880Z digest=sha256:b19b1f6da80de3f4d1c49ffec7cbc3b30cfcbefb763e14e45e5e3d5b4b6e5b45

Observation 3581a251-a555-4e0f-81f8-bd8810a09e88 · outbound

This paper cites MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-16T10:22:52.606264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:22:52.606264Z digest=sha256:770e7f7cbeadd17ac187a43363d34b1816674988de2d74d49a67b3568b80e0ee

Observation 8038171f-0e4d-4270-88fd-fe9b7f7436c3 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-16T10:22:52.613989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:22:52.613989Z digest=sha256:7c9bc1927cf0d37848776d15f3036737fa4cb808dbda8ee43d54c0e77eb0db58

Observation 99e3a955-7b77-43ed-a0d4-4de8baafce46 · outbound

This paper cites Divide, conquer and combine: A training-free framework for high-resolution im- age perception in multimodal large language models.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Divide, conquer and combine: A training-free framework for high-resolution im- age perception in multimodal large language models

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.258765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T10:22:52.617535Z digest=sha256:0a66b7d42af75b4d5f2e431975bfa850e027943423f36e2d26cfcbbefdcd7096

Observation 99dd5e4f-374c-4d32-bc37-8bb84f7c2773 · outbound

This paper cites Divide, Conquer and Combine: A Training-Free Framework for High-Resolution Image Perception in Multimodal Large Language Models.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Divide, Conquer and Combine: A Training-Free Framework for High-Resolution Image Perception in Multimodal Large Language Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-16T10:22:52.620905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:22:52.620905Z digest=sha256:d7373dce28f4a903a452b4a04606aa222ec28dbda11928117262a0900a6d52ca

Observation b4cab3b4-c321-46b0-9e4b-28dae3ee7bd5 · outbound

This paper cites Needle in a multimodal haystack.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Needle in a multimodal haystack

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.247840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T10:22:52.624325Z digest=sha256:aaaec0e8070d9be607ed7b461716abccaa6864d8a3989fde3d7c8f4c7c42f2d4

Observation eef52f33-274d-440e-b097-3312e6487ebe · outbound

This paper cites Panda: A gigapixel- level human-centric video dataset.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Panda: A gigapixel- level human-centric video dataset

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.236547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T10:22:52.627940Z digest=sha256:3a82460764cbafee22ac0087ffe5524a2ee6972219e0259ab6b38d62ab24476a

Observation 57526d2d-fb65-4948-b8c8-6d0d0a0fe354 · outbound

This paper cites Weiser, P.L.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Weiser, P.L

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.225206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T10:22:52.631117Z digest=sha256:4dfda6abe26065dffc570a4665f6d7378190ab9847b8eba2d60a6af668fc71f1

Observation 2c891846-3d79-4421-aae7-cdf56950113c · outbound

This paper cites V?: Guided visual search as a core mechanism in multimodal llms.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? V?: Guided visual search as a core mechanism in multimodal llms

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.203540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T10:22:52.637878Z digest=sha256:4e833f2d42c902294a4a448f6677e8532f3bbac3124b3c448c50a895b56b993b

Observation d38f6978-372a-4e46-914f-b76d9b486cff · outbound

This paper cites Visual Haystacks: A Vision-Centric Needle-In-A-Haystack Benchmark.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Visual Haystacks: A Vision-Centric Needle-In-A-Haystack Benchmark

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-16T10:22:52.641231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:22:52.641231Z digest=sha256:51b71a192399fb85ca93e0e9f3980f2f9bd6f960d7ca347a387fffdcb9af3a89

Observation 88fce6b9-61f5-46f8-b8c4-77343d86eccb · outbound

This paper cites DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-16T10:22:52.648406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:22:52.648406Z digest=sha256:37a33136ad4e16214dd243a2bc8753dadef108a984af6da406d02047f2d58987

Observation d8fdb897-c59f-4dad-b020-0499b7480244 · outbound

This paper cites Deep learning for detecting corona virus disease 2019 (covid-19) on high-resolution computed tomography: a pilot study.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Deep learning for detecting corona virus disease 2019 (covid-19) on high-resolution computed tomography: a pilot study

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.192182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T10:22:52.651683Z digest=sha256:927e46d683088d33099c4bdc8a7c5669de5b23e0b47aa74018b55151419510e2

Observation 319de3f4-0f78-4693-ad79-70d0235c154c · outbound

This paper cites Bdd100k: A diverse driving dataset for heterogeneous multitask learning.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Bdd100k: A diverse driving dataset for heterogeneous multitask learning

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.181330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T10:22:52.655316Z digest=sha256:f172f3f106fffa0834f5a6c737bbae5554a9d3ba97520025fc8e1863913a74d8

Observation a3770264-13b9-43f9-a15f-822ebf6caee3 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for ex- pert agi.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for ex- pert agi

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-16T10:22:52.659019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:22:52.659019Z digest=sha256:0f1b428b141b0b4f9f836d9c61e1bde3e71313cd8b16ab19e9f98d277f06001c

Observation a6e95b13-d20a-4042-9b3e-21ee90b4f30c · outbound

This paper cites Single-image crowd counting via multi-column convolutional neural network.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Single-image crowd counting via multi-column convolutional neural network

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.162943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T10:22:52.662419Z digest=sha256:87e405480edf7dd5ec8c47acc49e6db6af98034e167ae574202df945c82dfb7c

Observation 4ff3db0e-5e4a-45ea-8d17-2c8a00b98a54 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-16T10:22:52.665965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:22:52.665965Z digest=sha256:4f1902f6a6b9a8ebf3ab35190d954fe725b8b176f82e820fe1297947f2d44c0e

Observation ce0ae181-b1cf-4695-9ff0-773db659f6a4 · outbound

This paper cites MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-16T10:22:52.669629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:22:52.669629Z digest=sha256:fca26d0d295f43aa4de7e5ae1445935ee760b1b6c1aec25f9f24bdbd95edb5ea

Observation 6ad8a1ae-f2b4-4ff2-accb-87d5da2029f5 · outbound

This paper cites Monitoring Extracted from MME-Realworld, this dataset features images taken from public safety cameras in diverse scenarios.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Monitoring Extracted from MME-Realworld, this dataset features images taken from public safety cameras in diverse scenarios

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.151619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T10:22:52.673114Z digest=sha256:bbf8e15e2fcbb91bf37deea40af4c52f9ca7a83f48fbce0bbdd8b4079331d413

Observation f10ccd4a-4de5-4ce5-a358-e5474c129936 · outbound

This paper cites an unresolved cited work.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Unresolved cited work

Reference 82

Resolution
unresolved
raw_fallback, observed 2026-08-16T10:22:53.139626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T10:22:52.676865Z digest=sha256:1b3a9b68f4e40ee7704f7233db1cd7401e18a9486ed3fc8403cff396a19ae57e

Observation c18d8af0-22b6-4773-8268-b177c9265cec · outbound

This paper cites Phi- 3.5 [2] is a lightweight model designed for efficient lan- guage understanding and generation.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Phi- 3.5 [2] is a lightweight model designed for efficient lan- guage understanding and generation

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.128132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T10:22:52.680145Z digest=sha256:cb53f16d6c2a4772adcec11191feba6cff9b09ae6627295d1eecc4a0484bd1d7

Observation 4439af5e-c13a-4957-b8fe-e865bc6755f0 · outbound

This paper cites The scores are the average per- formance of all samples in val, test, testmini splits.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? The scores are the average per- formance of all samples in val, test, testmini splits

Reference 84

Resolution
malformed identifier
raw_fallback, observed 2026-08-16T10:22:53.003484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T10:22:52.684358Z digest=sha256:38355c1bc897340ce77d2074acf3ee464763f1696974b2d4ffff61c6e120f7ce

Observation ea35b337-831a-47a2-a900-74d0eeac204d · outbound

This paper cites question n Give an answer with this format: <ans>ANSWER</ans>, no redundant words. For example: <ans>A</ans>.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? question n Give an answer with this format: <ans>ANSWER</ans>, no redundant words. For example: <ans>A</ans>

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:52.992279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T10:22:52.688258Z digest=sha256:87cc9e896a517097d58e24258826e3f907601071b9ba531d3681bff61f6cd596

Observation 4b9f70eb-5177-4308-8a37-d78e5fd7e01a · outbound

This paper cites We compress the images to display them in the paper.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? We compress the images to display them in the paper

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:52.980470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T10:22:52.691595Z digest=sha256:a8e74642a010fb41ea951d5950f4f341d209b7c60c805bfd911f86d94ddc2c4a

Observation a0c9f485-4bef-49bc-b1b0-d659d904e87f · outbound

This paper cites Geological Survey data release, 2022.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Geological Survey data release, 2022

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.214714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T10:22:52.634550Z digest=sha256:1e93f01a8ff3134aadce40901ddde744027d1915d89177715b8ce72b0ba92ec4

Pith citing papers

No inbound Pith citation observations are available.