Pith. sign in

Paper Citation Record · LEDGER

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding?

As of 20 August 2026, this Paper Citation Record lists 84 of 84 outbound references and 0 inbound Pith citation observations for arXiv:2504.18406.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.18406 v2

Coverage vector

measured 84 of 84 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:22:52.691595Z

measured 84 of 84 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

84 of 84 outbound references displayed

  • verified exact0
  • verified fuzzy51
  • unresolved32
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f9b1899d-2417-4785-8d5f-28b113856ef8 · outbound

This paper cites 12 Kucherlapati Raju 13, Genome data analysis: Baylor College of Medicine Creighton Chad J.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? 12 Kucherlapati Raju 13, Genome data analysis: Baylor College of Medicine Creighton Chad J

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T10:22:52.389235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:22:52.389235Z digest=sha256:734da4b3bff79e3083dcff2a291e08da6064463840821520d81166e219137d8c

Observation 4d0c885f-e035-4a56-9bb4-5492d0030c35 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T10:22:52.393488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:22:52.393488Z digest=sha256:42eef140f57aea601b47b95a145db9ddfc195f19227c04bfc58d81b0dc3f12ee

Observation 1175f303-e6a7-4395-97ef-3c9cd122ee9e · outbound

This paper cites GPT-4 Technical Report.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? GPT-4 Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T10:22:52.397322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:22:52.397322Z digest=sha256:6b5b6c70432c9b916b6a85410b6602f807758b87d76a88526ba0cc57b3bdaab1

Observation cbc5567b-d8f4-429e-9f34-44a9c47784e2 · outbound

This paper cites The claude 3 model family: Opus, sonnet, haiku.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? The claude 3 model family: Opus, sonnet, haiku

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T10:22:52.401320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:22:52.401320Z digest=sha256:84cbba4ab6174906c01da7a12fc07bae939b801f3734ce8559345a55a8787ffb

Observation 15c957c5-8cd7-4cd8-bd1a-b94bde5faa64 · outbound

This paper cites Bach: Grand challenge on breast cancer histology im- ages.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Bach: Grand challenge on breast cancer histology im- ages

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T10:22:52.404867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:22:52.404867Z digest=sha256:b7153b5a7faad81a0ea8837dc43f4bf1d48a0f7dbb60948bc3bbb7fc4a818b93

Observation dd92b8e1-1e15-4332-b442-9ff0bc14226d · outbound

This paper cites Jauregui, and Juan Andr ´es Cardoso.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Jauregui, and Juan Andr ´es Cardoso

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T10:22:52.408735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:22:52.408735Z digest=sha256:8855e1cb503d59eaa3e06e15a0e07c784cf05d5d11f175978d20a170d16e78c3

Observation d3a4f4f7-944b-474c-995c-f7149645fcae · outbound

This paper cites Ef- ficient high-resolution deep learning: A survey.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Ef- ficient high-resolution deep learning: A survey

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.682253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:22:52.412513Z digest=sha256:ece1c97595e9fab3a90addaff731c6ae6d05e8332f88439c0db358e8e64a7537

Observation 19bb4625-5def-4fbf-9d05-ce0b6ba78736 · outbound

This paper cites Effi- cient high-resolution deep learning: A survey.ACM Comput.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Effi- cient high-resolution deep learning: A survey.ACM Comput

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.672630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:22:52.416326Z digest=sha256:d694d33a967fd6e772bb1a1ea1ae5add111d54b4af743f4885ff0ca74b8e7a55

Observation e944eaa2-510a-49db-9955-3c21580f39a6 · outbound

This paper cites Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T10:22:52.419841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:22:52.419841Z digest=sha256:63412a6f6e7b65f317ece799e9dc44887299423a704aa2f0215c13846e4f3db6

Observation 888b5071-0d94-441c-bd07-3b0a4ffc64c3 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T10:22:52.423560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:22:52.423560Z digest=sha256:bfd8c9134f888ddfdcde95f3a1a25042b31dc4d7076abaf6659e854c25037065

Observation e6070e0a-7b27-4f7e-a792-7a5bf964a236 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T10:22:52.427409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:22:52.427409Z digest=sha256:95898c4c5124473a1b036510f1a1d7cc00f82f7f1745c93c4fd3d2a6feb84405

Observation e97359c0-4da3-485b-8341-aa52329fbaa5 · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.663099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:22:52.431192Z digest=sha256:13b76feb21745f423698dd49ef869561bf04659a7c63309b16296be747f32d83

Observation 16647775-f89e-440e-b7ef-ec1659f534dd · outbound

This paper cites Functional map of the world.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Functional map of the world

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.652996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:22:52.434724Z digest=sha256:33b30f4dd4cf630fe6310ac661b382ecd01709d3949670494198441ef63119a1

Observation c6c8c29c-eb5e-4f8d-89e7-bfb21486adfd · outbound

This paper cites The cityscapes dataset for semantic urban scene understanding.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? The cityscapes dataset for semantic urban scene understanding

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.642930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:22:52.438441Z digest=sha256:5103f0c2ad6bb8a70278360b486289e9a17b686928989c03ac95d1d5aa8d79e0

Observation 57a333d4-255e-4b52-86fb-ddd5aea2bf8f · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T10:22:52.441852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:22:52.441852Z digest=sha256:6a69ddc855b59a20bb69a635bca19362164fa6ef95ccee2703dd831988c3a64f

Observation e34ee8a0-b7f5-476e-b3ed-e84af2704903 · outbound

This paper cites LungHist700: A dataset of histological images for deep learning in pulmonary pathology.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? LungHist700: A dataset of histological images for deep learning in pulmonary pathology

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.632512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:22:52.445363Z digest=sha256:617c8089aa85871fcbf8b333f3b5c54d5a8ad0fd77c4109b0790a7d9b9281e23

Observation b299ab0e-d64d-4bbd-87ba-8cadd9c49f01 · outbound

This paper cites LungHist700: A dataset of histological images for deep learning in pulmonary pathology.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? LungHist700: A dataset of histological images for deep learning in pulmonary pathology

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.621966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:22:52.448580Z digest=sha256:d400229bac2ea174ee39a2737139073908fef85ec821733e16a6af29052fb7eb

Observation ab042208-0d94-4f79-8ecc-e81ad17c7e34 · outbound

This paper cites Internlm-xcomposer2-4khd: A pioneer- ing large vision-language model handling resolutions from 336 pixels to 4k HD.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Internlm-xcomposer2-4khd: A pioneer- ing large vision-language model handling resolutions from 336 pixels to 4k HD

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.611682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:22:52.451866Z digest=sha256:742d0fcb5b0df684e90163602b7a6039ecddbd0cf83fa6223277ce580295ce77

Observation a263f2d1-568b-4dbf-b65f-3d6ffda9c112 · outbound

This paper cites The Llama 3 Herd of Models.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? The Llama 3 Herd of Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T10:22:52.455103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:22:52.455103Z digest=sha256:ac7e37a6f1610cd864d5d0cccefc6516ef777f7c15b8324eb79f344022811fce

Observation 3543c1bb-b91e-41ad-a5a0-6f8843f862db · outbound

This paper cites Describing differences in image sets with natural language.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Describing differences in image sets with natural language

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.601042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:22:52.458785Z digest=sha256:b688fc13e0f3716b3c7de5897ad00ae606e5e39199ca77c997a217f4557b2d7c

Observation f65c327e-f379-4aad-89a6-efc6604d2669 · outbound

This paper cites Floorplancad: A large-scale cad draw- ing dataset for panoptic symbol spotting.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Floorplancad: A large-scale cad draw- ing dataset for panoptic symbol spotting

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.590413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:22:52.462320Z digest=sha256:73e56e9f5c4681e4f882e8b7bbb7ffb4163900fd8e3ba33de468c86012ddfabd

Observation b33ecc19-b636-4868-807b-f9f156c0d2a3 · outbound

This paper cites Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T10:22:52.465714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:22:52.465714Z digest=sha256:2521c67d77e85b0abcdf1b3a20c958fcd85a4e4784c2638c3776820c9305ff73

Observation 67467734-d1fc-46cc-8cea-14e87c0e9ea4 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.579762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:22:52.469307Z digest=sha256:e65b378b8d0baeabf16131b42843fcf901659460f15fcae31fa399d3377999c9

Observation 5dae7f81-ff1f-428d-8965-120282bb8756 · outbound

This paper cites Llava-uhd: An LMM perceiving any aspect ratio and high- resolution images.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Llava-uhd: An LMM perceiving any aspect ratio and high- resolution images

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.569489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:22:52.472689Z digest=sha256:d172e67088dbdf0d86a5e8833c9263d8f5e414fa15b29c5f993c595a05b356cf

Observation a7cb86e0-1b89-49ee-8103-6ee52159a44f · outbound

This paper cites Cogagent: A visual lan- guage model for GUI agents.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Cogagent: A visual lan- guage model for GUI agents

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.558847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:22:52.476218Z digest=sha256:6ef3291c6d76c805c64d97505bc4b007b905c715513d4ff610aab0060df965cd

Observation e71e3d52-d4ac-49ac-bb8c-d59361ed0628 · outbound

This paper cites mplug-docowl 1.5: Unified structure learning for ocr-free document understanding.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? mplug-docowl 1.5: Unified structure learning for ocr-free document understanding

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.548052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:22:52.479535Z digest=sha256:978c37602a64eeb9d8ff12adb14467bc4d1442bcf685c51d5adcabe947d61ff2

Observation c82a385b-598a-4c8f-87ef-8d3fcbdbcdda · outbound

This paper cites mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T10:22:52.483040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:22:52.483040Z digest=sha256:ccf6e4f28f28c3e279154b8cd497afab01073929d37d56dd5c39bfaca678bdab

Observation 694fe1a0-1e68-4e9d-a306-5ab338b52b13 · outbound

This paper cites Novachart: A large- scale dataset towards chart understanding and generation of multimodal large language models.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Novachart: A large- scale dataset towards chart understanding and generation of multimodal large language models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.538004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:22:52.486969Z digest=sha256:8f963825554ce639d22ac91dd9798634b432d97f21b5bea43a7858c0238f1bf4

Observation 320bb115-4ebc-421b-a50b-ed91ba62d796 · outbound

This paper cites High reso- lution image quality database.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? High reso- lution image quality database

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.527667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:22:52.490267Z digest=sha256:a0952b57aae6fb09aeb437493c71d90258a752d31e759078e7a34442429ec8dc

Observation dbd2eaae-922f-4a31-b77a-d201cdb248a9 · outbound

This paper cites Mini-monkey: Alleviating the semantic saw- tooth effect for lightweight MLLMs via complementary im- age pyramid.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Mini-monkey: Alleviating the semantic saw- tooth effect for lightweight MLLMs via complementary im- age pyramid

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.517732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:22:52.493634Z digest=sha256:8237244d6dad958525efd53ca5b064e536b5536378cae7659f41c2bd29044794

Observation a3c99a96-fffb-4241-a0cb-e6cceaf6ca5c · outbound

This paper cites The apolloscape dataset for autonomous driving.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? The apolloscape dataset for autonomous driving

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.507439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:22:52.496876Z digest=sha256:6b0b5cd2fc7d6a8cb2aab7f5bcf8591aa1578b656acd0ebb58762a80cdf3a55b

Observation f5021136-94e3-4344-99ca-29f1255b301d · outbound

This paper cites GPT-4o System Card.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? GPT-4o System Card

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T10:22:52.500288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:22:52.500288Z digest=sha256:f1e4dfe14a9e47cf355cd5d0b839e702cb6998ca574a8599cba5ea0cbcd0aaf6

Observation 47e467f5-93b5-46f2-8aa4-3c00497d8f1d · outbound

This paper cites Multi-source multi-scale counting in extremely dense crowd images.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Multi-source multi-scale counting in extremely dense crowd images

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.496564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:22:52.503899Z digest=sha256:81e8557bd001ead960bdecf0087b5872cf70218fd33b2fa4b1bd3c99880e9942

Observation 7771275f-05a5-40d4-b808-9a8c39b60bd7 · outbound

This paper cites Composition loss for counting, density map estima- tion and localization in dense crowds.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Composition loss for counting, density map estima- tion and localization in dense crowds

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.486389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:22:52.507778Z digest=sha256:81572f1b666b4050f5e4eafeb3ce59da6b048d7fea9d0ef1a3fe052bf4bc58cb

Observation bbdc62ea-4066-464f-9852-4d5b711ee143 · outbound

This paper cites CosmoCLIP: Generalizing Large Vision-Language Models for Astronomical Imaging.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? CosmoCLIP: Generalizing Large Vision-Language Models for Astronomical Imaging

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T10:22:52.511131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:22:52.511131Z digest=sha256:b24007338813ee23a867d39b18d723e663d1f97d931b879be44da64cf620ee87

Observation af02aa70-aabd-4317-b138-8e2890980087 · outbound

This paper cites MMAD: A Comprehensive Benchmark for Multimodal Large Language Models in Industrial Anomaly Detection.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? MMAD: A Comprehensive Benchmark for Multimodal Large Language Models in Industrial Anomaly Detection

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T10:22:52.515246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:22:52.515246Z digest=sha256:b76cdc4a3ee9db0fcaadf113fc4e1fde7500a2413c55bd75d044ddc969470c0a

Observation f97d40ff-5c0e-4917-8a81-08ae1b554f3b · outbound

This paper cites A diagram is 10 worth a dozen images.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? A diagram is 10 worth a dozen images

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.475510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:22:52.518870Z digest=sha256:51458dc1e8398b6063500734e6e9ab5354d198868c547995dc1a74cdb6144c78

Observation ef01dd84-2979-419b-b91b-f2006978501c · outbound

This paper cites Berg, Wan-Yen Lo, Piotr Doll ´ar, and Ross B.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Berg, Wan-Yen Lo, Piotr Doll ´ar, and Ross B

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.463325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:22:52.522188Z digest=sha256:8948a3e0c6c297616515c9f144dce12c2c6ad3e4c51300f28b696e134d35db9d

Observation 652cbbda-5d7a-4ae4-b779-259e86ac5003 · outbound

This paper cites an unresolved cited work.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-16T10:22:53.452609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:22:52.525350Z digest=sha256:6132d8e780dfb7587cc9cfa4fbc692bcdafea297fae14f7a87d19a15f5f744d1

Observation ce38ad38-cf11-496e-b34a-23bafbf7087c · outbound

This paper cites A dataset of clinically generated visual questions and answers about radiology images.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? A dataset of clinically generated visual questions and answers about radiology images

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.441274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:22:52.528830Z digest=sha256:e081928cd22b242c65abef84dd1d51d493e2568f46c7fdae608e8f8c5e6ead70

Observation dd87e033-17dd-40d5-a7ac-b5ca145bb28e · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation, 2022.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation, 2022

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.431021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:22:52.532223Z digest=sha256:094f864cd368c28aa50e56a35ad7fca051dd82cd6e38868d0af6367d6c10be89

Observation 090ce644-3b98-469d-9077-51d19d9bb127 · outbound

This paper cites Hrvqa: A visual question answering benchmark for high-resolution aerial images.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Hrvqa: A visual question answering benchmark for high-resolution aerial images

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.420747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:22:52.535599Z digest=sha256:3de4174058eeef9c2821dc61539b7c31c46a6ed8ceb34491fa11d3018213c28d

Observation 5ea0ebcb-a443-4f56-bce3-67b7cd2542a7 · outbound

This paper cites To- wards streaming perception.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? To- wards streaming perception

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.410588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:22:52.538872Z digest=sha256:eac724d9571c48072cb21ddc308b7d8d6dc2721d86762d0ecdf8d2dc198ee41d

Observation e02da5d8-5072-4c2c-b1f7-49c0d86fa976 · outbound

This paper cites Mon- key: Image resolution and text label are important things for large multi-modal models.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Mon- key: Image resolution and text label are important things for large multi-modal models

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.399929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:22:52.542538Z digest=sha256:bc0c203928125c7acd8cc95da01e65540e1100059c89ef9a7221a9ebe84d4845

Observation 8834e4b0-29c1-49bb-ac17-2cf49e132b87 · outbound

This paper cites The ArtBench Dataset: Benchmarking Generative Models with Artworks.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? The ArtBench Dataset: Benchmarking Generative Models with Artworks

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-16T10:22:52.545955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:22:52.545955Z digest=sha256:4a7b6982166f357739b000734e4675afaa8182a54d6412cd120ba9cbf82e4e99

Observation 702d130a-c5fd-47a4-bb23-c923f4826bf1 · outbound

This paper cites Visual instruction tuning.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Visual instruction tuning

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.389280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:22:52.549625Z digest=sha256:714ca51eeba7de09589d83db684e1f14181099f040e4e8f763015182360b82ef

Observation b662ac6b-21e7-47e4-91a6-3b735978b878 · outbound

This paper cites Visual instruction tuning.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Visual instruction tuning

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.378720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:22:52.552983Z digest=sha256:da63c65fadf9ef53ba7940c7dce5fadab4e0a8bf2908f7d835543323ac567d7a

Observation 5c0f2fe2-da3d-4828-b7db-7357e9005fd2 · outbound

This paper cites Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.368266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:22:52.556413Z digest=sha256:f6c27577cc5f689328f9158c599ba2e9d397e9a8dc9a87ac6ed5035ea26f0e75

Observation f5d8e1cc-4b44-4e31-8afd-14b7e3d78322 · outbound

This paper cites A convnet for the 2020s.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? A convnet for the 2020s

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.357129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:22:52.560005Z digest=sha256:ea496fd4919fc921158321d684c069d22b2087582814a9d8479d295321e928cd

Observation f3b305db-fadf-4584-b15e-fa91c6b88242 · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-16T10:22:52.563373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:22:52.563373Z digest=sha256:a9e6d47bdfe004587436aa30a1537b5a2cfb77c0b583df58f61d792ef4341123

Observation 997a93e9-ecbe-44a5-9a1d-bf4f0c1d7d36 · outbound

This paper cites Feast your eyes: Mixture- of-resolution adaptation for multimodal large language mod- els.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Feast your eyes: Mixture- of-resolution adaptation for multimodal large language mod- els

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.344940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:22:52.567069Z digest=sha256:94486bcfedd52dc2b4625842d144bdb1cd70d422e4f13c945334d38f6d3ff799

Observation 432380a8-118f-4daf-b63d-c7788b99631b · outbound

This paper cites Infographicvqa.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Infographicvqa

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.333675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:22:52.570808Z digest=sha256:40f9dbb1de074680752c9b708a2847f020894dfa33e30b765f32f7a2b80ddf5d

Observation 304adae2-21fe-4aed-8c1e-cadeca871c33 · outbound

This paper cites MM1: methods, analysis and insights from multimodal LLM pre-training.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? MM1: methods, analysis and insights from multimodal LLM pre-training

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.322446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:22:52.574251Z digest=sha256:5ec86dc1aa9afe363f16224cd1be6b08af24ce946aff336e7ac92d9e5f758369

Observation 36d208a9-8be8-4387-b4e5-9c87de2704f5 · outbound

This paper cites The mame dataset: on the relevance of high resolution and variable shape image properties.Applied Intelligence, 52(10):11703–11724, 2022.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? The mame dataset: on the relevance of high resolution and variable shape image properties.Applied Intelligence, 52(10):11703–11724, 2022

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.311099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:22:52.577807Z digest=sha256:5874319c4f777234e65deb5cfcbb2e3e8470123bea736306503c078286ac00af

Observation 6d458e48-6f4a-4836-86b6-9d15bbb56e64 · outbound

This paper cites Learning transferable visual models from natural language supervision.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Learning transferable visual models from natural language supervision

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.300422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:22:52.581521Z digest=sha256:b41f4a6e44ccedc3b157b66bcadb8a1cc2d795379e352cf36dfdaf30b7e51b44

Observation 418f9234-1f9a-48e7-b5ec-48bed3326402 · outbound

This paper cites The synthia dataset: A large collection of synthetic images for semantic segmentation of urban scenes.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? The synthia dataset: A large collection of synthetic images for semantic segmentation of urban scenes

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-16T10:22:52.584823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:22:52.584823Z digest=sha256:896c16386a3c995e74ff8c64f68fe366b68a48940c0b34ab4ea9bb5b8f036ac1

Observation 6dbd83c3-87e4-46ae-bc16-6a5df342a570 · outbound

This paper cites Jhu-crowd++: Large-scale crowd counting dataset and a benchmark method.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Jhu-crowd++: Large-scale crowd counting dataset and a benchmark method

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.281672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:22:52.588328Z digest=sha256:15ce7842341836c4e8c1f44c5dad1d0b64c47680d389b89e6aae839188a9347e

Observation 2effe67c-8df2-4ba9-be1d-f9082195663a · outbound

This paper cites MileBench: Benchmarking MLLMs in Long Context.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? MileBench: Benchmarking MLLMs in Long Context

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-16T10:22:52.591664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:22:52.591664Z digest=sha256:d04ac40feb55737e7f33a1780b0e88b467cf8d30ce8f98174222aa449df722ed

Observation 2249752f-28cf-48b5-afd6-ee52df310d5b · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Gemini: A Family of Highly Capable Multimodal Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-16T10:22:52.595291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:22:52.595291Z digest=sha256:c6ef8da95c8b5077a93ea52b286e3ace33742a28000c3f9e8d75aaaf32798c27

Observation 0073b7ca-a1b8-42e1-bd00-913a855c079c · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-16T10:22:52.599306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:22:52.599306Z digest=sha256:6b86f0b1595840d899ae2e7cb9f61e8415568a8e4100d37c43232dfa436e62ef

Observation e8d37346-5381-4628-8b22-f894d82d57e9 · outbound

This paper cites Predicting breast tumor proliferation from whole-slide im- ages: the tupac16 challenge.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Predicting breast tumor proliferation from whole-slide im- ages: the tupac16 challenge

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.270358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:22:52.602880Z digest=sha256:7421bc5b4201ccc2707f0a89efb276ac7d801390a10a8fb175575796a6a88846

Observation 3581a251-a555-4e0f-81f8-bd8810a09e88 · outbound

This paper cites MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-16T10:22:52.606264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:22:52.606264Z digest=sha256:770e7f7cbeadd17ac187a43363d34b1816674988de2d74d49a67b3568b80e0ee

Observation 8038171f-0e4d-4270-88fd-fe9b7f7436c3 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-16T10:22:52.613989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:22:52.613989Z digest=sha256:7c9bc1927cf0d37848776d15f3036737fa4cb808dbda8ee43d54c0e77eb0db58

Observation 99e3a955-7b77-43ed-a0d4-4de8baafce46 · outbound

This paper cites Divide, conquer and combine: A training-free framework for high-resolution im- age perception in multimodal large language models.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Divide, conquer and combine: A training-free framework for high-resolution im- age perception in multimodal large language models

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.258765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:22:52.617535Z digest=sha256:23112983788f8a719d30292f137a6566cb492c9c2adcbec3347af1f4f19e8cc6

Observation 99dd5e4f-374c-4d32-bc37-8bb84f7c2773 · outbound

This paper cites Divide, Conquer and Combine: A Training-Free Framework for High-Resolution Image Perception in Multimodal Large Language Models.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Divide, Conquer and Combine: A Training-Free Framework for High-Resolution Image Perception in Multimodal Large Language Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-16T10:22:52.620905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:22:52.620905Z digest=sha256:d7373dce28f4a903a452b4a04606aa222ec28dbda11928117262a0900a6d52ca

Observation b4cab3b4-c321-46b0-9e4b-28dae3ee7bd5 · outbound

This paper cites Needle in a multimodal haystack.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Needle in a multimodal haystack

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.247840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:22:52.624325Z digest=sha256:1c2cc3ec9bc6c59dfcd6b097bfb89023016020c53321d4995d41f43cadb225c5

Observation eef52f33-274d-440e-b097-3312e6487ebe · outbound

This paper cites Panda: A gigapixel- level human-centric video dataset.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Panda: A gigapixel- level human-centric video dataset

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.236547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:22:52.627940Z digest=sha256:8c45745b1188f82dd42731d4c828dfd19dc666ea90987158bea826c3ccbc8258

Observation 57526d2d-fb65-4948-b8c8-6d0d0a0fe354 · outbound

This paper cites Weiser, P.L.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Weiser, P.L

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.225206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:22:52.631117Z digest=sha256:07752d4f7407a6c238f314e28cb1d7f01ee71f729e3cbd9314d7823315fb20d2

Observation 2c891846-3d79-4421-aae7-cdf56950113c · outbound

This paper cites V?: Guided visual search as a core mechanism in multimodal llms.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? V?: Guided visual search as a core mechanism in multimodal llms

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.203540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:22:52.637878Z digest=sha256:3dae0c9811da4f7f8644d38836b39dfaad824f2fb0a02cb9f7d98b99ae1b71b7

Observation d38f6978-372a-4e46-914f-b76d9b486cff · outbound

This paper cites Visual Haystacks: A Vision-Centric Needle-In-A-Haystack Benchmark.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Visual Haystacks: A Vision-Centric Needle-In-A-Haystack Benchmark

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-16T10:22:52.641231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:22:52.641231Z digest=sha256:51b71a192399fb85ca93e0e9f3980f2f9bd6f960d7ca347a387fffdcb9af3a89

Observation 88fce6b9-61f5-46f8-b8c4-77343d86eccb · outbound

This paper cites DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-16T10:22:52.648406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:22:52.648406Z digest=sha256:37a33136ad4e16214dd243a2bc8753dadef108a984af6da406d02047f2d58987

Observation d8fdb897-c59f-4dad-b020-0499b7480244 · outbound

This paper cites Deep learning for detecting corona virus disease 2019 (covid-19) on high-resolution computed tomography: a pilot study.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Deep learning for detecting corona virus disease 2019 (covid-19) on high-resolution computed tomography: a pilot study

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.192182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:22:52.651683Z digest=sha256:e28e930cf113b9b13a0d95d4807924a82fefaa8a8014c61ba8428653a65c3554

Observation 319de3f4-0f78-4693-ad79-70d0235c154c · outbound

This paper cites Bdd100k: A diverse driving dataset for heterogeneous multitask learning.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Bdd100k: A diverse driving dataset for heterogeneous multitask learning

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.181330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:22:52.655316Z digest=sha256:f0839000c44a16b9b797727b5d90dff58920c9c13b721e880f3493a97e05c3d4

Observation a3770264-13b9-43f9-a15f-822ebf6caee3 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for ex- pert agi.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for ex- pert agi

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-16T10:22:52.659019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:22:52.659019Z digest=sha256:0f1b428b141b0b4f9f836d9c61e1bde3e71313cd8b16ab19e9f98d277f06001c

Observation a6e95b13-d20a-4042-9b3e-21ee90b4f30c · outbound

This paper cites Single-image crowd counting via multi-column convolutional neural network.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Single-image crowd counting via multi-column convolutional neural network

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.162943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:22:52.662419Z digest=sha256:9a8fe8f9554a01d23fa6170ac22175a210e49d36384c653fa735e5ff62590571

Observation 4ff3db0e-5e4a-45ea-8d17-2c8a00b98a54 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-16T10:22:52.665965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:22:52.665965Z digest=sha256:4f1902f6a6b9a8ebf3ab35190d954fe725b8b176f82e820fe1297947f2d44c0e

Observation ce0ae181-b1cf-4695-9ff0-773db659f6a4 · outbound

This paper cites MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-16T10:22:52.669629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:22:52.669629Z digest=sha256:fca26d0d295f43aa4de7e5ae1445935ee760b1b6c1aec25f9f24bdbd95edb5ea

Observation 6ad8a1ae-f2b4-4ff2-accb-87d5da2029f5 · outbound

This paper cites Monitoring Extracted from MME-Realworld, this dataset features images taken from public safety cameras in diverse scenarios.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Monitoring Extracted from MME-Realworld, this dataset features images taken from public safety cameras in diverse scenarios

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.151619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:22:52.673114Z digest=sha256:c30832af6654d26c22ed3210765a03b514701df572a1b76dbe5c7d3b9c11146a

Observation f10ccd4a-4de5-4ce5-a358-e5474c129936 · outbound

This paper cites an unresolved cited work.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Unresolved cited work

Reference 82

Resolution
unresolved
raw_fallback, observed 2026-08-16T10:22:53.139626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:22:52.676865Z digest=sha256:78c919bf0ed2c7d3a7076878579ca648d0f9d51f3820e9782f4a62902dd28db3

Observation c18d8af0-22b6-4773-8268-b177c9265cec · outbound

This paper cites Phi- 3.5 [2] is a lightweight model designed for efficient lan- guage understanding and generation.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Phi- 3.5 [2] is a lightweight model designed for efficient lan- guage understanding and generation

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.128132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:22:52.680145Z digest=sha256:d81c96eb1ba81b639b1dba5d64b80746174a8b8881b9089ab81504aa2df7093d

Observation 4439af5e-c13a-4957-b8fe-e865bc6755f0 · outbound

This paper cites The scores are the average per- formance of all samples in val, test, testmini splits.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? The scores are the average per- formance of all samples in val, test, testmini splits

Reference 84

Resolution
malformed identifier
raw_fallback, observed 2026-08-16T10:22:53.003484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:22:52.684358Z digest=sha256:b56f94404c31008ffdca574f36f1656f67b76dd1468069860f40d29da959d418

Observation ea35b337-831a-47a2-a900-74d0eeac204d · outbound

This paper cites question n Give an answer with this format: <ans>ANSWER</ans>, no redundant words. For example: <ans>A</ans>.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? question n Give an answer with this format: <ans>ANSWER</ans>, no redundant words. For example: <ans>A</ans>

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:52.992279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:22:52.688258Z digest=sha256:5eea3115fcfe2e40cadcc21c0daa63a048bf1b41743469386cfe662e99976ee2

Observation 4b9f70eb-5177-4308-8a37-d78e5fd7e01a · outbound

This paper cites We compress the images to display them in the paper.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? We compress the images to display them in the paper

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:52.980470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:22:52.691595Z digest=sha256:a265beb8787586cdd792ec91a709938b80e60a833af16fe7762cf78fc55c0b65

Observation a0c9f485-4bef-49bc-b1b0-d659d904e87f · outbound

This paper cites Geological Survey data release, 2022.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Geological Survey data release, 2022

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:22:53.214714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:22:52.634550Z digest=sha256:3bfbc9ee1e2096ac269bc4dc090cd5999670d1ad351f5cddf58635ab1f6299a2

Pith citing papers

No inbound Pith citation observations are available.