Pith. sign in

Paper Citation Record · LEDGER

Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models

As of 19 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 2 inbound Pith citation observations for arXiv:2501.14276.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.14276 v1

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T15:18:49.974575Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:51:42.830659Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T11:38:38.471219Z

Reference resolution

48 of 48 outbound references displayed

  • verified exact0
  • verified fuzzy34
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b953dbc5-4c2c-4380-ab0a-4568c43acca0 · outbound

This paper cites Visual instruction tuning.

Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models Visual instruction tuning

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:18:50.729320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T15:18:49.624071Z digest=sha256:552e26938949d162a23295bfad67c33e07210c3f64575f714cc99ce1c112bade

Observation 793e77b5-ae36-4a4e-8708-d4d4d1ad1c04 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:18:50.714256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T15:18:49.629706Z digest=sha256:68d75b110d986fd8d186010dacacb37a84f8bee61af305db96d3f65322950422

Observation 06c17ce1-689b-4fa8-a059-5500e98a098d · outbound

This paper cites Sharegpt4v: Improving large multi- modal models with better captions.

Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models Sharegpt4v: Improving large multi- modal models with better captions

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:18:50.699395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T15:18:49.635118Z digest=sha256:0c23683cf7d8671911664c0a1c819c973e52221d87418d7be45ccb5832d31457

Observation a2eaa09e-434f-4f68-b289-9f04a195ccdd · outbound

This paper cites Monkey: Image resolution and text label are important things for large multi-modal models.

Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models Monkey: Image resolution and text label are important things for large multi-modal models

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:18:50.683185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T15:18:49.640278Z digest=sha256:9d7939a56bc74a86e30eabbccf8de06ef9455be35bd36329be441c22d620f7d0

Observation 4c8928aa-2cb5-4c60-94ad-ace2a76da0d9 · outbound

This paper cites InternLM-XComposer2-4KHD: A pioneering large vision-language model handling resolutions from 336 pixels to 4k HD.

Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models InternLM-XComposer2-4KHD: A pioneering large vision-language model handling resolutions from 336 pixels to 4k HD

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:18:50.667468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T15:18:49.645187Z digest=sha256:40c902f4523c1c68f539bb1760bf0c923880faa3490e7658d5627d8388d59c7a

Observation c450716a-35f9-47b7-9d62-b9aa6d260320 · outbound

This paper cites Salgan: Visual saliency prediction with adversarial networks.

Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models Salgan: Visual saliency prediction with adversarial networks

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:18:50.651828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T15:18:49.650579Z digest=sha256:d432011b44c12f7e009ad07db4997542fc73524f519074bf39e549b0beab2cab

Observation 83e6ff61-f14e-45ea-8bab-98341f2d92cb · outbound

This paper cites PaLI: A jointly-scaled multilingual language-image model.

Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models PaLI: A jointly-scaled multilingual language-image model

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:18:50.635363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T15:18:49.655951Z digest=sha256:7c2a07780f7fc4fdb5955f47458aa810cdac1a27dc9704e576e4c26d05c1fc6f

Observation 341bd568-5bef-4ff7-9209-2e8df930f74b · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T15:18:49.661345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:18:49.661345Z digest=sha256:1b1626628d561c01dff9ff634224ece90214b317d735dc000fe2130bfbbde502

Observation 4ac08dd5-2ae7-486a-b65d-a0f9052462d9 · outbound

This paper cites Cambrian-1: A fully open, vision-centric exploration of multimodal LLMs.

Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models Cambrian-1: A fully open, vision-centric exploration of multimodal LLMs

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:18:50.619723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T15:18:49.666160Z digest=sha256:5d7bc16e627efc6ed5367a08da80e8cb3dd6cfe1236952e5cfcf3158d511170e

Observation 0b98434d-d5c4-473f-ab85-df6b0825bfaf · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge, 2024.

Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models Llava-next: Improved reasoning, ocr, and world knowledge, 2024

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T15:18:49.670527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:18:49.670527Z digest=sha256:c8744f9afbe53fed6e8fb4ed5c29234b48fdae3d616ad58f9d6702fd0ddf40df

Observation 40ac4578-9d36-44c6-bc74-9aac4380cf8c · outbound

This paper cites How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites.

Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T15:18:49.674802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:18:49.674802Z digest=sha256:d5faa0b8da444437a4b11289b32f3675286cc491b1f2a1b8f9ac6e1c382a75ce

Observation 1365c077-116c-4169-a8d4-ffc36c18a8c9 · outbound

This paper cites Instruct- blip: Towards general-purpose vision-language models with instruction tuning.

Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models Instruct- blip: Towards general-purpose vision-language models with instruction tuning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:18:50.584820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T15:18:49.679115Z digest=sha256:b3f8edbf1be6d1f9a12e623231e9e14fef0d895c2a6105d37e1d3a1331782282

Observation 695af14c-d233-47bc-b4bd-f3520a045059 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models Learning transferable visual models from natural language supervision

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:18:50.569728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T15:18:49.683033Z digest=sha256:8f0bb3ba7f264420ff2bb78e22bc096ce1999021e689fa98822cc29dffcc0cd2

Observation 1602eccf-23fe-4268-b4c2-4910c6fce142 · outbound

This paper cites Dual modality prompt tuning for vision-language pre-trained model.

Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models Dual modality prompt tuning for vision-language pre-trained model

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:18:50.554579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T15:18:49.686957Z digest=sha256:b677e8d473f6b3fe8272b31025f75a9f954b94ef6509d9bf718d211870d8579b

Observation c149faca-a31e-419f-903f-fbdc5d9bfb1c · outbound

This paper cites Sgva-clip: Semantic-guided visual adapting of vision-language mod- els for few-shot image classification.

Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models Sgva-clip: Semantic-guided visual adapting of vision-language mod- els for few-shot image classification

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:18:50.539960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T15:18:49.690868Z digest=sha256:04ce697087584be1a2472a1fd5c0161edabe82d6d3e56ba316f60d33b91a6756

Observation f3181c5a-fa3f-4cb3-afc8-eff3fb1e2313 · outbound

This paper cites Gpt4ego: Unleashing the potential of pre-trained models for zero- shot egocentric action recognition.

Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models Gpt4ego: Unleashing the potential of pre-trained models for zero- shot egocentric action recognition

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:18:50.524452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T15:18:49.695889Z digest=sha256:8bc957cbfab6383386026cddd85bb1ce7601071740c986d39ce9ac79589c22a4

Observation c5de4ea7-eb2f-442a-bde9-424ba3531185 · outbound

This paper cites mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding.

Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T15:18:49.701778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:18:49.701778Z digest=sha256:3ea740d434e9f3b94863558981602f6a16c74afd4fd9ee8563a482499732f5cb

Observation 8ca60c17-f493-44a1-8a96-f9d0d7887b4b · outbound

This paper cites InternLM2 Technical Report.

Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models InternLM2 Technical Report

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T15:18:49.706713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:18:49.706713Z digest=sha256:0e0da9f95e24c51385168750e5068a7e8d91c5e5790d291b678147ec3a1bf7c7

Observation f95fd1fc-71b6-444c-bea8-964308349d3e · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T15:18:49.711643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:18:49.711643Z digest=sha256:a29304a7e9c3bdf17acd4bed95eabc432b9516d1cae4b9a6e52d5cff394c9ddf

Observation 3a553936-ea40-47d6-8922-850dd0ea4399 · outbound

This paper cites Ocrbench: on the hidden mystery of ocr in large multimodal models.

Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models Ocrbench: on the hidden mystery of ocr in large multimodal models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:18:50.510175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T15:18:49.717726Z digest=sha256:ac26ea4be35456ffd04c09b9df21d8ae24988bd8481bcf64e8c93034c44428f1

Observation 90c802b5-471a-4462-b9a4-4c44ce254abf · outbound

This paper cites Towards vqa models that can read.

Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models Towards vqa models that can read

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:18:50.496165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T15:18:49.722163Z digest=sha256:d3ccba30abd0e31c40de3a5ace58300b09d104229d6471925c263dce75e8631c

Observation fc766ca9-d6d0-4117-a0b4-122ad133a92f · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? In Proceed- ings of the European Conference on Computer Vision (ECCV) , pages 216–233, 2025.

Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models Mmbench: Is your multi-modal model an all-around player? In Proceed- ings of the European Conference on Computer Vision (ECCV) , pages 216–233, 2025

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:18:50.482251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T15:18:49.727363Z digest=sha256:f831161a77909913f987126ada19f2d9d4353d708153478948cba58a3cc3222c

Observation 22cb8890-990b-464b-9528-904eec590064 · outbound

This paper cites A diagram is worth a dozen images.

Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models A diagram is worth a dozen images

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:18:50.467391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T15:18:49.732028Z digest=sha256:eecbca4cad2ac43e36c27260ea985b7323f7a7ff689dd280c35b4ef231d60e96

Observation 45db2590-02e4-4fa6-ad30-f86564e89be3 · outbound

This paper cites Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models.

Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:18:50.453036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T15:18:49.736628Z digest=sha256:be05b286f0860430078280ed99813a44efa0d7b7873e2faf28159a1e4cc7c618

Observation 9fb2e153-f3fc-4ab9-a3de-3e38e67e7fb4 · outbound

This paper cites Improved baselines with visual instruction tuning.

Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models Improved baselines with visual instruction tuning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:18:50.438408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T15:18:49.740823Z digest=sha256:62142706fcd8bea50c675673ae39a81bf165a2f4cf4855c2ab782fa8fde261d3

Observation 92118df8-568f-47dc-bc9d-2b271f751bbf · outbound

This paper cites Decoupled weight decay regular- ization.

Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models Decoupled weight decay regular- ization

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:18:50.423092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T15:18:49.744780Z digest=sha256:484dcc46d38e6b5abbc3558cde543db463eeb24552c96abe09187ee5dcedddb8

Observation 54c31879-1f12-40ce-818f-3f1fbf90adc1 · outbound

This paper cites Dvqa: Understanding data visualizations via question answering.

Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models Dvqa: Understanding data visualizations via question answering

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:18:50.408507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T15:18:49.748916Z digest=sha256:e45676b042e712af6d61050065bdc2f6b3a5fc6cd2e757e9bc3dbdea1bb0ad1e

Observation 73927c6d-1325-4fb3-9fa0-b15755cc4b35 · outbound

This paper cites ChartQA: A benchmark for question answering about charts with visual and logical reasoning.

Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models ChartQA: A benchmark for question answering about charts with visual and logical reasoning

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:18:50.393775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T15:18:49.879875Z digest=sha256:5210d6b0114c56842ec129f1ab7dfac1f39658f168dd1b3976a088d21eb88dd9

Observation df0f00d9-5505-4a04-9084-806373e0ec1c · outbound

This paper cites Docvqa: A dataset for vqa on document images.

Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models Docvqa: A dataset for vqa on document images

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:18:50.379028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T15:18:49.884941Z digest=sha256:1b4723e97cef47895552871b84bad06d3c2ba1783d1671e81adf1dd7546943db

Observation 423dafa5-36b5-4243-b489-c6c85c98862a · outbound

This paper cites An augmented benchmark dataset for geometric question answering through dual parallel text encoding.

Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models An augmented benchmark dataset for geometric question answering through dual parallel text encoding

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:18:50.363884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T15:18:49.889819Z digest=sha256:8bc65ce74b1357c9dcc3e662329f89940726acef2e1e0682f8a03e688acfcbfe

Observation 3fea2abc-bf7f-49fa-877d-8fc5f3113e0d · outbound

This paper cites Ocr-free document understanding trans- former.

Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models Ocr-free document understanding trans- former

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:18:50.347991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T15:18:49.894446Z digest=sha256:885732ae849277d790bd4ec74e9d951c1adf618b182d9a39353a9300ed5a32c9

Observation 24fae8a2-ce73-4e34-b13b-0bee0ad822f5 · outbound

This paper cites Are we on the right way for evaluating large vision-language models? In Proceedings of the International Conference on Neural Information Processing Systems (NIPS) , 2024.

Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models Are we on the right way for evaluating large vision-language models? In Proceedings of the International Conference on Neural Information Processing Systems (NIPS) , 2024

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:18:50.332780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T15:18:49.898740Z digest=sha256:dcb1a8f1f14d4eca418367027d6b3503b3e3133b4698c66ddfb8fa4a72b2a0d1

Observation 5a44332a-664a-4c49-8cee-a2ec6c247fdc · outbound

This paper cites Mm-vet: evaluating large multimodal models for integrated capabilities.

Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models Mm-vet: evaluating large multimodal models for integrated capabilities

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:18:50.317294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T15:18:49.902995Z digest=sha256:5e08fdbfed3a885867c1e1a54bf4978ea4a4fabd4e3cc3e47c4d3e55160382bd

Observation 9906a31c-bea3-48b3-9336-8f37e7158ecd · outbound

This paper cites Seed-bench: Benchmarking multimodal large language models.

Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models Seed-bench: Benchmarking multimodal large language models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:18:50.302758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T15:18:49.907151Z digest=sha256:a2f101e806e3dfc4e2d6e1d59b292ee4203b411641e70e1b383ebd38aa07a51e

Observation c6c03810-9ee8-4a14-8e6e-a27d589e770a · outbound

This paper cites Grok-1.5 vision preview.

Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models Grok-1.5 vision preview

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:18:50.286691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T15:18:49.911148Z digest=sha256:29111f02cf4e4451451e6e49bcc3d4b4ffca2385654f794dddcbd3256457094c

Observation 6aef7a1f-6d98-4fb2-91b5-1776dc5c62f7 · outbound

This paper cites MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?.

Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T15:18:49.916042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:18:49.916042Z digest=sha256:461c1f37d2f858676f459406bb98a56824a7247fb2ebdd9d89989a1848b9c418

Observation 85204ccc-18ad-4c77-bee0-92926ec24000 · outbound

This paper cites Divide, Conquer and Combine: A Training-Free Framework for High-Resolution Image Perception in Multimodal Large Language Models.

Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models Divide, Conquer and Combine: A Training-Free Framework for High-Resolution Image Perception in Multimodal Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T15:18:49.921539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:18:49.921539Z digest=sha256:2636994265c04104f7bb4756bff9f0a73ad180190c8165324422336aa90abf7f

Observation d55b6464-93a5-4e79-8433-028d34009484 · outbound

This paper cites Infographicvqa.

Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models Infographicvqa

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:18:50.269531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T15:18:49.926739Z digest=sha256:c6536a3a357e2fa1765e20a2b6a5a21c058db75d485962ef3dbea865771f70d5

Observation 05ce2f96-c2e0-4b66-a5b8-5075adda0bfc · outbound

This paper cites Evaluating object hallucination in large vision-language models.

Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models Evaluating object hallucination in large vision-language models

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:18:50.253336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T15:18:49.931136Z digest=sha256:3dfd98f0636e27d09b86c3e0aa7d49a00590b2c46c022ae6e4f46b23370c123c

Observation dabf7dd9-6e3c-4415-ae52-f3c04b9c7b12 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:18:50.238145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T15:18:49.935488Z digest=sha256:497418c849bacb1b707153869dce07f5085f715f99570c817291dd7cd78d9e54

Observation b290609f-f6fa-4c5e-b999-c5d4635ae25f · outbound

This paper cites What matters when building vision-language models? In Proceedings of the International Conference on Neural Information Processing Systems (NIPS), 2024.

Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models What matters when building vision-language models? In Proceedings of the International Conference on Neural Information Processing Systems (NIPS), 2024

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:18:50.222525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T15:18:49.940272Z digest=sha256:50505ec286a9e7b5817f39c714731b64fe7434d6043cc3a68ef2d73f5f0a53e7

Observation cc3526a8-d06a-4df2-b2c6-bbb387b08b1c · outbound

This paper cites Vila: On pre-training for visual language models.

Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models Vila: On pre-training for visual language models

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:18:50.205851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T15:18:49.944967Z digest=sha256:7fd5ed75e935e48a28fd5e3a2b3e165df8b58ce0e5516171cbe012e6bf4c7b2f

Observation 5bf02c69-1611-4eb0-888a-7a11ed03c12f · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T15:18:49.949655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:18:49.949655Z digest=sha256:3eef7e9696b04852ca2e5554f3aa6606b4c3f0e6869fbc8a9fa774d44fd1f7ec

Observation a48a38b5-7dce-4027-9b86-642a409825f6 · outbound

This paper cites InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model.

Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T15:18:49.954946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:18:49.954946Z digest=sha256:2e07354ceaf3c5228bdd93ebd217e83033fc4c09891ffffb18b2c786e8da9973

Observation 26ee26dc-7886-4915-ab75-859a9ae4212c · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models PaliGemma: A versatile 3B VLM for transfer

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T15:18:49.959972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:18:49.959972Z digest=sha256:c6dd9d544669897157123105490f667ed95940672fbdebb954375e8b7cfcc398

Observation 140b4b66-190e-45ad-9a71-68564329d3f7 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models LLaVA-OneVision: Easy Visual Task Transfer

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T15:18:49.965049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:18:49.965049Z digest=sha256:4472e2b9fca30154d2ceee7651115c2f72c8884d403c2d785434fe736758b7dd

Observation de49709a-2c31-437b-9bed-827cca1ca6f6 · outbound

This paper cites Mini-internvl: a flexible-transfer pocket multi-modal model with 5% parameters and 90% performance.

Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models Mini-internvl: a flexible-transfer pocket multi-modal model with 5% parameters and 90% performance

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T15:18:49.970163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:18:49.970163Z digest=sha256:0bbb925a71845b562ffe05b4a7563aff5705c816c47f72725779bbe62d20bf7b

Observation 5fc3821e-0ff2-44bb-84c7-380649c8191e · outbound

This paper cites Ovis: Structural Embedding Alignment for Multimodal Large Language Model.

Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T15:18:49.974575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:18:49.974575Z digest=sha256:578644984f4afc4b90b999491aadf47dafc03a312d5cf3a3570cb5a2c4dce384

Pith citing papers

Observation 1d86c0e1-6bd8-4667-8769-8edaf95e3ad4 · inbound

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality cites this paper.

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T17:51:42.830659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:51:42.830659Z digest=sha256:c4cc7017ab62fed50f9fcfa9e10e3f4a66b5d4da79018101d27f45d2e7c056df

Observation 377603f8-1932-42d4-ac27-1cb2d7df6fed · inbound

HRVVS: A High-resolution Video Vasculature Segmentation Network via Hierarchical Autoregressive Residual Priors cites this paper.

HRVVS: A High-resolution Video Vasculature Segmentation Network via Hierarchical Autoregressive Residual Priors Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T11:38:38.548285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T11:38:36.041216Z digest=sha256:96a2dab40e7723f428172f55ee41fe6baac41396a0b0c29ae6fa2fefacb69392