Pith. sign in

Paper Citation Record · LEDGER

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning

As of 16 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 0 inbound Pith citation observations for arXiv:2411.12787.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.12787 v3

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T17:37:54.891321Z

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

38 of 38 outbound references displayed

  • verified exact0
  • verified fuzzy9
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d7f1bee5-c5be-4905-b8ca-a43edb0dd259 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.675226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.675226Z digest=sha256:45144f39b76024f37cfe07ae9c7dfef6fa7dd816304232f2cc71fb0011bda853

Observation a26616b5-c94c-419a-a46a-4fef259519f5 · outbound

This paper cites Layer Normalization.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning Layer Normalization

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.681289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.681289Z digest=sha256:5194e53f0090c273f0b1312f2c7cab25a9735b257e3583727b075f1188eb8604

Observation 2b986d2a-0096-41de-86b8-ffc475aaa73e · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.687801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.687801Z digest=sha256:59f80821811fb46c501e77fc4ab9b2dd1f2e990da9be525653244bef9455b1cb

Observation 5afec7ec-c18a-4f90-ae8a-94c76f4ef2bc · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.695363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.695363Z digest=sha256:5690b821ce928bd4dd28b84c9a47ec0619dbeaf0b3624b864d98afa847fb8794

Observation 66132ade-bf2d-4a37-b9ba-22066cb3ef22 · outbound

This paper cites LLaVA-MoLE: Sparse Mixture of LoRA Experts for Mitigating Data Conflicts in Instruction Finetuning MLLMs.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning LLaVA-MoLE: Sparse Mixture of LoRA Experts for Mitigating Data Conflicts in Instruction Finetuning MLLMs

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.701624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.701624Z digest=sha256:cfea3afcc97e2d4eb0083596b2450299ea424d42d651c145a146de8cf155d122

Observation 2f5db9c4-0d15-41d3-9a10-32912c279eef · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.709703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.709703Z digest=sha256:e8cb372910ae44560bcfc7a02d80c83e415acf9c20ef867b9cb908652aef1b7c

Observation 08921ef2-c25c-45a5-97d2-dde53a6983fd · outbound

This paper cites Instructblip: Towards general- purpose vision-language models with instruction tuning,.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning Instructblip: Towards general- purpose vision-language models with instruction tuning,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.715260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.715260Z digest=sha256:b5d9189353709cc96e1afb0770274c5094ffc48a1d7cb18ac4ffb30eb2dbb9ef

Observation 0da724d6-0a78-47c3-bf3b-d571344a3133 · outbound

This paper cites Parameter-efficient fine-tuning of large-scale pre-trained language models.Nature Machine In- telligence, 5(3):220–235, 2023.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning Parameter-efficient fine-tuning of large-scale pre-trained language models.Nature Machine In- telligence, 5(3):220–235, 2023

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.720473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.720473Z digest=sha256:1a6249ac26f376e2c31590c0f95b6dfb381c310f84c22153802d1c541d5ab0d5

Observation a6e169af-b8f8-4e5e-bdd5-e05a896e9cc3 · outbound

This paper cites MouSi: Poly-Visual-Expert Vision-Language Models.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning MouSi: Poly-Visual-Expert Vision-Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.726053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.726053Z digest=sha256:15acc3aaab4d4746eabbbd60823147cb16c527530064243c0126b4c8cf778fed

Observation 791b1a21-74a6-4ede-ba5f-e9c4af2a9656 · outbound

This paper cites Deep sparse rectifier neural networks.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning Deep sparse rectifier neural networks

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:37:55.718337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:37:54.732437Z digest=sha256:a6d1c5e4b6340e71b2fdd41368767a69846273802cbc157bf6b6133289da40f8

Observation 95eb8dfe-3c88-4319-944f-85912b3452dd · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning LoRA: Low-Rank Adaptation of Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.738181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.738181Z digest=sha256:bf37b38bb90bc893a412bd36b8ab96db75163d03d5b8f8414abc3df1e529d639

Observation acacf141-d1a7-492d-b5b9-3c6f8cf9c8e3 · outbound

This paper cites Harder tasks need more experts: Dynamic routing in moe models.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning Harder tasks need more experts: Dynamic routing in moe models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.743855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.743855Z digest=sha256:3ae25d173b7c6ed99cd9b1ec0e5d8757810959d4c0be4049131d101ea4ac5382

Observation 0f0c4c9c-b5a1-4e67-961f-33c073ec933b · outbound

This paper cites RoDE: Linear Rectified Mixture of Diverse Experts for Food Large Multi-Modal Models.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning RoDE: Linear Rectified Mixture of Diverse Experts for Food Large Multi-Modal Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.748872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.748872Z digest=sha256:503d862d88bc29033d21d67d6089619d3a70364c357292dcea3e96cf33409eec

Observation 2cae4764-886c-4fa6-92aa-ae68254158f4 · outbound

This paper cites Unlocking textual and visual wisdom: Open-vocabulary 3d object detection enhanced by comprehensive guidance from text and image.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning Unlocking textual and visual wisdom: Open-vocabulary 3d object detection enhanced by comprehensive guidance from text and image

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:37:55.699366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:37:54.754030Z digest=sha256:2303447076118cbd72c44250b25a832dbcb6e0a3aee052d5b631bfe987a3e58f

Observation 7ad139bc-c6db-4689-ac7b-37cdb056f429 · outbound

This paper cites Don't Deceive Me: Mitigating Gaslighting through Attention Reallocation in LMMs.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning Don't Deceive Me: Mitigating Gaslighting through Attention Reallocation in LMMs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.759113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.759113Z digest=sha256:a87017664163eee8a311a6969e6f641fffe1f934747d8bb8c54a58e97a04104b

Observation 8e8fdd22-c8c3-4c82-b9b2-96d3439bc3a4 · outbound

This paper cites Lumen: Unleashing versa- tile vision-centric capabilities of large multimodal models.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning Lumen: Unleashing versa- tile vision-centric capabilities of large multimodal models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:37:55.682222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:37:54.764630Z digest=sha256:730f5e60344001edd7cbb99452f2bd982dfcc02d3164208c29c31c4f9afd198e

Observation 7da838a1-9d43-4d58-99a5-22b63f4f3261 · outbound

This paper cites Imagenet classification with deep convolutional neural net- works.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning Imagenet classification with deep convolutional neural net- works

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.770220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.770220Z digest=sha256:e430bbfa0f0bab6da08a45f516545c69e6e48d3c124328b98452ccdc1234065c

Observation c7b873ba-d7bf-4ae4-b205-0edc508b4c29 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.775267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.775267Z digest=sha256:613968529f05a85bb0fa2c6aaad102c09ec7c6360ecc04fd0529f9fd349e5701

Observation eccca891-7eb6-4474-bd9e-d6cc548c56e1 · outbound

This paper cites Rouge: A package for automatic evaluation of summaries.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning Rouge: A package for automatic evaluation of summaries

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.780768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.780768Z digest=sha256:ec92dc881011480996dbab119acebf90bd7cb96017b68c8b179255be2819fc2d

Observation 95104aa3-f445-495f-93da-ecc27bb27c04 · outbound

This paper cites Visual instruction tuning, 2023.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning Visual instruction tuning, 2023

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:37:55.625708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:37:54.786295Z digest=sha256:207bd7bfdd83fda8ef5a5bc886acc6831e3be598586f56817bd4ca6f71e428f0

Observation 480ba212-1816-4dbd-9758-2dd3d696f467 · outbound

This paper cites Improved baselines with visual instruction tuning.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning Improved baselines with visual instruction tuning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:37:55.606297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:37:54.792350Z digest=sha256:b7989f71e2445af1ff9ea20b5a3ed0b1f5990d214f07d9ba2e3fe83aa5809714

Observation 9823dec1-cb94-42d7-a37d-f41e97f3ed15 · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.798888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.798888Z digest=sha256:c8b71181848585026abd8ddb243927285b6fa5e73a2d883f37e28f449a7ae015

Observation 0f5cf207-ebb8-4ff4-a6c4-54c25a6b83e6 · outbound

This paper cites AdaMoLE: Fine-Tuning Large Language Models with Adaptive Mixture of Low-Rank Adaptation Experts.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning AdaMoLE: Fine-Tuning Large Language Models with Adaptive Mixture of Low-Rank Adaptation Experts

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.805675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.805675Z digest=sha256:6cb0e8186c5b338e660d03892a6f49db3a354a955738e97c9e957ef5009f1305

Observation 2bf07059-d448-4a95-9510-58703b1ff33f · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.811609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.811609Z digest=sha256:e528371fcee0ebf054b00421a4f56d599c1b1d86807aefd5d25bd1190e3b97c3

Observation d93f0c8e-19ec-4a4a-b20d-2101feb765f4 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning DINOv2: Learning Robust Visual Features without Supervision

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.817637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.817637Z digest=sha256:0737bb15ad635b0dda1c94cab7a0cfac56928c196ab77961a85de120be8bd5f8

Observation 16286e8f-2876-4b89-9ad0-6c6ad8ab5879 · outbound

This paper cites Flickr30k entities: Collecting region-to-phrase corre- spondences for richer image-to-sentence models.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning Flickr30k entities: Collecting region-to-phrase corre- spondences for richer image-to-sentence models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:37:55.587874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:37:54.823531Z digest=sha256:7763055b64b63b01da0016f823ee142654574507bbbda7cdc07318e56c79ff03

Observation 4167a38d-3880-40f4-a5ef-164b61f5a131 · outbound

This paper cites A Call for Clarity in Reporting BLEU Scores.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning A Call for Clarity in Reporting BLEU Scores

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.828772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.828772Z digest=sha256:0185194be12c58e02839005b68ee735c5904401a90faaeeb6f4c86c8aa983259

Observation 61538d55-5b18-46f5-b59e-c707a540807b · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning Learning transferable visual models from natural language supervi- sion

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:37:55.570322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:37:54.834854Z digest=sha256:4b645483032f28bc2d137754feb11458ea6240807c8859b1b22d538a2de2b3e4

Observation 780a05aa-8369-4659-8f5a-6edb99520912 · outbound

This paper cites Scienceqa: A novel resource for question answering on scholarly articles.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning Scienceqa: A novel resource for question answering on scholarly articles

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:37:55.547356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:37:54.840309Z digest=sha256:e7ed1b51f1259694b0b1533c73463ca89dc52e5e3bbed4dda01bbdc6ae046cc6

Observation ecd51877-14f4-4d05-b97d-ea703d8943f3 · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.845427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.845427Z digest=sha256:f4e9a7ae00a6dfaa0fe80f49ae49219a092471ca58cdec177caf0982f186c237

Observation 031c52a2-023b-4533-841e-c6bbe7f3c87c · outbound

This paper cites Eyes wide shut? exploring the visual shortcomings of multimodal llms.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning Eyes wide shut? exploring the visual shortcomings of multimodal llms

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.850908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.850908Z digest=sha256:267317bffdd36651902c33b12055642982b88dc434a1ff2570a4662c135af924

Observation 264b8c67-1619-472e-9c4c-6c98e158f314 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning LLaMA: Open and Efficient Foundation Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.855723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.855723Z digest=sha256:12edad78f450944286a59c514ebf992cd22406a8b7d4456c5766d8a338d213d6

Observation 1a8cebf6-8ead-4290-9bfd-00fc64a31cf5 · outbound

This paper cites Mixture of LoRA Experts.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning Mixture of LoRA Experts

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.861350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.861350Z digest=sha256:ffd4ed756d91cfc29b72238d2397a3b5719dd4ab0698ef366a6c3176450c2ea8

Observation 9663fe63-948c-4bfd-9f35-dc7e1365bfa1 · outbound

This paper cites Vision transformer with deformable attention.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning Vision transformer with deformable attention

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:37:55.511682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:37:54.867653Z digest=sha256:67b9f807d64b851541f1c22bc38f95c744b620656e58428d197da5f3cda6cfe9

Observation 6aee29cb-5595-4836-be2c-d468dbc7c14c · outbound

This paper cites LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.873349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.873349Z digest=sha256:023eae48276fc35acc7d72a5a201dfb286da7634abaccf2d077485772d11e215

Observation c651cc23-de9b-469f-b8a4-48eb6bd89114 · outbound

This paper cites FoodLMM: A Versatile Food Assistant using Large Multi-modal Model.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning FoodLMM: A Versatile Food Assistant using Large Multi-modal Model

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.879662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.879662Z digest=sha256:ca26f3bc62875be6467daf297882719eb43099c67f450e7da3ae64ccfe4f4878

Observation bce9786f-9c3d-4cdd-8eeb-1345393caf8c · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.885564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.885564Z digest=sha256:371eb401aee8df7e6de05d928922fb3d929190e1b135d4e5cbd29dcba1ad6b00

Observation b787204d-0eae-4bab-a0a2-258eadc739e8 · outbound

This paper cites Deformable DETR: Deformable Transformers for End-to-End Object Detection.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning Deformable DETR: Deformable Transformers for End-to-End Object Detection

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.891321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.891321Z digest=sha256:ec9d5981c0235b5e15985508ffe0c2fc278b3b9e15684daba61640beffcd7196

Pith citing papers

No inbound Pith citation observations are available.