Pith. sign in

Paper Citation Record · LEDGER

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning

As of 16 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 0 inbound Pith citation observations for arXiv:2411.12787.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.12787 v3

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T17:37:54.891321Z

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

38 of 38 outbound references displayed

  • verified exact0
  • verified fuzzy9
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d7f1bee5-c5be-4905-b8ca-a43edb0dd259 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.675226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.675226Z digest=sha256:45144f39b76024f37cfe07ae9c7dfef6fa7dd816304232f2cc71fb0011bda853

Observation a26616b5-c94c-419a-a46a-4fef259519f5 · outbound

This paper cites Layer Normalization.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning Layer Normalization

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.681289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.681289Z digest=sha256:5194e53f0090c273f0b1312f2c7cab25a9735b257e3583727b075f1188eb8604

Observation 2b986d2a-0096-41de-86b8-ffc475aaa73e · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.687801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.687801Z digest=sha256:59f80821811fb46c501e77fc4ab9b2dd1f2e990da9be525653244bef9455b1cb

Observation 5afec7ec-c18a-4f90-ae8a-94c76f4ef2bc · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.695363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.695363Z digest=sha256:5690b821ce928bd4dd28b84c9a47ec0619dbeaf0b3624b864d98afa847fb8794

Observation 66132ade-bf2d-4a37-b9ba-22066cb3ef22 · outbound

This paper cites LLaVA-MoLE: Sparse Mixture of LoRA Experts for Mitigating Data Conflicts in Instruction Finetuning MLLMs.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning LLaVA-MoLE: Sparse Mixture of LoRA Experts for Mitigating Data Conflicts in Instruction Finetuning MLLMs

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.701624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.701624Z digest=sha256:9dbaaff725b69a31d68efb46a65883affcd0781879adb29497eee6c71fbf189b

Observation 2f5db9c4-0d15-41d3-9a10-32912c279eef · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.709703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.709703Z digest=sha256:e8cb372910ae44560bcfc7a02d80c83e415acf9c20ef867b9cb908652aef1b7c

Observation 08921ef2-c25c-45a5-97d2-dde53a6983fd · outbound

This paper cites Instructblip: Towards general- purpose vision-language models with instruction tuning,.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning Instructblip: Towards general- purpose vision-language models with instruction tuning,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.715260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.715260Z digest=sha256:b5d9189353709cc96e1afb0770274c5094ffc48a1d7cb18ac4ffb30eb2dbb9ef

Observation 0da724d6-0a78-47c3-bf3b-d571344a3133 · outbound

This paper cites Parameter-efficient fine-tuning of large-scale pre-trained language models.Nature Machine In- telligence, 5(3):220–235, 2023.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning Parameter-efficient fine-tuning of large-scale pre-trained language models.Nature Machine In- telligence, 5(3):220–235, 2023

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.720473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.720473Z digest=sha256:1a6249ac26f376e2c31590c0f95b6dfb381c310f84c22153802d1c541d5ab0d5

Observation a6e169af-b8f8-4e5e-bdd5-e05a896e9cc3 · outbound

This paper cites MouSi: Poly-Visual-Expert Vision-Language Models.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning MouSi: Poly-Visual-Expert Vision-Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.726053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.726053Z digest=sha256:e0640bced962446206c6051d2b5504e10c448150a92fc29b577132ad33e7165d

Observation 791b1a21-74a6-4ede-ba5f-e9c4af2a9656 · outbound

This paper cites Deep sparse rectifier neural networks.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning Deep sparse rectifier neural networks

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:37:55.718337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T17:37:54.732437Z digest=sha256:2f44d58bc0c273a6e13250245f8751b1ec75198f6410d78d69ef11dabb19fdd7

Observation 95eb8dfe-3c88-4319-944f-85912b3452dd · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning LoRA: Low-Rank Adaptation of Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.738181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.738181Z digest=sha256:bf37b38bb90bc893a412bd36b8ab96db75163d03d5b8f8414abc3df1e529d639

Observation acacf141-d1a7-492d-b5b9-3c6f8cf9c8e3 · outbound

This paper cites Harder tasks need more experts: Dynamic routing in moe models.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning Harder tasks need more experts: Dynamic routing in moe models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.743855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.743855Z digest=sha256:3ae25d173b7c6ed99cd9b1ec0e5d8757810959d4c0be4049131d101ea4ac5382

Observation 0f0c4c9c-b5a1-4e67-961f-33c073ec933b · outbound

This paper cites RoDE: Linear Rectified Mixture of Diverse Experts for Food Large Multi-Modal Models.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning RoDE: Linear Rectified Mixture of Diverse Experts for Food Large Multi-Modal Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.748872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.748872Z digest=sha256:9e7ed4f584ea9fbabaf38ca20dc3271c1b8daeee7d16f97c59216342b2f6fcd5

Observation 2cae4764-886c-4fa6-92aa-ae68254158f4 · outbound

This paper cites Unlocking textual and visual wisdom: Open-vocabulary 3d object detection enhanced by comprehensive guidance from text and image.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning Unlocking textual and visual wisdom: Open-vocabulary 3d object detection enhanced by comprehensive guidance from text and image

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:37:55.699366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T17:37:54.754030Z digest=sha256:38e3d03e21fe9065153827cb338fa4206c1a5477a759c5c59585dab105270aee

Observation 7ad139bc-c6db-4689-ac7b-37cdb056f429 · outbound

This paper cites Don't Deceive Me: Mitigating Gaslighting through Attention Reallocation in LMMs.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning Don't Deceive Me: Mitigating Gaslighting through Attention Reallocation in LMMs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.759113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.759113Z digest=sha256:56c31c37a92d0e1a580fb8434fd728cf812d6e94fb4a1cf3d9d629a9b9136af0

Observation 8e8fdd22-c8c3-4c82-b9b2-96d3439bc3a4 · outbound

This paper cites Lumen: Unleashing versa- tile vision-centric capabilities of large multimodal models.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning Lumen: Unleashing versa- tile vision-centric capabilities of large multimodal models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:37:55.682222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T17:37:54.764630Z digest=sha256:43932a4aec27ee46c63c9af727f0e6e423d5ec648b8753b3bffbadaf2d029a66

Observation 7da838a1-9d43-4d58-99a5-22b63f4f3261 · outbound

This paper cites Imagenet classification with deep convolutional neural net- works.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning Imagenet classification with deep convolutional neural net- works

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.770220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.770220Z digest=sha256:e430bbfa0f0bab6da08a45f516545c69e6e48d3c124328b98452ccdc1234065c

Observation c7b873ba-d7bf-4ae4-b205-0edc508b4c29 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.775267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.775267Z digest=sha256:613968529f05a85bb0fa2c6aaad102c09ec7c6360ecc04fd0529f9fd349e5701

Observation eccca891-7eb6-4474-bd9e-d6cc548c56e1 · outbound

This paper cites Rouge: A package for automatic evaluation of summaries.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning Rouge: A package for automatic evaluation of summaries

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.780768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.780768Z digest=sha256:ec92dc881011480996dbab119acebf90bd7cb96017b68c8b179255be2819fc2d

Observation 95104aa3-f445-495f-93da-ecc27bb27c04 · outbound

This paper cites Visual instruction tuning, 2023.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning Visual instruction tuning, 2023

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:37:55.625708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T17:37:54.786295Z digest=sha256:e7e38dbd9f2d42afb80c85c47015575c1e7530720f903cb7b095842d43d86bb1

Observation 480ba212-1816-4dbd-9758-2dd3d696f467 · outbound

This paper cites Improved baselines with visual instruction tuning.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning Improved baselines with visual instruction tuning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:37:55.606297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T17:37:54.792350Z digest=sha256:08e380065063e4161e2988da976e94b65256e002e402d192f105ae257088957b

Observation 9823dec1-cb94-42d7-a37d-f41e97f3ed15 · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.798888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.798888Z digest=sha256:c8b71181848585026abd8ddb243927285b6fa5e73a2d883f37e28f449a7ae015

Observation 0f5cf207-ebb8-4ff4-a6c4-54c25a6b83e6 · outbound

This paper cites AdaMoLE: Fine-Tuning Large Language Models with Adaptive Mixture of Low-Rank Adaptation Experts.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning AdaMoLE: Fine-Tuning Large Language Models with Adaptive Mixture of Low-Rank Adaptation Experts

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.805675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.805675Z digest=sha256:359f400535d8153f6eb7a2afff9cd6f8c3df902b651261dfbe2b28f9d52ad0fc

Observation 2bf07059-d448-4a95-9510-58703b1ff33f · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.811609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.811609Z digest=sha256:e528371fcee0ebf054b00421a4f56d599c1b1d86807aefd5d25bd1190e3b97c3

Observation d93f0c8e-19ec-4a4a-b20d-2101feb765f4 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning DINOv2: Learning Robust Visual Features without Supervision

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.817637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.817637Z digest=sha256:0737bb15ad635b0dda1c94cab7a0cfac56928c196ab77961a85de120be8bd5f8

Observation 16286e8f-2876-4b89-9ad0-6c6ad8ab5879 · outbound

This paper cites Flickr30k entities: Collecting region-to-phrase corre- spondences for richer image-to-sentence models.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning Flickr30k entities: Collecting region-to-phrase corre- spondences for richer image-to-sentence models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:37:55.587874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T17:37:54.823531Z digest=sha256:879f08b1c18509e30277f9daa2fba87813cee91e8729695b2d70faa1691e86fa

Observation 4167a38d-3880-40f4-a5ef-164b61f5a131 · outbound

This paper cites A Call for Clarity in Reporting BLEU Scores.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning A Call for Clarity in Reporting BLEU Scores

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.828772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.828772Z digest=sha256:0185194be12c58e02839005b68ee735c5904401a90faaeeb6f4c86c8aa983259

Observation 61538d55-5b18-46f5-b59e-c707a540807b · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning Learning transferable visual models from natural language supervi- sion

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:37:55.570322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T17:37:54.834854Z digest=sha256:142302c1507ab323bc1dcdde3a3cbe73b61cf2df6bdabd97ba81d3fe913bfc5b

Observation 780a05aa-8369-4659-8f5a-6edb99520912 · outbound

This paper cites Scienceqa: A novel resource for question answering on scholarly articles.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning Scienceqa: A novel resource for question answering on scholarly articles

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:37:55.547356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T17:37:54.840309Z digest=sha256:9836cfe886e87d407de4423d6863dd2919998e80ee2405c537cbea67555629fc

Observation ecd51877-14f4-4d05-b97d-ea703d8943f3 · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.845427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.845427Z digest=sha256:f4e9a7ae00a6dfaa0fe80f49ae49219a092471ca58cdec177caf0982f186c237

Observation 031c52a2-023b-4533-841e-c6bbe7f3c87c · outbound

This paper cites Eyes wide shut? exploring the visual shortcomings of multimodal llms.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning Eyes wide shut? exploring the visual shortcomings of multimodal llms

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.850908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.850908Z digest=sha256:267317bffdd36651902c33b12055642982b88dc434a1ff2570a4662c135af924

Observation 264b8c67-1619-472e-9c4c-6c98e158f314 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning LLaMA: Open and Efficient Foundation Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.855723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.855723Z digest=sha256:12edad78f450944286a59c514ebf992cd22406a8b7d4456c5766d8a338d213d6

Observation 1a8cebf6-8ead-4290-9bfd-00fc64a31cf5 · outbound

This paper cites Mixture of LoRA Experts.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning Mixture of LoRA Experts

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.861350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.861350Z digest=sha256:7d10277781d7d561b81fd90e2618de62199384f7d78419427f287a19a61807d1

Observation 9663fe63-948c-4bfd-9f35-dc7e1365bfa1 · outbound

This paper cites Vision transformer with deformable attention.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning Vision transformer with deformable attention

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:37:55.511682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T17:37:54.867653Z digest=sha256:6276e4298373c374147de72339222870b020e931c3766624dcc5aceba4ea8e4e

Observation 6aee29cb-5595-4836-be2c-d468dbc7c14c · outbound

This paper cites LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.873349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.873349Z digest=sha256:4c9a2e22d1bb6d3fe4f6ef3bc029da9c06cf3d4ab133de70c39f38a360ce5960

Observation c651cc23-de9b-469f-b8a4-48eb6bd89114 · outbound

This paper cites FoodLMM: A Versatile Food Assistant using Large Multi-modal Model.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning FoodLMM: A Versatile Food Assistant using Large Multi-modal Model

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.879662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.879662Z digest=sha256:61d810f7bea80d608d68774ed08fb5d7f49c209f5d06862edae7cec38f64189b

Observation bce9786f-9c3d-4cdd-8eeb-1345393caf8c · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.885564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.885564Z digest=sha256:371eb401aee8df7e6de05d928922fb3d929190e1b135d4e5cbd29dcba1ad6b00

Observation b787204d-0eae-4bab-a0a2-258eadc739e8 · outbound

This paper cites Deformable DETR: Deformable Transformers for End-to-End Object Detection.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning Deformable DETR: Deformable Transformers for End-to-End Object Detection

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.891321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.891321Z digest=sha256:ec9d5981c0235b5e15985508ffe0c2fc278b3b9e15684daba61640beffcd7196

Pith citing papers

No inbound Pith citation observations are available.