Pith. sign in

Paper Citation Record · LEDGER

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion

As of 14 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 2 inbound Pith citation observations for arXiv:2412.01289.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.01289 v2

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T04:35:52.480788Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T22:40:00.510742Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T16:27:09.401525Z

Reference resolution

56 of 56 outbound references displayed

  • verified exact1
  • verified fuzzy23
  • unresolved32
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e5020a5b-2820-412e-922d-8a01dcd07960 · outbound

This paper cites Llama 3 model card.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Llama 3 model card

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:53.926092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:35:52.161969Z digest=sha256:c679bf4f2e93557643c22ee75f2859df03ea408aff121bd6b563dac5503f98f9

Observation 579021a5-be40-4b14-b897-eb5b2b030ef5 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Flamingo: a visual language model for few-shot learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:52.169246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:52.169246Z digest=sha256:028d965456a9a888319cd1d6dc2450d58d8a61208772ebca74b06db9f2361d6c

Observation 93f85b83-c205-419c-af0b-3cd0ddf64249 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:52.174871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:52.174871Z digest=sha256:57dca38b2d59a74f0f50bbb59d59c1b8d46e949f053ee9d4b3cb078e6b85b589

Observation c84bf1a3-e5d7-4fe6-a94e-c2360674bb86 · outbound

This paper cites Token Merging: Your ViT But Faster.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Token Merging: Your ViT But Faster

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:52.180236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:52.180236Z digest=sha256:88652ab9ad34eecc55c7897ee1951eada7a1031a7f00a1931de4b1d7943b9f61

Observation ffb5b431-b27b-4176-90e4-5d67fb9a08d4 · outbound

This paper cites Honeybee: Locality-enhanced projector for multimodal llm.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Honeybee: Locality-enhanced projector for multimodal llm

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:53.891160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:35:52.185358Z digest=sha256:1ea3c810b4c13d72a24bbd5fd50aee12a9a6d66f7fda257a353c9c5fe5342cd2

Observation 578e2d59-2d4a-4796-888a-adb393df60e7 · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:52.190470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:52.190470Z digest=sha256:25a500edd70ca640c41de0f306cb6557a248a39913e901e38f32416193625905

Observation 89f70929-9fb2-4836-8a7c-a31a4843be58 · outbound

This paper cites An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models, 2024.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models, 2024

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:53.873075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:35:52.196430Z digest=sha256:b6e7361d8e321412da080b1bac043226e063323b9d23cad3c6d9f436e5ac304c

Observation 3fd7e438-4b00-4406-ab26-73712da16cb6 · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, march.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, march

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:52.201755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:52.201755Z digest=sha256:878c1dbe0c55337a7ac38a45f87e35ad690de8f2016a9b2c46f02cba874b7b8f

Observation 55d157bb-c303-4343-a09c-d07f1ad997d5 · outbound

This paper cites Fusing finetuned models for better pretraining.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Fusing finetuned models for better pretraining

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:52.206726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:52.206726Z digest=sha256:65e8de4a3d63bc3ec925fd7b5960ae241aad791d1640bec9198b9613d77df658

Observation e3c46c07-88a7-4f99-a008-8230dc87b68f · outbound

This paper cites MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:52.211989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:52.211989Z digest=sha256:93d87f644be4420166359a5519bf7b5728dee183b7adc28abad0054f55673079

Observation 1961de4f-e606-4566-904a-1bb5a3654107 · outbound

This paper cites Instructblip: Towards general- purpose vision-language models with instruction tuning,.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Instructblip: Towards general- purpose vision-language models with instruction tuning,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:52.217289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:52.217289Z digest=sha256:7effade58ae3a8705f1f8941455f4115e10cbf31f9414cf76d35c3ad66628dd1

Observation 8dc888e9-7e0c-426d-ba46-4932d7dcc54f · outbound

This paper cites Epsilon Sampling Rocks: Investigating Sampling Strategies for Minimum Bayes Risk Decoding for Machine Translation.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Epsilon Sampling Rocks: Investigating Sampling Strategies for Minimum Bayes Risk Decoding for Machine Translation

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-12T04:35:53.110880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:35:52.222575Z digest=sha256:73fed3ca9468e2187d7cd5fe21dee7965ae0588f2d5d91919ff57e5835592f50

Observation 3736654f-5dd5-4490-a3db-223588607acf · outbound

This paper cites Mme: A compre- hensive evaluation benchmark for multimodal large language models, 2024.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Mme: A compre- hensive evaluation benchmark for multimodal large language models, 2024

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:53.823105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:35:52.229333Z digest=sha256:a0b0618ab47085bd5c4317f4b851593e2326faf09655cfbcec33e56b605c95c9

Observation 432abb0b-d338-4279-8de6-313248a04d2b · outbound

This paper cites ConvLLaVA: Hierarchical Backbones as Visual Encoder for Large Multimodal Models.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion ConvLLaVA: Hierarchical Backbones as Visual Encoder for Large Multimodal Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:52.235865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:52.235865Z digest=sha256:5c01fba8b3defdbcd37061e2de416c1edbaed67c7849984852dc252e0385cbd9

Observation 932ccbd9-163f-4711-8a14-80f66a95fb4a · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:53.800117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:35:52.241613Z digest=sha256:e7d78788b6746413b7142554fe547d53cac31d6b6627f0fa4df58b575b8de47a

Observation 0765a90e-1a92-4eb8-8f00-fd0f3af35a10 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:53.774006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:35:52.247353Z digest=sha256:206f81e4bf83c249b19d3caee233e6a28eb24d619561c6da0346013e007e017b

Observation 0014ba9e-0cad-4347-bf10-699e3e4c6b3c · outbound

This paper cites Vizwiz grand challenge: Answering visual questions from blind people.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Vizwiz grand challenge: Answering visual questions from blind people

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:53.753151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:35:52.252591Z digest=sha256:17a46b3a9de08d986cfe512d5f5535dda575f7083376d75184f9003b389a3016

Observation 36af12b2-109f-46e6-84e8-0ac063c7138f · outbound

This paper cites Onellm: One framework to align all modalities with language.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Onellm: One framework to align all modalities with language

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:53.734724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:35:52.257413Z digest=sha256:b74438ff45756fcebe5abeab7ff4cb94128ce976eea8c7f54616a5e6dfd45032

Observation 83921769-1cbc-4a52-a3ad-df5b70759392 · outbound

This paper cites Editing models with task arithmetic.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Editing models with task arithmetic

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:53.714807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:35:52.262529Z digest=sha256:fad01b4c88a018c60520215543977ead297de7e7fd60af4dbe34bfa1a6705805

Observation 11de87cf-3771-4388-8e95-300b97321891 · outbound

This paper cites Editing models with task arithmetic.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Editing models with task arithmetic

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:53.695652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:35:52.267261Z digest=sha256:545e5b4c4efebdfb9acae45d36b125b7e62cb67c283e2bcbfad1b2a8666ce2a0

Observation 9d214fd1-cbe8-4607-b7d8-87a6f45498ff · outbound

This paper cites LLM-Blender: Ensembling Large Language Models with Pairwise Ranking and Generative Fusion.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion LLM-Blender: Ensembling Large Language Models with Pairwise Ranking and Generative Fusion

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:52.272424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:52.272424Z digest=sha256:712706abb4ac586338e13f36efec9abfad7a716ca5aedd35f145351661057afe

Observation 802b26ba-707d-41ff-8c6b-aaeabe7dd343 · outbound

This paper cites Dataless knowledge fusion by merging weights of language models.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Dataless knowledge fusion by merging weights of language models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:53.675911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:35:52.277765Z digest=sha256:812c3bd214b5188323a7760b40c2b5d52ce4fd021ba51ec289a6f7126f5b4824

Observation 7cd60cc8-56bf-4ac8-a293-760a515e2c89 · outbound

This paper cites A diagram is worth a dozen images.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion A diagram is worth a dozen images

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:52.282873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:52.282873Z digest=sha256:52a28ccd735cdd3cab6ca315d3aa49e548d85dd8af6f123f9f14727fe2aebe99

Observation 332759f4-3009-434c-955d-cdd30957e077 · outbound

This paper cites In- troducing idefics: An open reproduction of state-of-the-art visual language model, 2023.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion In- troducing idefics: An open reproduction of state-of-the-art visual language model, 2023

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:53.646125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:35:52.296757Z digest=sha256:08b41e73485bccc85a213e502b6b70525e6de5cc43fe3359807dadfcc50cfa37

Observation e9980917-69d3-430b-9f95-603a1244ca12 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:52.302694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:52.302694Z digest=sha256:edf65f50027c242eaeba0330594a5867194be6a22079ec04f38697d7720b72bf

Observation 5e915f62-38f3-4445-8129-ccf5dee9a37b · outbound

This paper cites LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:52.307747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:52.307747Z digest=sha256:1416e824a858ef453ea1b4616e901c843acf5804833a06a0c26d283568334bb6

Observation 3f797669-e749-44a5-bdc2-08af6022129b · outbound

This paper cites Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:52.313670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:52.313670Z digest=sha256:cdb3cfe7a358372b7e09cf6c20b3d862cebfba12f7f096e33594c93133bd5267

Observation 2998fcce-700e-4121-b0b2-7cd9735a91ba · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:52.321158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:52.321158Z digest=sha256:f8d86ed48ebe4099fad66496bad5e3b09400e865fe1dae3a4411b7998ca7b97e

Observation 031776af-ef46-435a-b03c-2f7c4bfe9e85 · outbound

This paper cites Vila: On pre-training for visual language models.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Vila: On pre-training for visual language models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:53.611660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:35:52.326881Z digest=sha256:747c14269c7955dee76a458e7d5e175865d9c602c5762f21516bd882bee65387

Observation 8bc5a57e-f4ea-4e02-8cd2-d4d1b5d11ab2 · outbound

This paper cites Microsoft coco: Common objects in context.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Microsoft coco: Common objects in context

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:52.332753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:52.332753Z digest=sha256:ee3f0c2c00b3ee66c90a67ff937b787b31fac77eaa46499f14bfc54731e250de

Observation 96d2cefd-2f49-4e5a-a952-a525211d076b · outbound

This paper cites Mitigating hallucination in large multi-modal models via robust instruction tuning.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Mitigating hallucination in large multi-modal models via robust instruction tuning

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:53.577821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:35:52.339639Z digest=sha256:29b58aa6025a696e13c39c3b2619b488f7596268e7e69167a03898705b7f1a0d

Observation e449a05d-4bb9-40ba-b126-78193aa34fc0 · outbound

This paper cites Improved baselines with visual instruction tuning.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Improved baselines with visual instruction tuning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:53.559725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:35:52.346028Z digest=sha256:028fa81057f499f7dd48ebab2416488d999ec4face83aba31e259a5c391faf9c

Observation 6d7f0d48-151c-4fb3-a4e6-ce8c87379131 · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:53.400194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:35:52.350752Z digest=sha256:5a1c3974c747705367a4fd052874e7a011b8860e6140f1cdfd417095febfd81f

Observation 0073817c-eee9-4e9e-bb05-9a1b029e7253 · outbound

This paper cites Visual instruction tuning.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Visual instruction tuning

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:53.380245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:35:52.356248Z digest=sha256:370ff7cee0b3142c87b023973e26ad028416436d5fbdf28628cad662802bf68d

Observation 300eb4da-51a2-4cc8-b6dd-493e0edca866 · outbound

This paper cites BitDelta: Your Fine-Tune May Only Be Worth One Bit.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion BitDelta: Your Fine-Tune May Only Be Worth One Bit

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:52.361915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:52.361915Z digest=sha256:6ce75ca7b9bd352a196cf8baa601976caf347e1928a1bdd4843f2101b5566d9d

Observation 99d3f12d-3706-4a72-8c75-50782528f299 · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion MMBench: Is Your Multi-modal Model an All-around Player?

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:52.368597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:52.368597Z digest=sha256:69747ac8960e4ff68e07c333819f01058b13422fa7c1821495d2dd147af308c9

Observation 2df3300e-0170-4a34-a7f2-adac24513433 · outbound

This paper cites OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:52.374335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:52.374335Z digest=sha256:c117a3e2147fe8b6515d6e0358a058a8e31f2281105587f2a9de2458b559432a

Observation d97a5506-0c4f-4577-87f1-af88335367f4 · outbound

This paper cites An Empirical Study of Scaling Instruct-Tuned Large Multimodal Models.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion An Empirical Study of Scaling Instruct-Tuned Large Multimodal Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:52.380002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:52.380002Z digest=sha256:c4f46ef39d14dd82c8ef3c1d00cf4b4d13d251d3e9455d13eede432acb78df85

Observation d6c2f7b5-19f5-46c5-aa8e-c34aa7e9cb31 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Learn- ing transferable visual models from natural language super- vision

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:52.385690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:52.385690Z digest=sha256:c248e571da9f29310246daf529f1afa3d1c00c2266f7972f7df9103f7352e7dc

Observation a6cc9117-80e7-49c2-b2af-c2da7fac13b2 · outbound

This paper cites Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:52.390853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:52.390853Z digest=sha256:feaff8a89a6eca862fa9d8cbb54d7a1f65ca920bb4c3135405f6baf391b55890

Observation b82fd9ad-d49f-40bb-b243-f5d0d3416c85 · outbound

This paper cites Animating rotation with quaternion curves.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Animating rotation with quaternion curves

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:53.340341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:35:52.396328Z digest=sha256:be4b15cf9752e59da96f28e73e5b513fb2fc6e5ccd6117b708cde92a618329eb

Observation 4d0d2f25-67e1-4bf3-bd8f-22e13d25057e · outbound

This paper cites Towards vqa models that can read.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Towards vqa models that can read

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:53.312247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:35:52.403028Z digest=sha256:8a7d3d130b0aa71523a3aa77dc09815ea92bc3fb2f2c93dedd360e87ba3a77db

Observation b4bce915-5b0a-4100-9b43-9b89da4f8f55 · outbound

This paper cites Visualizing data using t-sne.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Visualizing data using t-sne

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:53.292775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:35:52.408981Z digest=sha256:f436c433e0260932d5f38c3fcf4984b7b3ea665b66f08bacfed315e63fed1211

Observation 174fae8c-ad11-4a6c-895d-cc5a8550ca61 · outbound

This paper cites Knowledge Fusion of Large Language Models.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Knowledge Fusion of Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:52.415298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:52.415298Z digest=sha256:99f2354929443723496ea877c8f540facb92ab4f45b23990f50ad7f9db83a1ed

Observation 56d56766-116b-49ec-b54f-7c8135a8682b · outbound

This paper cites Grok-1.5 Vision Preview, 2024.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Grok-1.5 Vision Preview, 2024

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:53.268461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:35:52.421218Z digest=sha256:b4fb07ce10d59eb066f8861835537facfa97b825fa14481651763ff1113e2beb

Observation d522b4b3-3f60-491e-a3bd-ff52039a38aa · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:52.426778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:52.426778Z digest=sha256:6be32acf6b84e0ca7fcdc3835b872c7571651e13e5831b6cf1fb08ca2bbe80ef

Observation 6429dc0b-6a57-4882-b451-55182ee5fb1a · outbound

This paper cites TIES-Merging: Resolving Interference When Merging Models.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion TIES-Merging: Resolving Interference When Merging Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:52.432601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:52.432601Z digest=sha256:0b4c3fa3d1d321b1ce03789e729ece9f7ee5a16c72de70a2e3372811f7f5038b

Observation d6bd0b00-1e4d-4529-9caf-f9b31182c94c · outbound

This paper cites Law of vision represen- tation in mllms.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Law of vision represen- tation in mllms

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:52.438373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:52.438373Z digest=sha256:3194d7dbb64b9b5a2221885415308fc6ffa05f4cb27a9eeae4e39943eb5fb17b

Observation b2605e8c-5450-46e0-a93d-1563c818a0da · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:52.444166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:52.444166Z digest=sha256:97b3df4a49dbff304b323f995a3a724343adb7525bf80d398f5669023aa9561c

Observation 5a448751-c2b9-4eca-96e4-baba4f97c04c · outbound

This paper cites A Survey on Multimodal Large Language Models.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion A Survey on Multimodal Large Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:52.450553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:52.450553Z digest=sha256:69b8c306466732206fe9f5a0a75ada7fba94c2efd32bd2f81fc3d4757f2ca031

Observation 5486dcd6-2b10-486f-9432-c81a8ef32a2d · outbound

This paper cites Language models are super mario: Absorbing abilities from homologous models as a free lunch.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Language models are super mario: Absorbing abilities from homologous models as a free lunch

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:53.250350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:35:52.455827Z digest=sha256:afa149d2b3e307211cdaddc32f3abf027a80ae1cf57d35a9e042e220c945898e

Observation ba3966b2-3d3e-4c1e-922a-3b6331e1f0c1 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for ex- pert agi.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for ex- pert agi

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:53.232358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:35:52.460798Z digest=sha256:0011d1b9c2ac96bdc5a14373646991cdb119ffb36a3164f4417ea257816d1eed

Observation 558f1a50-6dac-4dea-b57e-712b593cedeb · outbound

This paper cites LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:52.465577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:52.465577Z digest=sha256:d70d88b9ec08d4ae7f745a83d27e067f107302ba2ed8c3b23cc201ecb2b1414b

Observation 2293816a-e9c4-45b7-9785-2041aec0a79d · outbound

This paper cites Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:52.470510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:52.470510Z digest=sha256:52cdccc977c5c0aaa87374bdb34a32ac959309bcc1a0e943d8d3d16699f3729c

Observation af8029f1-3b68-4ea6-857e-e5fb67d7643a · outbound

This paper cites SVIT: Scaling up Visual Instruction Tuning.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion SVIT: Scaling up Visual Instruction Tuning

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:52.475515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:52.475515Z digest=sha256:ad3d83acd72241fcb4bc2ba4f538e8b4502dec049dbc65851afd9ce27bb8d4fa

Observation 35fc70fc-8ca2-46a9-8ed4-77b774c000ba · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:52.480788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:52.480788Z digest=sha256:17661b085870db161733bf30f7d2e926a3807f52ac30479e03c01fc444514d5c

Pith citing papers

Observation 81cc936d-1ccb-4918-be52-a8ad35fac335 · inbound

PivotMerge: Bridging Heterogeneous Multimodal Pre-training via Post-Alignment Model Merging cites this paper.

PivotMerge: Bridging Heterogeneous Multimodal Pre-training via Post-Alignment Model Merging Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:06:52.904006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T07:04:54.909073Z digest=sha256:c2f3fb0b48db0d5ccdee0879324364abc403aa213c10dd3b7338531f8ad82993

Observation d395fd3b-2d8b-4ac5-9735-002f65d9c99b · inbound

Closed-Form Spectral Regularization for Multi-Task Model Merging cites this paper.

Closed-Form Spectral Regularization for Multi-Task Model Merging Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-02T16:27:09.403073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-27T22:40:00.510742Z digest=sha256:4351288f781294224a8ab149ad2ba96c85793c54e9f5cfc05a3b16fced9329bf