Pith. sign in

Paper Citation Record · LEDGER

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion

As of 15 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 2 inbound Pith citation observations for arXiv:2412.01289.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.01289 v2

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T04:35:52.480788Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T22:40:00.510742Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T16:27:09.401525Z

Reference resolution

56 of 56 outbound references displayed

  • verified exact1
  • verified fuzzy23
  • unresolved32
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e5020a5b-2820-412e-922d-8a01dcd07960 · outbound

This paper cites Llama 3 model card.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Llama 3 model card

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:53.926092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:35:52.161969Z digest=sha256:1ffb762e70174f79bfd806b704c988a243b1da7b7181870a975f1aec0352d008

Observation 579021a5-be40-4b14-b897-eb5b2b030ef5 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Flamingo: a visual language model for few-shot learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:52.169246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:52.169246Z digest=sha256:38678509241224b28620c6220ba7bf21a90c1325adaf047d1d07798e2db8113b

Observation 93f85b83-c205-419c-af0b-3cd0ddf64249 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:52.174871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:52.174871Z digest=sha256:0ab89520d6567e5deb7bf4c0c4a8e7f66d87a5f9a24e8dade928d9ebebef21fe

Observation c84bf1a3-e5d7-4fe6-a94e-c2360674bb86 · outbound

This paper cites Token Merging: Your ViT But Faster.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Token Merging: Your ViT But Faster

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:52.180236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:52.180236Z digest=sha256:6b2a356184a7690e2fa982c25a08de83ee7b74958b2af999b5ea77ad94a8e803

Observation ffb5b431-b27b-4176-90e4-5d67fb9a08d4 · outbound

This paper cites Honeybee: Locality-enhanced projector for multimodal llm.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Honeybee: Locality-enhanced projector for multimodal llm

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:53.891160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:35:52.185358Z digest=sha256:12165ea793bf8824504ffa907c915d664d6043c023ba6c4f3585defca5bdcac7

Observation 578e2d59-2d4a-4796-888a-adb393df60e7 · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:52.190470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:52.190470Z digest=sha256:b3a264ff9cf579ecec39b82dfbc5255e8d8843ed2f4dfcf1c0901fd74fa19160

Observation 89f70929-9fb2-4836-8a7c-a31a4843be58 · outbound

This paper cites An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models, 2024.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models, 2024

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:53.873075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:35:52.196430Z digest=sha256:cb42e76a049024f54d6335498a53060237757d1079a2a767bc520443eaec3db6

Observation 3fd7e438-4b00-4406-ab26-73712da16cb6 · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, march.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, march

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:52.201755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:52.201755Z digest=sha256:f0c5912028938c1cf6c7e60272ef333266f0f1d4dd0ca230e92cf39df2bba822

Observation 55d157bb-c303-4343-a09c-d07f1ad997d5 · outbound

This paper cites Fusing finetuned models for better pretraining.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Fusing finetuned models for better pretraining

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:52.206726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:52.206726Z digest=sha256:33a35da684120107af434916ccff36fa23f4cfbe0214f4ce62e14f4370f1b927

Observation e3c46c07-88a7-4f99-a008-8230dc87b68f · outbound

This paper cites MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:52.211989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:52.211989Z digest=sha256:68a6a3dfe4ce05c71fcf6f2a5d54f0938cfa030310f38a3acc23971239102f38

Observation 1961de4f-e606-4566-904a-1bb5a3654107 · outbound

This paper cites Instructblip: Towards general- purpose vision-language models with instruction tuning,.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Instructblip: Towards general- purpose vision-language models with instruction tuning,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:52.217289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:52.217289Z digest=sha256:56356dcf31adcea03e48f67141c830a5516105748cfc829bab933b62e4559fab

Observation 8dc888e9-7e0c-426d-ba46-4932d7dcc54f · outbound

This paper cites Epsilon Sampling Rocks: Investigating Sampling Strategies for Minimum Bayes Risk Decoding for Machine Translation.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Epsilon Sampling Rocks: Investigating Sampling Strategies for Minimum Bayes Risk Decoding for Machine Translation

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-12T04:35:53.110880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:35:52.222575Z digest=sha256:0a0cd2721e05fd3d512bf6b6b6077f4eef742f88482fe61638cf27efa50933a1

Observation 3736654f-5dd5-4490-a3db-223588607acf · outbound

This paper cites Mme: A compre- hensive evaluation benchmark for multimodal large language models, 2024.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Mme: A compre- hensive evaluation benchmark for multimodal large language models, 2024

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:53.823105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:35:52.229333Z digest=sha256:faa247798413333954a57eba2ac8b46068922d485a3f6004718179b3b47b5179

Observation 432abb0b-d338-4279-8de6-313248a04d2b · outbound

This paper cites ConvLLaVA: Hierarchical Backbones as Visual Encoder for Large Multimodal Models.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion ConvLLaVA: Hierarchical Backbones as Visual Encoder for Large Multimodal Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:52.235865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:52.235865Z digest=sha256:9c0fa1507a2933adde3e658aab4aeeea8e48c52339dd4899399fba39e457d28a

Observation 932ccbd9-163f-4711-8a14-80f66a95fb4a · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:53.800117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:35:52.241613Z digest=sha256:3d15e243f6e530a6154924362fba8c4bdac3bdc2d0db8246593bf54159d010db

Observation 0765a90e-1a92-4eb8-8f00-fd0f3af35a10 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:53.774006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:35:52.247353Z digest=sha256:be313100f9f5d63c3c8b86d50301977b6391b8891401cb2375388ea638738a33

Observation 0014ba9e-0cad-4347-bf10-699e3e4c6b3c · outbound

This paper cites Vizwiz grand challenge: Answering visual questions from blind people.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Vizwiz grand challenge: Answering visual questions from blind people

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:53.753151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:35:52.252591Z digest=sha256:2d35de517cb811c1b6c84170c0a203905b01a1296124c64b5f06e2d419028f1f

Observation 36af12b2-109f-46e6-84e8-0ac063c7138f · outbound

This paper cites Onellm: One framework to align all modalities with language.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Onellm: One framework to align all modalities with language

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:53.734724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:35:52.257413Z digest=sha256:d7e01803a8d95c584a14d2b4d01b08deec743ed5715ec64aef1cd894ad83d549

Observation 83921769-1cbc-4a52-a3ad-df5b70759392 · outbound

This paper cites Editing models with task arithmetic.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Editing models with task arithmetic

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:53.714807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:35:52.262529Z digest=sha256:067da70e721030225b771ce0d2bcc4664b629d0457391e968e0b53540773f9b1

Observation 11de87cf-3771-4388-8e95-300b97321891 · outbound

This paper cites Editing models with task arithmetic.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Editing models with task arithmetic

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:53.695652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:35:52.267261Z digest=sha256:2dbb1930730d06070caea22144839cae899c37a9cf2aa52d70c39f1e1f87f3b7

Observation 9d214fd1-cbe8-4607-b7d8-87a6f45498ff · outbound

This paper cites LLM-Blender: Ensembling Large Language Models with Pairwise Ranking and Generative Fusion.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion LLM-Blender: Ensembling Large Language Models with Pairwise Ranking and Generative Fusion

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:52.272424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:52.272424Z digest=sha256:1b8ad2908f1bc12a1a719fa2c2e54f321bfcf80068eca9ce996c1c215b7d623c

Observation 802b26ba-707d-41ff-8c6b-aaeabe7dd343 · outbound

This paper cites Dataless knowledge fusion by merging weights of language models.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Dataless knowledge fusion by merging weights of language models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:53.675911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:35:52.277765Z digest=sha256:cdf6e5b23b14a9637830d9184aba06c3e2a2ec73577bf78cf4c0f2effd308c0a

Observation 7cd60cc8-56bf-4ac8-a293-760a515e2c89 · outbound

This paper cites A diagram is worth a dozen images.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion A diagram is worth a dozen images

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:52.282873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:52.282873Z digest=sha256:f490748762492e52e78cc424bd0a756aebb8c706fa210f64ce50aab3a5cf55e7

Observation 332759f4-3009-434c-955d-cdd30957e077 · outbound

This paper cites In- troducing idefics: An open reproduction of state-of-the-art visual language model, 2023.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion In- troducing idefics: An open reproduction of state-of-the-art visual language model, 2023

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:53.646125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:35:52.296757Z digest=sha256:f3c8c65a94a4f77bff79ea3bbe2d72f711ac36bc37332a2cf2b2a3f829386169

Observation e9980917-69d3-430b-9f95-603a1244ca12 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:52.302694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:52.302694Z digest=sha256:1dfdffa53f4f235c649333222af57fed0d85dcfa37712ea0c5265c3c3a3a3581

Observation 5e915f62-38f3-4445-8129-ccf5dee9a37b · outbound

This paper cites LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:52.307747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:52.307747Z digest=sha256:cde17f71bd395552648484550c99e7dd076bda7f31d2dccbd8807255a0e7dc9e

Observation 3f797669-e749-44a5-bdc2-08af6022129b · outbound

This paper cites Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:52.313670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:52.313670Z digest=sha256:52d1f9a9ab5474eb40cd70c10d128ed33af868a736e78d9ea78f7965360a103f

Observation 2998fcce-700e-4121-b0b2-7cd9735a91ba · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:52.321158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:52.321158Z digest=sha256:5c4d6ef7e7f9d2bb38b4ffd21661d57ad62f8bdc9f6441b6f8366b2ad7c6e220

Observation 031776af-ef46-435a-b03c-2f7c4bfe9e85 · outbound

This paper cites Vila: On pre-training for visual language models.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Vila: On pre-training for visual language models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:53.611660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:35:52.326881Z digest=sha256:0f5dd0d68ac22b028ea2b16f2ec681db7c04c3978a495ece47d408b19315199a

Observation 8bc5a57e-f4ea-4e02-8cd2-d4d1b5d11ab2 · outbound

This paper cites Microsoft coco: Common objects in context.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Microsoft coco: Common objects in context

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:52.332753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:52.332753Z digest=sha256:d903395d5cb02e1da936873a1c61b7f330103e8dfaa8f479fbf75dd5d1f4aa7e

Observation 96d2cefd-2f49-4e5a-a952-a525211d076b · outbound

This paper cites Mitigating hallucination in large multi-modal models via robust instruction tuning.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Mitigating hallucination in large multi-modal models via robust instruction tuning

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:53.577821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:35:52.339639Z digest=sha256:0c9dcee8e20366c36a5c98f2211c091e7e1c9362af6d9ad9c91a175ce7ab643b

Observation e449a05d-4bb9-40ba-b126-78193aa34fc0 · outbound

This paper cites Improved baselines with visual instruction tuning.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Improved baselines with visual instruction tuning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:53.559725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:35:52.346028Z digest=sha256:4935c04338da2fe109e7c6a59f56fda020f07fe745a755a05d05fa508ecec5fd

Observation 6d7f0d48-151c-4fb3-a4e6-ce8c87379131 · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:53.400194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:35:52.350752Z digest=sha256:79d8dffb58f7995e880e5dceed10a0bdc708bd844a42c05870efd93c69f93b0c

Observation 0073817c-eee9-4e9e-bb05-9a1b029e7253 · outbound

This paper cites Visual instruction tuning.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Visual instruction tuning

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:53.380245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:35:52.356248Z digest=sha256:ac79cc40cfdacecfbc2c1d31ca4ace720af32a5dc274ae690f8aca6520c87406

Observation 300eb4da-51a2-4cc8-b6dd-493e0edca866 · outbound

This paper cites BitDelta: Your Fine-Tune May Only Be Worth One Bit.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion BitDelta: Your Fine-Tune May Only Be Worth One Bit

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:52.361915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:52.361915Z digest=sha256:49d8821a08db06e4de6197cbcc9acff3f8f51632eea30b9113f430622b2dcdf4

Observation 99d3f12d-3706-4a72-8c75-50782528f299 · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion MMBench: Is Your Multi-modal Model an All-around Player?

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:52.368597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:52.368597Z digest=sha256:587db570a7a7bae33945fb3e9723cc312c6b0ef707e0b55529555321914437fa

Observation 2df3300e-0170-4a34-a7f2-adac24513433 · outbound

This paper cites OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:52.374335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:52.374335Z digest=sha256:c0c4380472a3f6407a425cd1696e92fdeee1b60e23bb9aa400069aa9ff4b36ce

Observation d97a5506-0c4f-4577-87f1-af88335367f4 · outbound

This paper cites An Empirical Study of Scaling Instruct-Tuned Large Multimodal Models.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion An Empirical Study of Scaling Instruct-Tuned Large Multimodal Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:52.380002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:52.380002Z digest=sha256:1a8079f75ab38f219248501bdf27164dd6f9050ec9c91aeac4fc2647495005b3

Observation d6c2f7b5-19f5-46c5-aa8e-c34aa7e9cb31 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Learn- ing transferable visual models from natural language super- vision

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:52.385690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:52.385690Z digest=sha256:31e3e1fea4a07b9ea09cf4c97436d9c4128db26c2a9dac287b829bf026cf9e92

Observation a6cc9117-80e7-49c2-b2af-c2da7fac13b2 · outbound

This paper cites Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:52.390853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:52.390853Z digest=sha256:226688b4ac2d0c0481290b0e298c3c7bbdcfe8b679706bd3c897fb2696434afc

Observation b82fd9ad-d49f-40bb-b243-f5d0d3416c85 · outbound

This paper cites Animating rotation with quaternion curves.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Animating rotation with quaternion curves

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:53.340341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:35:52.396328Z digest=sha256:7eddb11f17c231a158b40546adb021cdaccb5aafc109ba3cbc1281ca44c8c7a6

Observation 4d0d2f25-67e1-4bf3-bd8f-22e13d25057e · outbound

This paper cites Towards vqa models that can read.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Towards vqa models that can read

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:53.312247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:35:52.403028Z digest=sha256:6fe18e08bd650b8cd772435813857ea2d898b6399f7e52e95a496472d5741c37

Observation b4bce915-5b0a-4100-9b43-9b89da4f8f55 · outbound

This paper cites Visualizing data using t-sne.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Visualizing data using t-sne

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:53.292775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:35:52.408981Z digest=sha256:8fafa1568ac88c2c8c76505ceab6cc5a79c35d4fb17ee315c8875463bcf0e176

Observation 174fae8c-ad11-4a6c-895d-cc5a8550ca61 · outbound

This paper cites Knowledge Fusion of Large Language Models.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Knowledge Fusion of Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:52.415298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:52.415298Z digest=sha256:26cc93b9de7fd451f8d2a27a01fdd6a735045c92ff37ae8def5b853439e8dee7

Observation 56d56766-116b-49ec-b54f-7c8135a8682b · outbound

This paper cites Grok-1.5 Vision Preview, 2024.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Grok-1.5 Vision Preview, 2024

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:53.268461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:35:52.421218Z digest=sha256:ed7b82c1ff0cca7077addf10c6c22699cf0374bf018920c702e18e86c19b02bf

Observation d522b4b3-3f60-491e-a3bd-ff52039a38aa · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:52.426778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:52.426778Z digest=sha256:df2febfd92a71098799af99fe961ce596399065e57d5ec75d275fecea5fe47df

Observation 6429dc0b-6a57-4882-b451-55182ee5fb1a · outbound

This paper cites TIES-Merging: Resolving Interference When Merging Models.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion TIES-Merging: Resolving Interference When Merging Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:52.432601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:52.432601Z digest=sha256:1442df8d0705897eca253fee9652f73edb5f35edea59fb59d8dfc7631d08821c

Observation d6bd0b00-1e4d-4529-9caf-f9b31182c94c · outbound

This paper cites Law of vision represen- tation in mllms.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Law of vision represen- tation in mllms

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:52.438373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:52.438373Z digest=sha256:21678be4d5fdf6a54ca3a4f037ab87e494317a1a4fc84e67ec998e2efd2c1ca1

Observation b2605e8c-5450-46e0-a93d-1563c818a0da · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:52.444166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:52.444166Z digest=sha256:d4bcfe3851cacf12358661e3fa3289e20d745813dbad38c68c047c94c3912eb1

Observation 5a448751-c2b9-4eca-96e4-baba4f97c04c · outbound

This paper cites A Survey on Multimodal Large Language Models.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion A Survey on Multimodal Large Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:52.450553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:52.450553Z digest=sha256:76126951cbfb9380722138666e329c102ca4a5db72b594554dcfe1897293e714

Observation 5486dcd6-2b10-486f-9432-c81a8ef32a2d · outbound

This paper cites Language models are super mario: Absorbing abilities from homologous models as a free lunch.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Language models are super mario: Absorbing abilities from homologous models as a free lunch

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:53.250350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:35:52.455827Z digest=sha256:12dca1817bce9365324dc41f4a1722a04b8f6f84450c1ccaa45a8304fd9af083

Observation ba3966b2-3d3e-4c1e-922a-3b6331e1f0c1 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for ex- pert agi.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for ex- pert agi

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:53.232358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:35:52.460798Z digest=sha256:b69267623420cb5a3909bb77fa2982e2ff457212e9560310eee64d5829a94825

Observation 558f1a50-6dac-4dea-b57e-712b593cedeb · outbound

This paper cites LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:52.465577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:52.465577Z digest=sha256:9a9a398ad19490966c185c9c6a6cbeaae65354d67996fa189c9038169cb27e7a

Observation 2293816a-e9c4-45b7-9785-2041aec0a79d · outbound

This paper cites Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:52.470510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:52.470510Z digest=sha256:4ac72f0c4f68add4ebc36908aacf28ece95860ee117b688cddb91a93296bba80

Observation af8029f1-3b68-4ea6-857e-e5fb67d7643a · outbound

This paper cites SVIT: Scaling up Visual Instruction Tuning.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion SVIT: Scaling up Visual Instruction Tuning

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:52.475515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:52.475515Z digest=sha256:2d5722643bbb6a42711823149ffda25c9a9d6c365478f5dac7d4a31d7cbefcf1

Observation 35fc70fc-8ca2-46a9-8ed4-77b774c000ba · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:52.480788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:52.480788Z digest=sha256:2559ee2c1f41bc3d34a5a640de14d8ff5297cf8b057ccaed49b78c5f6267098b

Pith citing papers

Observation 81cc936d-1ccb-4918-be52-a8ad35fac335 · inbound

PivotMerge: Bridging Heterogeneous Multimodal Pre-training via Post-Alignment Model Merging cites this paper.

PivotMerge: Bridging Heterogeneous Multimodal Pre-training via Post-Alignment Model Merging Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:06:52.904006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T07:04:54.909073Z digest=sha256:a6709b70aedd5583917dff765bd17a87e7cfa3fc3d9db3a751f8f37707206ecb

Observation d395fd3b-2d8b-4ac5-9735-002f65d9c99b · inbound

Closed-Form Spectral Regularization for Multi-Task Model Merging cites this paper.

Closed-Form Spectral Regularization for Multi-Task Model Merging Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-02T16:27:09.403073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T22:40:00.510742Z digest=sha256:7220b80e73ded5681d05a8c1b2568b60b7285025d0e14d6571fab9ab9f3f3fe7