Pith. sign in

Paper Citation Record · LEDGER

Decoupled Visual Processing: Efficient Multimodal Adaptation via Modality-Specific Transformer Substitution

As of 9 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 0 inbound Pith citation observations for arXiv:2607.26596.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.26596 v1

Coverage vector

measured 15 of 15 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T12:40:56.954792Z

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

15 of 15 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1f88dbb6-2687-4aa1-9f9a-997a3ba183fc · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

Decoupled Visual Processing: Efficient Multimodal Adaptation via Modality-Specific Transformer Substitution MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T12:40:56.926574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:40:56.926574Z digest=sha256:4be103c8a48a8492030e5964f176dec0d71f4f885839465928d94c498f9c56a8

Observation 5d9c196b-f4a7-4bbd-9d2f-052b3c22c11b · outbound

This paper cites Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi.

Decoupled Visual Processing: Efficient Multimodal Adaptation via Modality-Specific Transformer Substitution Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T12:40:56.935423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:40:56.935423Z digest=sha256:3ee57e7b4421829d5d14977abbefa349b20eda896ed0cc18c647ca35e9a12cc2

Observation 5a61cdb0-9ed5-417f-95a6-a18341fcdc43 · outbound

This paper cites VILA: On Pre-training for Visual Language Models.

Decoupled Visual Processing: Efficient Multimodal Adaptation via Modality-Specific Transformer Substitution VILA: On Pre-training for Visual Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T12:40:56.938008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:40:56.938008Z digest=sha256:dea1b7109eb50c4f6be3b5f45c53d1ecdacb27ab15504571bee802915877db12

Observation 830ae66c-781d-41e5-9326-2c2e2ff03646 · outbound

This paper cites Decoupled Weight Decay Regularization.

Decoupled Visual Processing: Efficient Multimodal Adaptation via Modality-Specific Transformer Substitution Decoupled Weight Decay Regularization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T12:40:56.940805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:40:56.940805Z digest=sha256:9fa0f43c2e9ff98598714d67ce3958b2af4e83421804583431d3ac5bbf9435c5

Observation 3100d93d-404d-4521-a8ec-ab1fecd3994b · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

Decoupled Visual Processing: Efficient Multimodal Adaptation via Modality-Specific Transformer Substitution ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T12:40:56.943390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:40:56.943390Z digest=sha256:3722c23bbb745620de334880f538886408b29a37bada3938d450f8d8e2cfcc48

Observation 790da517-9b0a-423b-a228-e98ba5390d4b · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

Decoupled Visual Processing: Efficient Multimodal Adaptation via Modality-Specific Transformer Substitution Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T12:40:56.945948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:40:56.945948Z digest=sha256:be843ebb6a7c20b2650c7e25852ee234f1a06ebddaa0a7af6f60c6ff262c967f

Observation 25be7dd9-6dfe-479e-a1ac-7ed684938b88 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Decoupled Visual Processing: Efficient Multimodal Adaptation via Modality-Specific Transformer Substitution Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T12:40:56.949309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:40:56.949309Z digest=sha256:aa9ef8d0f9f50151570b2388b40a89eb241397db29637846ea0ace7cad05dba9

Observation 00355b0e-bfab-4df1-8d82-d3624c9629dc · outbound

This paper cites FLGo: A Fully Customizable Federated Learning Platform.

Decoupled Visual Processing: Efficient Multimodal Adaptation via Modality-Specific Transformer Substitution FLGo: A Fully Customizable Federated Learning Platform

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T12:40:56.951963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:40:56.951963Z digest=sha256:cee7722c6dcd9291f4ec05ce9109f47042bc314ea0133d3b45435ed3ca03809c

Observation abbab7a4-eef6-4abe-8d74-a2efe25f6549 · outbound

This paper cites mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration.

Decoupled Visual Processing: Efficient Multimodal Adaptation via Modality-Specific Transformer Substitution mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T12:40:56.954792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:40:56.954792Z digest=sha256:e997591c863b498663cec0c541ff92073394a13eacbecca3f7658b2ad7769698

Observation a2d46f7d-6730-4e90-80dc-490bfbf64241 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Decoupled Visual Processing: Efficient Multimodal Adaptation via Modality-Specific Transformer Substitution BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-01T12:40:56.887890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:40:56.887890Z digest=sha256:53607f5721055ca8cf7c18e9e3dabe5d0dac456f3237a0e4fc264bc4e20ceeb9

Observation 6facfb72-b996-4426-bca3-9d5f5b3d3a2b · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Decoupled Visual Processing: Efficient Multimodal Adaptation via Modality-Specific Transformer Substitution LoRA: Low-Rank Adaptation of Large Language Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-01T12:40:56.929777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:40:56.929777Z digest=sha256:a24effde7aa369e0366ae6f66f4b239beeadcde3a1991b0093c22c37ae12c24e

Observation 4e6b62ce-890d-4902-a46d-f722d15d2379 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Decoupled Visual Processing: Efficient Multimodal Adaptation via Modality-Specific Transformer Substitution Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-01T12:40:56.858304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:40:56.858304Z digest=sha256:36622a58efc787e6ec71781348fbe85bbd8c808c148f2dfad819acf6d2ff2f58

Observation 7f387456-358b-4518-b435-c8a1f63056a1 · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

Decoupled Visual Processing: Efficient Multimodal Adaptation via Modality-Specific Transformer Substitution InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-01T12:40:56.871061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:40:56.871061Z digest=sha256:a39490c3220111db5c582589072f7b5de10d8444ad7bcb75c016e3cdd635f34c

Observation 2c30a087-2d53-4d67-ad03-b22958e7e088 · outbound

This paper cites DReSS: Data-driven Regularized Structured Streamlining for Large Language Models.

Decoupled Visual Processing: Efficient Multimodal Adaptation via Modality-Specific Transformer Substitution DReSS: Data-driven Regularized Structured Streamlining for Large Language Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-01T12:40:56.907968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:40:56.907968Z digest=sha256:cd54cd285a2a3b1892e356cbfc732006b939f00d062dbcecc9e6560030a95fbe

Observation b1902258-08e3-4dc7-bbb3-d3a19e07e521 · outbound

This paper cites Exploring Knowledge Purification in Multi-Teacher Knowledge Distillation for LLMs.

Decoupled Visual Processing: Efficient Multimodal Adaptation via Modality-Specific Transformer Substitution Exploring Knowledge Purification in Multi-Teacher Knowledge Distillation for LLMs

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-01T12:40:56.932940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:40:56.932940Z digest=sha256:b939e567c392839f422f4741f73c6f002bf87695a93b774043cf1ac2ec0bcdc7

Pith citing papers

No inbound Pith citation observations are available.