Pith. sign in

Paper Citation Record · LEDGER

Towards LLM-Centric Multimodal Fusion: A Survey on Integration Strategies and Techniques

As of 10 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2506.04788.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.04788 v1

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:36:56.237018Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

13 of 13 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 426e823d-bf67-4157-8c7e-5a3c82d199f3 · outbound

This paper cites GPT-4 Technical Report.

Towards LLM-Centric Multimodal Fusion: A Survey on Integration Strategies and Techniques GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T10:36:56.180523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:36:56.180523Z digest=sha256:f7f811fab65eb7cecc004607322df2e2f16c9d87e7d40eb1504b748cf3394b2d

Observation eded0f9a-9741-4aa3-b944-1b04c6b765ce · outbound

This paper cites wav2vec: Unsupervised Pre-training for Speech Recognition.

Towards LLM-Centric Multimodal Fusion: A Survey on Integration Strategies and Techniques wav2vec: Unsupervised Pre-training for Speech Recognition

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T10:36:56.191276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:36:56.191276Z digest=sha256:b0c418de1ae3c514fe027e86d812eb90c2a2414a06ce7ab7fec45e7bb86a733e

Observation 500868d0-b501-4c3c-965b-4714d8c7625b · outbound

This paper cites InIEEE/CVF Confer- ence on Computer Vision and Pattern Recognition, CVPR 2024, Seattle, WA, USA, June 16-22, 2024, pages 15120–15130.

Towards LLM-Centric Multimodal Fusion: A Survey on Integration Strategies and Techniques InIEEE/CVF Confer- ence on Computer Vision and Pattern Recognition, CVPR 2024, Seattle, WA, USA, June 16-22, 2024, pages 15120–15130

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:36:56.754264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T10:36:56.196880Z digest=sha256:22e89cecdc9b07dca032cca8f16fa6734b3684d4d71361330a330ede2bc26ea6

Observation 935f29c8-819d-44e0-9ed8-0b5acf5b550a · outbound

This paper cites Jailbreak in pieces: Compositional Adversarial Attacks on Multi-Modal Language Models.

Towards LLM-Centric Multimodal Fusion: A Survey on Integration Strategies and Techniques Jailbreak in pieces: Compositional Adversarial Attacks on Multi-Modal Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T10:36:56.201686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:36:56.201686Z digest=sha256:2b4b0197a1dc6d70e2c86db291264a671bc82f833c64a61c0a3d1f14d2e6ada3

Observation 7d6180a7-6901-41b0-a5dc-fbc53b65ccf8 · outbound

This paper cites Changli Tang, Wenyi Yu, Guangzhi Sun, Xianzhao Chen, Tian Tan, Wei Li, Lu Lu, Zejun Ma, and Chao Zhang.

Towards LLM-Centric Multimodal Fusion: A Survey on Integration Strategies and Techniques Changli Tang, Wenyi Yu, Guangzhi Sun, Xianzhao Chen, Tian Tan, Wei Li, Lu Lu, Zejun Ma, and Chao Zhang

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T10:36:56.212001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:36:56.212001Z digest=sha256:317c6618e78c9ca76fe8b64409f6bbfce983d15c08e2e3b4f1c7f2b433d0ad22

Observation 0f0e7a96-b62c-437c-9140-88118de914b4 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Towards LLM-Centric Multimodal Fusion: A Survey on Integration Strategies and Techniques Gemini: A Family of Highly Capable Multimodal Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T10:36:56.216528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:36:56.216528Z digest=sha256:190f683a4f7ac14c38c25c3e9e470e1a52a3764ba7a6e2dc8c7752418f209851

Observation 3514720f-5be1-4c11-a1c7-5f55dd0e539f · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Towards LLM-Centric Multimodal Fusion: A Survey on Integration Strategies and Techniques LLaMA: Open and Efficient Foundation Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T10:36:56.221847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:36:56.221847Z digest=sha256:0baa63022b3c761a8bf40a4ca2ad16e9a9e4d07268ba7107ed17f1593949f59f

Observation 21ca545f-01e0-475d-80f4-ac9c32598554 · outbound

This paper cites Wenhai Wang, Zhe Chen, Xiaokang Chen, Jiannan Wu, Xizhou Zhu, Gang Zeng, Ping Luo, Tong Lu, Jie Zhou, Yu Qiao, and Jifeng Dai.

Towards LLM-Centric Multimodal Fusion: A Survey on Integration Strategies and Techniques Wenhai Wang, Zhe Chen, Xiaokang Chen, Jiannan Wu, Xizhou Zhu, Gang Zeng, Ping Luo, Tong Lu, Jie Zhou, Yu Qiao, and Jifeng Dai

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T10:36:56.226890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:36:56.226890Z digest=sha256:9b3003db208a3e6f367cc775bc5b7df22e131f43a3caeea35c58363d9e5bf71f

Observation 67d5af40-f86a-4030-a61a-90de85db71b9 · outbound

This paper cites MM-LLMs: Recent Advances in MultiModal Large Language Models.

Towards LLM-Centric Multimodal Fusion: A Survey on Integration Strategies and Techniques MM-LLMs: Recent Advances in MultiModal Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T10:36:56.232186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:36:56.232186Z digest=sha256:6866fd71857e09345dd5f66576ebe5387e6e69c14a6fd6211693aaff83d343b1

Observation 8d7b984d-99c0-441e-ba9a-dbb4663cdf44 · outbound

This paper cites an unresolved cited work.

Towards LLM-Centric Multimodal Fusion: A Survey on Integration Strategies and Techniques Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:36:56.738577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T10:36:56.237018Z digest=sha256:dfafb6d54e18045d4d9a4239a8f07b1cb9b83eb3a30f04af9c9b7ec251fd4d54

Observation 48eff9f7-c9bc-4a51-b885-8cc48c7e8ec1 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Towards LLM-Centric Multimodal Fusion: A Survey on Integration Strategies and Techniques DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T10:36:56.173390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:36:56.173390Z digest=sha256:7822b6035cc70b29ef6a104213b56aa17c38f28630a84a04e48623cf93079b4b

Observation 3e11c5f6-141d-4a64-b401-f753d12fb7db · outbound

This paper cites X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning.

Towards LLM-Centric Multimodal Fusion: A Survey on Integration Strategies and Techniques X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T10:36:56.185722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:36:56.185722Z digest=sha256:026dda7349da7e3062fed9d1bb9ef4ee8c396a2638bcc843edca611d291595a4

Observation d8a8afa8-b18b-4ba7-93ce-417e3287ed56 · outbound

This paper cites How to Bridge the Gap between Modalities: Survey on Multimodal Large Language Model.

Towards LLM-Centric Multimodal Fusion: A Survey on Integration Strategies and Techniques How to Bridge the Gap between Modalities: Survey on Multimodal Large Language Model

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T10:36:56.207237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:36:56.207237Z digest=sha256:99ca2e7550319a6878b91cfe61d8ecef2d4942127fa40b27bef68d379e34a95c

Pith citing papers

No inbound Pith citation observations are available.