Pith. sign in

Paper Citation Record · LEDGER

X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning

As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 19 inbound Pith citation observations for arXiv:2311.18799.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.18799 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 19 of 19 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:48:15.981135Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T08:09:41.279523Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 6f02d1db-a0d8-40b6-b7b2-d537c641e7b9 · inbound

VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs cites this paper.

VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T02:44:53.605377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-11T02:44:53.284345Z digest=sha256:1476a14d6b6661874acc5a0feb33d25c0aac49eaef535eec188d899ff96c14fd

Observation 2017dba7-a8e0-4ae1-85a3-df5a1fdfb783 · inbound

LongVALE: Vision-Audio-Language-Event Benchmark Towards Time-Aware Omni-Modal Perception of Long Videos cites this paper.

LongVALE: Vision-Audio-Language-Event Benchmark Towards Time-Aware Omni-Modal Perception of Long Videos X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T05:53:35.802575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:53:35.802575Z digest=sha256:ee030598dc309205d071d3693fabde21cc39d9e62bb08198d9072a701ec7c98a

Observation c7d9c0ca-fa84-4fcc-9c2c-24f50449ee61 · inbound

Investigating the Scaling Effect of Instruction Templates for Training Multimodal Language Model cites this paper.

Investigating the Scaling Effect of Instruction Templates for Training Multimodal Language Model X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T18:03:22.912520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:03:22.912520Z digest=sha256:5cccd2e8376ba326cdab69fa6d22aba549fead490496c6f0238a5e0fc5ea395e

Observation af7ff0c7-ffa8-4aa4-919c-5a1ac56f26e4 · inbound

Modality-Inconsistent Continual Learning of Multimodal Large Language Models cites this paper.

Modality-Inconsistent Continual Learning of Multimodal Large Language Models X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-23T06:52:40.170982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-23T06:50:12.315919Z digest=sha256:527886434d9b7327685e6566b2b43416e0df5dbae57ee0dbe07393712c2c9b56

Observation 2c48730d-aea6-4791-95e7-6cf2672d3188 · inbound

A Review of Multimodal Explainable Artificial Intelligence: Past, Present and Future cites this paper.

A Review of Multimodal Explainable Artificial Intelligence: Past, Present and Future X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning

Reference 280

Resolution
unresolved
no resolver link, observed 2026-08-11T12:33:41.961306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:33:41.961306Z digest=sha256:27644e2a3f239bdca911aebd3c907c3ff2d95a6521079a89902ca504088fdbac

Observation 9ee2955f-56a1-4c70-a346-44c9a35a680e · inbound

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs cites this paper.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:53.456596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:53.456596Z digest=sha256:66cdb4224e2e8f2c0550ee967071cb5514d24517df48d9c3ea2665d3fe4325fe

Observation 54232287-b629-49b0-9e68-f757043a9a7e · inbound

3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding cites this paper.

3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:12.501353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:12.501353Z digest=sha256:35f4c61bf1486789d0497ae45394c476178e5514917cd81bc9514b6081b3a4bb

Observation e3f1422e-6964-464e-9a20-bf4909b6da1c · inbound

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding cites this paper.

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning

Reference 160

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:20:00.174028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-11T01:19:59.603343Z digest=sha256:b513df3e14c89121e523af721cd47a9a719c1912c9be9f69f2a68b6843a17751

Observation 2089806d-d4d8-4faf-9925-14246f765018 · inbound

Foundational Models for 3D Point Clouds: A Survey and Outlook cites this paper.

Foundational Models for 3D Point Clouds: A Survey and Outlook X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning

Reference 176

Resolution
unresolved
no resolver link, observed 2026-08-09T22:54:24.931481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T22:54:24.931481Z digest=sha256:a357dfb2d4d64f5a25ca701ec5e6b91030e5e11581bb6acd3a41e32b8700b75a

Observation cfebbb1a-da40-42bf-b345-d6b3fb941bfb · inbound

WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs cites this paper.

WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:53:26.404419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-17T05:53:26.066674Z digest=sha256:70498177e7e4d0afae8758df52be85400ef3838376313c47b2ca52b7051a6463

Observation a17bc5c5-59a7-4933-b169-1d602e105c20 · inbound

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning cites this paper.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.981135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.981135Z digest=sha256:d6d1614cebf0a0d24019a3e1620e3db6f0845d6889bfcbb3d13309e2bcf0b495

Observation 7b9bed74-cfd5-43d2-81ed-4a1fdc659e68 · inbound

AuthGuard: Generalizable Deepfake Detection via Language Guidance cites this paper.

AuthGuard: Generalizable Deepfake Detection via Language Guidance X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:46.925326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:45:46.925326Z digest=sha256:b8cd6cc23b03abef8b2fe2b06e76bc3c868f8129847333bf72d1be2b2d60ff4e

Observation 3e11c5f6-141d-4a64-b401-f753d12fb7db · inbound

Towards LLM-Centric Multimodal Fusion: A Survey on Integration Strategies and Techniques cites this paper.

Towards LLM-Centric Multimodal Fusion: A Survey on Integration Strategies and Techniques X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T10:36:56.185722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:36:56.185722Z digest=sha256:89f14cf7f448e0f969701a7ac772e8f06e0663ece0439fa528e376c85b84b3c9

Observation 191b7aef-7ad7-4a45-8d33-809fc4b8aa07 · inbound

SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes cites this paper.

SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T01:05:44.701187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:05:44.701187Z digest=sha256:be3b45a8fa91ab161b7aea6812c289f1ee62293e017dcc43af3b9e0ad51be7b0

Observation 74b74614-b132-4d43-ad93-f59553198c90 · inbound

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning cites this paper.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T17:51:16.687278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:51:16.687278Z digest=sha256:73f04dc488e530d3ced310eaac33f165959b345d902023fffb0007f410315f08

Observation bee4125a-b824-4a19-a701-e96c733a7aa2 · inbound

Mitigating Hallucinations in Multimodal LLMs via Object-aware Preference Optimization cites this paper.

Mitigating Hallucinations in Multimodal LLMs via Object-aware Preference Optimization X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T15:18:35.610779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:18:35.610779Z digest=sha256:298dd69e35507c063721c1d5b2e8d2affb83bbe4b69a8282e996deab1a574087

Observation 1fdfb7a6-5a81-47c0-a436-b1e74a46124e · inbound

Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs cites this paper.

Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:02:06.328981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T02:00:02.786195Z digest=sha256:e63b92826087a23418b1efc4fbc9cb2d2f1bdeeb43a9b3d0c667262d897fbd66

Observation acb3a1a4-d757-4d89-a3fd-b84f2b0363c0 · inbound

Closed-Form Spectral Regularization for Multi-Task Model Merging cites this paper.

Closed-Form Spectral Regularization for Multi-Task Model Merging X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning

Reference 97

Resolution
verified exact
arxiv_id, observed 2026-07-02T16:27:09.382888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T22:40:00.510742Z digest=sha256:f45b033e0b90c3243a8dce7dbf62a0ca1e25ddb1d166db317bd9adb7a42f29c4

Observation 04d39992-225f-4a31-be9b-a27bbdc21052 · inbound

CapRiCorn-1K: A Comprehensive Benchmark for Video Captioning and Subject Referential Consistency Across Temporal Scales cites this paper.

CapRiCorn-1K: A Comprehensive Benchmark for Video Captioning and Subject Referential Consistency Across Temporal Scales X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T08:09:41.281286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T12:09:01.026544Z digest=sha256:f37ae816f8033ee1127383b2441109b8411587866bd68221a24fae9b51c0899e