Pith. sign in

Paper Citation Record · LEDGER

X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2311.18799.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.18799 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:45:46.925326Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T08:09:41.279523Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 6f02d1db-a0d8-40b6-b7b2-d537c641e7b9 · inbound

VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs cites this paper.

VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T02:44:53.605377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T02:44:53.284345Z digest=sha256:f9e67450a1d19bdf129d90008a8e9da24e904be5f71a2fb9f27a7c4504959a44

Observation af7ff0c7-ffa8-4aa4-919c-5a1ac56f26e4 · inbound

Modality-Inconsistent Continual Learning of Multimodal Large Language Models cites this paper.

Modality-Inconsistent Continual Learning of Multimodal Large Language Models X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-23T06:52:40.170982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-23T06:50:12.315919Z digest=sha256:afaa6cd7e5ae3d7a2be949c334a5ec61eb4c6fc476a06cad92964f26d1d28650

Observation e3f1422e-6964-464e-9a20-bf4909b6da1c · inbound

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding cites this paper.

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning

Reference 160

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:20:00.174028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T01:19:59.603343Z digest=sha256:0104a32d12d386ee9379be94787b14bc81a432b4299c1cd26eba51a3fdfa99bb

Observation cfebbb1a-da40-42bf-b345-d6b3fb941bfb · inbound

WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs cites this paper.

WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:53:26.404419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T05:53:26.066674Z digest=sha256:adc9397d521c26259b88836cfd71d46923a9dd6e90d1965f59170a047fd777f4

Observation 7b9bed74-cfd5-43d2-81ed-4a1fdc659e68 · inbound

AuthGuard: Generalizable Deepfake Detection via Language Guidance cites this paper.

AuthGuard: Generalizable Deepfake Detection via Language Guidance X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:46.925326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:45:46.925326Z digest=sha256:022d5113e65a862c5d915ad394e4f0872498b528515464129c081a8e76bf7de0

Observation 3e11c5f6-141d-4a64-b401-f753d12fb7db · inbound

Towards LLM-Centric Multimodal Fusion: A Survey on Integration Strategies and Techniques cites this paper.

Towards LLM-Centric Multimodal Fusion: A Survey on Integration Strategies and Techniques X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T10:36:56.185722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:36:56.185722Z digest=sha256:94cd9ab6e14dc7ce3fe83ba90710e7bc5c1aa4beda54d323f89d87e308d16b59

Observation 191b7aef-7ad7-4a45-8d33-809fc4b8aa07 · inbound

SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes cites this paper.

SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T01:05:44.701187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:05:44.701187Z digest=sha256:6e96c787d80f29755333e7d4dc7e1db63bf89cae2085bebafb2daca7987c053a

Observation bee4125a-b824-4a19-a701-e96c733a7aa2 · inbound

Mitigating Hallucinations in Multimodal LLMs via Object-aware Preference Optimization cites this paper.

Mitigating Hallucinations in Multimodal LLMs via Object-aware Preference Optimization X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T15:18:35.610779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:18:35.610779Z digest=sha256:1a7394489dca5ac84ec6a0c198e98b75b8fb130f44f1bc4967593d0ac6d618c6

Observation 1fdfb7a6-5a81-47c0-a436-b1e74a46124e · inbound

Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs cites this paper.

Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:02:06.328981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T02:00:02.786195Z digest=sha256:ccde06a44ea1ff012b73322234194c129a3247fae36f03bc6678a87b2dd617b4

Observation acb3a1a4-d757-4d89-a3fd-b84f2b0363c0 · inbound

Closed-Form Spectral Regularization for Multi-Task Model Merging cites this paper.

Closed-Form Spectral Regularization for Multi-Task Model Merging X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning

Reference 97

Resolution
verified exact
arxiv_id, observed 2026-07-02T16:27:09.382888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T22:40:00.510742Z digest=sha256:f16b45cbfd694620f2faf1a64acb9aea636aef73d73a9b791db12b428e9749a9

Observation 04d39992-225f-4a31-be9b-a27bbdc21052 · inbound

CapRiCorn-1K: A Comprehensive Benchmark for Video Captioning and Subject Referential Consistency Across Temporal Scales cites this paper.

CapRiCorn-1K: A Comprehensive Benchmark for Video Captioning and Subject Referential Consistency Across Temporal Scales X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T08:09:41.281286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T12:09:01.026544Z digest=sha256:fc63c01dc4fbbf2f4479eacdcfef0ae3de8de078104c3ddf1706d15592070d68