Pith. sign in

Paper Citation Record · LEDGER

The Evolution of Multimodal Model Architectures

As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2405.17927.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.17927 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T20:06:00.937328Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T23:45:08.065399Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 4d6ce7cc-7417-4886-94d8-b9146d79aa76 · inbound

Granite Vision: a lightweight, open-source multimodal model for enterprise Intelligence cites this paper.

Granite Vision: a lightweight, open-source multimodal model for enterprise Intelligence The Evolution of Multimodal Model Architectures

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-07T20:06:00.937328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T20:06:00.937328Z digest=sha256:18ad5170bc33f9a0fc000873f3012d59b6bcb0ef56146cf3af6d33b0528996bc

Observation 03b7c11b-16a0-4b97-8a9c-c4fa9e3062d9 · inbound

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning cites this paper.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning The Evolution of Multimodal Model Architectures

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T00:33:50.954145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:ecf06eefd5b7446ed4e9f7afd8587c9a37df790496db4f6590a96b00f3399354

Observation 83ca2660-bd83-4b12-8151-efb3a4872338 · inbound

ViFusion: In-Network Tensor Fusion for Scalable Video Feature Indexing cites this paper.

ViFusion: In-Network Tensor Fusion for Scalable Video Feature Indexing The Evolution of Multimodal Model Architectures

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:27.380440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:49:27.380440Z digest=sha256:421649a3f1f0fd322ece1ff0c67cb4a47920b0e2dcc1c448e66ea28c54fdffe7

Observation 14335849-9096-477c-bd25-60bc4fe18231 · inbound

When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models cites this paper.

When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models The Evolution of Multimodal Model Architectures

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-12T00:51:15.302376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-12T00:50:06.234399Z digest=sha256:49992e62e193277f81cea6b4e17f379d831b7547974ce045396db5b9ee84a742

Observation 2a0004dc-de41-4e7f-ae82-9c1ce8d5f34a · inbound

When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models cites this paper.

When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models The Evolution of Multimodal Model Architectures

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:22:58.955860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-14T21:21:37.185232Z digest=sha256:0e9a80f3de04c75be0f94a60aae5ad76888224e41753c67265d9f2d388f1874a

Observation a79c9ba7-19ee-41d0-acee-fcdca699fd3d · inbound

When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models cites this paper.

When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models The Evolution of Multimodal Model Architectures

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-19T17:07:41.708283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-19T17:06:03.483247Z digest=sha256:29184b04781382713c43f404e363858b4289bd247c3f272014e25408f8f021d9

Observation b4684228-826d-40d9-add3-cb758875e017 · inbound

When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models cites this paper.

When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models The Evolution of Multimodal Model Architectures

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-06-30T23:45:08.068208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-30T23:41:36.199099Z digest=sha256:d08e03863dc9166598498451ecb596c29e3f952864a15233a514ccaa0f047a8f