Pith. sign in

Paper Citation Record · LEDGER

The Evolution of Multimodal Model Architectures

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2405.17927.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.17927 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T20:06:00.937328Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T23:45:08.065399Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 4d6ce7cc-7417-4886-94d8-b9146d79aa76 · inbound

Granite Vision: a lightweight, open-source multimodal model for enterprise Intelligence cites this paper.

Granite Vision: a lightweight, open-source multimodal model for enterprise Intelligence The Evolution of Multimodal Model Architectures

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-07T20:06:00.937328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T20:06:00.937328Z digest=sha256:36e4cdd4f7063fb516e4ab16a953c4a4f7e72593e85fce12dc7b21336366970d

Observation 03b7c11b-16a0-4b97-8a9c-c4fa9e3062d9 · inbound

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning cites this paper.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning The Evolution of Multimodal Model Architectures

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T00:33:50.954145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:cf08d9714a395f8c8fb7bb7cf0f2565e45e322065932fc0da1d756412d9576ac

Observation 83ca2660-bd83-4b12-8151-efb3a4872338 · inbound

ViFusion: In-Network Tensor Fusion for Scalable Video Feature Indexing cites this paper.

ViFusion: In-Network Tensor Fusion for Scalable Video Feature Indexing The Evolution of Multimodal Model Architectures

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:27.380440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:49:27.380440Z digest=sha256:ef683af6d80ba0f7c902027c38f015b720a437f6108076f34047217fe580e976

Observation 14335849-9096-477c-bd25-60bc4fe18231 · inbound

When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models cites this paper.

When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models The Evolution of Multimodal Model Architectures

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-12T00:51:15.302376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T00:50:06.234399Z digest=sha256:c84dde6598227a935facccb810dd53a29ce2b8a2a34676f953c436a5e8cd35fe

Observation 2a0004dc-de41-4e7f-ae82-9c1ce8d5f34a · inbound

When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models cites this paper.

When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models The Evolution of Multimodal Model Architectures

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:22:58.955860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-14T21:21:37.185232Z digest=sha256:ad326b41d58a86bbf745ef7f0b3c95764401b1c4405aac5aa44ee1dd463827d4

Observation a79c9ba7-19ee-41d0-acee-fcdca699fd3d · inbound

When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models cites this paper.

When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models The Evolution of Multimodal Model Architectures

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-19T17:07:41.708283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T17:06:03.483247Z digest=sha256:225f4d2032352e4b4128d466505af9a3c0c42a49b13077cfa0957ab01cb4cf7b

Observation b4684228-826d-40d9-add3-cb758875e017 · inbound

When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models cites this paper.

When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models The Evolution of Multimodal Model Architectures

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-06-30T23:45:08.068208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T23:41:36.199099Z digest=sha256:8ac1ceddc13833e267db16dd3bff2b85546f4e9e3e19d510b6e2bf7ae5fef5c7