Pith. sign in

Paper Citation Record · LEDGER

A Review of Multi-Modal Large Language and Vision Models

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2404.01322.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2404.01322 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T18:22:10.238771Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

9
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 08e50ad5-4808-4a8c-8b16-f3d5365b226e · inbound

TOKON: TOKenization-Optimized Normalization for time series analysis with a large language model cites this paper.

TOKON: TOKenization-Optimized Normalization for time series analysis with a large language model A Review of Multi-Modal Large Language and Vision Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T18:22:10.238771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:22:10.238771Z digest=sha256:051c7d3f9085b61cec40b95a77fa8a70c37b12a2f3fb5a8b916e68aaf5a52076

Observation 71f318e2-1579-4182-b768-4d14b299a69d · inbound

Uncertainty-o: One Model-agnostic Framework for Unveiling Uncertainty in Large Multimodal Models cites this paper.

Uncertainty-o: One Model-agnostic Framework for Unveiling Uncertainty in Large Multimodal Models A Review of Multi-Modal Large Language and Vision Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T05:36:15.479733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:36:15.479733Z digest=sha256:8bf784f9dc90246d7491fb24e607eca07674a4236a85ad491a0c2be07a0c1c78

Observation 9830cf20-e016-45ed-bdd3-d6c84b3cfadd · inbound

HKD4VLM: A Progressive Hybrid Knowledge Distillation Framework for Robust Multimodal Hallucination and Factuality Detection in VLMs cites this paper.

HKD4VLM: A Progressive Hybrid Knowledge Distillation Framework for Robust Multimodal Hallucination and Factuality Detection in VLMs A Review of Multi-Modal Large Language and Vision Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T00:41:41.097111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:41:41.097111Z digest=sha256:555bc971d194b9856078be8111c746031e322b632a3f806cbf19ad45a64f9edd

Observation 473f1fd4-803c-41fd-8dc1-524eea4a2891 · inbound

LLMs for LLMs: A Structured Prompting Methodology for Long Legal Documents cites this paper.

LLMs for LLMs: A Structured Prompting Methodology for Long Legal Documents A Review of Multi-Modal Large Language and Vision Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T11:48:43.288927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:48:43.288927Z digest=sha256:71cc4445d4e6197fc6cab5fb5300717a07b16e01d943798648934c04a402e491

Observation a4878e1b-08ea-4c31-b93e-a8c591c66d13 · inbound

Spatiotemporal Knowledge Graphs as Persistent Scene Memory for Embodied Question Answering cites this paper.

Spatiotemporal Knowledge Graphs as Persistent Scene Memory for Embodied Question Answering A Review of Multi-Modal Large Language and Vision Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T12:55:35.364167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:55:35.364167Z digest=sha256:2e391a5eade336343836f9de955db0acbc377ca364f2cdb115cc8bfee6450f7c

Observation bf1612b3-304e-483a-84bf-ea8147f3fb1d · inbound

WildFireVQA: A Large-Scale Radiometric Thermal VQA Benchmark for Aerial Wildfire Monitoring cites this paper.

WildFireVQA: A Large-Scale Radiometric Thermal VQA Benchmark for Aerial Wildfire Monitoring A Review of Multi-Modal Large Language and Vision Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-10T01:10:09.177122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T01:08:34.674972Z digest=sha256:8ad1a1f675aecafb6f41286eb6e10440b2446d0862390d79156acc0b58c0fff3

Observation b10e73ee-6c9e-4744-9d6e-7b6a957cd787 · inbound

Cross-Layer Energy Analysis of Multimodal Training on Grace Hopper Superchips cites this paper.

Cross-Layer Energy Analysis of Multimodal Training on Grace Hopper Superchips A Review of Multi-Modal Large Language and Vision Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:36:08.685148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T15:58:52.425995Z digest=sha256:16b64e204838295a8228e7dcf2cceb0cef8cd7e1fc0bfa433d1aa927e9f95e59

Observation ef9fd723-9775-4162-b6dd-aa42fe10423c · inbound

Cross-Layer Energy Analysis of Multimodal Training on Grace Hopper Superchips cites this paper.

Cross-Layer Energy Analysis of Multimodal Training on Grace Hopper Superchips A Review of Multi-Modal Large Language and Vision Models

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T21:58:47.248139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T15:58:52.425995Z digest=sha256:cc4758cf0a95ac26ac7c875bf9d66ea85096ef31f4441dcfb4f0c4549f3ed6c8

Observation 57d959cc-61c0-4e1c-9a52-f17246a94efe · inbound

Rethinking Video-Language Model from the Language Input Perspective cites this paper.

Rethinking Video-Language Model from the Language Input Perspective A Review of Multi-Modal Large Language and Vision Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:33:28.475273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-29T13:24:46.360149Z digest=sha256:f765932609c5a4e3b645403d3cf5cf99f722cc2b4ddcaee2fc19938dd305ed58