Pith. sign in

Paper Citation Record · LEDGER

LLaVA-Phi: Efficient Multi-Modal Assistant with Small Language Model

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 22 inbound Pith citation observations for arXiv:2401.02330.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2401.02330 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 22 of 22 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T19:14:24.357070Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T06:39:37.481871Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 238f4ae0-43af-44c0-ac1d-6f6f576eac31 · inbound

MobileVLM V2: Faster and Stronger Baseline for Vision Language Model cites this paper.

MobileVLM V2: Faster and Stronger Baseline for Vision Language Model LLaVA-Phi: Efficient Multi-Modal Assistant with Small Language Model

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-05-18T15:27:52.145811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T15:27:51.839171Z digest=sha256:f5371f458e3b05c849fb02a327fcbb046e65d48e34db0139ceeb092a4282c9a2

Observation efa3fa37-7825-49ae-ae74-70459f1c9d1b · inbound

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models cites this paper.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models LLaVA-Phi: Efficient Multi-Modal Assistant with Small Language Model

Reference 137

Resolution
verified exact
arxiv_id, observed 2026-05-23T22:20:21.568598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:6fd0c3ffb8e1fb36d99d1aa5c2d9f0a7234981ee2206d2d9a7ab711321303963

Observation 1c655613-c311-4e0f-9bf6-b1980c9068c9 · inbound

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training cites this paper.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training LLaVA-Phi: Efficient Multi-Modal Assistant with Small Language Model

Reference 135

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T04:09:36.430260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:a472726d034d463e3d5d58598b1f8e9abf6277b9225bda88e469a808d8464365

Observation e83d3dbb-328f-4248-b67d-2a7d1d376388 · inbound

SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation cites this paper.

SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation LLaVA-Phi: Efficient Multi-Modal Assistant with Small Language Model

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:48:36.288113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T22:48:36.010306Z digest=sha256:a0d499b0be90a702b2b58e36c1cb94f6f26edbf698cc06c7ec8e7c3561989351

Observation 01874c40-64b1-4a60-ad83-df77265fd892 · inbound

TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation cites this paper.

TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation LLaVA-Phi: Efficient Multi-Modal Assistant with Small Language Model

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-17T16:12:26.041647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-17T16:12:25.980853Z digest=sha256:a73766c613c2cad81d9bb2ea39a87da504b120d3790fad16ae8ce833341a17d6

Observation 064ebc94-c7b9-47c6-956c-da541c1e6655 · inbound

Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation cites this paper.

Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation LLaVA-Phi: Efficient Multi-Modal Assistant with Small Language Model

Reference 96

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:09:16.073820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T22:09:16.001309Z digest=sha256:97b7945ca308d986e4ba2f4b3b7ca2880ca8be45438b0a05bd9fd606b361eafb

Observation bfa17e6d-eb0d-460e-9494-c54ee7549810 · inbound

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning cites this paper.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning LLaVA-Phi: Efficient Multi-Modal Assistant with Small Language Model

Reference 111

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:24.357070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:24.357070Z digest=sha256:99cdcaf7fe983ddbe5ad7edb35c44fd226a1757d6b25193b0c65e801f7e1d00e

Observation 350fbf87-738c-4120-bf83-c1e3e53f517d · inbound

Generalist Virtual Agents: A Survey on Autonomous Agents Across Digital Platforms cites this paper.

Generalist Virtual Agents: A Survey on Autonomous Agents Across Digital Platforms LLaVA-Phi: Efficient Multi-Modal Assistant with Small Language Model

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T19:10:14.361152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:10:14.361152Z digest=sha256:d3fd43292b4ca08a1d495f3fcd54e097897759298497f60c01455154380f9e20

Observation 8433939e-def4-4455-b368-f57453b87aea · inbound

Understanding Museum Exhibits using Vision-Language Reasoning cites this paper.

Understanding Museum Exhibits using Vision-Language Reasoning LLaVA-Phi: Efficient Multi-Modal Assistant with Small Language Model

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:39.788524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:39.788524Z digest=sha256:2d98c03b01e285e719b3bafd87cdab7d04bbaaf16aeab68ae49b6a7a69649037

Observation 01a51fb4-2b19-467e-95e9-18ad1872f3f2 · inbound

Olympus: A Universal Task Router for Computer Vision Tasks cites this paper.

Olympus: A Universal Task Router for Computer Vision Tasks LLaVA-Phi: Efficient Multi-Modal Assistant with Small Language Model

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:06.818547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:06.818547Z digest=sha256:7fed013c2088985766b4816d6d7db69b901c0fe1c8f941ae21a7c1a29aa7c9a3

Observation bf603baa-01b3-46dd-9ca7-8a40f53975ee · inbound

AlzheimerRAG: Multimodal Retrieval Augmented Generation for Clinical Use Cases using PubMed articles cites this paper.

AlzheimerRAG: Multimodal Retrieval Augmented Generation for Clinical Use Cases using PubMed articles LLaVA-Phi: Efficient Multi-Modal Assistant with Small Language Model

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T10:25:34.251640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:25:34.251640Z digest=sha256:a185cd60ff15e1b5a72d198e47d2f1cbd5302c24ba42719d4ef474fb845cc1a1

Observation 066c069c-f4d9-4b05-9906-c9891d41afa8 · inbound

WalkVLM:Aid Visually Impaired People Walking by Vision Language Model cites this paper.

WalkVLM:Aid Visually Impaired People Walking by Vision Language Model LLaVA-Phi: Efficient Multi-Modal Assistant with Small Language Model

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-10T23:12:19.104157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:12:19.104157Z digest=sha256:196e48dd5af980e92f66aa259c55f8db193cd43b1720efce6ea60d62dec79cbc

Observation b47da43e-68bb-4219-90ff-0a9e8e9856fe · inbound

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts cites this paper.

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts LLaVA-Phi: Efficient Multi-Modal Assistant with Small Language Model

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:12.664774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:42:12.664774Z digest=sha256:6d0d2df499d7eaa629c392f07d227290a2844e6adeab4f6f0de0514fc51796a4

Observation 069d95ed-154c-49f9-8fb8-b65a17131abb · inbound

LLMQuoter: Enhancing RAG Capabilities Through Efficient Quote Extraction From Large Contexts cites this paper.

LLMQuoter: Enhancing RAG Capabilities Through Efficient Quote Extraction From Large Contexts LLaVA-Phi: Efficient Multi-Modal Assistant with Small Language Model

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T21:16:23.985198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:16:23.985198Z digest=sha256:2828ffdb35bac3f78d7b4e60fdba98bdb3d69b128a3054c66cbed91288c90d55

Observation ee3c0e73-0e11-4f6a-9d7f-ae60a25d8028 · inbound

Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling cites this paper.

Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling LLaVA-Phi: Efficient Multi-Modal Assistant with Small Language Model

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:14:53.057371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-11T08:14:52.890145Z digest=sha256:4c874f42b92cc250a3a1dfc639b3c2e03560240fc7b34a68df8ed64a9bd570b7

Observation 22b36fa5-7f81-4cef-a74a-5cb88c287ac6 · inbound

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding cites this paper.

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding LLaVA-Phi: Efficient Multi-Modal Assistant with Small Language Model

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-08T19:32:32.091286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:32:32.091286Z digest=sha256:e87814c7c40014a6ae694f31ada86af01880e66f07cec3594e84ec6beca5075d

Observation c94ca341-21c4-4a6b-970f-82d1e6a2e98f · inbound

Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation cites this paper.

Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation LLaVA-Phi: Efficient Multi-Modal Assistant with Small Language Model

Reference 99

Resolution
verified exact
arxiv_id, observed 2026-05-17T07:24:04.756472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-17T07:24:04.460276Z digest=sha256:7d70a6eccedca7f394629d93ef3ee36e9677a9d9c1f9b93f8c2ee819c0ddf91f

Observation 62bdc963-0009-484b-8998-f6b6db471a92 · inbound

FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocities cites this paper.

FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocities LLaVA-Phi: Efficient Multi-Modal Assistant with Small Language Model

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:59.785169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:04:59.785169Z digest=sha256:1c87862df2a1fbb83a8b829fd24ab24c9fc723e49b9a1a9eee6aeb0f028522d2

Observation e3d9b7ca-2369-4474-9e1e-20e63ef9d394 · inbound

Boosting Embodied AI Agents through Perception-Generation Disaggregation and Asynchronous Pipeline Execution cites this paper.

Boosting Embodied AI Agents through Perception-Generation Disaggregation and Asynchronous Pipeline Execution LLaVA-Phi: Efficient Multi-Modal Assistant with Small Language Model

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-04T18:58:13.064497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:58:13.064497Z digest=sha256:9e97113c02ac3b108250004476da03f1802533d2a8a6c8542401c11392b762c2

Observation 5f37af42-41e5-4448-9e3e-cb67becaef07 · inbound

From Plausibility to Verifiability: Risk-Controlled Generative OCR with Vision-Language Models cites this paper.

From Plausibility to Verifiability: Risk-Controlled Generative OCR with Vision-Language Models LLaVA-Phi: Efficient Multi-Modal Assistant with Small Language Model

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T08:55:19.436731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T08:53:18.268970Z digest=sha256:68a5b622c94702ba233ceb1e86c876ad584ea64029359061cf39092e4855219f

Observation e441337d-b762-42cb-9e2c-65302fb3803d · inbound

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning cites this paper.

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning LLaVA-Phi: Efficient Multi-Modal Assistant with Small Language Model

Reference 285

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T06:39:37.483246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-26T14:19:53.450263Z digest=sha256:3d5ed788f776eeb5d3baa74dbdd9301110fe149ea4c89d214c1e32c8796ac875

Observation 1f0fdd25-f117-41c8-8184-f0f80beee9a8 · inbound

SCOPE and SCION: A Benchmark and an Auditable Reference Pipeline for Schema Induction and Fusion from Text cites this paper.

SCOPE and SCION: A Benchmark and an Auditable Reference Pipeline for Schema Induction and Fusion from Text LLaVA-Phi: Efficient Multi-Modal Assistant with Small Language Model

Reference 165

Resolution
unresolved
no resolver link, observed 2026-08-02T13:37:02.165763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T13:37:02.165763Z digest=sha256:aa3b95d32c102acbc4aacadc8efd3e0b620b8dada3b216ef6afc2c09b9cd6a2b