Pith. sign in

Paper Citation Record · LEDGER

VILA: On Pre-training for Visual Language Models

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 33 inbound Pith citation observations for arXiv:2312.07533.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.07533 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 33 of 33 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T12:20:08.455261Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T12:09:48.871684Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 5aa228ac-1987-4675-bf6f-e0fe34447313 · inbound

MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI cites this paper.

MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI VILA: On Pre-training for Visual Language Models

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:37:41.639085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T05:37:41.401736Z digest=sha256:547925a23fcdf29c56d89f86b929ef485d77cfeeb1b4344dc4725c00dce86951

Observation 75b31be1-6db3-41b0-bb1b-15da1e94e251 · inbound

Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception cites this paper.

Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception VILA: On Pre-training for Visual Language Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:19:28.026823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T00:19:27.965902Z digest=sha256:01fa41dec23eccf711bdacfc04170fea29a454204390a5b2849d74c453c178d0

Observation e62c1d3d-15ac-4a15-ab2e-037fa7173fb2 · inbound

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training cites this paper.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training VILA: On Pre-training for Visual Language Models

Reference 71

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T04:09:36.503622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:2b278738924ce6f18b6f696177f3ea806a250b6f5ccafa2b417ef6183035bd2c

Observation 84ad2ef8-c196-4b86-b7ff-8a212b9abe75 · inbound

PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning cites this paper.

PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning VILA: On Pre-training for Visual Language Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:21:57.992998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T20:21:57.873354Z digest=sha256:ead2e8267232f95e4270fef10ab9226d0d9255828234da35477243fad799a83e

Observation 61e5d36f-238b-4510-8d74-148bc6e0df71 · inbound

OpenVLA: An Open-Source Vision-Language-Action Model cites this paper.

OpenVLA: An Open-Source Vision-Language-Action Model VILA: On Pre-training for Visual Language Models

Reference 88

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:46:36.498450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T14:46:35.942338Z digest=sha256:825a3a4b1b19621e0721a323bfeb0d51a3051c37a49bca80d7efe55d81f2e434

Observation 5a8e2401-59df-4aae-86f0-e0c8011f7464 · inbound

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models cites this paper.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models VILA: On Pre-training for Visual Language Models

Reference 227

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:20:36.442753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:204783c8dfd9687e8270601389f68085af4d76c769858e982bbecc53d82cfd1e

Observation 588643b7-9054-49ca-8bae-8de9a918f330 · inbound

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos cites this paper.

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos VILA: On Pre-training for Visual Language Models

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-14T00:32:41.131367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T00:32:41.059558Z digest=sha256:e349d9da2e64f7d0362b1a5ede0dfad0707f03478d3aadcf04b92d5def5c624e

Observation 229832af-0b80-48e4-9f71-cbd0c6446fa8 · inbound

Vision-Language Models for Edge Networks: A Comprehensive Survey cites this paper.

Vision-Language Models for Edge Networks: A Comprehensive Survey VILA: On Pre-training for Visual Language Models

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-08T12:20:08.455261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:20:08.455261Z digest=sha256:494a696ef788ba6fb70063c1c47266237400bce69685067f3e2da1af43618b2a

Observation 64b238b0-36c6-48f8-a9b6-8fcd4efb3e7f · inbound

Affordance Benchmark for MLLMs cites this paper.

Affordance Benchmark for MLLMs VILA: On Pre-training for Visual Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:53.733092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:59:53.733092Z digest=sha256:615c76a3af6104be3ede6da07f2ef5f7bbeefc03c2990cc15489a4c5112a61a3

Observation 7893db75-165d-4948-a513-86e9cacc1672 · inbound

SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics cites this paper.

SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics VILA: On Pre-training for Visual Language Models

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:22:37.461989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T21:22:36.902119Z digest=sha256:40a6533ba95d22f4d2c59216e91c44608f916ccc96a1129ac75b673faba5f3d4

Observation bbeed42a-e77b-437f-b766-638f6d431e27 · inbound

GenRecal: Generation after Recalibration from Large to Small Vision-Language Models cites this paper.

GenRecal: Generation after Recalibration from Large to Small Vision-Language Models VILA: On Pre-training for Visual Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T23:57:24.185844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:57:24.185844Z digest=sha256:19f9ad4be046f2515251b82d46ea92cad78b5dfdd9a5acfe4a4c8779fba2cc29

Observation 3c15058f-a19e-4f63-9807-89b51860cc96 · inbound

A Survey on Autonomy-Induced Security Risks in Large Model-Based Agents cites this paper.

A Survey on Autonomy-Induced Security Risks in Large Model-Based Agents VILA: On Pre-training for Visual Language Models

Reference 139

Resolution
unresolved
no resolver link, observed 2026-08-06T21:34:45.150075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:34:45.150075Z digest=sha256:ad8aadaf397792cd288641157061b91130c55b81f255553cd7dcbe3bf2026e63

Observation 91d85d33-4de6-46c1-9662-ac35ff987eda · inbound

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding cites this paper.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding VILA: On Pre-training for Visual Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:52.368568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:52.368568Z digest=sha256:b8b7872429af2b2725bf641a1d681102b57ddceb0dbc2315a7a811e7f2b9e375

Observation 08768bc4-df80-43f8-a3c7-0cdc7de55359 · inbound

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model cites this paper.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model VILA: On Pre-training for Visual Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:10.522038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:10.522038Z digest=sha256:81caf522dd05c285ba0fe661f019a0e182133eb3beb85e8c028d413aa1d0347e

Observation 5abc18a0-57c1-41e4-9b75-a1c7165fae67 · inbound

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning cites this paper.

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning VILA: On Pre-training for Visual Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T05:36:18.864386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:36:18.864386Z digest=sha256:d226629c511ebb30e49c17dd3c97bf33875d0f500e8e0188f0a9291129c4bf29

Observation 1354d313-04a2-4e86-a48a-d63ce58b7a6e · inbound

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models cites this paper.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models VILA: On Pre-training for Visual Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.622377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.622377Z digest=sha256:a72cd38331c52dc1339d35bbd299552ecc2675f226377e3873f86a1d83636aa6

Observation fb3d657b-2e0a-4e38-abd9-018bed994755 · inbound

BcQLM: Efficient Vision-Language Understanding with Distilled Q-Gated Cross-Modal Fusion cites this paper.

BcQLM: Efficient Vision-Language Understanding with Distilled Q-Gated Cross-Modal Fusion VILA: On Pre-training for Visual Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T20:20:31.990067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T20:20:31.990067Z digest=sha256:dd18240ca0f9558c6240a3634f755d7aa8822df5403796a4abc658af59358426

Observation 564b90a6-2952-4def-9ddf-883dc07b490e · inbound

Estimating the Empowerment of Language Model Agents cites this paper.

Estimating the Empowerment of Language Model Agents VILA: On Pre-training for Visual Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T14:53:17.092139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:53:17.092139Z digest=sha256:a53a04333973579ec9d43dfb803cde4937a83e3cd86c9c91320d0be773b9f102

Observation 1b5e48f3-1ec3-40f7-91cf-92a9edad9b97 · inbound

Video Reasoning without Training cites this paper.

Video Reasoning without Training VILA: On Pre-training for Visual Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T09:12:08.258287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T09:12:08.258287Z digest=sha256:e14cf1f2c0db63555ca2ab80f7112653e238749939f33a90b40612c6d9d1fc97

Observation cfa2b176-0507-42e7-9483-4c2c5a486be4 · inbound

SpaceTools: Tool-Augmented Spatial Reasoning via Double Interactive RL cites this paper.

SpaceTools: Tool-Augmented Spatial Reasoning via Double Interactive RL VILA: On Pre-training for Visual Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T18:42:11.790946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:42:11.790946Z digest=sha256:f2d73323f99250816b466e68dfd5abb9966ba73a2465ab9484e749aff922e0c2

Observation 03fd888e-7c74-439c-9582-e6d7ba631392 · inbound

HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video Understanding cites this paper.

HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video Understanding VILA: On Pre-training for Visual Language Models

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:57:53.836328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T12:55:04.564442Z digest=sha256:06da575edaa29e4627031cdff99ec3dadba788304e7964aa1f0e34cebbc02cdf

Observation 86d6886c-7282-433f-990e-073e791dc387 · inbound

See Fair, Speak Truth: Equitable Attention Improves Grounding and Reduces Hallucination in Vision-Language Alignment cites this paper.

See Fair, Speak Truth: Equitable Attention Improves Grounding and Reduces Hallucination in Vision-Language Alignment VILA: On Pre-training for Visual Language Models

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T08:26:00.660703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T16:38:51.797197Z digest=sha256:9c5ad1dd48718c20621688c933e2c211ef40785dd172bb5be17b9e7180af2ce3

Observation bf203738-4503-43dc-b8d6-5d69041428fb · inbound

Back to the Barn with LLAMAs: Evolving Pretrained LLM Backbones in Finetuning Vision Language Models cites this paper.

Back to the Barn with LLAMAs: Evolving Pretrained LLM Backbones in Finetuning Vision Language Models VILA: On Pre-training for Visual Language Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:20:59.239922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T16:07:13.367863Z digest=sha256:df90d9eeb30447acdaa85c4cc0dc445f135b5c531a4fa2e618084e5ff4a873e0

Observation d3678fd2-7160-450c-8434-235fd90be38d · inbound

GoClick: Lightweight Element Grounding Model for Autonomous GUI Interaction cites this paper.

GoClick: Lightweight Element Grounding Model for Autonomous GUI Interaction VILA: On Pre-training for Visual Language Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:36:17.209504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T04:46:22.676901Z digest=sha256:3a6aced934b76dea1886b8082d3149d8482c531c17fcd35a0535a8124b9f5ad0

Observation 1e138bcd-5b29-4938-bf6d-82a014ab1786 · inbound

What Limits Vision-and-Language Navigation ? cites this paper.

What Limits Vision-and-Language Navigation ? VILA: On Pre-training for Visual Language Models

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:02:32.831408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T17:59:58.891151Z digest=sha256:e83350cd7f77d22bb715c3e3a9ad87b75f2dcee0296d8077c28b9f6f41f8045d

Observation 49b95122-0aff-4e75-a057-6ad370c89577 · inbound

Q-GeoMem: Question-Guided Geometric Memory for Video Spatial Reasoning cites this paper.

Q-GeoMem: Question-Guided Geometric Memory for Video Spatial Reasoning VILA: On Pre-training for Visual Language Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:23:50.990383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T18:15:28.261185Z digest=sha256:70bef244bfbd3dced3d4d12a4691c1ff708388be6f4c7540c20b8c008ef02be6

Observation 8b8304d8-1782-40b8-a388-98484736660b · inbound

Q-GeoMem: Question-Guided Geometric Memory for Video Spatial Reasoning cites this paper.

Q-GeoMem: Question-Guided Geometric Memory for Video Spatial Reasoning VILA: On Pre-training for Visual Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-12T15:54:12.909974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T15:54:12.909974Z digest=sha256:b386bfef89da25105ce7de81a22241167b73f7b3f4db04e5ded5607e367f0cfd

Observation 118b7e57-d14b-4d7a-8c91-97878b70e66a · inbound

Reinforcing Dual-Path Reasoning in Spatial Vision Language Models cites this paper.

Reinforcing Dual-Path Reasoning in Spatial Vision Language Models VILA: On Pre-training for Visual Language Models

Reference 104

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T20:08:55.533025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T01:42:30.005911Z digest=sha256:696f4a2fd4e34aef61a9f204cd5515bf2d14e42cb9216b09862eb11a47a04dbe

Observation 5f1e41ae-a87e-4c4d-89e3-d6cd7dc11e84 · inbound

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning cites this paper.

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning VILA: On Pre-training for Visual Language Models

Reference 290

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T06:39:37.446577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T14:19:53.450263Z digest=sha256:6b9aa499c89b5ddc3b9cfc17ba63e7832470a343cb98d8768a550b86b997a69c

Observation 13fe428f-8c6a-4bb8-a1c8-1a4f4c257cc0 · inbound

Dense Reward for Multi-View 3D Reasoning with Global Maps and Local Views cites this paper.

Dense Reward for Multi-View 3D Reasoning with Global Maps and Local Views VILA: On Pre-training for Visual Language Models

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T10:39:45.368680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T08:38:46.044079Z digest=sha256:f4720c06744b3aa66a2ce3c57d98c265cd8290d8bddfa4126aeaacdd8e850b4a

Observation 9c355570-673b-4952-9a46-c4e119daec73 · inbound

Kamera: Unified Position-Invariant Multimodal KV Cache for Training-Free Reuse cites this paper.

Kamera: Unified Position-Invariant Multimodal KV Cache for Training-Free Reuse VILA: On Pre-training for Visual Language Models

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T12:09:48.873083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T07:13:19.593093Z digest=sha256:de455338d1485f9ec793674835742142f748cca9d8ee083bc4a1f72801dba8bb

Observation d1a90f2f-7dbe-4c44-974c-2ec2accbd19b · inbound

ESC: Emotional Self-Correction for Reliable Vision-Language Models cites this paper.

ESC: Emotional Self-Correction for Reliable Vision-Language Models VILA: On Pre-training for Visual Language Models

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-07-03T21:28:58.423966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-03T21:20:00.041277Z digest=sha256:541206e9f325ee816384aefe1cf1cfd7bb5fb72207cde466ac819d72d0973d37

Observation 5a61cdb0-9ed5-417f-95a6-a18341fcdc43 · inbound

Decoupled Visual Processing: Efficient Multimodal Adaptation via Modality-Specific Transformer Substitution cites this paper.

Decoupled Visual Processing: Efficient Multimodal Adaptation via Modality-Specific Transformer Substitution VILA: On Pre-training for Visual Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T12:40:56.938008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:40:56.938008Z digest=sha256:dea1b7109eb50c4f6be3b5f45c53d1ecdacb27ab15504571bee802915877db12