Pith. sign in

Paper Citation Record · LEDGER

SOLO: A Single Transformer for Scalable Vision-Language Modeling

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2407.06438.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.06438 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:49:27.734866Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T16:50:10.233691Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b0c3328d-b256-43ab-91ea-11c8c42329ab · inbound

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding cites this paper.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding SOLO: A Single Transformer for Scalable Vision-Language Modeling

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:08.613191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:08.613191Z digest=sha256:40d79029b762db33c3655de14f38930153954a93f5b83ad564df83f98a057049

Observation ae2dcaa5-bf61-40e2-8bce-da9ee0cf63ce · inbound

LLaVA Steering: Visual Instruction Tuning with 500x Fewer Parameters through Modality Linear Representation-Steering cites this paper.

LLaVA Steering: Visual Instruction Tuning with 500x Fewer Parameters through Modality Linear Representation-Steering SOLO: A Single Transformer for Scalable Vision-Language Modeling

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T14:13:24.434647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:13:24.434647Z digest=sha256:3a384dcc73380edd4e4d57640ce1e0706fb5639a2dd0b2c338e4223bf69b8d62

Observation 6d27094e-ac80-4478-9cc7-737b3866e3d9 · inbound

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer cites this paper.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer SOLO: A Single Transformer for Scalable Vision-Language Modeling

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.454414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.454414Z digest=sha256:8231ec0ca6b459228de32a2755f33269c9bbf91e9c5865d8aed19db7f167019d

Observation 3b1b00db-8e89-435d-aa31-4194798f1758 · inbound

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding cites this paper.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding SOLO: A Single Transformer for Scalable Vision-Language Modeling

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:07.975874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:07.975874Z digest=sha256:5b034dfae7cab24554f1f5fd8473fb261be464a3624e0d7ac3cc07757f760739

Observation 75fee6e8-3154-41aa-9617-1ab5e12fcb06 · inbound

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey cites this paper.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey SOLO: A Single Transformer for Scalable Vision-Language Modeling

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.379180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.379180Z digest=sha256:fcb6c170130dce714016f554f27522944a58e852de7a5dc1b9fdb9b56c77e4e4

Observation d783e8fa-cc1c-4ef8-b4d6-8b90a045d9f8 · inbound

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models cites this paper.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models SOLO: A Single Transformer for Scalable Vision-Language Modeling

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.590338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.590338Z digest=sha256:317c6bae1d89a7656d9c3d5a7221ae10dec66ba0f4f168ebd4b9927ef38c4a33

Observation f7a0b0cc-42e7-4957-8077-93ec3a6cb291 · inbound

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training cites this paper.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training SOLO: A Single Transformer for Scalable Vision-Language Modeling

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.734866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.734866Z digest=sha256:12542f79c3dfad4fcf5c367ffc8915c6a2c090f0e9c794559fb40c0bca75c19d

Observation 545f009c-c476-415b-8c41-ec954171c74e · inbound

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs cites this paper.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs SOLO: A Single Transformer for Scalable Vision-Language Modeling

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:43.915177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:43.915177Z digest=sha256:a6d2957e903da7aecb5c8c245fe070617bc2929e44a26e7672a30ffe96e0f697

Observation c1b79421-ce28-4c47-b15b-bc68e211d958 · inbound

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs cites this paper.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs SOLO: A Single Transformer for Scalable Vision-Language Modeling

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:09.614748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:09.614748Z digest=sha256:8887f546e90cf37061538f7b7b7cdfbac2b8eb0d6530268308ba63da74a4f72c

Observation 1c525bd1-5ec4-4b4f-9005-06ab4929f066 · inbound

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models cites this paper.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models SOLO: A Single Transformer for Scalable Vision-Language Modeling

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:50:10.351624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T16:49:56.373596Z digest=sha256:1060f7441d98f73e0b9c9bf74bf8c3cddb7b88433326c99c66d23dda12061617

Observation b02216c0-a23f-4b29-9223-7e859dbb5a5b · inbound

NarrativeTrack: Evaluating Entity-Centric Reasoning for Narrative Understanding cites this paper.

NarrativeTrack: Evaluating Entity-Centric Reasoning for Narrative Understanding SOLO: A Single Transformer for Scalable Vision-Language Modeling

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T13:01:57.577357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:01:57.577357Z digest=sha256:b2c32ac489f2a196f0a1e48cc520b0621c2115b1d860f867b44488c6e5cc2fbd

Observation 41d5c32b-d851-49e6-af21-ed1bce506d36 · inbound

NarrativeTrack: Evaluating Entity-Centric Reasoning for Narrative Understanding cites this paper.

NarrativeTrack: Evaluating Entity-Centric Reasoning for Narrative Understanding SOLO: A Single Transformer for Scalable Vision-Language Modeling

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T06:36:33.779965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:36:33.779965Z digest=sha256:5628f580149b90b6ebf7361f09d7e781ffeaa7fa3ed0d98341e2bb24b9150a17