Pith. sign in

Paper Citation Record · LEDGER

Liquid: Language Models are Scalable and Unified Multi-modal Generators

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 22 inbound Pith citation observations for arXiv:2412.04332.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.04332 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 22 of 22 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T17:08:56.594805Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T00:04:22.401157Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 8dd074de-e5b9-4644-b9b5-bd5ae90f6f6a · inbound

UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths cites this paper.

UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths Liquid: Language Models are Scalable and Unified Multi-modal Generators

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-08T15:24:46.446643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:24:46.446643Z digest=sha256:8978f4fe03827b95f901faf93b46c58f60de8ace368dfb3420ba665a83b2e7bf

Observation 9bbaf9aa-4ec2-4012-b95c-27c78871a41e · inbound

WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation cites this paper.

WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation Liquid: Language Models are Scalable and Unified Multi-modal Generators

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-15T16:24:27.491335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T16:24:27.407376Z digest=sha256:27cec092a03429fc7e4241b9fe63fe519f7caf7ad0d61d49f656f5225575840c

Observation ce410087-09ce-4ded-974b-f0731aa4b796 · inbound

Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation cites this paper.

Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation Liquid: Language Models are Scalable and Unified Multi-modal Generators

Reference 85

Resolution
verified exact
arxiv_id, observed 2026-05-17T07:24:04.698301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T07:24:04.460276Z digest=sha256:22a1352a54654d0e0588ec26cd7d7393a00a594bebe41e3bc7e1ce8e869180b5

Observation 67e2ff16-9551-4c35-a2cc-7fc008b4854b · inbound

Emerging Properties in Unified Multimodal Pretraining cites this paper.

Emerging Properties in Unified Multimodal Pretraining Liquid: Language Models are Scalable and Unified Multi-modal Generators

Reference 85

Resolution
verified exact
arxiv_id, observed 2026-05-10T16:23:42.041939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:23:41.854132Z digest=sha256:c016aaec8e7f136b1713b84eef6abe109b57707c35886661b9e1d8708c1b084e

Observation 547639b4-fe7c-403f-be0c-81d633721b56 · inbound

TokBench: Evaluating Your Visual Tokenizer before Visual Generation cites this paper.

TokBench: Evaluating Your Visual Tokenizer before Visual Generation Liquid: Language Models are Scalable and Unified Multi-modal Generators

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:43.966306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:43.966306Z digest=sha256:755219dbaaa341d4d7f5eca214dc96be2bfbdd13a9b085fa005f6b6df3c15f00

Observation 88636981-0a87-40b5-98da-fadb43d3c281 · inbound

Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation cites this paper.

Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation Liquid: Language Models are Scalable and Unified Multi-modal Generators

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:58.581514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:58.581514Z digest=sha256:d8200806f4592e2a0e8e5c6a4a94e88bbc2d650fed9b5912a5f1d6c402705f69

Observation e308dcb7-501d-44bf-9a0f-d4708432b45f · inbound

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning cites this paper.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning Liquid: Language Models are Scalable and Unified Multi-modal Generators

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:26.366317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:26.366317Z digest=sha256:483f7ad9dfad275ab8290367cf034810d543a44510c3da6a54ede85c218b8734

Observation 09c1505d-c155-403f-a089-bbac8913b2f5 · inbound

Show-o2: Improved Native Unified Multimodal Models cites this paper.

Show-o2: Improved Native Unified Multimodal Models Liquid: Language Models are Scalable and Unified Multi-modal Generators

Reference 118

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:51:15.779964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T18:51:15.428692Z digest=sha256:a85d8eeed2e625a84b596a21c610856a67de974ca36edf030eb21d25b15bd36e

Observation fa027142-f6a4-4ae2-bfcd-7fbd7e199ab2 · inbound

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation cites this paper.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Liquid: Language Models are Scalable and Unified Multi-modal Generators

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.850509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.850509Z digest=sha256:06bc241279001ad19f3853c994f4cfd6d6ca61e9a80de32e8b08fae879d6f17b

Observation e7393e11-46e1-4417-b8ef-3e01f8a137b6 · inbound

Instella-T2I: Pushing the Limits of 1D Discrete Latent Space Image Generation cites this paper.

Instella-T2I: Pushing the Limits of 1D Discrete Latent Space Image Generation Liquid: Language Models are Scalable and Unified Multi-modal Generators

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T22:43:04.510524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:43:04.510524Z digest=sha256:06d1b962d9881c3345d86358257ae11aaa91a8fa7e0663d4d384d4a410ff8d29

Observation 5afcd07c-6d9f-464d-8890-d84d5122f57b · inbound

Reconstruction Alignment Improves Unified Multimodal Models cites this paper.

Reconstruction Alignment Improves Unified Multimodal Models Liquid: Language Models are Scalable and Unified Multi-modal Generators

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-04T22:36:08.238282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T22:36:08.238282Z digest=sha256:63ed7954f31f30485a07d3be145c40582bb3172977b1ffe07ab88392acd8daad

Observation b540e1e8-a5c9-400e-a37f-a14975165cb7 · inbound

AutoNeural: Co-Designing Vision-Language Models for NPU Inference cites this paper.

AutoNeural: Co-Designing Vision-Language Models for NPU Inference Liquid: Language Models are Scalable and Unified Multi-modal Generators

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-03T18:57:00.890684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:57:00.890684Z digest=sha256:fb8acbef30d31bc7978a685753ed6b0aece491cfc2c91191f6e7275a5a0aff99

Observation ed2cb239-2d97-4609-b8e9-b737b8539c91 · inbound

Pseudo-Unification: Entropy Probing Reveals Divergent Information Patterns in Unified Multimodal Models cites this paper.

Pseudo-Unification: Entropy Probing Reveals Divergent Information Patterns in Unified Multimodal Models Liquid: Language Models are Scalable and Unified Multi-modal Generators

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:50:59.197308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:27:14.491492Z digest=sha256:0cb47ac25ee0ec5ec8d9d47c877ffa012e5584148b2606cc46354f13093a3ec3

Observation 45accae2-06e0-437a-a0c2-3dc9140cdc29 · inbound

WinTok: A Win-Win Hybrid Tokenizer via Decomposing Visual Understanding and Generation with Transferable Tokens cites this paper.

WinTok: A Win-Win Hybrid Tokenizer via Decomposing Visual Understanding and Generation with Transferable Tokens Liquid: Language Models are Scalable and Unified Multi-modal Generators

Reference 91

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:08:15.782804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:04:19.761430Z digest=sha256:0e51e2157db6d15ee8f2dd07dc5f647e1297092aba07afe6f1902fc4110c8e24

Observation 510d55b9-c96c-44f8-8248-8422ac28ebf1 · inbound

On the Limits of Token Reduction for Efficient Unified Vision Language Training cites this paper.

On the Limits of Token Reduction for Efficient Unified Vision Language Training Liquid: Language Models are Scalable and Unified Multi-modal Generators

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-06-28T17:02:24.295561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:8ab4558fa5ce626f1b017588533aac4eccc0282cd6f2cc6e026510517ffb2ca6

Observation 0c738ec6-7764-4a1d-a5fd-22cfe2530c27 · inbound

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers cites this paper.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers Liquid: Language Models are Scalable and Unified Multi-modal Generators

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:38:28.909508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:c07d4174e166e7277e5fbedf97aae1ab0e5fb872face662c1147be9192e20aab

Observation 8f2c5e0c-af6d-4e79-b120-7b19fba7d402 · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence Liquid: Language Models are Scalable and Unified Multi-modal Generators

Reference 137

Resolution
verified exact
local_arxiv, observed 2026-07-08T00:04:22.402454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-07T23:59:38.702609Z digest=sha256:20f6dc14ef07008e10ae8bc89a62449d246b19c9bb9de677ecf049a8d8a6afb3

Observation 4b2703cb-d4f9-48e6-a5af-b518d68a27a4 · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence Liquid: Language Models are Scalable and Unified Multi-modal Generators

Reference 137

Resolution
unresolved
no resolver link, observed 2026-07-11T07:46:49.059192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T07:46:49.059192Z digest=sha256:7793833fc6c9dfaf91c6e5cabcd556284145d33dcd083ea622c8db84b08fefb9

Observation 385d2136-e0a2-4319-94e1-3412150f6711 · inbound

Twins: Learn to Predict Unified Representations with Focal Loss cites this paper.

Twins: Learn to Predict Unified Representations with Focal Loss Liquid: Language Models are Scalable and Unified Multi-modal Generators

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-01T04:29:51.029704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T04:29:51.029704Z digest=sha256:40e97c9aed403af97f1bb372dd66a43306df413228e45e90798e5ef9b947515e

Observation 2ef15d77-cc04-4d18-80d0-a35ec502ee27 · inbound

Argus-Unified: Towards A Compact and Economical Unified Model for Image Understanding and Generation cites this paper.

Argus-Unified: Towards A Compact and Economical Unified Model for Image Understanding and Generation Liquid: Language Models are Scalable and Unified Multi-modal Generators

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T02:14:09.687868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:14:09.687868Z digest=sha256:d12583ac3f18e7c96509c624b94104a044d808a0fecdcecda31287c48192fabd

Observation 06fe37aa-195e-4ebb-95c9-68e02a64175f · inbound

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes cites this paper.

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes Liquid: Language Models are Scalable and Unified Multi-modal Generators

Reference 138

Resolution
unresolved
no resolver link, observed 2026-08-06T11:55:28.167608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:55:28.167608Z digest=sha256:7eeacb196a8461b91c624ff3c04eef9536c7d7bbb9f63ddbbd903de278ccdae2

Observation b87b291e-b14d-41af-a597-767d3bd51e63 · inbound

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes cites this paper.

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes Liquid: Language Models are Scalable and Unified Multi-modal Generators

Reference 138

Resolution
unresolved
no resolver link, observed 2026-08-08T17:08:56.594805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:08:56.594805Z digest=sha256:5f74f08a7f8a562c30890a674d486fa8211f1f5e1b96bdfb9f418cd9aa485950