Pith. sign in

Paper Citation Record · LEDGER

LMFusion: Adapting Pretrained Language Models for Multimodal Generation

As of 16 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 51 inbound Pith citation observations for arXiv:2412.15188.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.15188 v4

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T11:37:11.041300Z

measured 75 of 75 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 51 of 51 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:09:10.901259Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

24 of 24 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 2357e8fc-cbdd-4e4d-9800-22e685329399 · outbound

This paper cites Jointly Training Large Autoregressive Multimodal Models.

LMFusion: Adapting Pretrained Language Models for Multimodal Generation Jointly Training Large Autoregressive Multimodal Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:10.936279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:10.936279Z digest=sha256:944c75d4bad10a109f49e556d2f914642d3bb2ba8494b09116f36838e9b94f29

Observation 28098a00-1df7-4891-8855-94f5ea28dac9 · outbound

This paper cites The Llama 3 Herd of Models.

LMFusion: Adapting Pretrained Language Models for Multimodal Generation The Llama 3 Herd of Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:10.951261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:10.951261Z digest=sha256:1f69669b1db8706726d2eee7835ae2e34ec7e7451e33fa0dd694722f3edcd97e

Observation e3684a2f-3ec4-403a-95b6-7fad20f51d8b · outbound

This paper cites SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation.

LMFusion: Adapting Pretrained Language Models for Multimodal Generation SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:10.960187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:10.960187Z digest=sha256:cc50f664cc431b7e51afb59637a1b8196ec1da05a9640e42f15674bd1c911e89

Observation 942c7eae-ddf2-43b5-a237-6a86330fb483 · outbound

This paper cites Auto-Encoding Variational Bayes.

LMFusion: Adapting Pretrained Language Models for Multimodal Generation Auto-Encoding Variational Bayes

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:10.969318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:10.969318Z digest=sha256:5142107ab95659bdf8b5ad222ec5a7f2549e5eaf63730e8135283778f1e3940b

Observation b3e75af6-e46b-4bd4-965c-d3c1240024cd · outbound

This paper cites GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding.

LMFusion: Adapting Pretrained Language Models for Multimodal Generation GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:10.973378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:10.973378Z digest=sha256:c20ff9278746593eaaa6a072b476743528af934f5f3b798675cfdfa224ca65f4

Observation e0417250-e716-417a-8e28-60b4761ef14a · outbound

This paper cites MoE-LLaVA: Mixture of Experts for Large Vision-Language Models.

LMFusion: Adapting Pretrained Language Models for Multimodal Generation MoE-LLaVA: Mixture of Experts for Large Vision-Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:10.977006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:10.977006Z digest=sha256:a4132bd3aaa60f86635ceca5b32e84af313037db90772b97f4a47251cebad33e

Observation 33177bf8-3934-4482-ac5f-5bcc1fedcd81 · outbound

This paper cites Improved baselines with visual instruction tuning.

LMFusion: Adapting Pretrained Language Models for Multimodal Generation Improved baselines with visual instruction tuning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:37:11.345509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T11:37:10.986658Z digest=sha256:ef8dc6b48f93e64959d39561c32e5ecffb6f02f70dc45d0dd48529d71de9202d

Observation fed01ad0-89a5-4f7b-a333-0bf396f2cffb · outbound

This paper cites ChartQA: A benchmark for question answering about charts with visual and logical reasoning.

LMFusion: Adapting Pretrained Language Models for Multimodal Generation ChartQA: A benchmark for question answering about charts with visual and logical reasoning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:10.991130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:10.991130Z digest=sha256:41ee378727ca2ca1cc023aa1de346685eba805243454ee4e0931dcbac9f42173

Observation 23855cfc-c50c-455d-818f-6ea21c9317fe · outbound

This paper cites OLMoE: Open Mixture-of-Experts Language Models.

LMFusion: Adapting Pretrained Language Models for Multimodal Generation OLMoE: Open Mixture-of-Experts Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:10.995277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:10.995277Z digest=sha256:5033db51415b70eb91e50d70856422b0c839a5575c79a70b6d37781274b8f38a

Observation f60c53a6-26f1-49d9-8de5-35c162883f37 · outbound

This paper cites Social iqa: Commonsense reasoning about social interactions.

LMFusion: Adapting Pretrained Language Models for Multimodal Generation Social iqa: Commonsense reasoning about social interactions

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:37:11.311650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T11:37:11.003822Z digest=sha256:945e8b22a437f97f4de1d08203217943efa5b6b81ef7bba1617422eabe86e45a

Observation 44a38423-3e55-448a-99fa-045b3e9612ad · outbound

This paper cites Emu: Generative Pretraining in Multimodality.

LMFusion: Adapting Pretrained Language Models for Multimodal Generation Emu: Generative Pretraining in Multimodality

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:11.018372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:11.018372Z digest=sha256:2b9a9611db59e7ddcd7afd14ebd69f8162cdcec0031f277804a4cef300edc62f

Observation 138570ee-9ce3-4384-8038-6d3b2b51818d · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

LMFusion: Adapting Pretrained Language Models for Multimodal Generation Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:11.023935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:11.023935Z digest=sha256:8d1f434b2f092c5af4d32f0d84f9d0b47160816fcba4441fa6c0bceb5916637b

Observation 58c6be49-250f-4c1c-bb91-a452eda8afb6 · outbound

This paper cites MetaMorph: Multimodal Understanding and Generation via Instruction Tuning.

LMFusion: Adapting Pretrained Language Models for Multimodal Generation MetaMorph: Multimodal Understanding and Generation via Instruction Tuning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:11.028592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:11.028592Z digest=sha256:f4bf9b552088ad01903ba10df39db0fc0be15d693b629677b3fce426fb9133f1

Observation 4236d180-888b-408d-bf85-976a0421a142 · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

LMFusion: Adapting Pretrained Language Models for Multimodal Generation Emu3: Next-Token Prediction is All You Need

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:11.037017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:11.037017Z digest=sha256:693882e36b4437c4e5a71b3027674c054a0ea6af1fc10732fe59f67ec3e35475

Observation 292f1f6e-ef4a-4b3a-a88b-4bf399200f4a · outbound

This paper cites Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model.

LMFusion: Adapting Pretrained Language Models for Multimodal Generation Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:11.041300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:11.041300Z digest=sha256:8586fde2524269774980ad0db5da7b29ed7f70871608f20365cd4cb31046881b

Observation 65164a7a-6899-474a-967b-4f3cf16dca0f · outbound

This paper cites MoMa: Efficient Early-Fusion Pre-training with Mixture of Modality-Aware Experts.

LMFusion: Adapting Pretrained Language Models for Multimodal Generation MoMa: Efficient Early-Fusion Pre-training with Mixture of Modality-Aware Experts

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:10.981853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:10.981853Z digest=sha256:ef4ab3f47ad177aec110f0361cf15f1b7724018a100a4c934ea66337a0bee636

Observation cb3bc181-11ff-40b6-b434-7e5282954bce · outbound

This paper cites VLMo: Unified Vision-Language Pre-Training with Mixture-of-Modality-Experts.

LMFusion: Adapting Pretrained Language Models for Multimodal Generation VLMo: Unified Vision-Language Pre-Training with Mixture-of-Modality-Experts

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:11.032673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:11.032673Z digest=sha256:85e8108f922aafc9aab09252e2bb64bf24bc9b85b1d09448b3bd46e7a745f6e0

Observation f2447408-82ca-41d8-a704-11e2131eea0b · outbound

This paper cites Scaling Vision-Language Models with Sparse Mixture of Experts.

LMFusion: Adapting Pretrained Language Models for Multimodal Generation Scaling Vision-Language Models with Sparse Mixture of Experts

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:11.014189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:11.014189Z digest=sha256:c5ec10c5bdd0753102b86304393687f16dc5b9f4b30dec655c78d9e9e2113353

Observation 676349cd-9188-40ac-badf-ff5c019aea5f · outbound

This paper cites Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer.

LMFusion: Adapting Pretrained Language Models for Multimodal Generation Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:11.008771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:11.008771Z digest=sha256:ab1123053b310141bc0cabd28f9ed7c856f15060d848184da4290beff8441ed5

Observation c6deebb6-705d-4903-bc10-e97a54c52f78 · outbound

This paper cites EVE: Efficient Vision-Language Pre-training with Masked Prediction and Modality-Aware MoE.

LMFusion: Adapting Pretrained Language Models for Multimodal Generation EVE: Efficient Vision-Language Pre-training with Masked Prediction and Modality-Aware MoE

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:10.941520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:10.941520Z digest=sha256:65baab7f5731871c0360d2c15b6edcc6dfbaa56d7cb648e7b0a457cf173cff75

Observation e992d828-5c0b-4ca8-b2de-11554d5e041c · outbound

This paper cites U-net: Convolutional networks for biomedical image segmentation.

LMFusion: Adapting Pretrained Language Models for Multimodal Generation U-net: Convolutional networks for biomedical image segmentation

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:37:11.324797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T11:37:10.999642Z digest=sha256:8c985f25f3a110dd8fc16d389ebbbde788196ecfe5710d435cd11b2b000a163b

Observation eab11fe3-250f-456b-b8b5-3c52585ccb0f · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

LMFusion: Adapting Pretrained Language Models for Multimodal Generation MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:10.955844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:10.955844Z digest=sha256:43fee223db50e8a2e97115e15d095ad35e84bc8718377bd48e533307f9e87d5f

Observation f151846d-3be3-4942-b2b2-6a9a7eaed538 · outbound

This paper cites DreamLLM: Synergistic Multimodal Comprehension and Creation.

LMFusion: Adapting Pretrained Language Models for Multimodal Generation DreamLLM: Synergistic Multimodal Comprehension and Creation

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:10.946403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:10.946403Z digest=sha256:64ac56caae7447fa3a8ca68967e3eeaeadcdb145128bbb64f25c32209efc9b18

Observation 80703a0e-fd71-4e46-9967-6a9dceae9847 · outbound

This paper cites MARS: Mixture of Auto-Regressive Models for Fine-grained Text-to-image Synthesis.

LMFusion: Adapting Pretrained Language Models for Multimodal Generation MARS: Mixture of Auto-Regressive Models for Fine-grained Text-to-image Synthesis

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:10.964681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:10.964681Z digest=sha256:a453110bd28ebe53bc6dc3e74480fe406e42efb9afa6a31f5cfa46ab6196dbb8

Pith citing papers

Observation 4281b462-436e-40f7-899c-6eae8f1ce502 · inbound

UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths cites this paper.

UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T15:24:46.407707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:24:46.407707Z digest=sha256:a37ead77172ede1a464a7e42c752b7fc982b91d8b1265c2fea659a255dcbf6c7

Observation ee30dcd7-c7ae-46b0-88c6-2244f47e3098 · inbound

Diffusion Instruction Tuning cites this paper.

Diffusion Instruction Tuning LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-09T11:20:02.915960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:20:02.915960Z digest=sha256:b16cecd69794c7f09480f70b4edde4734d8a0583e9fc3c4e846dec9c7eca62fb

Observation 82064ae3-1537-41f5-954b-ed80812f159e · inbound

I Think, Therefore I Diffuse: Enabling Multimodal In-Context Reasoning in Diffusion Models cites this paper.

I Think, Therefore I Diffuse: Enabling Multimodal In-Context Reasoning in Diffusion Models LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-08T10:24:41.573614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:24:41.573614Z digest=sha256:daf2286496500baba7f614c145ce64e92decafa7c64a017955fcba071ec48a15

Observation df39cbdd-3381-441e-a123-c2e256285cdc · inbound

WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation cites this paper.

WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T16:24:27.719941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T16:24:27.407376Z digest=sha256:b34e08c4d5283213ecb67ad55e96572e5ec7eeeec03594183c1bf1e5cdfc658a

Observation b00e0bf0-cad8-4fb7-9243-39b51f257aa3 · inbound

Transfer between Modalities with MetaQueries cites this paper.

Transfer between Modalities with MetaQueries LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T22:49:23.256298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-14T22:49:23.074271Z digest=sha256:6ff72f2aebf6a91a7744cd83e3304963e661d593398097b530df00cfa37ec296

Observation fe53db0e-c508-4791-9d1c-6ea7850c202a · inbound

TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation cites this paper.

TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-15T23:09:10.901259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:09:10.901259Z digest=sha256:b6f108949f3cacccbd617fd4aba0f5d250cd3e214dce122aca914fd39d13902f

Observation f1969a78-60b6-4416-b4b6-16c8b833dcbf · inbound

BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset cites this paper.

BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T23:34:27.059549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T23:34:26.878354Z digest=sha256:d642ece2363a4a0a66220ec83f3f9a87236c537beb925f4039407a7340e093c3

Observation 04671a1f-5f16-4c92-8f25-e16a2ac660be · inbound

Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis cites this paper.

Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T21:21:22.481319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:21:22.481319Z digest=sha256:7c8b45037dab30dd19937254a80f5201c0b41bac2991e1f4887a90a035a44c8d

Observation 51b2df35-9b9c-4a5f-92ae-eaf04c263e21 · inbound

Emerging Properties in Unified Multimodal Pretraining cites this paper.

Emerging Properties in Unified Multimodal Pretraining LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 66

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T16:23:41.942313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T16:23:41.854132Z digest=sha256:2056153c8e7a55a3544ce521960941c776bbd54d2dbfd84e17a0015bae0f9fbf

Observation 49da6413-21c2-4efa-906c-0995830cdede · inbound

Embedding-to-Prefix: Parameter-Efficient Personalization for Pre-Trained Large Language Models cites this paper.

Embedding-to-Prefix: Parameter-Efficient Personalization for Pre-Trained Large Language Models LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T21:02:03.202470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:02:03.202470Z digest=sha256:2bd16fb9c68e0c1354e9c22e1a3b3f6af3d1f559883bff1f9bf7f48d087421d6

Observation 05dac12b-2ec1-4114-b49e-baa9bdda25cf · inbound

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning cites this paper.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:25.463377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:25.463377Z digest=sha256:593f347786e86a585bd2b51f89220cba5128e9d1d39d18aea016f686e202e242

Observation 4e7650c6-7a75-4f80-a7ee-22d476024b72 · inbound

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation cites this paper.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:17.119373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:17.119373Z digest=sha256:4b8dad2ac40fc3b6274a221d6ca63c5b36a2e777a29d191cb641fe72445a06a9

Observation 31f5f6e2-7a61-446a-af72-fef9c92bf846 · inbound

Knowledge Insulating Vision-Language-Action Models: Train Fast, Run Fast, Generalize Better cites this paper.

Knowledge Insulating Vision-Language-Action Models: Train Fast, Run Fast, Generalize Better LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:23.846395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:23.846395Z digest=sha256:e04da9a502857d63bf30f2bfb84d82b8901b268bf16cc9cdf1acf64e4aa81dd1

Observation a98f8182-ac5b-424c-b2ac-401da3970f5a · inbound

LaTtE-Flow: Layerwise Timestep-Expert Flow-based Transformer cites this paper.

LaTtE-Flow: Layerwise Timestep-Expert Flow-based Transformer LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T05:51:42.915375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:51:42.915375Z digest=sha256:ee666867149f161e0c2a97b963a1324fe503c1711531c9433ba0ca6fa04a4396

Observation 58e5e446-5684-49ed-a9fb-a57566674917 · inbound

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation cites this paper.

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:02.245217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:26:02.245217Z digest=sha256:e1b18b07d8bd5e1c170777a06d02cb1c05633eb1532c96dcec72390e2a7426db

Observation 4a9b2bc1-764f-4d42-96fb-7ae08c69dbe3 · inbound

Dreamland: Controllable World Creation with Simulator and Generative Models cites this paper.

Dreamland: Controllable World Creation with Simulator and Generative Models LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:16.360980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:26:16.360980Z digest=sha256:0028d0a3d894e4e8b6f604faf5e5d6c2e12d9f0e66a2ebb5084d83abc02457ce

Observation 73a3cffc-4bfb-4dea-88c1-15afa9fa889f · inbound

Show-o2: Improved Native Unified Multimodal Models cites this paper.

Show-o2: Improved Native Unified Multimodal Models LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 95

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T18:51:15.590013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T18:51:15.428692Z digest=sha256:c035cde0a0566a1a182d7495ba9d1308c2699b350f9abca475c818dac786c6e3

Observation 301bff5f-135a-4629-bd06-220f1bcd1fea · inbound

OmniGen2: Towards Instruction-Aligned Multimodal Generation cites this paper.

OmniGen2: Towards Instruction-Aligned Multimodal Generation LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 68

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T07:52:10.858881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:2e6d984f8a3004e202311adfdd91591b50e84db6c505b839f2f54772f9435c36

Observation cd09d3a2-d46b-49ef-be2a-5baa8f857fa3 · inbound

WordCon: Word-level Typography Control in Scene Text Rendering cites this paper.

WordCon: Word-level Typography Control in Scene Text Rendering LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T22:34:16.184346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:34:16.184346Z digest=sha256:f075370f635ffb028148d78d4cb5bd470a25995682db6f4bf86189ae73fc967f

Observation 72d32ee3-09c9-486d-a573-65be95ab85cf · inbound

Rethinking Discrete Tokens: Treating Them as Conditions for Continuous Autoregressive Image Synthesis cites this paper.

Rethinking Discrete Tokens: Treating Them as Conditions for Continuous Autoregressive Image Synthesis LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T20:49:42.028479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:49:42.028479Z digest=sha256:58174eecd3fb04c6274930b0659006df38a995aefab454663d2d2bd218230019

Observation d2c473d0-10b9-4e03-938d-35c50e31c5ef · inbound

FreeLoRA: Enabling Training-Free LoRA Fusion for Autoregressive Multi-Subject Personalization cites this paper.

FreeLoRA: Enabling Training-Free LoRA Fusion for Autoregressive Multi-Subject Personalization LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T20:46:25.122139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:46:25.122139Z digest=sha256:50adb69474382d08e9bf0be1bdfc538c73ceee1f43361d35f3cb7e8d6ae7092b

Observation 594a0dbc-5e29-48b0-89c5-62155ea44e51 · inbound

X-Omni: Reinforcement Learning Makes Discrete Autoregressive Image Generative Models Great Again cites this paper.

X-Omni: Reinforcement Learning Makes Discrete Autoregressive Image Generative Models Great Again LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T12:10:08.289255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:10:08.289255Z digest=sha256:97a1fe11f49606d72a083d71230f0f25ba9cee2aa398eb9ec91f261f0a427a39

Observation c0b5789b-2dbd-4c63-8b2f-c2f0f5f01f4b · inbound

OmniHuman-1.5: Instilling an Active Mind in Avatars via Cognitive Simulation cites this paper.

OmniHuman-1.5: Instilling an Active Mind in Avatars via Cognitive Simulation LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T16:57:29.686053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:57:29.686053Z digest=sha256:16672badbebf3a248593c094e782f0a4ac02e8cc7118f50f57b157b599decb30

Observation b17a1063-74ee-4200-8211-e01266c29ff6 · inbound

Galaxea Open-World Dataset and G0 Dual-System VLA Model cites this paper.

Galaxea Open-World Dataset and G0 Dual-System VLA Model LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T13:31:09.919265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:31:09.919265Z digest=sha256:dd614384fb756a57357fa49076c00c1ab6d8835325f962c3e1bce5a39a058cc4

Observation fc05f3a1-bb1c-445f-8eda-3a47119e9884 · inbound

OccVLA: Vision-Language-Action Model with Implicit 3D Occupancy Supervision cites this paper.

OccVLA: Vision-Language-Action Model with Implicit 3D Occupancy Supervision LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:15.644541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:26:15.644541Z digest=sha256:c3ddf620372fe5e1681565e04d46919f249fe94f266aaee9e53581fb38d7bad6

Observation 93875204-2588-4310-b563-7bab02756f4c · inbound

Interleaving Reasoning for Better Text-to-Image Generation cites this paper.

Interleaving Reasoning for Better Text-to-Image Generation LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:44.950531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:44.950531Z digest=sha256:c85090aeba0326d98633bbb87aa6ec8d9b00ca429922fff64cde8c310b4343fb

Observation b562f445-e176-474e-9c2d-9b13fbd5371d · inbound

Reconstruction Alignment Improves Unified Multimodal Models cites this paper.

Reconstruction Alignment Improves Unified Multimodal Models LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-04T22:36:08.140951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T22:36:08.140951Z digest=sha256:e813ac96c2defeb5e80c65d579882e31715db0063da47ecc4a585329d7976549

Observation ef364fbd-f8fd-4b0c-bf37-103f5d74a52b · inbound

UniVideo: Unified Understanding, Generation, and Editing for Videos cites this paper.

UniVideo: Unified Understanding, Generation, and Editing for Videos LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T10:49:49.019944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:49:49.019944Z digest=sha256:efc061f5e2fcde243529bc4a9fbd6c29686dcb321f3f8cf05e3e8464fb283633

Observation 527a0908-9afd-40e7-91e4-9cdb4c2ca112 · inbound

SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models cites this paper.

SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T09:54:26.449060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:54:26.449060Z digest=sha256:cdc6a1426fe85a03f5a455e8d637c77be9a6396bd5867308ec87535218a675d6

Observation 70e91f46-7163-4878-a42c-6e83fdf87912 · inbound

FasterVAR: Plug-and-Play Acceleration for Visual Autoregressive Models cites this paper.

FasterVAR: Plug-and-Play Acceleration for Visual Autoregressive Models LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T15:36:51.177913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T15:36:51.177913Z digest=sha256:d0cb539b77250113dad2b2457310975f3ad28b9a3754a2759d0f0ece701f143e

Observation 71837f0d-f365-4e9a-814e-146d94d08b85 · inbound

CG-MLLM: Captioning and Generating 3D content via Multi-modal Large Language Models cites this paper.

CG-MLLM: Captioning and Generating 3D content via Multi-modal Large Language Models LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-21T14:50:14.718128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-21T14:48:21.787919Z digest=sha256:798cb1e6983e748e65db74bdd2d78286b8c697aa72666bad95407007a84f3716

Observation 8854c87b-93c1-45fc-926c-f848d8caded8 · inbound

ChatUMM: Robust Context Tracking for Conversational Interleaved Generation cites this paper.

ChatUMM: Robust Context Tracking for Conversational Interleaved Generation LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T03:57:32.062558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:57:32.062558Z digest=sha256:5613223f75289737f95dbc30470cdd0165bb4ac35afabe596ccaf461583bc460

Observation c89b769c-3ee7-42e6-8247-0d57e1dcebd4 · inbound

LLaMo: Scaling Pretrained Language Models for Unified Motion Understanding and Generation with Continuous Autoregressive Tokens cites this paper.

LLaMo: Scaling Pretrained Language Models for Unified Motion Understanding and Generation with Continuous Autoregressive Tokens LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 58

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T05:02:19.618705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T05:01:11.880003Z digest=sha256:e9f0414277a627c90738d05db194e6cc404dc4841c4fdd8345b0f2281a7796a7

Observation aa3453fd-dc74-4795-95b3-f947d0d1f384 · inbound

Demystifying Video Reasoning cites this paper.

Demystifying Video Reasoning LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-07-13T23:27:11.006580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:27:11.006580Z digest=sha256:5056860293d7e6ace12ac8743bb83620d38196b4e64a60eb56a0b797a5117e22

Observation 8f54f030-7751-43f7-9d1a-645e1db41827 · inbound

Demystifying Video Reasoning cites this paper.

Demystifying Video Reasoning LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-03T02:33:57.876583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:33:57.876583Z digest=sha256:645510cb8bef9c890040bc306aca132931e85d7f48cb334ad1dafa8111897fc7

Observation 10617339-23d3-4317-89c2-12d670fc9617 · inbound

LMGenDrive: Bridging Multimodal Understanding and Generative World Modeling for End-to-End Driving cites this paper.

LMGenDrive: Bridging Multimodal Understanding and Generative World Modeling for End-to-End Driving LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:31:01.398245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T17:36:38.627415Z digest=sha256:686e7d34ce2ec7acd2d8e4a5c9df1af77730a0989544806fb7924e51e609fc07

Observation b0800302-88df-4745-bf7b-dd720fc3f116 · inbound

Free Lunch for Unified Multimodal Models: Enhancing Generation via Reflective Rectification with Inherent Understanding cites this paper.

Free Lunch for Unified Multimodal Models: Enhancing Generation via Reflective Rectification with Inherent Understanding LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T14:15:28.953805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T14:15:02.723774Z digest=sha256:ec2d11abfe7a93e9aa3ffabb9753dd84468c2ca6bcf3e86e2935db90291a24d6

Observation 06ef9691-e5ef-45bd-8e01-14279bf16a8f · inbound

Co-generation of Layout and Shape from Text via Autoregressive 3D Diffusion cites this paper.

Co-generation of Layout and Shape from Text via Autoregressive 3D Diffusion LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T09:23:37.044329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T09:22:10.651922Z digest=sha256:71c81fe2b46f53b975fce494a2aceb09d997d3742b9b61ad05d8e2299cb713f0

Observation c235ae80-05e8-4e10-8822-a65b8dec7d74 · inbound

MMCORE: MultiModal COnnection with Representation Aligned Latent Embeddings cites this paper.

MMCORE: MultiModal COnnection with Representation Aligned Latent Embeddings LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T03:29:21.669427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T03:27:50.144706Z digest=sha256:9cf2f554538cab057b179b388b38d8098ed6c97f51f737b0b0b3dccc69645d48

Observation 66ee8cee-fdc4-41e5-ab3a-66281e56cb1d · inbound

Meta-CoT: Enhancing Granularity and Generalization in Image Editing cites this paper.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 52

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:46:14.691378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:6f24d4ed915bbec6c97191fc64aa9b3c475c7686c565992bbe0f1e88ef8a9c32

Observation 51bd0ddd-055f-465b-a524-64859693240b · inbound

SpatialFusion: Endowing Unified Image Generation with Intrinsic 3D Geometric Awareness cites this paper.

SpatialFusion: Endowing Unified Image Generation with Intrinsic 3D Geometric Awareness LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:46:26.816740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-07T13:45:53.346402Z digest=sha256:ca934290f63120e86d1fd937c9e1ce9aa10ad1272b4df3173345bb230b29c30d

Observation 73068788-af9b-4a9b-89b8-001858cd0320 · inbound

STARFlow2: Bridging Language Models and Normalizing Flows for Unified Multimodal Generation cites this paper.

STARFlow2: Bridging Language Models and Normalizing Flows for Unified Multimodal Generation LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T02:45:57.156891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T02:44:59.644215Z digest=sha256:abb831f894b3ffed3975d5f029882e1f9adaa38f9dff93e2198fa571a9c174d8

Observation c63185a6-3916-499f-8a62-05a744587d3b · inbound

Reversing the Flow: Generation-to-Understanding Synergy in Large Multimodal Models cites this paper.

Reversing the Flow: Generation-to-Understanding Synergy in Large Multimodal Models LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T18:48:53.450985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-20T18:44:54.835575Z digest=sha256:05a478a8ea97ffb0971e798ce3dc2f3d1bada87c38bc7b0abaaa4b6cb882a2a0

Observation 8fe1fa5e-1efa-4d0a-a3b9-8971c93eae2c · inbound

Semantic Generative Tuning for Unified Multimodal Models cites this paper.

Semantic Generative Tuning for Unified Multimodal Models LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:33:14.202728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-20T11:32:24.007847Z digest=sha256:bab6de5a423acfe9672ec3e4e00a352cc6d2c7e164fbebe7b9c6b69a9cb77a06

Observation e554ff05-2944-404f-b193-242c8dc4740e · inbound

Semantic Generative Tuning for Unified Multimodal Models cites this paper.

Semantic Generative Tuning for Unified Multimodal Models LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:35:00.282413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:674a8dfdc9ab20d34ea8705c6f5faee3a73403f96bb008a683375bdd3d5e9f48

Observation 86284e46-aff6-4509-b9a9-68049f183b84 · inbound

Where to Refine, When to Stop: Rethinking Redundancy via Latent Discrepancy for Efficient Visual Autoregressive Generation cites this paper.

Where to Refine, When to Stop: Rethinking Redundancy via Latent Discrepancy for Efficient Visual Autoregressive Generation LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T22:42:46.709383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-28T22:38:26.147074Z digest=sha256:5dca51d0d59f50dc1df9ef8ee108ea2f7dbc9560d312df5f40b20d3538bfbb09

Observation a6e04f3e-7839-49c8-b13a-3a891a308702 · inbound

Polaris: Scaling Up Instruction-Guided Image Generation Towards Millions of Personalized Style Needs cites this paper.

Polaris: Scaling Up Instruction-Guided Image Generation Towards Millions of Personalized Style Needs LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 97

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T22:46:18.726597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-28T15:07:03.439928Z digest=sha256:dda72baaa11fb461a9898934b1fadf192a1d0637897defffbc3211e8980314f3

Observation 74f42b3d-7ebf-4053-b30b-be90316c1160 · inbound

UniTac: A Unified Multimodal Model for Cross-Sensor Tactile Understanding and Generation cites this paper.

UniTac: A Unified Multimodal Model for Cross-Sensor Tactile Understanding and Generation LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-01T10:05:40.624385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-01T05:56:03.597839Z digest=sha256:2ddafd2d194a233dfe8d63018a378eb8490b3ba5b2bda3159278aeaa96f1e2a1

Observation 583ff840-3bba-428a-894e-007a3b67114d · inbound

Streaming Multi-Agent Autoregressive Diffusion Model with World State Registers cites this paper.

Streaming Multi-Agent Autoregressive Diffusion Model with World State Registers LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 126

Resolution
unresolved
no resolver link, observed 2026-08-01T07:02:51.493005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:02:51.493005Z digest=sha256:3da6c44b0217a9b1d9925e4f89f9caffdcb488706eac0830c2628fc18a9b490e

Observation 9c8c05c8-dcfb-46f2-961d-2d9147766e61 · inbound

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes cites this paper.

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-06T11:55:28.069606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:55:28.069606Z digest=sha256:725b7eee1ee47feefdf26f91f87850d3d9ddd14148d21e921f8b16a387bfc07b

Observation 1b47279d-1775-4e41-80f8-d3ed0ac8850b · inbound

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes cites this paper.

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-08T17:08:55.803846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:08:55.803846Z digest=sha256:9c49c4938f817c99f4c6012eea9d72d737ae16556aca169e0c63d363c282184e