Pith. sign in

Paper Citation Record · LEDGER

LMFusion: Adapting Pretrained Language Models for Multimodal Generation

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 45 inbound Pith citation observations for arXiv:2412.15188.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.15188 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 45 of 45 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T17:08:55.803846Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 4281b462-436e-40f7-899c-6eae8f1ce502 · inbound

UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths cites this paper.

UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T15:24:46.407707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:24:46.407707Z digest=sha256:268b95c6b8720430815ad5509087c194219fa70d4af9101a627e644cc9edb8ae

Observation 82064ae3-1537-41f5-954b-ed80812f159e · inbound

I Think, Therefore I Diffuse: Enabling Multimodal In-Context Reasoning in Diffusion Models cites this paper.

I Think, Therefore I Diffuse: Enabling Multimodal In-Context Reasoning in Diffusion Models LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-08T10:24:41.573614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:24:41.573614Z digest=sha256:0420e6716eb012bbebb818870a6fca599f8a162ffd46d4071863ddd282aaf1ee

Observation df39cbdd-3381-441e-a123-c2e256285cdc · inbound

WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation cites this paper.

WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T16:24:27.719941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T16:24:27.407376Z digest=sha256:1f131454d2db922618dcd7c70562bedb68545eb7d86be7d152a6d3c5a11f774b

Observation b00e0bf0-cad8-4fb7-9243-39b51f257aa3 · inbound

Transfer between Modalities with MetaQueries cites this paper.

Transfer between Modalities with MetaQueries LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T22:49:23.256298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-14T22:49:23.074271Z digest=sha256:223f6c2600d1dc68e06dbbc5c19ff712785a07f60003395b7c61d52cb8c7141e

Observation f1969a78-60b6-4416-b4b6-16c8b833dcbf · inbound

BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset cites this paper.

BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T23:34:27.059549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T23:34:26.878354Z digest=sha256:a2f5cdc258246e207cc939450f33601d27aa1408b9a271b57bf0ccbc3ce28b45

Observation 51b2df35-9b9c-4a5f-92ae-eaf04c263e21 · inbound

Emerging Properties in Unified Multimodal Pretraining cites this paper.

Emerging Properties in Unified Multimodal Pretraining LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 66

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T16:23:41.942313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T16:23:41.854132Z digest=sha256:379210a3689c6c280e9ffaca477ca52bd819f0700cf067c04fa9a5c77e857ac1

Observation 05dac12b-2ec1-4114-b49e-baa9bdda25cf · inbound

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning cites this paper.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:25.463377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:25.463377Z digest=sha256:233dced5990abb6aac34e05fe211ded96f92517db45a65b89802d9b2427ea342

Observation 4e7650c6-7a75-4f80-a7ee-22d476024b72 · inbound

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation cites this paper.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:17.119373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:17.119373Z digest=sha256:4e7e57b0af3df14c6291bc7d7d6392d4eb0ec5cc62b78be923385f8c04d9a51a

Observation 31f5f6e2-7a61-446a-af72-fef9c92bf846 · inbound

Knowledge Insulating Vision-Language-Action Models: Train Fast, Run Fast, Generalize Better cites this paper.

Knowledge Insulating Vision-Language-Action Models: Train Fast, Run Fast, Generalize Better LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:23.846395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:23.846395Z digest=sha256:65e9b333414d9fe6783ec5bb2e426c1e5faf82a457eebb95ab1bb1c1c92cba28

Observation a98f8182-ac5b-424c-b2ac-401da3970f5a · inbound

LaTtE-Flow: Layerwise Timestep-Expert Flow-based Transformer cites this paper.

LaTtE-Flow: Layerwise Timestep-Expert Flow-based Transformer LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T05:51:42.915375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:51:42.915375Z digest=sha256:90c123bd89c6b20bb27d914beab82f1f5dcfbddd4aca8bf9c52cbf1a0abb16a6

Observation 58e5e446-5684-49ed-a9fb-a57566674917 · inbound

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation cites this paper.

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:02.245217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:26:02.245217Z digest=sha256:db8e7e79573c1016b51936429b9cf49fb8dc715a381690ef981c7188fc817fe9

Observation 4a9b2bc1-764f-4d42-96fb-7ae08c69dbe3 · inbound

Dreamland: Controllable World Creation with Simulator and Generative Models cites this paper.

Dreamland: Controllable World Creation with Simulator and Generative Models LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:16.360980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:26:16.360980Z digest=sha256:7e9ecbbb06ee40493a5db009ce8e5af0918ac22736f9ceb63b9b83729a29b0ef

Observation 73a3cffc-4bfb-4dea-88c1-15afa9fa889f · inbound

Show-o2: Improved Native Unified Multimodal Models cites this paper.

Show-o2: Improved Native Unified Multimodal Models LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 95

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T18:51:15.590013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T18:51:15.428692Z digest=sha256:aa7f78e89d15ada943dfd790fdf16497bff94e5c3d235bce7ab3564dd0b65310

Observation 301bff5f-135a-4629-bd06-220f1bcd1fea · inbound

OmniGen2: Towards Instruction-Aligned Multimodal Generation cites this paper.

OmniGen2: Towards Instruction-Aligned Multimodal Generation LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 68

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T07:52:10.858881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:ee1dd7980e8e304e8625be4d1a518904d2a8d11f61933257b6d8cbb647a57e2e

Observation cd09d3a2-d46b-49ef-be2a-5baa8f857fa3 · inbound

WordCon: Word-level Typography Control in Scene Text Rendering cites this paper.

WordCon: Word-level Typography Control in Scene Text Rendering LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T22:34:16.184346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:34:16.184346Z digest=sha256:9c92c450225a3d2312f1cd107ab31b4a0df34b626a8eb7b4cb74b2d9feb04a85

Observation 72d32ee3-09c9-486d-a573-65be95ab85cf · inbound

Rethinking Discrete Tokens: Treating Them as Conditions for Continuous Autoregressive Image Synthesis cites this paper.

Rethinking Discrete Tokens: Treating Them as Conditions for Continuous Autoregressive Image Synthesis LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T20:49:42.028479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:49:42.028479Z digest=sha256:4d31deac31ca97895605eb34d9aabe8116e1e1966c82431a6c78f96ab2e64466

Observation d2c473d0-10b9-4e03-938d-35c50e31c5ef · inbound

FreeLoRA: Enabling Training-Free LoRA Fusion for Autoregressive Multi-Subject Personalization cites this paper.

FreeLoRA: Enabling Training-Free LoRA Fusion for Autoregressive Multi-Subject Personalization LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T20:46:25.122139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:46:25.122139Z digest=sha256:0f4553bf9d1a9abe99f55c6dff362aba90a0d65334512b97e634fe0ecde9901e

Observation 594a0dbc-5e29-48b0-89c5-62155ea44e51 · inbound

X-Omni: Reinforcement Learning Makes Discrete Autoregressive Image Generative Models Great Again cites this paper.

X-Omni: Reinforcement Learning Makes Discrete Autoregressive Image Generative Models Great Again LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T12:10:08.289255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:10:08.289255Z digest=sha256:7d027373705e553360cb7ac23f8aabeabe1de711c865dbcab04338022e07af5b

Observation b17a1063-74ee-4200-8211-e01266c29ff6 · inbound

Galaxea Open-World Dataset and G0 Dual-System VLA Model cites this paper.

Galaxea Open-World Dataset and G0 Dual-System VLA Model LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T13:31:09.919265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:31:09.919265Z digest=sha256:d8672a16410b051ddcbee48c39595d867c640d182ccf80cf46047bbfe577dfe2

Observation 93875204-2588-4310-b563-7bab02756f4c · inbound

Interleaving Reasoning for Better Text-to-Image Generation cites this paper.

Interleaving Reasoning for Better Text-to-Image Generation LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:44.950531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:44.950531Z digest=sha256:55eb5d2f957701ad6e858ada29948acd71d25f23dc02e3d0d7db43ff8daf015e

Observation b562f445-e176-474e-9c2d-9b13fbd5371d · inbound

Reconstruction Alignment Improves Unified Multimodal Models cites this paper.

Reconstruction Alignment Improves Unified Multimodal Models LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-04T22:36:08.140951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T22:36:08.140951Z digest=sha256:06b7aceef103f766bc469fabb071d5054844b5c49d8d3f0f52b98d6b15160dba

Observation ef364fbd-f8fd-4b0c-bf37-103f5d74a52b · inbound

UniVideo: Unified Understanding, Generation, and Editing for Videos cites this paper.

UniVideo: Unified Understanding, Generation, and Editing for Videos LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T10:49:49.019944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:49:49.019944Z digest=sha256:8681fe2c40a0b3618665e85ebe09de5cbcdb36ee6da14f8ca346368fc77d7346

Observation 527a0908-9afd-40e7-91e4-9cdb4c2ca112 · inbound

SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models cites this paper.

SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T09:54:26.449060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:54:26.449060Z digest=sha256:02fe8ee15e99a148da7d3331d59c9a3ba478c5a3487ccca8ee753968af4b21c8

Observation 70e91f46-7163-4878-a42c-6e83fdf87912 · inbound

FasterVAR: Plug-and-Play Acceleration for Visual Autoregressive Models cites this paper.

FasterVAR: Plug-and-Play Acceleration for Visual Autoregressive Models LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T15:36:51.177913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T15:36:51.177913Z digest=sha256:c8681ae52a7daf9957e877abffd4495df2f6ea844ede33841aa8b1236f5f223d

Observation 71837f0d-f365-4e9a-814e-146d94d08b85 · inbound

CG-MLLM: Captioning and Generating 3D content via Multi-modal Large Language Models cites this paper.

CG-MLLM: Captioning and Generating 3D content via Multi-modal Large Language Models LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-21T14:50:14.718128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T14:48:21.787919Z digest=sha256:a4a63389e75f2637867a0c8ecd64f01e8bf2a01cc1ec62e6eebb3180c9fa31c5

Observation 8854c87b-93c1-45fc-926c-f848d8caded8 · inbound

ChatUMM: Robust Context Tracking for Conversational Interleaved Generation cites this paper.

ChatUMM: Robust Context Tracking for Conversational Interleaved Generation LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T03:57:32.062558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:57:32.062558Z digest=sha256:bd66505a7d007f62d3092d9e637fbc476cf35b18e13c5a7615367ddc42c209be

Observation c89b769c-3ee7-42e6-8247-0d57e1dcebd4 · inbound

LLaMo: Scaling Pretrained Language Models for Unified Motion Understanding and Generation with Continuous Autoregressive Tokens cites this paper.

LLaMo: Scaling Pretrained Language Models for Unified Motion Understanding and Generation with Continuous Autoregressive Tokens LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 58

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T05:02:19.618705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T05:01:11.880003Z digest=sha256:3bb91b69e2e7ef4163ab8b67dbc214bbcd88043b12d8446254ddf2a578393a7a

Observation aa3453fd-dc74-4795-95b3-f947d0d1f384 · inbound

Demystifying Video Reasoning cites this paper.

Demystifying Video Reasoning LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-07-13T23:27:11.006580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:27:11.006580Z digest=sha256:6e1d6f1c87fce881399297f9f2d6b07042795cbb3d5e574ea76d18d5f4443550

Observation 8f54f030-7751-43f7-9d1a-645e1db41827 · inbound

Demystifying Video Reasoning cites this paper.

Demystifying Video Reasoning LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-03T02:33:57.876583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:33:57.876583Z digest=sha256:bfa97083422fcc367e16615d6ff2eefabe3b096dbf81b4a024bb869ed18039f8

Observation 10617339-23d3-4317-89c2-12d670fc9617 · inbound

LMGenDrive: Bridging Multimodal Understanding and Generative World Modeling for End-to-End Driving cites this paper.

LMGenDrive: Bridging Multimodal Understanding and Generative World Modeling for End-to-End Driving LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:31:01.398245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T17:36:38.627415Z digest=sha256:17dff4ea77ec944099c14071d0236468e649cf698d17c5b70a169b17405e5cbf

Observation b0800302-88df-4745-bf7b-dd720fc3f116 · inbound

Free Lunch for Unified Multimodal Models: Enhancing Generation via Reflective Rectification with Inherent Understanding cites this paper.

Free Lunch for Unified Multimodal Models: Enhancing Generation via Reflective Rectification with Inherent Understanding LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T14:15:28.953805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T14:15:02.723774Z digest=sha256:17ce8db9f282bba85078ac682d3ce5a03090bb4eb991235949fbce5e14bb565e

Observation 06ef9691-e5ef-45bd-8e01-14279bf16a8f · inbound

Co-generation of Layout and Shape from Text via Autoregressive 3D Diffusion cites this paper.

Co-generation of Layout and Shape from Text via Autoregressive 3D Diffusion LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T09:23:37.044329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T09:22:10.651922Z digest=sha256:8eb26fcfcbed6665218c8f9bddc26ef5ec574204e31d5d9339225a21b6790f0a

Observation c235ae80-05e8-4e10-8822-a65b8dec7d74 · inbound

MMCORE: MultiModal COnnection with Representation Aligned Latent Embeddings cites this paper.

MMCORE: MultiModal COnnection with Representation Aligned Latent Embeddings LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T03:29:21.669427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T03:27:50.144706Z digest=sha256:2e465bec943045f36d0ccedfe302e0123b037540479b48889dcad098e91393ae

Observation 66ee8cee-fdc4-41e5-ab3a-66281e56cb1d · inbound

Meta-CoT: Enhancing Granularity and Generalization in Image Editing cites this paper.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 52

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:46:14.691378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:c52888cfa261db98f935b90a6166d992d13678a78efdba9d9224b6af73c3c4da

Observation 51bd0ddd-055f-465b-a524-64859693240b · inbound

SpatialFusion: Endowing Unified Image Generation with Intrinsic 3D Geometric Awareness cites this paper.

SpatialFusion: Endowing Unified Image Generation with Intrinsic 3D Geometric Awareness LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:46:26.816740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-07T13:45:53.346402Z digest=sha256:e5f71729ee626bf4295977aff63cdb402167aa01b4e53d2617dd1d83b9818444

Observation 73068788-af9b-4a9b-89b8-001858cd0320 · inbound

STARFlow2: Bridging Language Models and Normalizing Flows for Unified Multimodal Generation cites this paper.

STARFlow2: Bridging Language Models and Normalizing Flows for Unified Multimodal Generation LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T02:45:57.156891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T02:44:59.644215Z digest=sha256:a97ec043053c29a74dc122f1979305c7b0374a4086e022545c9c8606ce8f00ea

Observation c63185a6-3916-499f-8a62-05a744587d3b · inbound

Reversing the Flow: Generation-to-Understanding Synergy in Large Multimodal Models cites this paper.

Reversing the Flow: Generation-to-Understanding Synergy in Large Multimodal Models LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T18:48:53.450985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T18:44:54.835575Z digest=sha256:ea9d335138aed1c26a533daa45adcff81e1ebc2f252a5f1cf158f9a4c29ddc1f

Observation 8fe1fa5e-1efa-4d0a-a3b9-8971c93eae2c · inbound

Semantic Generative Tuning for Unified Multimodal Models cites this paper.

Semantic Generative Tuning for Unified Multimodal Models LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:33:14.202728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T11:32:24.007847Z digest=sha256:ed976df149291e7086b64a7541e354561de9e5d315cce2253626d2c636481980

Observation e554ff05-2944-404f-b193-242c8dc4740e · inbound

Semantic Generative Tuning for Unified Multimodal Models cites this paper.

Semantic Generative Tuning for Unified Multimodal Models LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:35:00.282413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:89b8483b3dcf96baaa259c26e2f73ee1c8697d600a297527b277ea131d859645

Observation 86284e46-aff6-4509-b9a9-68049f183b84 · inbound

Where to Refine, When to Stop: Rethinking Redundancy via Latent Discrepancy for Efficient Visual Autoregressive Generation cites this paper.

Where to Refine, When to Stop: Rethinking Redundancy via Latent Discrepancy for Efficient Visual Autoregressive Generation LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T22:42:46.709383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T22:38:26.147074Z digest=sha256:72f606979d7ab3a14f7ee55adebb9f8cf18564d22528b77048744a79f2492ef6

Observation a6e04f3e-7839-49c8-b13a-3a891a308702 · inbound

Polaris: Scaling Up Instruction-Guided Image Generation Towards Millions of Personalized Style Needs cites this paper.

Polaris: Scaling Up Instruction-Guided Image Generation Towards Millions of Personalized Style Needs LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 97

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T22:46:18.726597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T15:07:03.439928Z digest=sha256:b1dbfb8b1ace45c15d8a37dcf627287491ac7db15bac6db3060954b4f4f7fc24

Observation 74f42b3d-7ebf-4053-b30b-be90316c1160 · inbound

UniTac: A Unified Multimodal Model for Cross-Sensor Tactile Understanding and Generation cites this paper.

UniTac: A Unified Multimodal Model for Cross-Sensor Tactile Understanding and Generation LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-01T10:05:40.624385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-01T05:56:03.597839Z digest=sha256:5a3d427eb444b98fd3fd53674ad1f1d96598352e5dd6170652b26271030479f4

Observation 583ff840-3bba-428a-894e-007a3b67114d · inbound

Streaming Multi-Agent Autoregressive Diffusion Model with World State Registers cites this paper.

Streaming Multi-Agent Autoregressive Diffusion Model with World State Registers LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 126

Resolution
unresolved
no resolver link, observed 2026-08-01T07:02:51.493005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:02:51.493005Z digest=sha256:9db034ad0972ab369d0aa3321cef89a1a49bad9d542d4ca2ea8c94ad14ef53b7

Observation 9c8c05c8-dcfb-46f2-961d-2d9147766e61 · inbound

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes cites this paper.

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-06T11:55:28.069606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:55:28.069606Z digest=sha256:7cce56158ea46f4b43937c158ec07d02fcec243c0186d8559d5ab3ee68e8330f

Observation 1b47279d-1775-4e41-80f8-d3ed0ac8850b · inbound

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes cites this paper.

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-08T17:08:55.803846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:08:55.803846Z digest=sha256:29e0704c3d8ba57a73ed68328be226436ad4a79a2ac49d2f840b4d5642ca0114