Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T10:35:22.738851Z
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 15 inbound Pith citation observations for arXiv:2412.00127.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T10:35:22.738851Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-08T19:32:31.857081Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T14:38:28.877985Z
31 of 31 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 77f9cd7a-ca2c-4e1e-bb54-c6cf3fc751d1 · outbound
Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads ShareGPT4V: Improving Large Multi-Modal Models with Better Captions
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2faf6238-7963-4c50-b184-6752ca9e28b1 · outbound
Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb481a48-4e06-490c-a8f4-d22a1838b1e0 · outbound
Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads Fluid: Scaling Autoregressive Text-to-image Generative Models with Continuous Tokens
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b51af1a8-c7c0-4991-adff-233ae7fbd36d · outbound
Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads Examples on Image Editing Figure 8 shows random examples of image editing by Orthus-base post-trained on Instruct-Pix2Pix (Brooks et al., 2023)
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 03bc8b68-f2f6-4488-8751-3efce6b991e8 · outbound
Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads Compared to editing-specific diffusion models (Brooks et al., 2023), Orthus demonstrates better fidelity to the original image in regions where no editing is required
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d7c9f7c3-362e-4eb6-bc45-719ee0b9fe62 · outbound
Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads Autoregressive Image Generation without Vector Quantization
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6577574b-44c9-47d8-a068-36bed90e6a6a · outbound
Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads Lumina-mGPT: Illuminate Flexible Photorealistic Text-to-Image Generation with Multimodal Generative Pretraining
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ea3f318-639e-4d1c-a982-a341bcc5a1f2 · outbound
Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 663075ca-af0b-4dd9-bc35-b862a59ddcbe · outbound
Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads Hierarchical Text-Conditional Image Generation with CLIP Latents
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8107d99-ece7-4f05-9e8b-d07871c352bb · outbound
Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads Chameleon: Mixed-Modal Early-Fusion Foundation Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7480c06-2858-41b9-b804-ae373aafffad · outbound
Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads Gemini: A Family of Highly Capable Multimodal Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da0c46fc-4781-4e30-96e8-de54f22446d3 · outbound
Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb904d04-f6e1-4b8b-a33b-32e619b773ee · outbound
Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads LLaMA: Open and Efficient Foundation Language Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4100c0a2-6b2b-4069-86af-f268553dee02 · outbound
Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads MMAR: Towards Lossless Multi-Modal Auto-Regressive Probabilistic Modeling
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28c93f00-1ad4-409b-aa98-adf39a934615 · outbound
Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads CoCa: Contrastive Captioners are Image-Text Foundation Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb563a88-11f2-43a1-9378-7cfce7584d53 · outbound
Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54dcf4e4-b158-45fa-a0a1-da7ed3b8474c · outbound
Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads An Image is Worth 32 Tokens for Reconstruction and Generation
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f13eff42-636d-47d8-a49f-ce3539e4e9b0 · outbound
Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads MonoFormer: One Transformer for Both Diffusion and Autoregression
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e846a60-51c6-45f3-9b08-7f7277b55f92 · outbound
Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf9bdcb4-59ff-4804-8ddb-ea0a45e477ce · outbound
Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69645efd-1cd3-4cd5-8036-0b024d2da2ac · outbound
Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads Model PSNR ↑ SSIM (Wang et al., 2004)↑ VQ-V AE (Team,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation aa690c79-08b4-4c7d-b4e5-3905f454d345 · outbound
Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation
Reference 2004
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4a17b2e-fb69-485d-972a-8685a303db16 · outbound
Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation
Reference 2013
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f5864f6-d848-4371-92e4-7f52b941f206 · outbound
Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads Classifier-Free Diffusion Guidance
Reference 2014
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b16c8a1-6ec1-48dc-a21b-f5254a7efcfe · outbound
Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads Denoising Diffusion Implicit Models
Reference 2015
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30edef1b-8bd3-4da5-88d3-c47361526a14 · outbound
Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads Emu3: Next-Token Prediction is All You Need
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2edca8ae-4b00-437a-baa4-03e82cb034fc · outbound
Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads Auto-Encoding Variational Bayes
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3906f80b-f5b0-4a9b-be18-83ce207dc9f5 · outbound
Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads PaLM-E: An Embodied Multimodal Language Model
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f00ceab9-be16-4132-baea-276fd707cd52 · outbound
Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a169182-9c08-43e3-a73d-67faa7fefa57 · outbound
Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4a5b56f-d725-4dd4-9c01-f3cc7ffd1e86 · outbound
Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 526fbf8b-d74f-470b-a624-48c9600b9ca7 · inbound
UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26df7ed4-86df-468d-a80a-8949774e4c65 · inbound
UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73d7de18-61b2-4ae5-81af-4e0664d211cb · inbound
WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d0d52794-3986-4ec9-92e2-92310a2bf8d1 · inbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f82f7973-daeb-4c91-8e21-88af5fb17ed7 · inbound
ComfyMind: Toward General-Purpose Generation via Tree-Based Planning and Reactive Feedback Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 668be6c0-d07f-4c16-b854-b3f4c42bb6ef · inbound
UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a043893-9d9b-4be9-99db-27387a32bbd3 · inbound
Show-o2: Improved Native Unified Multimodal Models Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation adc791ca-bce5-4791-b3ac-43dec2d35008 · inbound
X-Omni: Reinforcement Learning Makes Discrete Autoregressive Image Generative Models Great Again Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8736375b-45de-4cca-b905-c90cdf22826f · inbound
A Survey on Diffusion Language Models Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads
Reference 156
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1527e277-ac0f-4309-a7ab-9b8de4a040af · inbound
LongCat-Image Technical Report Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 33f1e5e0-1c2c-4ce2-9aab-a9e9fa1a6100 · inbound
EduIllustrate: Towards Scalable Automated Generation Of Multimodal Educational Content Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 23f05fc9-1090-448d-a91b-d4d2b495e287 · inbound
MAR-GRPO: Stabilized GRPO for AR-diffusion Hybrid Image Generation Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a608ae32-7f47-482e-a371-970ab1f6c834 · inbound
ProductWebGen: Benchmarking Multimodal Product Webpage Generation Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 0f0ae40e-923c-4008-b8ac-c0f24f7e5690 · inbound
HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 62437193-a12a-4169-9511-6066d7701a30 · inbound
Twins: Learn to Predict Unified Representations with Focal Loss Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.