Pith. sign in

Paper Citation Record · LEDGER

Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads

As of 12 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 15 inbound Pith citation observations for arXiv:2412.00127.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.00127 v2

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T10:35:22.738851Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T19:32:31.857081Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T14:38:28.877985Z

Reference resolution

31 of 31 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 77f9cd7a-ca2c-4e1e-bb54-c6cf3fc751d1 · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:22.607462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:22.607462Z digest=sha256:b930661abe7683566aafef05dad1113d648f1f5f0607b65ed9975c897751ea47

Observation 2faf6238-7963-4c50-b184-6752ca9e28b1 · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:22.617670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:22.617670Z digest=sha256:4aaedfff5e18007a292b68da1d2f6a6c9ee9f1e59728b43d09684e9c19f39567

Observation bb481a48-4e06-490c-a8f4-d22a1838b1e0 · outbound

This paper cites Fluid: Scaling Autoregressive Text-to-image Generative Models with Continuous Tokens.

Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads Fluid: Scaling Autoregressive Text-to-image Generative Models with Continuous Tokens

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:22.626842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:22.626842Z digest=sha256:a567fe70d9a9dac79bd939a304c764bc038275f83f2601e8f2484ba1152acd90

Observation b51af1a8-c7c0-4991-adff-233ae7fbd36d · outbound

This paper cites Examples on Image Editing Figure 8 shows random examples of image editing by Orthus-base post-trained on Instruct-Pix2Pix (Brooks et al., 2023).

Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads Examples on Image Editing Figure 8 shows random examples of image editing by Orthus-base post-trained on Instruct-Pix2Pix (Brooks et al., 2023)

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:35:23.108225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:35:22.734378Z digest=sha256:e2e0e988c3f2e487500a8f6173c2576ca6246431e25b3879636d293294aabe6d

Observation 03bc8b68-f2f6-4488-8751-3efce6b991e8 · outbound

This paper cites Compared to editing-specific diffusion models (Brooks et al., 2023), Orthus demonstrates better fidelity to the original image in regions where no editing is required.

Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads Compared to editing-specific diffusion models (Brooks et al., 2023), Orthus demonstrates better fidelity to the original image in regions where no editing is required

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:35:23.095658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:35:22.738851Z digest=sha256:2a82ddaa3617c8062875b4e28d51f4c59d52d8f914303511746937df6ffa6f2d

Observation d7c9f7c3-362e-4eb6-bc45-719ee0b9fe62 · outbound

This paper cites Autoregressive Image Generation without Vector Quantization.

Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads Autoregressive Image Generation without Vector Quantization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:22.647039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:22.647039Z digest=sha256:d375a4f18f3f8771c7325a3c8ab24cea2e035f821196659f1e5b4a741b3fe7b9

Observation 6577574b-44c9-47d8-a068-36bed90e6a6a · outbound

This paper cites Lumina-mGPT: Illuminate Flexible Photorealistic Text-to-Image Generation with Multimodal Generative Pretraining.

Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads Lumina-mGPT: Illuminate Flexible Photorealistic Text-to-Image Generation with Multimodal Generative Pretraining

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:22.651398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:22.651398Z digest=sha256:32be2b62f30b9ee433d74d01d4ef17ed12e4b4d6dcb5d77395cf83fa5f714b9b

Observation 0ea3f318-639e-4d1c-a982-a341bcc5a1f2 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:22.655674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:22.655674Z digest=sha256:6a6b9e50c045630075e353cbb6ea856ac8e5dfe5f94d44f1254deb40e695cf2b

Observation 663075ca-af0b-4dd9-bc35-b862a59ddcbe · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:22.660086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:22.660086Z digest=sha256:2d96784cb18d4c4a95462ceab1d2af850dc5b59d0e5ff1ce4abb9c479b5e97fd

Observation b8107d99-ece7-4f05-9e8b-d07871c352bb · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:22.668589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:22.668589Z digest=sha256:3238449b35db0a4d908868b54a45de15e209e6ba1deb937198117c417e9966a6

Observation c7480c06-2858-41b9-b804-ae373aafffad · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads Gemini: A Family of Highly Capable Multimodal Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:22.673180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:22.673180Z digest=sha256:a8acc48c754e7de2d80afc37450550a4585773de257bfa94041ff8f75074b762

Observation da0c46fc-4781-4e30-96e8-de54f22446d3 · outbound

This paper cites Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction.

Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:22.677127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:22.677127Z digest=sha256:e0ff50b889fb2a71ce0231c0534b2c3bf55ba5c4e57a59ef8bd8df7b5fbbc851

Observation bb904d04-f6e1-4b8b-a33b-32e619b773ee · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads LLaMA: Open and Efficient Foundation Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:22.681090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:22.681090Z digest=sha256:6cecdd9fb282df03d6b5b8a16638e107ddcf5e09f4ea8f06b651f59cdebe2d7c

Observation 4100c0a2-6b2b-4069-86af-f268553dee02 · outbound

This paper cites MMAR: Towards Lossless Multi-Modal Auto-Regressive Probabilistic Modeling.

Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads MMAR: Towards Lossless Multi-Modal Auto-Regressive Probabilistic Modeling

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:22.698742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:22.698742Z digest=sha256:a8b4d13be62759ea75a828f61796f10819cf730b22ce47cc0a99686a87ca02d0

Observation 28c93f00-1ad4-409b-aa98-adf39a934615 · outbound

This paper cites CoCa: Contrastive Captioners are Image-Text Foundation Models.

Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads CoCa: Contrastive Captioners are Image-Text Foundation Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:22.702855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:22.702855Z digest=sha256:d597fdd2296faec444c859b4870c70abec0cfab6441deb216299333a5e65bbf3

Observation fb563a88-11f2-43a1-9378-7cfce7584d53 · outbound

This paper cites Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation.

Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:22.707405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:22.707405Z digest=sha256:d3fc9bd1f8973363102afda0b162498084249c82752f5e382788f8c872ad5864

Observation 54dcf4e4-b158-45fa-a0a1-da7ed3b8474c · outbound

This paper cites An Image is Worth 32 Tokens for Reconstruction and Generation.

Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads An Image is Worth 32 Tokens for Reconstruction and Generation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:22.711558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:22.711558Z digest=sha256:4e615eb9b9d4bc9bb5d9643f78217ebfda6f99c2c60905e8d1d624103a1eaa49

Observation f13eff42-636d-47d8-a49f-ce3539e4e9b0 · outbound

This paper cites MonoFormer: One Transformer for Both Diffusion and Autoregression.

Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads MonoFormer: One Transformer for Both Diffusion and Autoregression

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:22.716088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:22.716088Z digest=sha256:ffd360ee6bdb5564427049cc8812d57adc6d371b712cf23760e86a05c8409104

Observation 5e846a60-51c6-45f3-9b08-7f7277b55f92 · outbound

This paper cites Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model.

Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:22.720644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:22.720644Z digest=sha256:54bb1098cc03cd6bfbedac7db4e72c90c2a2bec6c885b7b7e84f24f1141bc4f0

Observation cf9bdcb4-59ff-4804-8ddb-ea0a45e477ce · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:22.725169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:22.725169Z digest=sha256:56d873fedaed7b21a3f6f840f86f8ba393cf5980d77dee87a4d211f36940fcec

Observation 69645efd-1cd3-4cd5-8036-0b024d2da2ac · outbound

This paper cites Model PSNR ↑ SSIM (Wang et al., 2004)↑ VQ-V AE (Team,.

Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads Model PSNR ↑ SSIM (Wang et al., 2004)↑ VQ-V AE (Team,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:35:23.120585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:35:22.729824Z digest=sha256:3099f76609b8c9ed315490ce9b70a68dc9007c54924f19f071aa84f14dc97505

Observation aa690c79-08b4-4c7d-b4e5-3905f454d345 · outbound

This paper cites Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation.

Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation

Reference 2004

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:22.690206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:22.690206Z digest=sha256:c4c18c3623259a4ceaa3bfbf33234a279b55f9df2ca1233600f2d10863703a3c

Observation a4a17b2e-fb69-485d-972a-8685a303db16 · outbound

This paper cites VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation.

Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation

Reference 2013

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:22.694652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:22.694652Z digest=sha256:616a3f6e708839d9e435eb1a803ab01e3542cf51fd75897a440ef9b10b0c1fbb

Observation 8f5864f6-d848-4371-92e4-7f52b941f206 · outbound

This paper cites Classifier-Free Diffusion Guidance.

Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads Classifier-Free Diffusion Guidance

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:22.637735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:22.637735Z digest=sha256:61297bb5423370b68c67267b67836059690145fd397c633614b253116dcfdf61

Observation 8b16c8a1-6ec1-48dc-a21b-f5254a7efcfe · outbound

This paper cites Denoising Diffusion Implicit Models.

Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads Denoising Diffusion Implicit Models

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:22.664209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:22.664209Z digest=sha256:0aa8f0dc9599c26492f6c41e3bb69c8b9e8e88c32496ddc2c7eac46d49a5c04c

Observation 30edef1b-8bd3-4da5-88d3-c47361526a14 · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads Emu3: Next-Token Prediction is All You Need

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:22.685832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:22.685832Z digest=sha256:aede30d2134bf2ee7ff7b59cef0323d969425832f08ca51038cfb03daca630de

Observation 2edca8ae-4b00-437a-baa4-03e82cb034fc · outbound

This paper cites Auto-Encoding Variational Bayes.

Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads Auto-Encoding Variational Bayes

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:22.642410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:22.642410Z digest=sha256:4079c5faafd0cb9c342c9e5f6d84f02d828737a0e446bd3293d82a0806ddb2fc

Observation 3906f80b-f5b0-4a9b-be18-83ce207dc9f5 · outbound

This paper cites PaLM-E: An Embodied Multimodal Language Model.

Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads PaLM-E: An Embodied Multimodal Language Model

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:22.622443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:22.622443Z digest=sha256:59dac2fc356cde351d0404ef493141d0a434ad720a040911a6a7f7be1c55f85e

Observation f00ceab9-be16-4132-baea-276fd707cd52 · outbound

This paper cites SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation.

Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:22.632927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:22.632927Z digest=sha256:a2f4e4bdfc5c5e24cffad15e097ee1033eaee045b7ddf97c6660952b8928c20e

Observation 5a169182-9c08-43e3-a73d-67faa7fefa57 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:22.602310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:22.602310Z digest=sha256:e8b3548df7312690cbcb40d52ffbd3a2f4bffdfee54feb9fad50fa80ad9aa557

Observation f4a5b56f-d725-4dd4-9c01-f3cc7ffd1e86 · outbound

This paper cites ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation.

Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:22.612319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:22.612319Z digest=sha256:0c9cfbe86188c588fdd8837366de7ab8038c708c4adb2f519cfd6daef64f5553

Pith citing papers

Observation 526fbf8b-d74f-470b-a624-48c9600b9ca7 · inbound

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding cites this paper.

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T19:32:31.857081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:32:31.857081Z digest=sha256:b01a99aa198547bf6a18d9816cf3c458129e6f8da63a1219c63ccb09cb38e38d

Observation 26df7ed4-86df-468d-a80a-8949774e4c65 · inbound

UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths cites this paper.

UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T15:24:46.353838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:24:46.353838Z digest=sha256:fb1e6a2074f38edfa18146272206f1bc2f88fd877494abdb94f9c67fde4abffe

Observation 73d7de18-61b2-4ae5-81af-4e0664d211cb · inbound

WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation cites this paper.

WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-15T16:24:27.626031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T16:24:27.407376Z digest=sha256:a349a3ce10b5e11895f099334134da365f2afd956546dbadc42ce62c00b7c9e5

Observation d0d52794-3986-4ec9-92e2-92310a2bf8d1 · inbound

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning cites this paper.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:46:06.155454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:dead0584cb0e44c73ba1e99b1e28e36f549ca46339be7873e87e3f6f181aad42

Observation f82f7973-daeb-4c91-8e21-88af5fb17ed7 · inbound

ComfyMind: Toward General-Purpose Generation via Tree-Based Planning and Reactive Feedback cites this paper.

ComfyMind: Toward General-Purpose Generation via Tree-Based Planning and Reactive Feedback Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:15.410916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:15.410916Z digest=sha256:8f5aa7cea7ce6d60f30b6c36ae6c1238918122e08d99bcedfe18dbd3d6a1f234

Observation 668be6c0-d07f-4c16-b854-b3f4c42bb6ef · inbound

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning cites this paper.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:23.148041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:23.148041Z digest=sha256:1a41d62c4809e10f2040b3b651b5693cf228455550d1ab97e78c8210f7022bc6

Observation 5a043893-9d9b-4be9-99db-27387a32bbd3 · inbound

Show-o2: Improved Native Unified Multimodal Models cites this paper.

Show-o2: Improved Native Unified Multimodal Models Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:51:15.820751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-12T18:51:15.428692Z digest=sha256:fedf4d59146d81f0743197ee01e359ecb5b3233b5a4f741a33e37eef606f1a01

Observation adc791ca-bce5-4791-b3ac-43dec2d35008 · inbound

X-Omni: Reinforcement Learning Makes Discrete Autoregressive Image Generative Models Great Again cites this paper.

X-Omni: Reinforcement Learning Makes Discrete Autoregressive Image Generative Models Great Again Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T12:10:08.223636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:10:08.223636Z digest=sha256:b1c989ea7730687b1565ed5f12b524831c18306425fa902d88abc5412248031e

Observation 8736375b-45de-4cca-b905-c90cdf22826f · inbound

A Survey on Diffusion Language Models cites this paper.

A Survey on Diffusion Language Models Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads

Reference 156

Resolution
unresolved
no resolver link, observed 2026-08-05T20:15:27.240749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:15:27.240749Z digest=sha256:ec6a200d561e2da44e36ab50d0561ce85c6bfc4cfa2d26e4b278b7642dc60132

Observation 1527e277-ac0f-4309-a7ab-9b8de4a040af · inbound

LongCat-Image Technical Report cites this paper.

LongCat-Image Technical Report Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T08:04:13.052686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T08:04:12.949075Z digest=sha256:a7b7925ca5d7a1bf454cd3373afaebbce90e819fd0716c0004e4eeca44b30b3a

Observation 33f1e5e0-1c2c-4ce2-9aab-a9e9fa1a6100 · inbound

EduIllustrate: Towards Scalable Automated Generation Of Multimodal Educational Content cites this paper.

EduIllustrate: Towards Scalable Automated Generation Of Multimodal Educational Content Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:20:48.532209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T19:57:46.438440Z digest=sha256:f6057a8390d66220cef34c36fecd164e8630f0720866e4c3a810dc00a0acd21c

Observation 23f05fc9-1090-448d-a91b-d4d2b495e287 · inbound

MAR-GRPO: Stabilized GRPO for AR-diffusion Hybrid Image Generation cites this paper.

MAR-GRPO: Stabilized GRPO for AR-diffusion Hybrid Image Generation Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:40:58.246847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T18:00:50.105629Z digest=sha256:3607cb87d271460418fdff42e6cd56f9c6fd77fb56408cfe5f407d767c228d82

Observation a608ae32-7f47-482e-a371-970ab1f6c834 · inbound

ProductWebGen: Benchmarking Multimodal Product Webpage Generation cites this paper.

ProductWebGen: Benchmarking Multimodal Product Webpage Generation Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-01T21:06:14.727715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-28T17:26:44.297446Z digest=sha256:429b9195a9083d4c35bf3432f99034c24f21b99e286009f95bf11f939eb24860

Observation 0f0ae40e-923c-4008-b8ac-c0f24f7e5690 · inbound

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers cites this paper.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads

Reference 66

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:38:28.879242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:a463d858b11e3af0eba0952a919211f1df44bd2394629035d7c9aad110a2577f

Observation 62437193-a12a-4169-9511-6066d7701a30 · inbound

Twins: Learn to Predict Unified Representations with Focal Loss cites this paper.

Twins: Learn to Predict Unified Representations with Focal Loss Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-01T04:29:49.264057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T04:29:49.264057Z digest=sha256:6eb3f19f6a62318ca6ec9f59c6337e7b8bf98e6184268d7cef43a8bae6f43206