Pith. sign in

Paper Citation Record · LEDGER

Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 15 inbound Pith citation observations for arXiv:2412.00127.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.00127 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T19:32:31.857081Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T14:38:28.877985Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 526fbf8b-d74f-470b-a624-48c9600b9ca7 · inbound

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding cites this paper.

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T19:32:31.857081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:32:31.857081Z digest=sha256:7f0ec800e2546da1cbc191b17830ee1fbfd94048fc1d54ac15d7cf100187d5fe

Observation 26df7ed4-86df-468d-a80a-8949774e4c65 · inbound

UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths cites this paper.

UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T15:24:46.353838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:24:46.353838Z digest=sha256:5966112ae0ea69a55fa3a58bc04a3448809eaf6f6ffc2ae39ccf7986874d8c0c

Observation 73d7de18-61b2-4ae5-81af-4e0664d211cb · inbound

WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation cites this paper.

WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-15T16:24:27.626031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T16:24:27.407376Z digest=sha256:6ba455a41ee5f2b5f5a8d371b5833cb8136c95b87b4623dbe24b633be3c923eb

Observation d0d52794-3986-4ec9-92e2-92310a2bf8d1 · inbound

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning cites this paper.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:46:06.155454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:168f4b1bc266aa707a4efe7397a604bdeb3cffe087e03dbd4e471a6810d666ed

Observation f82f7973-daeb-4c91-8e21-88af5fb17ed7 · inbound

ComfyMind: Toward General-Purpose Generation via Tree-Based Planning and Reactive Feedback cites this paper.

ComfyMind: Toward General-Purpose Generation via Tree-Based Planning and Reactive Feedback Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:15.410916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:15.410916Z digest=sha256:691d514c2dc16ee706f5622fd5e10404de9babd19f58bf703e17cb0aa7909446

Observation 668be6c0-d07f-4c16-b854-b3f4c42bb6ef · inbound

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning cites this paper.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:23.148041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:23.148041Z digest=sha256:40bc78e926d4bfea9f64060242c641ae7f0f0ac07e71680a935a886c0d2ee8f2

Observation 5a043893-9d9b-4be9-99db-27387a32bbd3 · inbound

Show-o2: Improved Native Unified Multimodal Models cites this paper.

Show-o2: Improved Native Unified Multimodal Models Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:51:15.820751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T18:51:15.428692Z digest=sha256:66e823748da54470a9698196b54c875164665401ace7e06e22ea5a7c76473e03

Observation adc791ca-bce5-4791-b3ac-43dec2d35008 · inbound

X-Omni: Reinforcement Learning Makes Discrete Autoregressive Image Generative Models Great Again cites this paper.

X-Omni: Reinforcement Learning Makes Discrete Autoregressive Image Generative Models Great Again Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T12:10:08.223636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:10:08.223636Z digest=sha256:21815e98661fbe9dfc91cf910c6b8438fd06fdb6d62524c469cd8497d0c6ec61

Observation 8736375b-45de-4cca-b905-c90cdf22826f · inbound

A Survey on Diffusion Language Models cites this paper.

A Survey on Diffusion Language Models Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads

Reference 156

Resolution
unresolved
no resolver link, observed 2026-08-05T20:15:27.240749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:15:27.240749Z digest=sha256:762ccc971d5445a430b7a9df36fb2d68dd13fc834c55e6b24db5f56d0dea7573

Observation 1527e277-ac0f-4309-a7ab-9b8de4a040af · inbound

LongCat-Image Technical Report cites this paper.

LongCat-Image Technical Report Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T08:04:13.052686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T08:04:12.949075Z digest=sha256:e491f9dfe48603c99293011a8d1cf452a14faf28c96e7924d4df02fc17f5af9f

Observation 33f1e5e0-1c2c-4ce2-9aab-a9e9fa1a6100 · inbound

EduIllustrate: Towards Scalable Automated Generation Of Multimodal Educational Content cites this paper.

EduIllustrate: Towards Scalable Automated Generation Of Multimodal Educational Content Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:20:48.532209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T19:57:46.438440Z digest=sha256:ddb73904a71e546d8ae0d889cf4eb74753fca60bb9c8679de8ed9a0dab4f752b

Observation 23f05fc9-1090-448d-a91b-d4d2b495e287 · inbound

MAR-GRPO: Stabilized GRPO for AR-diffusion Hybrid Image Generation cites this paper.

MAR-GRPO: Stabilized GRPO for AR-diffusion Hybrid Image Generation Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:40:58.246847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T18:00:50.105629Z digest=sha256:556ff98f87361d340e01cfa14e275df24c8d9a7b8333c2c65d827c2aa8e67886

Observation a608ae32-7f47-482e-a371-970ab1f6c834 · inbound

ProductWebGen: Benchmarking Multimodal Product Webpage Generation cites this paper.

ProductWebGen: Benchmarking Multimodal Product Webpage Generation Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-01T21:06:14.727715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T17:26:44.297446Z digest=sha256:e4f3f3ebb807439a7e5700e57994da957fbb12684aad645081016d0daf1fa521

Observation 0f0ae40e-923c-4008-b8ac-c0f24f7e5690 · inbound

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers cites this paper.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads

Reference 66

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:38:28.879242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:de04ad94fb84a98ac8ba320c473c94fac9224662974d57108a10183c1fa074a3

Observation 62437193-a12a-4169-9511-6066d7701a30 · inbound

Twins: Learn to Predict Unified Representations with Focal Loss cites this paper.

Twins: Learn to Predict Unified Representations with Focal Loss Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-01T04:29:49.264057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T04:29:49.264057Z digest=sha256:31080df9d2602eadcce8b843504a79e9d9e889c9db0ed56ea720fdea8a504acd