Pith. sign in

Paper Citation Record · LEDGER

ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 30 inbound Pith citation observations for arXiv:2407.06135.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.06135 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 30 of 30 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T20:21:30.373363Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T16:29:56.712695Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d549e1aa-5df8-4520-9333-471a6814f482 · inbound

Imagine while Reasoning in Space: Multimodal Visualization-of-Thought cites this paper.

Imagine while Reasoning in Space: Multimodal Visualization-of-Thought ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-16T23:09:34.853178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T23:09:34.805552Z digest=sha256:d2b893c81d4cce2fadb2f5fa61d1cdaa893aa9beff0fa1f221ca12e5a1b9713c

Observation 58d5da58-a9ab-4936-a875-026441128f46 · inbound

HealthGPT: A Medical Large Vision-Language Model for Unifying Comprehension and Generation via Heterogeneous Knowledge Adaptation cites this paper.

HealthGPT: A Medical Large Vision-Language Model for Unifying Comprehension and Generation via Heterogeneous Knowledge Adaptation ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T20:21:30.373363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T20:21:30.373363Z digest=sha256:17aa9ac041ae5bb142842f69f533a6447ed8c4d2345dc9c3076d72064dabea89

Observation fad3175f-07c1-4792-bee1-036d26231f4b · inbound

MuDoC: An Interactive Multimodal Document-grounded Conversational AI System cites this paper.

MuDoC: An Interactive Multimodal Document-grounded Conversational AI System ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T20:20:13.163709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T20:20:13.163709Z digest=sha256:83b8237ee45fcdc30c7c9dac8ba1c80104f5b6ed6513f9871f2eb6fca295871a

Observation 2ac5ac4d-1524-4b64-9514-e256ac64affa · inbound

StyleAR: Customizing Multimodal Autoregressive Model for Style-Aligned Text-to-Image Generation cites this paper.

StyleAR: Customizing Multimodal Autoregressive Model for Style-Aligned Text-to-Image Generation ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:46.064937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:09:46.064937Z digest=sha256:48d53333f8648b7488404cfd0d99fdfe3acdfd0c2d244a90dfa09dd30da5bf61

Observation e4b8cd6b-a9a1-4815-9017-881ab97394c8 · inbound

Are Any-to-Any Models More Consistent Across Modality Transfers Than Specialists? cites this paper.

Are Any-to-Any Models More Consistent Across Modality Transfers Than Specialists? ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:36:37.945412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:36:37.945412Z digest=sha256:022ab997e864977e00f5d268f1d909a4e6475afd27fd68e367fae345ea4b80b7

Observation e3caf459-b8b1-403e-a755-12a927bd2ff1 · inbound

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics cites this paper.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:24.262152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:24.262152Z digest=sha256:e2fd97ab84a8573f072e64e640d822011241f5a916cbbdb1868c3265537c23ee

Observation 5b1cc51a-8109-4340-8a18-1e8590487aa4 · inbound

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation cites this paper.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.618223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.618223Z digest=sha256:f151ba20228ccd9666f2139bc569cf4ed011f509cc41ab10433e5f08aca91328

Observation d9b48a2b-da9c-441b-806f-bbcc140f388d · inbound

Extracting Visual Facts from Intermediate Layers for Mitigating Hallucinations in Multimodal Large Language Models cites this paper.

Extracting Visual Facts from Intermediate Layers for Mitigating Hallucinations in Multimodal Large Language Models ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T15:32:34.455945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:32:34.455945Z digest=sha256:a0c58067c6942b9e48d68dbaabeaa28a29422c421e74e4208fb37ae12c5662c5

Observation b8a06caf-109b-4c80-9cdf-df37b84d8c32 · inbound

GoViG: Goal-Conditioned Visual Navigation Instruction Generation via Multimodal Reasoning cites this paper.

GoViG: Goal-Conditioned Visual Navigation Instruction Generation via Multimodal Reasoning ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T23:02:52.704228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T23:02:14.913189Z digest=sha256:b05854158f9cdba9301445622f35a0f590ba1776e4e2053c8444b088ffc8e3ff

Observation ac75b562-6416-467f-9d6e-b42ac1b682c3 · inbound

VVS: Accelerating Speculative Decoding for Visual Autoregressive Generation via Partial Verification Skipping cites this paper.

VVS: Accelerating Speculative Decoding for Visual Autoregressive Generation via Partial Verification Skipping ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:35:17.562297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T21:34:14.371914Z digest=sha256:faa2fce61d5d1c6737a0e3c2405798cb2130ea10661ba4b3bfaf65dc6c03d518

Observation bce8e8b5-a923-4bc2-981c-763879253328 · inbound

Mull-Tokens: Modality-Agnostic Latent Thinking cites this paper.

Mull-Tokens: Modality-Agnostic Latent Thinking ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:58:38.638012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T22:57:04.802871Z digest=sha256:041fd9607f83b311a95ef3c3edbadccd9f775febd7a1b1f3fc9659a12ae37687

Observation bc8e3f01-2f13-4856-871e-90ddc180c11f · inbound

Omni-R1: Towards the Unified Generative Paradigm for Multimodal Reasoning cites this paper.

Omni-R1: Towards the Unified Generative Paradigm for Multimodal Reasoning ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-16T14:37:59.986665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T14:37:05.402850Z digest=sha256:527b9751fd90cc46d97326edd8c272fba5e106eb0778d96bd4dd3f410b6526b0

Observation 171a5348-85c7-4a3a-ae57-0299994caaad · inbound

A Survey on Evaluating Quality and Trustworthiness in LLM-Generated Data cites this paper.

A Survey on Evaluating Quality and Trustworthiness in LLM-Generated Data ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T08:15:15.601866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T08:15:15.601866Z digest=sha256:3b4eeae58d85173a47cc50129e9ed82a662a13112d02b931aa5d604b8b5b42cb

Observation 0eab7f02-e4e2-47e3-9b72-36d1e19959e2 · inbound

SJD-PAC: Accelerating Speculative Jacobi Decoding via Proactive Drafting and Adaptive Continuation cites this paper.

SJD-PAC: Accelerating Speculative Jacobi Decoding via Proactive Drafting and Adaptive Continuation ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-13T22:30:39.049531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:30:39.049531Z digest=sha256:e443ca562ba44aa345d3504cd3a57f8030098df4d159cc94a3fc244b461d1d17

Observation 15e0d7f2-b3f3-455d-a874-08855ac0ac8c · inbound

CASCADE: Context-Aware Relaxation for Speculative Image Decoding cites this paper.

CASCADE: Context-Aware Relaxation for Speculative Image Decoding ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:55:55.103143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T02:08:27.374066Z digest=sha256:4b2e970ed85c043f97d40265dfd96c9708e36b654b8e6f435b2e44bc1da39391

Observation 9f5be935-48d2-433d-8137-702321f38ce5 · inbound

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture cites this paper.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:17:18.811647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:dd4d2ff0bafddab56aefc229e210c4e30e43744f316c3369aa208a4be4eff6b0

Observation 7cb32506-9195-49ad-b539-69d0342741ec · inbound

ATLAS: Agentic or Latent Visual Reasoning? One Word is Enough for Both cites this paper.

ATLAS: Agentic or Latent Visual Reasoning? One Word is Enough for Both ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-15T03:14:53.499825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T03:09:58.411261Z digest=sha256:e78442d0fc94ead81142b087a6eb1e295faee17acae2bc6a597271526858b826

Observation 96ae87d1-6052-41fc-82c2-a29f92ff9eba · inbound

ETCHR: Editing To Clarify and Harness Reasoning cites this paper.

ETCHR: Editing To Clarify and Harness Reasoning ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:25:19.378172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-25T04:21:50.862780Z digest=sha256:bb517732fdf9534f4100abd85d30291674395af4a0a0b7ed3f52da6d06292783

Observation 7a85336c-e2ca-4c2c-bd00-c87b7ded61f7 · inbound

DeepLatent: Think with Images via Parallel Latent Visual Reasoning cites this paper.

DeepLatent: Think with Images via Parallel Latent Visual Reasoning ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 89

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T19:52:36.081851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T18:44:39.545911Z digest=sha256:68d4668788a5d890d5ec1fc24724cdeecaded218b99b1248604ab2ceea1a8baa

Observation bfe18290-46dd-4369-b218-fcd1ebcb5253 · inbound

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation cites this paper.

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T03:06:29.571000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T10:23:43.501656Z digest=sha256:5a37a39d950586e3446694820de3d5f58b4991af42d5d06d9be575b70263ef69

Observation 079e8fc5-6c9e-4622-b814-4b922cf8c71e · inbound

Parallel Jacobi Decoding for Fast Autoregressive Image Generation cites this paper.

Parallel Jacobi Decoding for Fast Autoregressive Image Generation ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:26:57.122683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T02:05:58.714029Z digest=sha256:b8af183c563d17abb45532b4a570638d2d75cc31ab0a27205d1092d401562a74

Observation c973eaab-4ad4-424d-89a3-fea897dc6c04 · inbound

SVoT: State-aware Visualization-of-Thought for Spatial Reasoning via Reinforcement Learning cites this paper.

SVoT: State-aware Visualization-of-Thought for Spatial Reasoning via Reinforcement Learning ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:27:56.953878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T09:59:02.899488Z digest=sha256:17ed3269fcbdc985538fa7e7114d25ea9a5979438124390e27660fa0d9c7d21c

Observation 54943f0c-8254-425d-832c-86245aee7d58 · inbound

Bridging Modal Isolation in Interleaved Thinking: Supervising Modality Transitions via Stepwise Reinforcement cites this paper.

Bridging Modal Isolation in Interleaved Thinking: Supervising Modality Transitions via Stepwise Reinforcement ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-03T13:38:19.352383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T07:43:09.538833Z digest=sha256:b691b8f281bcee1b47641349598a7364bcf3e90a7274f7fd316a482a26dd5c7a

Observation c0d3e570-86ff-4ddb-9fba-afdfc2198749 · inbound

Bridging Modal Isolation in Interleaved Thinking: Supervising Modality Transitions via Stepwise Reinforcement cites this paper.

Bridging Modal Isolation in Interleaved Thinking: Supervising Modality Transitions via Stepwise Reinforcement ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-12T14:14:01.413340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T14:14:01.413340Z digest=sha256:4303f5c6de3437a3a18001f97c743e23f90118a400073d1e78d3ae53161e6365

Observation 33f4139f-20bd-4da1-8c6c-4c4b48745748 · inbound

NavWM: A Unified Navigation World Model for Foresight-Driven Planning cites this paper.

NavWM: A Unified Navigation World Model for Foresight-Driven Planning ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:29:56.714946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T00:42:30.237545Z digest=sha256:32ceb1312308bce52ee894d5809207ef9f3a2a00a2322696a8f6146df344039c

Observation aed47806-a2fe-43b5-a94e-1cebd9dee218 · inbound

Illuminating Unified Multimodal Model for Free-form Interleaved Text-Image Generation cites this paper.

Illuminating Unified Multimodal Model for Free-form Interleaved Text-Image Generation ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T06:04:20.971346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T06:04:07.327934Z digest=sha256:1d13ea30a80c6a0d07b98cfa81b0025bca5fce20daef4b71801f2a7f5bf51985

Observation dfeaea4e-7ee4-4034-859c-520013f7dba4 · inbound

ProLaViT: Learning Progressive Latent Visual Thoughts in Structured Latent Space cites this paper.

ProLaViT: Learning Progressive Latent Visual Thoughts in Structured Latent Space ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-12T06:13:16.894125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:13:16.894125Z digest=sha256:d01fa5687de86dbb5b68661902dfb255ed1bfbe263f2ad6d31a9c5c5e3ad441d

Observation 41a213a5-11cd-4b02-b507-9ce1f963f5ab · inbound

DynaVieW: Schema-Guided World Modeling for Understanding Hierarchical Visual Dynamics cites this paper.

DynaVieW: Schema-Guided World Modeling for Understanding Hierarchical Visual Dynamics ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-11T21:36:42.378149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T21:36:42.378149Z digest=sha256:655240d8d9d374b0faa4e63e3c9dc03ca2931edd8912c75666619a99c4e64385

Observation f0351427-741e-4ed7-a6ee-07d7e6805aad · inbound

Improving Sample Diversity in Autoregressive Text-to-Image Generation via Cluster Truncation cites this paper.

Improving Sample Diversity in Autoregressive Text-to-Image Generation via Cluster Truncation ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-14T10:59:43.137854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:59:43.137854Z digest=sha256:de7bcbfc87d36dd75d82bb828a24681587ef424c021566aeac0a8aaacb0965f5

Observation 114e8193-2dd6-4389-9d07-a0d85a9da3be · inbound

VIG-RL: Learning to Search and Insert for Verified Image Grounding cites this paper.

VIG-RL: Learning to Search and Insert for Verified Image Grounding ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-07-31T19:16:58.978808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T19:16:58.978808Z digest=sha256:67bec1fab648edf206356d0eab1f7b5d18181ad7479a40ca986dc15b16958f03