Pith. sign in

Paper Citation Record · LEDGER

Think in Strokes, Not Pixels: Process-Driven Image Generation via Interleaved Reasoning

As of 5 August 2026, this Paper Citation Record lists 12 of 12 outbound references and 1 inbound Pith citation observation for arXiv:2604.04746.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.04746 v3

Coverage vector

measured 12 of 12 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T19:16:58.323955Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T21:07:28.254518Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

12 of 12 outbound references displayed

  • verified exact5
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch7

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c48b0112-00ed-4818-9745-dc8d6737318e · outbound

This paper cites Selection-Inference: Exploiting Large Language Models for Interpretable Logical Reasoning.

Think in Strokes, Not Pixels: Process-Driven Image Generation via Interleaved Reasoning Selection-Inference: Exploiting Large Language Models for Interpretable Logical Reasoning

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:10:53.823287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:16:58.323955Z digest=sha256:a381e754c9e89efcb2d606e1db286e650e3c8b4ea534f3c0d7f4faeff523e7b6

Observation 122e0cea-d14a-401c-9c97-0662210c0192 · outbound

This paper cites DreamLLM: Synergistic Multimodal Comprehension and Creation.

Think in Strokes, Not Pixels: Process-Driven Image Generation via Interleaved Reasoning DreamLLM: Synergistic Multimodal Comprehension and Creation

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:10:53.745000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:16:58.323955Z digest=sha256:6f6e452761e229422fd3eb008878e690b2d2f27068dcc4164785e6e1fdf9b73f

Observation f87889a0-6e2c-4247-a13e-21de85e6ef2e · outbound

This paper cites Scaling Rectified Flow Transformers for High-Resolution Image Synthesis.

Think in Strokes, Not Pixels: Process-Driven Image Generation via Interleaved Reasoning Scaling Rectified Flow Transformers for High-Resolution Image Synthesis

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:27:54.173931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:16:58.323955Z digest=sha256:3ecbb79cfaa1972cacf2fd255e65613f50404017eade4cd16e0129243637ec75

Observation fe7b2ad6-500a-4de1-bdef-44db87f0268f · outbound

This paper cites GenEval: An Object-Focused Framework for Evaluating Text-to-Image Alignment.

Think in Strokes, Not Pixels: Process-Driven Image Generation via Interleaved Reasoning GenEval: An Object-Focused Framework for Evaluating Text-to-Image Alignment

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:10:53.793691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:16:58.323955Z digest=sha256:39250659650b90c0a575e1ce3378bc0dda7fb96ce24ed8cea49fdb866eb7a315

Observation 9c2a2f93-6c03-460c-b878-df31db549271 · outbound

This paper cites T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT.

Think in Strokes, Not Pixels: Process-Driven Image Generation via Interleaved Reasoning T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:10:53.770009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:16:58.323955Z digest=sha256:9a4c195ba3644a406db9b02b51b66e94c60da8cf473a650c0fd401d38c52cb63

Observation 390e76f5-04cc-4123-aafa-e0252ab906b1 · outbound

This paper cites Is 'Right' Right? Enhancing Object Orientation Understanding in Multimodal Large Language Models through Egocentric Instruction Tuning.

Think in Strokes, Not Pixels: Process-Driven Image Generation via Interleaved Reasoning Is 'Right' Right? Enhancing Object Orientation Understanding in Multimodal Large Language Models through Egocentric Instruction Tuning

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:10:53.787696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:16:58.323955Z digest=sha256:c27248c88cfe431988a586e771509b3a7ee516ed1e043b42bee66dfc5e650e05

Observation cf9c15bd-dda4-4d12-bb6b-89ed2ba72ab1 · outbound

This paper cites JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation.

Think in Strokes, Not Pixels: Process-Driven Image Generation via Interleaved Reasoning JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:10:53.779102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:16:58.323955Z digest=sha256:75a50dcabad4e403363b8f4d290f2b39d7cef6573d2152d7bc7df279c2aa0205

Observation b922f013-e444-4eac-bea0-709163224b4c · outbound

This paper cites EmbodiedGPT: Vision-Language Pre-Training via Embodied Chain of Thought.

Think in Strokes, Not Pixels: Process-Driven Image Generation via Interleaved Reasoning EmbodiedGPT: Vision-Language Pre-Training via Embodied Chain of Thought

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:10:53.736351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:16:58.323955Z digest=sha256:b0a445a0b8e8ce6ef454011f7a08c9fb0c65e34385e474edccc4db55baaf6f8c

Observation 4243c24a-2608-44c6-bb20-c55d35749a40 · outbound

This paper cites WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation.

Think in Strokes, Not Pixels: Process-Driven Image Generation via Interleaved Reasoning WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-15T16:24:27.819650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:16:58.323955Z digest=sha256:4f9d2053458434b1bd6888a653ef6c8d0fb3e39a49e96e1d16bb57219393005c

Observation 211c994b-f3f4-45f2-a9ff-72e6f011aac3 · outbound

This paper cites TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation.

Think in Strokes, Not Pixels: Process-Driven Image Generation via Interleaved Reasoning TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:10:53.714152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:16:58.323955Z digest=sha256:1280e30ffe12d69608356d182989d5610f122fe1fbb35101495ab6430aff7c8d

Observation c9bb5c41-7094-4727-b2ce-cf082f638fbd · outbound

This paper cites MetaMorph: Multimodal Understanding and Generation via Instruction Tuning.

Think in Strokes, Not Pixels: Process-Driven Image Generation via Interleaved Reasoning MetaMorph: Multimodal Understanding and Generation via Instruction Tuning

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T07:51:13.882527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:16:58.323955Z digest=sha256:f9d94f7b3f7b665b8970b143f49bc7b4269f8d898fc0bad221d79f7a42056d4f

Observation 1f819ba7-250d-49e9-8d67-6f4ce8e40496 · outbound

This paper cites Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model.

Think in Strokes, Not Pixels: Process-Driven Image Generation via Interleaved Reasoning Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:57:27.005672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:16:58.323955Z digest=sha256:00ffc9b5b4e70941959a5161da1f8980f3700210bd2974ed843e5c47aaf3b6a7

Pith citing papers

Observation 7281173a-0a6c-4164-a079-ad1a89d19ed4 · inbound

Apple-$\pi$: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence cites this paper.

Apple-$\pi$: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence Think in Strokes, Not Pixels: Process-Driven Image Generation via Interleaved Reasoning

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-01T21:07:28.254518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:07:28.254518Z digest=sha256:4109de0af379d76f2ffb5bc8104e0afbe7b7f7eb2252d1f14f6f8e5b46e2db48