Pith. sign in

Paper Citation Record · LEDGER

Taming Transformers for High-Resolution Image Synthesis

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 27 inbound Pith citation observations for arXiv:2012.09841.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2012.09841 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 27 of 27 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T21:34:57.381834Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-10T06:15:00.866473Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c9e1fa55-7d71-401f-b199-a0f536a99d2c · inbound

High-Resolution Image Synthesis with Latent Diffusion Models cites this paper.

High-Resolution Image Synthesis with Latent Diffusion Models Taming Transformers for High-Resolution Image Synthesis

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T22:02:10.049564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T22:02:09.899347Z digest=sha256:743e34c4762e1df8cde49b85d3c441c78d5822ad4f8801eb2f9315ec1c68a4f0

Observation 8688c964-81cb-44ee-bc32-acf0cc50e4bf · inbound

Hierarchical Text-Conditional Image Generation with CLIP Latents cites this paper.

Hierarchical Text-Conditional Image Generation with CLIP Latents Taming Transformers for High-Resolution Image Synthesis

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T16:55:57.706155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T16:55:57.612364Z digest=sha256:ffb819255422cf3e5a978887edf55bf8baac768c5ade3757534ee8d25998525f

Observation 627d82d3-d55d-4f65-beb1-a76d206ff8b7 · inbound

CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers cites this paper.

CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers Taming Transformers for High-Resolution Image Synthesis

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:24:30.169637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T12:24:30.071822Z digest=sha256:257dc797876a584a9c1a4dfa9d7f02423212db81080b428b7fc9b6976dd78c0a

Observation d992ae0f-c056-43e8-8e4b-a13887ea7fe6 · inbound

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets cites this paper.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Taming Transformers for High-Resolution Image Synthesis

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:58:52.084829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:01f945a109070ccbc3a25653c060da37ea43eac7ac7e357a2ac9cddfe62afb6d

Observation e8f7e586-273a-4e01-97e8-6a4102a7f999 · inbound

Emu3: Next-Token Prediction is All You Need cites this paper.

Emu3: Next-Token Prediction is All You Need Taming Transformers for High-Resolution Image Synthesis

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T10:56:07.899835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T10:56:06.418360Z digest=sha256:199b53c16de35a9aa208513bb00f209bb4e56bdae83642be37a3fb913358196c

Observation cd4dac47-1368-453e-9956-66214302bc43 · inbound

Towards Virtual Clinical Trials of Radiology AI with Conditional Generative Modeling cites this paper.

Towards Virtual Clinical Trials of Radiology AI with Conditional Generative Modeling Taming Transformers for High-Resolution Image Synthesis

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T21:34:57.381834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T21:34:57.381834Z digest=sha256:fd7762e731dbe4374a951e311d42cce79efc32521b336c0f8dc5efea83635e6e

Observation 577a4d28-e0a1-4008-9215-cc6894d99b66 · inbound

Tokenizing Electron Cloud in Protein-Ligand Interaction Learning cites this paper.

Tokenizing Electron Cloud in Protein-Ligand Interaction Learning Taming Transformers for High-Resolution Image Synthesis

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:26:30.195838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:26:30.195838Z digest=sha256:1876180a98cb1eaa058aad1977cfde0505bb313992225cefb9aa3fec772b58a2

Observation 0e9060c8-496a-4dba-989b-441152cb205f · inbound

LeDiFlow: Learned Distribution-guided Flow Matching to Accelerate Image Generation cites this paper.

LeDiFlow: Learned Distribution-guided Flow Matching to Accelerate Image Generation Taming Transformers for High-Resolution Image Synthesis

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:33.616788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:33.616788Z digest=sha256:cef5e1fb5aa9c5e386e7fd6f66d50122e6c88b634401c9f9d5cc97da8f3a2798

Observation 72e589fa-862b-4016-b3a5-4f6c6c8f6cea · inbound

Exploring the Effectiveness of Deep Features from Domain-Specific Foundation Models in Retinal Image Synthesis cites this paper.

Exploring the Effectiveness of Deep Features from Domain-Specific Foundation Models in Retinal Image Synthesis Taming Transformers for High-Resolution Image Synthesis

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:11.354540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:11.354540Z digest=sha256:6223b5572cbb8bdb4a80bddb595f0d9d83dae5669dffa8e85a4f6a0079e03a8f

Observation 39f2225e-a259-4729-818e-3aae54b1f9ce · inbound

Transition Matching: Scalable and Flexible Generative Modeling cites this paper.

Transition Matching: Scalable and Flexible Generative Modeling Taming Transformers for High-Resolution Image Synthesis

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T21:46:28.121890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:46:28.121890Z digest=sha256:bc1dcc17912ccf4d68fc91495aef0ab5b5621eb1ebf0c3ac77e835148201684c

Observation fda2a2e6-294f-44e9-9362-250e07a6b7ec · inbound

Latent Sensor Fusion: Multimedia Learning of Physiological Signals for Resource-Constrained Devices cites this paper.

Latent Sensor Fusion: Multimedia Learning of Physiological Signals for Resource-Constrained Devices Taming Transformers for High-Resolution Image Synthesis

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T17:58:36.667446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:58:36.667446Z digest=sha256:821b1b3666837af45e6697c168dcf3645d731a0fbdaac567005e92c2920bf5ee

Observation 25999249-f595-49b7-89a0-68e4bd432bcb · inbound

StorySync: Training-Free Subject Consistency in Text-to-Image Generation via Region Harmonization cites this paper.

StorySync: Training-Free Subject Consistency in Text-to-Image Generation via Region Harmonization Taming Transformers for High-Resolution Image Synthesis

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T10:47:00.547306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:47:00.547306Z digest=sha256:463b1161eb3242bbf3769a7e2ce9dd2025a2db91b9b0fd047bf74d39e930a691

Observation 898cb73b-271a-4ba7-830e-301c5af7812a · inbound

Controllable Skin Synthesis via Lesion-Focused Vector Autoregression Model cites this paper.

Controllable Skin Synthesis via Lesion-Focused Vector Autoregression Model Taming Transformers for High-Resolution Image Synthesis

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T15:41:16.150346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:41:16.150346Z digest=sha256:58e4dbea71ee70541e737c03cd0920a98ab97806f26220fba061b9e3da3b5317

Observation 64746141-28d3-40d5-b6fa-56f1b65a1650 · inbound

Analysis of Speaker Verification Performance Trade-offs with Neural Audio Codec Transmission cites this paper.

Analysis of Speaker Verification Performance Trade-offs with Neural Audio Codec Transmission Taming Transformers for High-Resolution Image Synthesis

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T11:29:03.656854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:29:03.656854Z digest=sha256:3bd9d5df87c6003fb02b3cb68155006d5e7ee25dc2e783a2d1cf5f14e2ed1b7d

Observation a5b1424c-630d-4465-8998-9f1c458fb51b · inbound

InertialAR: Autoregressive 3D Molecule Generation with Inertial Frames cites this paper.

InertialAR: Autoregressive 3D Molecule Generation with Inertial Frames Taming Transformers for High-Resolution Image Synthesis

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T07:22:12.652710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:22:12.652710Z digest=sha256:c9d8ea0e5e6e5611012849dee1cacb3aaaebd441d7bc137927cf0f0bbdcf3b05

Observation 5b12fbaa-9ad2-42d8-b177-119cdc997842 · inbound

What Matters for Diffusion-Friendly Latent Manifold? Prior-Aligned Autoencoders for Latent Diffusion cites this paper.

What Matters for Diffusion-Friendly Latent Manifold? Prior-Aligned Autoencoders for Latent Diffusion Taming Transformers for High-Resolution Image Synthesis

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:05:57.417886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T01:57:24.033068Z digest=sha256:33807109a8185626c4e36d8deff2fb6fafb1df4a8116be7c761b2380d4b5eb9e

Observation 734f5076-98ec-43e1-8c66-86eb436b1fa8 · inbound

CaloArt: Large-Patch x-Prediction Diffusion Transformers for High-Granularity Calorimeter Shower Generation cites this paper.

CaloArt: Large-Patch x-Prediction Diffusion Transformers for High-Granularity Calorimeter Shower Generation Taming Transformers for High-Resolution Image Synthesis

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-13T04:42:15.836286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T04:40:25.725622Z digest=sha256:5865243abddb514e5d87b6493e857b615aeb38d3fd3f8354bf8287c9317309dc

Observation 03843ae1-6ea9-475b-a954-43b6d8fe80ca · inbound

Learning from Semantic Dictionaries: Discriminative Codebook Contrastive Learning for Unified Visual Representation and Generation cites this paper.

Learning from Semantic Dictionaries: Discriminative Codebook Contrastive Learning for Unified Visual Representation and Generation Taming Transformers for High-Resolution Image Synthesis

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T12:34:38.743266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T12:28:33.099835Z digest=sha256:6e0534c1878b378508840e5109250a4ce7c3f445855a9019847057d131deb1a3

Observation aa44a522-961e-4bbc-942b-d6f2520129dc · inbound

GPIC: A Giant Permissive Image Corpus for Visual Generation cites this paper.

GPIC: A Giant Permissive Image Corpus for Visual Generation Taming Transformers for High-Resolution Image Synthesis

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T07:43:14.161246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T07:36:21.262064Z digest=sha256:9b81ce9569b8a8f8891ad130c881bdeab40533124b9e82d9ee0f86d114c63267

Observation 1dc90202-abb1-4242-bb93-fc2e14f65420 · inbound

Consistency Training while Mitigating Obfuscation via Rate Matching cites this paper.

Consistency Training while Mitigating Obfuscation via Rate Matching Taming Transformers for High-Resolution Image Synthesis

Reference 141

Resolution
verified exact
arxiv_id, observed 2026-06-28T14:32:18.108266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T14:25:43.147442Z digest=sha256:68ca4b0db58ca237edb2cba86464f026e106983544d4511a3be3d1b2ef3aa8dc

Observation e0631fdc-ce03-4906-ba42-240e4f13806d · inbound

Mapping Whisper Representations to Human ECoG Responses with Interpretable Time-Resolved Neural Encoding cites this paper.

Mapping Whisper Representations to Human ECoG Responses with Interpretable Time-Resolved Neural Encoding Taming Transformers for High-Resolution Image Synthesis

Reference 128

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T11:42:04.140355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T11:39:30.563754Z digest=sha256:d53a1abda6513937acfe617d373c68712ddbc6831b575b667db02270108d1fd8

Observation 7aa61616-3ae2-417a-b70b-fa8eb1c68ca9 · inbound

Variational Proximal Policy Optimization cites this paper.

Variational Proximal Policy Optimization Taming Transformers for High-Resolution Image Synthesis

Reference 52

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T21:57:26.059020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T19:22:23.768249Z digest=sha256:816389929bf30442563b25d77d8a275bd0be134c491d2ff7428e2f097083bff8

Observation 542bbc52-ab6e-49ae-9736-e90bc4695b2c · inbound

NSVQ: Mitigating Codebook Collapse by Stabilizing Encoder Drift in Vector Quantization cites this paper.

NSVQ: Mitigating Codebook Collapse by Stabilizing Encoder Drift in Vector Quantization Taming Transformers for High-Resolution Image Synthesis

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T05:27:39.708579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T13:18:47.472178Z digest=sha256:92fb25869b699e96d3fe9201145bb0278bb8c7a08c92568cdd5232aaa5c51eb2

Observation 70464072-1d41-4a12-98c3-acceea611229 · inbound

The Market in the Model: Latent Diffusion as Neural Economy cites this paper.

The Market in the Model: Latent Diffusion as Neural Economy Taming Transformers for High-Resolution Image Synthesis

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T02:49:24.758638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T19:05:25.386543Z digest=sha256:c5814063273d6f98edfe9cbe4383a27c618a6a18f6c6b5106862d0dd3f4230ab

Observation 9f2441b9-0330-4bcb-b0a0-5d1c596fbad7 · inbound

Moebius: 0.2B Lightweight Image Inpainting Framework with 10B-Level Performance cites this paper.

Moebius: 0.2B Lightweight Image Inpainting Framework with 10B-Level Performance Taming Transformers for High-Resolution Image Synthesis

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:39:17.090513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T21:03:43.962738Z digest=sha256:6b1fd3043132f5d0d6a60fb9942539415f2514d321da51e5f469de87ef586fc6

Observation d0707cb4-3fa8-4053-8994-3dca0f5a748d · inbound

Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing cites this paper.

Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing Taming Transformers for High-Resolution Image Synthesis

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-01T13:39:00.040780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:39:00.040780Z digest=sha256:8664006d1cf17e002d54837a088d429b521548f8c2e6606edcddac33f807c061

Observation 01433ba0-a80d-4a11-b89c-51a8c8fefd1a · inbound

FairDiffuseVQVAE: Sampling-Time Fairness in Tabular Diffusion via Conditional Refinement of Vector-Quantized Latents cites this paper.

FairDiffuseVQVAE: Sampling-Time Fairness in Tabular Diffusion via Conditional Refinement of Vector-Quantized Latents Taming Transformers for High-Resolution Image Synthesis

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T16:50:35.363661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T16:50:35.363661Z digest=sha256:1a62e1cda5bf141bd4eb1bb1c588f4aa04a1789767dbbcfb0ec94a9a8b77840c