Pith. sign in

Paper Citation Record · LEDGER

Muse: Text-To-Image Generation via Masked Generative Transformers

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 65 inbound Pith citation observations for arXiv:2301.00704.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2301.00704 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 65 of 65 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 65 of 65 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T05:54:18.435628Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

119
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 9232872e-81f2-4cf9-80d4-dadf982e014b · inbound

Scaling Robot Learning with Semantically Imagined Experience cites this paper.

Scaling Robot Learning with Semantically Imagined Experience Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 78

Resolution
verified exact
arxiv_id, observed 2026-05-17T18:59:10.586456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T18:59:10.352342Z digest=sha256:96c4c0e7830df03c5976087577f1975ee85e32d4b60102a6a45e5ebb7969abd3

Observation 0f796b13-7c6f-4551-8c62-4e22c30ac520 · inbound

DragNUWA: Fine-grained Control in Video Generation by Integrating Text, Image, and Trajectory cites this paper.

DragNUWA: Fine-grained Control in Video Generation by Integrating Text, Image, and Trajectory Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 129

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:03:57.973110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-20T13:03:57.828598Z digest=sha256:d882ac8dacabec3d48ad6166786a7b8a832b1dbe6f3c7501343cb78377a20f39

Observation a694b0ae-382a-41c6-b127-2f87bba00510 · inbound

Finite Scalar Quantization: VQ-VAE Made Simple cites this paper.

Finite Scalar Quantization: VQ-VAE Made Simple Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-16T18:34:13.478920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T18:34:13.439993Z digest=sha256:f8071690fda309ee9e839a796fe37d9945a5fa9300bedba6ef55d7f9841a09c7

Observation c77dcd03-48dd-492f-9de5-fdf451d90738 · inbound

Learning Interactive Real-World Simulators cites this paper.

Learning Interactive Real-World Simulators Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 173

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T02:15:18.633063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-16T02:15:18.265190Z digest=sha256:cb2e832288c15f5d2aa249daa0ae64c997815cae9f9cc94228e302b724ba407e

Observation a07e5ce6-1751-4c76-9b09-d8f18e2bd385 · inbound

VideoCrafter1: Open Diffusion Models for High-Quality Video Generation cites this paper.

VideoCrafter1: Open Diffusion Models for High-Quality Video Generation Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:40:44.012282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T21:40:43.956642Z digest=sha256:ab18f95542ed697006a5ec04f58bdafa8882c799d582972e224898f9d69e3844

Observation c40f8024-858c-4265-a131-657112c79235 · inbound

VideoPoet: A Large Language Model for Zero-Shot Video Generation cites this paper.

VideoPoet: A Large Language Model for Zero-Shot Video Generation Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T17:51:05.648073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T17:51:05.465548Z digest=sha256:ac28b8d29508977aa43d82798bfd66780933ac8307b73c589ee3168dda805ab9

Observation 5b3eb7f1-a16f-4980-8df3-66b9a26b42c7 · inbound

Controllable Image Generation with Composed Parallel Token Prediction cites this paper.

Controllable Image Generation with Composed Parallel Token Prediction Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-24T00:53:40.560109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-24T00:53:36.830943Z digest=sha256:b5077124666b384b70594460e7d0d493471adb715ec4dccd6f347f1c2fcc3d3d

Observation 53a8d595-7b89-4e86-9456-b2e176e8dd26 · inbound

Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation cites this paper.

Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:09:16.842300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T22:09:16.622717Z digest=sha256:3325f1a570aabf180d9bd61accc666237b2d408e9cd5274c83d825835b8a977d

Observation 131f433b-66db-4afd-a722-86e083c996f1 · inbound

Show-o: One Single Transformer to Unify Multimodal Understanding and Generation cites this paper.

Show-o: One Single Transformer to Unify Multimodal Understanding and Generation Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:03:33.528653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T21:03:33.427939Z digest=sha256:2663a1a1432d2e179209f29831adcc7cebbf61efb04d40451e7ae24e4cd36728

Observation 80dc1c22-9878-45d2-9c2a-0b0e873d3836 · inbound

Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation cites this paper.

Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:09:16.400212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T22:09:16.001309Z digest=sha256:2b9de4a92b67d0ad7028b99cff9e8c58e4d85e95042389e9d59a37cf8e457714

Observation 66e0d9ea-dd01-44ae-baca-dd2d23af5fab · inbound

Autoregressive Video Generation without Vector Quantization cites this paper.

Autoregressive Video Generation without Vector Quantization Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-17T15:07:39.770391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T15:07:39.718555Z digest=sha256:f478578b9b6416ce087503b18ae1eb060c7c0755243911838a7747f06ea2904a

Observation 8ab7713c-374a-45dc-a085-e8a845ca2ad3 · inbound

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training cites this paper.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.435628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.435628Z digest=sha256:2d9772efea15e8164ae03488202e59c44a39dc11af738a63db58cd17b550a7e5

Observation b86abbaa-4d4c-4d29-acc8-624c0b490d9d · inbound

A Comprehensive Review on Noise Control of Diffusion Model cites this paper.

A Comprehensive Review on Noise Control of Diffusion Model Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T21:59:05.452842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:59:05.452842Z digest=sha256:4183856535335dbc002b3def58b7e3b9954498b98fb7da2a0c16698d837f652b

Observation 24e8bcc5-a1d0-4987-94ad-3c0953077f90 · inbound

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding cites this paper.

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T19:32:31.782645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:32:31.782645Z digest=sha256:f2a1b10eb80b93e21209c64b156c5a15033c25d774c559c978596a537e734f05

Observation e226a789-8427-4973-bcc9-0d1358c5bb95 · inbound

UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths cites this paper.

UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-08T15:24:46.297104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:24:46.297104Z digest=sha256:03c1e0506b460ef8bfce74b1232ceba049cce57cadee233bc78adab57d60162a

Observation a882e97a-118b-4887-abc5-24d56748b228 · inbound

HDT: Hierarchical Discrete Transformer for Multivariate Time Series Forecasting cites this paper.

HDT: Hierarchical Discrete Transformer for Multivariate Time Series Forecasting Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T05:44:46.441310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T05:44:46.441310Z digest=sha256:6e2c62f72c3e0dc816d473e465e3cf534fb91c2eaf238c10f34c9e328dbb973e

Observation 75741d7b-c34c-41a6-a371-b8e7c276f273 · inbound

E-MD3C: Taming Masked Diffusion Transformers for Efficient Zero-Shot Object Customization cites this paper.

E-MD3C: Taming Masked Diffusion Transformers for Efficient Zero-Shot Object Customization Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T22:29:07.156820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T22:29:07.156820Z digest=sha256:86bbd97d9f3d2df4faef08f68f5eedb6e25c4775d6e4431d127e73c2d47a9fea

Observation e4c62dd1-97f8-43f1-b85b-6c6e62cec95a · inbound

Large Language Diffusion Models cites this paper.

Large Language Diffusion Models Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:42:54.605300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T01:42:54.279353Z digest=sha256:ebcff088a630350c6e5dd03c49bab47b6f445a899eb6ee999b538f38b4fe3064

Observation 2a1143d1-7a46-454e-9f27-222bada2c9f2 · inbound

MSDformer: Multi-scale Discrete Transformer For Time Series Generation cites this paper.

MSDformer: Multi-scale Discrete Transformer For Time Series Generation Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-22T13:44:52.815486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T13:44:08.891339Z digest=sha256:95cec135512aaf4e6a8b3adec1c46302f0b65de5e8477c4e812b9508f0961d4e

Observation 25db7335-c11d-44ce-a2bb-695de5c292dd · inbound

MMaDA: Multimodal Large Diffusion Language Models cites this paper.

MMaDA: Multimodal Large Diffusion Language Models Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 107

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:50:59.789654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:4289bdae75de83fc544be7528ac2c58d626794763d0519115c1c9d2461602ea3

Observation e77c68c5-45da-461b-aeba-416f1234010b · inbound

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet cites this paper.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:57.486525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:57.486525Z digest=sha256:4fa5f01915e56c35baa8c92601d0a4677fdc73a449f2c4d66af0287b9d10c551

Observation 9e99710c-9d77-427f-8763-c4e96e5bdcd7 · inbound

LaViDa: A Large Diffusion Language Model for Multimodal Understanding cites this paper.

LaViDa: A Large Diffusion Language Model for Multimodal Understanding Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:59:32.902329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:59:32.902329Z digest=sha256:dab4e4866e60a7d7ad49330b8ac60b452d5eeae82bf48467b732e869385b1692

Observation f51b3af5-c512-4de2-977c-a437fe25ccf4 · inbound

Semantics-Aware Human Motion Generation from Audio Instructions cites this paper.

Semantics-Aware Human Motion Generation from Audio Instructions Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:35.069080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:48:35.069080Z digest=sha256:34f8f958a178b96e0bd3a1482dde72c6fbd7ebf359558e9ed9b776c43ec20462

Observation 20a3d936-6f55-46db-9e63-6e5bf523bf11 · inbound

IMPACT: Iterative Mask-based Parallel Decoding for Text-to-Audio Generation with Diffusion Modeling cites this paper.

IMPACT: Iterative Mask-based Parallel Decoding for Text-to-Audio Generation with Diffusion Modeling Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:23.180725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:03:23.180725Z digest=sha256:0982d4153dfafde88ae9fab7d41482227fd5efdf4e5e270ff4e39ce78a924ea7

Observation 36887104-8b08-42f6-9869-9f24112971e3 · inbound

Native-Resolution Image Synthesis cites this paper.

Native-Resolution Image Synthesis Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:15:21.740312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:15:21.740312Z digest=sha256:7adca159f35d54d85c9f1f3da9ed22edf57602f66d45d1c7bab388293aa68243

Observation 4b2b60c7-2598-43c8-ad79-1de3049e13ba · inbound

HMAR: Efficient Hierarchical Masked Auto-Regressive Image Generation cites this paper.

HMAR: Efficient Hierarchical Masked Auto-Regressive Image Generation Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T10:47:52.273933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:47:52.273933Z digest=sha256:9978ef1a09f6e5c3f40c26bfd0fe25c0a8249228b7bfb365fad27c25a1dd8aca

Observation c3712f10-5504-470b-a3ff-dff63c008640 · inbound

Noise Consistency Regularization for Improved Subject-Driven Image Synthesis cites this paper.

Noise Consistency Regularization for Improved Subject-Driven Image Synthesis Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:59:47.443309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:59:47.443309Z digest=sha256:0d49951f920f0abbc8ab1495fa67fb8bc9a9ed4a013bbe48f9273032510e9101

Observation c9c346ec-e6fd-49ab-adf8-2a271f52af36 · inbound

MapBERT: Bitwise Masked Modeling for Real-Time Semantic Mapping Generation cites this paper.

MapBERT: Bitwise Masked Modeling for Real-Time Semantic Mapping Generation Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:23.832417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:23.832417Z digest=sha256:3de6c6e4dac96bddb88e6fe09766b1fcdf8131c6386da4388f947a0fc3463f14

Observation ff38e35f-58a8-4028-8e43-25e4f3c590c4 · inbound

Output Scaling: YingLong-Delayed Chain of Thought in a Large Pretrained Time Series Forecasting Model cites this paper.

Output Scaling: YingLong-Delayed Chain of Thought in a Large Pretrained Time Series Forecasting Model Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-07T15:39:24.622157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:39:24.622157Z digest=sha256:aeb9c63fb8c1ef8b83592531891b2fac24883b6c1c0e4f0429defae3d0aa1112

Observation 93898f86-93a1-4007-bbe9-fd4ddcd0c70e · inbound

MARch\'e: Fast Masked Autoregressive Image Generation with Cache-Aware Attention cites this paper.

MARch\'e: Fast Masked Autoregressive Image Generation with Cache-Aware Attention Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:51:36.281876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:51:36.281876Z digest=sha256:4144009a243e6ca31066d8c36a8d1c0006005740909cf3a8880f2e7d8758007b

Observation 7f962a37-9a2e-400b-9a7d-5a0dbce28b9d · inbound

Learning golf swing signatures from a single wrist-worn inertial sensor cites this paper.

Learning golf swing signatures from a single wrist-worn inertial sensor Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T23:35:33.442154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:35:33.442154Z digest=sha256:f79a87f7ed6e51ecbd0d1e51aa957f868fbf7c75906b5351433bd2cd9183c996

Observation e8f6cd43-d7dc-455d-a05e-11b2a34633df · inbound

MVGBench: Comprehensive Benchmark for Multi-view Generation Models cites this paper.

MVGBench: Comprehensive Benchmark for Multi-view Generation Models Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T04:52:09.670357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:52:09.670357Z digest=sha256:a2895d3ea0c2d726cc8c6f562227c5eeab459827075d45ef461340490bfa79b2

Observation 53d434c6-451e-4728-bb7a-2559e37af9e4 · inbound

Is Visual in-Context Learning for Compositional Medical Tasks within Reach? cites this paper.

Is Visual in-Context Learning for Compositional Medical Tasks within Reach? Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T21:12:05.800608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:12:05.800608Z digest=sha256:e30a963890557b357d4ee70b6261205908d63d2ead1213983708a3c633e18e0f

Observation dc6d1286-0035-4aa1-9ff1-5e25eab8a452 · inbound

CI-VID: A Coherent Interleaved Text-Video Dataset cites this paper.

CI-VID: A Coherent Interleaved Text-Video Dataset Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T20:43:53.858179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:43:53.858179Z digest=sha256:8a6e74c3abca0ae03eac2a01f4d304d1538d91fe15c9844167b5cba4d62cef21

Observation 0757d32b-5882-4aac-a7ab-e83195aef7f6 · inbound

DC-AR: Efficient Masked Autoregressive Image Generation with Deep Compression Hybrid Tokenizer cites this paper.

DC-AR: Efficient Masked Autoregressive Image Generation with Deep Compression Hybrid Tokenizer Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T19:41:23.621281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:41:23.621281Z digest=sha256:9fb6ab3dffe1f97a6e9f9f123ef7a113d5718a5c4a73c090cf244e5a3415e7b4

Observation ed205e6c-4d82-4db4-bc82-fb3e72b207a6 · inbound

Room Impulse Response Generation Conditioned on Acoustic Parameters cites this paper.

Room Impulse Response Generation Conditioned on Acoustic Parameters Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T16:58:28.623018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:58:28.623018Z digest=sha256:0dfa00cd4e8fbefa272d41d5e5f6b670460fbc48b47d62eaff2ac534958f30ec

Observation 07764351-64d3-4d07-9710-39ec82cf5df2 · inbound

MADI: Masking-Augmented Diffusion with Inference-Time Scaling for Visual Editing cites this paper.

MADI: Masking-Augmented Diffusion with Inference-Time Scaling for Visual Editing Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T16:48:29.960911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:48:29.960911Z digest=sha256:0f1d8d701435a3289e9ab6914456331172e19b6745c5d55cfb5c3b8e64a5df6b

Observation 3469a9a0-887f-4059-ad72-6d7dc63615e5 · inbound

Learning neuro-symbolic convergent term rewriting systems cites this paper.

Learning neuro-symbolic convergent term rewriting systems Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T14:28:45.675804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:28:45.675804Z digest=sha256:0bfe36fd1f70cdad13cb0d4a8f17d0c57f5cbc0a24fd6837fcb0347af3dd6f7b

Observation 7235f01d-ba09-49de-a38a-6496680f10d2 · inbound

Zero-Residual Concept Erasure via Progressive Alignment in Text-to-Image Model cites this paper.

Zero-Residual Concept Erasure via Progressive Alignment in Text-to-Image Model Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T00:04:32.077010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:04:32.077010Z digest=sha256:59710ae8048aa2d686c991c93232c20e63180aef79c2f7dbab35241e359ec372

Observation c0233a12-2300-4dcf-8ad9-faf649f90181 · inbound

FBI: Learning Dexterous In-hand Manipulation with Dynamic Visuotactile Shortcut Policy cites this paper.

FBI: Learning Dexterous In-hand Manipulation with Dynamic Visuotactile Shortcut Policy Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T18:37:17.815731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:37:17.815731Z digest=sha256:2ee4c9b21c2ec03ae3a669540337ab7b1979286755e893d5bcf59a401c27baa1

Observation 03f319a0-c123-4471-9679-1616fb66726a · inbound

Layout-Conditioned Autoregressive Text-to-Image Generation via Structured Masking cites this paper.

Layout-Conditioned Autoregressive Text-to-Image Generation via Structured Masking Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T16:43:49.262777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T16:43:49.262777Z digest=sha256:7254efb723681402588429dceec26a3022a6d6fdb749f76b6c67aba416e51647

Observation a3eeabf2-40d5-4655-b1f5-f0bdfb7aee7c · inbound

Lavida-O: Elastic Large Masked Diffusion Models for Unified Multimodal Understanding and Generation cites this paper.

Lavida-O: Elastic Large Masked Diffusion Models for Unified Multimodal Understanding and Generation Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T15:39:34.022747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T15:39:34.022747Z digest=sha256:c0a0ea7b7ad60f40fe851e7b50f94a06656bfa1e6935593ac46eb708b1c55615

Observation 5b9524c7-760c-40a4-b39b-46c5de90dd0f · inbound

IAR2: Improving Autoregressive Visual Generation with Semantic-Detail Associated Token Prediction cites this paper.

IAR2: Improving Autoregressive Visual Generation with Semantic-Detail Associated Token Prediction Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T11:08:02.386459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:08:02.386459Z digest=sha256:eb37ebd6d999d5a7af1361170689a7641740801cd4c0cb5f44cc6da563cc99b3

Observation 71063a29-1ae8-45a6-87fa-202e586ea3d0 · inbound

RubricRL: Simple Generalizable Rewards for Text-to-Image Generation cites this paper.

RubricRL: Simple Generalizable Rewards for Text-to-Image Generation Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T20:15:25.700535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:15:25.700535Z digest=sha256:fd9b36bed04c82207a8070ce6d254a20a0614a827740fc59c671f66cd0428cd2

Observation ab85a340-a5c4-4404-a48d-c2a01f7d38e5 · inbound

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models cites this paper.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:31.725029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:31.725029Z digest=sha256:45b774df2971f8ce4711f03f9541f2bda7c59d12a77870a327d4a75e8bff757d

Observation e823a537-1fdc-4c00-bb45-09ecb3b07f47 · inbound

The Script is All You Need: An Agentic Framework for Long-Horizon Dialogue-to-Cinematic Video Generation cites this paper.

The Script is All You Need: An Agentic Framework for Long-Horizon Dialogue-to-Cinematic Video Generation Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-03T08:16:22.354776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:16:22.354776Z digest=sha256:e49973efc4e9b6a5a61b495ac05ddc5b420d2f15505f34acc795ff3fdac7f202

Observation e8362087-437c-4cbf-a03b-83b1d4e73e9b · inbound

Improving Text-to-Image Generation with Intrinsic Self-Confidence Rewards cites this paper.

Improving Text-to-Image Generation with Intrinsic Self-Confidence Rewards Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:41:28.773088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T18:40:18.940889Z digest=sha256:4e71e23f43da14a9efca1e00ba307ce2c1639b9271913d99faa7b352c993d879

Observation a83873ff-e267-4133-8856-ff170a96f15b · inbound

Unifying Contrastive and Generative Objectives for Visual Understanding and Text-to-Image Generation cites this paper.

Unifying Contrastive and Generative Objectives for Visual Understanding and Text-to-Image Generation Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-21T12:20:06.998166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T12:15:57.019720Z digest=sha256:bb668bc25d8a6d8d24b810f9af8f18a4e024190e0b71645a10a2bdc71ecbaf75

Observation ef4784c5-9a87-4b1d-9891-5fb6a6b4e38f · inbound

Controllable Image Generation with Composed Parallel Token Prediction cites this paper.

Controllable Image Generation with Composed Parallel Token Prediction Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:50:48.709020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T19:33:03.066564Z digest=sha256:a838f15ea2bd71121742af1eda4f991e3ae1cdaf513a8f53e6824cc0f3bdde54

Observation 14212172-9d6b-4810-9261-eae613d889da · inbound

UDM-GRPO: Stable and Efficient Group Relative Policy Optimization for Uniform Discrete Diffusion Models cites this paper.

UDM-GRPO: Stable and Efficient Group Relative Policy Optimization for Uniform Discrete Diffusion Models Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T05:36:01.840743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T05:32:54.235578Z digest=sha256:279f9b3cf38b7bc89887e6590e0e320e9d318915dfaec2fa44d49c855c54f1f6

Observation 7d97494e-1a26-412c-a643-e19525e82873 · inbound

UDM-GRPO: Stable and Efficient Group Relative Policy Optimization for Uniform Discrete Diffusion Models cites this paper.

UDM-GRPO: Stable and Efficient Group Relative Policy Optimization for Uniform Discrete Diffusion Models Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-07-05T11:41:02.733153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-05T11:32:38.636335Z digest=sha256:053a24baa1595c4cdad8db25e5492a520b8352b93074bcf2bb6521485b74f5bd

Observation 83ed988c-ef11-45c2-a602-6aad93d637dc · inbound

Coupling Models for One-Step Discrete Generation cites this paper.

Coupling Models for One-Step Discrete Generation Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 59

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T03:00:55.161047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-11T02:58:10.909499Z digest=sha256:597af9f29ae42c01796c67bc5123204980cb5203739d2785a9cf0d41c8e33f60

Observation 9d6ce223-0082-4686-ae56-d033d04f4f12 · inbound

VAGS: Velocity Adaptive Guidance Scale for Image Editing and Generation cites this paper.

VAGS: Velocity Adaptive Guidance Scale for Image Editing and Generation Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-20T18:33:53.079392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T18:33:47.573164Z digest=sha256:61695df32a0c299e5d8b72e988a3475dedfbefe2ae471d949cb69555d2180090

Observation 609ca211-2bb8-4d53-b578-5bbb2630e701 · inbound

Dimension-Free Convergence of Discrete Diffusion Models: Adjoint Equations Induce the Right Space cites this paper.

Dimension-Free Convergence of Discrete Diffusion Models: Adjoint Equations Induce the Right Space Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:48:23.247806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-20T14:47:33.474443Z digest=sha256:6e439abb02c7114973f1fe59590bde88652d35b9deb61f9982e91914cb5cbdf8

Observation 04ea7e06-9606-4226-9f00-13106dcbd614 · inbound

Dimension-Free Convergence of Discrete Diffusion Models: Adjoint Equations Induce the Right Space cites this paper.

Dimension-Free Convergence of Discrete Diffusion Models: Adjoint Equations Induce the Right Space Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-06-30T19:15:00.171208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-30T19:13:35.794787Z digest=sha256:af8c84e421bf753de2f35e597e14c8bc35fb4975761a38b9021c0ccce1ae4ca3

Observation 281bbec6-d233-4b26-9316-e46384d8f824 · inbound

Dimension-Free Convergence of Discrete Diffusion Models: Adjoint Equations Induce the Right Space cites this paper.

Dimension-Free Convergence of Discrete Diffusion Models: Adjoint Equations Induce the Right Space Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T13:53:01.557172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T13:53:01.557172Z digest=sha256:bd0d4d6036c14208a2046e924c8eec0390d13afd9e56571df02262e337dec9ab

Observation bc82891e-2ef2-4cef-886d-9807b393be94 · inbound

Going PLACES: Participatory Localized Red Teaming for Text-to-Image Safety in the Global South cites this paper.

Going PLACES: Participatory Localized Red Teaming for Text-to-Image Safety in the Global South Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:08:07.000756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T07:06:57.555070Z digest=sha256:d68270f0c4208a45e4d37fbd8989481613ecc5e5f43045e34caf7742b372c0b1

Observation 41295a1b-dadb-4525-8144-73c64613c3ba · inbound

SplitAvatar: One-shot Head Avatar with Autoregressive Gaussian Splitting cites this paper.

SplitAvatar: One-shot Head Avatar with Autoregressive Gaussian Splitting Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T23:04:00.978197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T23:02:13.380306Z digest=sha256:3aa1b114efef8652fbfd00bd18900ec66a6767792d881bc3732322f2b2cb1a9a

Observation 98eb01a1-65d2-4d36-b016-935ee963d365 · inbound

Diffusing in the Right Space: A Systematic Study of Latent Diffusability cites this paper.

Diffusing in the Right Space: A Systematic Study of Latent Diffusability Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 78

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T02:36:27.646872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T10:44:24.318786Z digest=sha256:6b14822806997558d4e2f178832f65c70894315ffb75f992d9d4117f1a3c4d4f

Observation b7772da1-d62b-4175-b2b5-9909ce03c38b · inbound

What Type of Inference is Active Inference? cites this paper.

What Type of Inference is Active Inference? Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-06-28T06:51:44.836825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T06:45:53.233544Z digest=sha256:6f90358029916290f980d1cab8153a00bbe5fadab9d7fcdd6f8c0006beed3cdc

Observation 05d80a24-c1eb-495f-8d8c-cf224c716547 · inbound

ARM: An AutoRegressive Large Multimodal Model with Unified Discrete Representations cites this paper.

ARM: An AutoRegressive Large Multimodal Model with Unified Discrete Representations Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:07:38.441534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T13:29:11.526106Z digest=sha256:c5379dc9cd4abd8a5840f1fcc79585e2373d10a4b34cdd49f3ecaf6bd3055a9a

Observation 5a832f26-c88c-40e4-81b6-d31f4a969b93 · inbound

Expected Free Energy-based Planning as Variational Inference cites this paper.

Expected Free Energy-based Planning as Variational Inference Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-06-27T13:40:57.096595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T13:39:28.078342Z digest=sha256:2c4d19cd74c47661dcb6f7099be0b29be4363e06486636534c233587a50cbe07

Observation 4e6b0e0c-9da9-4d51-bef2-c9129619da42 · inbound

Co-occurring associated retained concepts in Diffusion Unlearning cites this paper.

Co-occurring associated retained concepts in Diffusion Unlearning Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T16:09:57.475269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T00:50:58.689718Z digest=sha256:c7761923682ce58bb06c7bde0c99dc23299aa2fbf24c8ec3d0e9d004e5aec87f

Observation 7d4417cc-c09e-40ff-9e0e-29344603c31c · inbound

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model cites this paper.

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 153

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T11:55:42.259206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T03:50:26.873406Z digest=sha256:cc8679213866877b9deafdaf919cbe2f83d2dbf31974acf45b69f625df71c8d9

Observation bda92eb2-a1ff-4c83-8505-db6d71b2ede1 · inbound

MentalThink: Shaping Thoughts in Mental SVG World cites this paper.

MentalThink: Shaping Thoughts in Mental SVG World Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-12T01:50:59.184754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T01:50:59.184754Z digest=sha256:fc0dd580ac9c4d295d40f1646a7d5032078ce29f0d8eaefc1eea5c5cdc259076