Pith. sign in

Paper Citation Record · LEDGER

Make-A-Scene: Scene-Based Text-to-Image Generation with Human Priors

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2203.13131.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2203.13131 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T21:03:23.615737Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T23:04:01.962523Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 17a15ade-d7d2-41dd-9c9f-51699fc5eea9 · inbound

High-Resolution Image Synthesis with Latent Diffusion Models cites this paper.

High-Resolution Image Synthesis with Latent Diffusion Models Make-A-Scene: Scene-Based Text-to-Image Generation with Human Priors

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T22:02:10.163577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T22:02:09.899347Z digest=sha256:8456566c5fea0ed3b3e0b8db5bf1647977bf8adce8afb461b3da3c019434c01f

Observation 8628b0d2-3651-4406-8016-f1f5052a6b1e · inbound

Hierarchical Text-Conditional Image Generation with CLIP Latents cites this paper.

Hierarchical Text-Conditional Image Generation with CLIP Latents Make-A-Scene: Scene-Based Text-to-Image Generation with Human Priors

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T16:55:57.716963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T16:55:57.612364Z digest=sha256:c98baa276102aae7f5dfdcfcbefcc67467163debfb484f4af8b257fe07101087

Observation 42d8bd7a-bb9e-48c7-8223-a78aec1d9517 · inbound

Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding cites this paper.

Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding Make-A-Scene: Scene-Based Text-to-Image Generation with Human Priors

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:38:53.399476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T07:38:53.056362Z digest=sha256:aa890aad4d1d1a2d290bb431d302ca41ef62088b11dd1e584ca8d17ee0070f05

Observation 69afec39-e1a1-4aad-8f38-061e7d1ca95f · inbound

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation cites this paper.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation Make-A-Scene: Scene-Based Text-to-Image Generation with Human Priors

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T04:49:31.051437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:bc37c013a88a3cfd1c8a9e8d3f3ffac9c3232f3dd7726d6291b2e973805b319e

Observation 34a4a8bd-fd6e-488b-9688-f024a5810660 · inbound

Prompt-to-Prompt Image Editing with Cross Attention Control cites this paper.

Prompt-to-Prompt Image Editing with Cross Attention Control Make-A-Scene: Scene-Based Text-to-Image Generation with Human Priors

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T07:00:01.514550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T07:00:01.154743Z digest=sha256:6c79ec5799702dcfb6e9ec0c867be6f6ed82071d524158e3d38ffc285b86c3c1

Observation e1def9de-ccba-4988-b46f-8057a4b47479 · inbound

Make-A-Video: Text-to-Video Generation without Text-Video Data cites this paper.

Make-A-Video: Text-to-Video Generation without Text-Video Data Make-A-Scene: Scene-Based Text-to-Image Generation with Human Priors

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:13:03.268349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T01:13:03.213358Z digest=sha256:0b5fef35124400d36f192b7718c52281c3f40c161721e792e8eac6c94a952c8b

Observation ed82296c-576b-4c67-870b-f8a6d783e25f · inbound

eDiff-I: Text-to-Image Diffusion Models with an Ensemble of Expert Denoisers cites this paper.

eDiff-I: Text-to-Image Diffusion Models with an Ensemble of Expert Denoisers Make-A-Scene: Scene-Based Text-to-Image Generation with Human Priors

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T01:44:22.771350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T01:44:22.710205Z digest=sha256:e7b0a250df265c8ea5af65e2a5766aa019447b6accf9c1db7a6acd33282e1659

Observation b7ab1f25-669e-4a4d-937f-4e6bcb000f4c · inbound

Visual Instruction Tuning cites this paper.

Visual Instruction Tuning Make-A-Scene: Scene-Based Text-to-Image Generation with Human Priors

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T08:22:03.728483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T08:22:03.403362Z digest=sha256:cfa66fd580eb0064e6ae304aa2f28ab9eb1c491c75080739390e621d7c4c0f58

Observation 79d8d55d-a167-445d-ac29-5518b73e014d · inbound

Shap-E: Generating Conditional 3D Implicit Functions cites this paper.

Shap-E: Generating Conditional 3D Implicit Functions Make-A-Scene: Scene-Based Text-to-Image Generation with Human Priors

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T15:32:06.740762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T15:32:06.563955Z digest=sha256:2b823d063918161df09ca758c19bb5cc2ea65df87527018ec122a30f930fa72e

Observation b20c9735-a372-460f-a71d-1069a8275e84 · inbound

DragNUWA: Fine-grained Control in Video Generation by Integrating Text, Image, and Trajectory cites this paper.

DragNUWA: Fine-grained Control in Video Generation by Integrating Text, Image, and Trajectory Make-A-Scene: Scene-Based Text-to-Image Generation with Human Priors

Reference 244

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:03:58.150297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-20T13:03:57.828598Z digest=sha256:8a714732cdaaa3c939ad9a37936fc191bde7db78731bb20fa0b9868d6f1cfb75

Observation 9bf17676-6f02-4ced-8729-26629bfa3760 · inbound

Chameleon: Mixed-Modal Early-Fusion Foundation Models cites this paper.

Chameleon: Mixed-Modal Early-Fusion Foundation Models Make-A-Scene: Scene-Based Text-to-Image Generation with Human Priors

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T10:03:28.051369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T10:03:27.919346Z digest=sha256:e6d8499c1237b92d99b62a7a832b6c0eb87e6a9ec1be8f689aefb9d469b35c87

Observation 64cde2e2-49af-4cb3-9f2a-7d0a9de71ac5 · inbound

Cached Multi-Lora Composition for Multi-Concept Image Generation cites this paper.

Cached Multi-Lora Composition for Multi-Concept Image Generation Make-A-Scene: Scene-Based Text-to-Image Generation with Human Priors

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T21:03:23.615737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T21:03:23.615737Z digest=sha256:3112585c9c68202a1fa19960db5b1db8cf54203d3723f78a53f6726f3501fba6

Observation 5a7b3cea-6c0e-4375-a1a1-21d554c2d08c · inbound

STARFlow: Scaling Latent Normalizing Flows for High-resolution Image Synthesis cites this paper.

STARFlow: Scaling Latent Normalizing Flows for High-resolution Image Synthesis Make-A-Scene: Scene-Based Text-to-Image Generation with Human Priors

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:59.560240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:59.560240Z digest=sha256:20f57ef57ce7ea792a4f296cdc6f59a3bc37bfd27aeab36ecf02245b3a4ff090

Observation 6e88afe7-aa0c-466e-b77f-134e9213febc · inbound

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models cites this paper.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Make-A-Scene: Scene-Based Text-to-Image Generation with Human Priors

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:52.182184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:52.182184Z digest=sha256:e3718f21542577e0b66356dcead71205cf379d2d51c35fea6916fca1f141c96d

Observation 2412b760-f18d-47cd-b1b8-09ebfe517144 · inbound

TextPixs: Glyph-Conditioned Diffusion with Character-Aware Attention and OCR-Guided Supervision cites this paper.

TextPixs: Glyph-Conditioned Diffusion with Character-Aware Attention and OCR-Guided Supervision Make-A-Scene: Scene-Based Text-to-Image Generation with Human Priors

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:49.855080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:18:49.855080Z digest=sha256:9e3e164c251dfad83442762d455b220267d3bad41b13113285a6bc24a26fe490

Observation 6885816a-f5e8-4510-bb7f-51db2d534a36 · inbound

Toward Native Multimodal Modeling: A Roadmap cites this paper.

Toward Native Multimodal Modeling: A Roadmap Make-A-Scene: Scene-Based Text-to-Image Generation with Human Priors

Reference 181

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T23:04:01.964697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T22:58:38.610609Z digest=sha256:9202cb818601583b908f798fb0b51e45174c8e39e57d2114733506d30086d413