Pith. sign in

Paper Citation Record · LEDGER

Multimodal Benchmarking and Recommendation of Text-to-Image Generation Models

As of 17 August 2026, this Paper Citation Record lists 12 of 12 outbound references and 1 inbound Pith citation observation for arXiv:2505.04650.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.04650 v1

Coverage vector

measured 12 of 12 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:44:59.549750Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T13:05:48.204639Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

12 of 12 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation eff095ca-04c5-4fb1-b87d-9bf9f0f9b34e · outbound

This paper cites High- resolution image synthesis with latent diffusion models,.

Multimodal Benchmarking and Recommendation of Text-to-Image Generation Models High- resolution image synthesis with latent diffusion models,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T23:44:59.497854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:44:59.497854Z digest=sha256:0d0abbc03f3f7f88b753fcf82aec664987a0eed03741406852ac2ab1fbb12798

Observation cd30383c-a6e2-4730-9dac-c0fee0d51553 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Multimodal Benchmarking and Recommendation of Text-to-Image Generation Models SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T23:44:59.502937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:44:59.502937Z digest=sha256:ec51bad620e6d1b072a938df1c2b841426d9b317bf59c45ecf63aef874f3c87b

Observation b5f2185e-abfa-4acf-96d2-c3d55b3c93de · outbound

This paper cites CogView2: Faster and Better Text-to-Image Generation via Hierarchical Transformers.

Multimodal Benchmarking and Recommendation of Text-to-Image Generation Models CogView2: Faster and Better Text-to-Image Generation via Hierarchical Transformers

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T23:44:59.507648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:44:59.507648Z digest=sha256:1995f138b6f9d34da510a2a011a88caa977e68d91173e9f531c54c46a1e940b2

Observation 8549cf46-f4f9-4de6-9ac6-66d6f3ff25fc · outbound

This paper cites PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis.

Multimodal Benchmarking and Recommendation of Text-to-Image Generation Models PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T23:44:59.512313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:44:59.512313Z digest=sha256:c106569a906c037cfe1d5442802475dde1867e40553ffb878018aab9540aa5c9

Observation b1832c37-db70-49ef-b119-cdf91bf0b014 · outbound

This paper cites Group Diffusion Transformers are Unsupervised Multitask Learners.

Multimodal Benchmarking and Recommendation of Text-to-Image Generation Models Group Diffusion Transformers are Unsupervised Multitask Learners

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T23:44:59.517083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:44:59.517083Z digest=sha256:adb692255dec7e64113554be81694ce7bf823c878cbe888fa55105bc5c486e20

Observation 8dc0a084-2c5e-47ff-87ad-0422bf85f232 · outbound

This paper cites Scaling Rectified Flow Transformers for High-Resolution Image Synthesis.

Multimodal Benchmarking and Recommendation of Text-to-Image Generation Models Scaling Rectified Flow Transformers for High-Resolution Image Synthesis

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T23:44:59.522047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:44:59.522047Z digest=sha256:ab78c989b3a3e6f8d8bb373db5bb4cb0246d23891dcde8b09b9e22e14f0e96a0

Observation 1cfc0751-e660-48b2-bd59-4a75126eef82 · outbound

This paper cites Lumina-Image 2.0: A Unified and Efficient Image Generative Framework.

Multimodal Benchmarking and Recommendation of Text-to-Image Generation Models Lumina-Image 2.0: A Unified and Efficient Image Generative Framework

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T23:44:59.527474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:44:59.527474Z digest=sha256:3028c5caa32794a5ae9165433e4ac4ba3b0d6b42820090deb5876c073ba375d7

Observation 01d7214c-0757-4641-8d36-634174458edd · outbound

This paper cites Sana-sprint: One-step diffusion with continuous-time consistency distillation,.

Multimodal Benchmarking and Recommendation of Text-to-Image Generation Models Sana-sprint: One-step diffusion with continuous-time consistency distillation,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T23:44:59.532017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:44:59.532017Z digest=sha256:394545443db5ca3938565f1061606aaeff020bdcf75384209c661bb2ffdfa2a0

Observation 0e1c56b7-e3fa-4f45-9eaa-b003e38442b7 · outbound

This paper cites KOALA: Empirical Lessons Toward Memory-Efficient and Fast Diffusion Models for Text-to-Image Synthesis.

Multimodal Benchmarking and Recommendation of Text-to-Image Generation Models KOALA: Empirical Lessons Toward Memory-Efficient and Fast Diffusion Models for Text-to-Image Synthesis

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T23:44:59.536379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:44:59.536379Z digest=sha256:381fa1b7cd279d49b52ab53fd99a51c64b9b82bc3981b5b0a2409233ef2f65d5

Observation ed86a1c1-4b52-4eef-bd6d-103a9ab80c6e · outbound

This paper cites Text2human: Text-driven controllable human image generation,.

Multimodal Benchmarking and Recommendation of Text-to-Image Generation Models Text2human: Text-driven controllable human image generation,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:44:59.900920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T23:44:59.540939Z digest=sha256:845522fa888c5091abf6e5f9ccc52eb02f38033f1434d2840a87b3794d4d7ef2

Observation 834167f6-aaf1-4942-aaa3-9c1af783f156 · outbound

This paper cites Latent consistency models: Synthesizing high-resolution images with few-step inference,.

Multimodal Benchmarking and Recommendation of Text-to-Image Generation Models Latent consistency models: Synthesizing high-resolution images with few-step inference,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:44:59.883830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T23:44:59.545382Z digest=sha256:6c823af452a693a5badd51c5c8a25f62bea0519025c768600363b9889615658d

Observation 6a8e3309-a013-4be8-b534-a9d137003873 · outbound

This paper cites IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models.

Multimodal Benchmarking and Recommendation of Text-to-Image Generation Models IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T23:44:59.549750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:44:59.549750Z digest=sha256:9fd37762e214192fd80531b5b4e22db01448d54707775e6221f504901a1c57bc

Pith citing papers

Observation 29f3b2d8-7c37-4c42-af6c-99c131c4ff04 · inbound

A Comprehensive Dataset for Human vs. AI Generated Image Detection cites this paper.

A Comprehensive Dataset for Human vs. AI Generated Image Detection Multimodal Benchmarking and Recommendation of Text-to-Image Generation Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T13:05:48.204639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:05:48.204639Z digest=sha256:3095099bbefe0cb7877a4867799766566380a6b388444509bcd512d8d52ddc71