Pith. sign in

Paper Citation Record · LEDGER

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions

As of 15 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 0 inbound Pith citation observations for arXiv:2506.16679.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.16679 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:42:44.693479Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

45 of 45 outbound references displayed

  • verified exact0
  • verified fuzzy19
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b29b55f1-1eae-4431-8287-c32b58950ab5 · outbound

This paper cites eDiff-I: Text-to-Image Diffusion Models with an Ensemble of Expert Denoisers.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions eDiff-I: Text-to-Image Diffusion Models with an Ensemble of Expert Denoisers

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:39.205574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:39.205574Z digest=sha256:4c95e870e642f584b54094ef03d50648babac6e21ed38fc7285626d13480e61c

Observation b7ccddc5-04e7-49b8-b370-194bce0d1623 · outbound

This paper cites Improving image generation with better captions, 2023.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions Improving image generation with better captions, 2023

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:39.318251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:39.318251Z digest=sha256:f1e9bcba854c988d76523cf5daf1d3c6fa4459b0a50178ad09c338892f6dca62

Observation e9c07c11-e186-4cf8-810b-83d1fd08c5fa · outbound

This paper cites Easily Ac- cessible Text-to-Image Generation Amplifies Demographic Stereotypes at Large Scale.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions Easily Ac- cessible Text-to-Image Generation Amplifies Demographic Stereotypes at Large Scale

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:39.446715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:39.446715Z digest=sha256:06553af1187db1262f619f15db2cc791f85036dde282eabfe2313c04932dfb9a

Observation 558f82cb-ca6b-4513-bf14-22cd904bf406 · outbound

This paper cites Se- mantics derived automatically from language corpora con- tain human-like biases.Science, 356(6334):183–186, 2017.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions Se- mantics derived automatically from language corpora con- tain human-like biases.Science, 356(6334):183–186, 2017

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:39.584832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:39.584832Z digest=sha256:c9aad22d5dd514a24e5c07aa67bdd58c5da85977f6364005600567314fcb90eb

Observation 2155edb3-d6f1-4785-8840-a235f3b9fe30 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:39.794485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:39.794485Z digest=sha256:6ccaab554c6bbd1ff1de3e04fd38da7b55a0990037c11f2c5de6fe3b77a2371a

Observation 6624f839-a4f3-4b3a-8d99-8774cb2628d5 · outbound

This paper cites Lmdeploy: A toolkit for com- pressing, deploying, and serving llm.https://github.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions Lmdeploy: A toolkit for com- pressing, deploying, and serving llm.https://github

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:39.904443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:39.904443Z digest=sha256:17c288f04456acdf0a8c912b94ca04977f6b9ff95334a983a27a0f74475ab1a8

Observation 6a27f1dc-6ac3-4360-9fe3-16dc6ee37eab · outbound

This paper cites DeepFloyd-IF-I-XL-v1.0: DeepFloyd’s Image Generation Model, 2023.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions DeepFloyd-IF-I-XL-v1.0: DeepFloyd’s Image Generation Model, 2023

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:40.102939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:40.102939Z digest=sha256:fa9f908466a0b19de63623cd0f5117b6b29e19f6453aa101c3e18efe3c02485c

Observation d0220b11-71d3-4c05-8397-1d2acf8f482f · outbound

This paper cites Scaling rectified flow trans- formers for high-resolution image synthesis.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions Scaling rectified flow trans- formers for high-resolution image synthesis

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:40.255286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:40.255286Z digest=sha256:b46e3b5c84b2805b5b7ca701335f84da0b7f6bb81c2b3851c85ef4412544a898

Observation 347fc60b-73ce-4c28-b156-dd2933804507 · outbound

This paper cites Auditing and instructing text-to-image gener- ation models on fairness.AI and Ethics, 2024.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions Auditing and instructing text-to-image gener- ation models on fairness.AI and Ethics, 2024

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:40.405376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:40.405376Z digest=sha256:3b5207e0f434d86b17045cfd91d20b558544b07d81cf8c650a7cf27340fe0475

Observation a1c12695-5946-44e0-8b8d-c2ddea93614f · outbound

This paper cites Multilingual text-to-image generation magnifies gender stereotypes and prompt engineering may not help you, 2024.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions Multilingual text-to-image generation magnifies gender stereotypes and prompt engineering may not help you, 2024

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:40.495626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:40.495626Z digest=sha256:539a3831464b734da9136fa70fc1438a831381bad96949d7e4bb52ffde92d631

Observation 22dec8b3-0204-4ef4-89e8-31bd5a06cafd · outbound

This paper cites Datacomp: In search of the next generation of multimodal datasets.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions Datacomp: In search of the next generation of multimodal datasets

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:40.624921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:40.624921Z digest=sha256:4cf7f26dc71d79be0a0774a6498f46e055a8ac872055c063221ea1fff9c39f92

Observation 61b54fcd-4a5f-434d-944c-ae818c146561 · outbound

This paper cites CommonCanvas: An Open Diffusion Model Trained with Creative-Commons Images.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions CommonCanvas: An Open Diffusion Model Trained with Creative-Commons Images

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:40.734956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:40.734956Z digest=sha256:6eabaf2a722ea5a72d5a106338626408729c84169a5d603ee460707b60f459ab

Observation 54200460-ec09-4521-bc29-7146b74c3f7b · outbound

This paper cites Hal- lusionbench: An advanced diagnostic suite for entangled language hallucination and visual illusion in large vision- language models.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions Hal- lusionbench: An advanced diagnostic suite for entangled language hallucination and visual illusion in large vision- language models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:40.846001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:40.846001Z digest=sha256:d5803e53471474df8683042175313fa2214e27e1247733ec9f05091598500d8f

Observation 4b7918e8-bed1-449e-b904-6cf61361700b · outbound

This paper cites Benchmarking of deep architectures for segmentation of medical images.Trans.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions Benchmarking of deep architectures for segmentation of medical images.Trans

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:40.966618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:40.966618Z digest=sha256:1c0aa869c3163c8b000db5bf698a80986e4b6e28327369bdf6a192bbd1418320

Observation fc7b575b-4bc4-4a28-8285-15bc81136e46 · outbound

This paper cites CLIPScore: a reference-free evaluation met- ric for image captioning.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions CLIPScore: a reference-free evaluation met- ric for image captioning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:41.194882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:41.194882Z digest=sha256:d951eb18756fa39b103635045dafc7b3f5a8abf0166dda01887dbfde161bcc69

Observation c571d370-70d8-47c8-b310-bf761d777dbf · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilib- rium.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions Gans trained by a two time-scale update rule converge to a local nash equilib- rium

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:41.290350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:41.290350Z digest=sha256:68482588ada7aa45b1ebb26dc13280f8a47c898a92328934cad113992be31c0b

Observation b095a204-ce0b-497f-80f7-b008bf11c865 · outbound

This paper cites Classifier-Free Diffusion Guidance.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions Classifier-Free Diffusion Guidance

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:41.417383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:41.417383Z digest=sha256:64e27d6b0aad78ac1f3668f97d3e1731a462397fc251eca02712f9ddb0483dc3

Observation 81b22947-d3cf-4904-9aca-c6978eecb2f8 · outbound

This paper cites Fairface: Face at- tribute dataset for balanced race, gender, and age for bias measurement and mitigation.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions Fairface: Face at- tribute dataset for balanced race, gender, and age for bias measurement and mitigation

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:51.905389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T23:42:41.602235Z digest=sha256:c30e80195ebd40602aa22956a2137e83e25b29b8bcd8fc83a5baf73799389f96

Observation 58cfc80b-e862-46cc-a7cb-63029ed7ea07 · outbound

This paper cites Transformers are minimax optimal nonparametric in-context learners.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions Transformers are minimax optimal nonparametric in-context learners

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:51.555707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T23:42:41.764753Z digest=sha256:31fa965c584d4f90881d248a57469f8814db2ebf7b0da17b65e691efe39f9375

Observation 0a02f6ce-6ae5-48ba-a14f-a37192251143 · outbound

This paper cites Pick-a-pic: An open dataset of user preferences for text-to-image generation.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions Pick-a-pic: An open dataset of user preferences for text-to-image generation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:51.178767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T23:42:41.934753Z digest=sha256:905b19e8c5da55b3f317b66c7a43250a44c611237b892f727123be5e3ba3e00b

Observation 08b35d18-f633-4871-9784-f9114dbe2b77 · outbound

This paper cites Genai-bench: Evaluat- ing and improving compositional text-to-visual generation,.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions Genai-bench: Evaluat- ing and improving compositional text-to-visual generation,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:50.974970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T23:42:42.064986Z digest=sha256:211e1973bdb8416a7a6b9d2d55dc9881ec53d6b8b8f416475772c7e85582c140

Observation da717749-599d-4916-9ecb-f93dbace5d54 · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:42.178165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:42.178165Z digest=sha256:86931802c271e423d9e2d0aff6311de346580ce6cc13274014a971f81e3e7903

Observation 6e38cb9e-9dfb-4235-9150-5a65b0a9f6fb · outbound

This paper cites Manmatha, Ashwin Swaminathan, Zhuowen Tu, Stefano Ermon, and Stefano Soatto.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions Manmatha, Ashwin Swaminathan, Zhuowen Tu, Stefano Ermon, and Stefano Soatto

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:50.455089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T23:42:42.375184Z digest=sha256:23653c440d43638b83e5ac954a23183b945e7d738f64d3fe50efd40f9f706151

Observation 1b97e62d-a13d-48ba-a05f-7af49f06fea8 · outbound

This paper cites an unresolved cited work.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:42:50.064635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T23:42:42.455111Z digest=sha256:c04de0e3df1fb60fdea1c107cdb8b7a33b9431164f55bcb6431052448b4a1897

Observation 908c64c8-763d-4626-b7b0-54265600e425 · outbound

This paper cites MVPTR: multi- level semantic alignment for vision-language pre-training via multi-stage learning.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions MVPTR: multi- level semantic alignment for vision-language pre-training via multi-stage learning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:49.814499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T23:42:42.544752Z digest=sha256:b90d0d54987225e125182e296c0c60413acb18f3dd62fc2896e925a2f30eb4d4

Observation d98e2388-d1ae-4244-977f-c4bbd7a40305 · outbound

This paper cites Playground v3: Im- proving text-to-image alignment with deep-fusion large lan- guage models, 2024.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions Playground v3: Im- proving text-to-image alignment with deep-fusion large lan- guage models, 2024

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:49.498510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T23:42:42.629743Z digest=sha256:9ba842fba9925cc511a61903918db44354c41fdf0f1d61d333c5e1662f988d8a

Observation 990f61d5-5c46-4105-ba7e-0679f2c7a42d · outbound

This paper cites Decoupled weight decay regularization.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions Decoupled weight decay regularization

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:49.261266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T23:42:42.715356Z digest=sha256:f9eb3e3dab511be38a370cd89780581f340c9427808463b55a0e422fae2db1c4

Observation 7b80c2bb-db38-4371-8bc0-d7b1e7857db7 · outbound

This paper cites DataDecide: How to Predict Best Pretraining Data with Small Experiments.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions DataDecide: How to Predict Best Pretraining Data with Small Experiments

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:42.799200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:42.799200Z digest=sha256:85703e28eabe391062ec3c61c28345a00973ad173273fe7150bdb864d1d11c68

Observation 5583e7a2-d656-4ab7-9bce-9fadce231b7f · outbound

This paper cites On aliased resizing and surprising subtleties in GAN evaluation.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions On aliased resizing and surprising subtleties in GAN evaluation

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:49.045221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T23:42:42.975038Z digest=sha256:eccb0cfd67d7fdc751b3600294314263c2ed90d1915b8d8f17315ea9c6f49149

Observation a0f40c02-ffab-472f-9a2c-777a7742a149 · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:43.044829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:43.044829Z digest=sha256:8f0f0c24aa4e669d0a715fb0032a4fc235fafa59ecb7aa0b2c44b5e581dfca92

Observation 99aa93f8-c68d-42aa-87c6-b454a4a6f723 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions High-resolution image synthesis with latent diffusion models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:48.820391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T23:42:43.133552Z digest=sha256:0474a7a28d69f820eff91ba1cbedf04734f3a4355f447a785501e099f92bb7d9

Observation 547bb8b8-6e95-42d6-a8f6-2dd81afe9eec · outbound

This paper cites Sara Mahdavi, Rapha Gontijo Lopes, Tim Salimans, Jonathan Ho, David J.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions Sara Mahdavi, Rapha Gontijo Lopes, Tim Salimans, Jonathan Ho, David J

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:48.515166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T23:42:43.199837Z digest=sha256:4b8ac83e71a19fd9410c06834be22e6aed678c0ac95cb98d70ca2a2209884f0e

Observation 4dd22a3e-4cda-4a02-9453-0aa14946d9a1 · outbound

This paper cites Fast High-Resolution Image Synthesis with Latent Adversarial Diffusion Distillation.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions Fast High-Resolution Image Synthesis with Latent Adversarial Diffusion Distillation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:43.346847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:43.346847Z digest=sha256:df94346a38dad53c378455cb9b74fa2dccb5e8aad8b0fcdca20b5c032eaf5d3d

Observation 901e554f-1db5-4b46-a477-311ee7b789d0 · outbound

This paper cites Safe latent diffusion: Mitigating inappro- priate degeneration in diffusion models.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions Safe latent diffusion: Mitigating inappro- priate degeneration in diffusion models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:48.294749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T23:42:43.454840Z digest=sha256:90b7ef1189578d04dd4fa2de08d3ee9ab2c4696255d7add904a83d0802ae7cb3

Observation fb5fa092-cb34-4479-b4f0-4d7e0a19c085 · outbound

This paper cites LAION- 400M: open dataset of clip-filtered 400 million image-text pairs.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions LAION- 400M: open dataset of clip-filtered 400 million image-text pairs

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:48.074836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T23:42:43.546387Z digest=sha256:40b7e3beabf1204eec1bb110f9d9dc45a9aa4fa5acb18fd9f7ce91265a93e949

Observation e9f375a3-88be-4dab-bf85-3746ad49ef47 · outbound

This paper cites Laion-5b: An open large-scale dataset for training next gen- eration image-text models.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions Laion-5b: An open large-scale dataset for training next gen- eration image-text models

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:47.803600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T23:42:43.727063Z digest=sha256:6d37093d0c57f5ed3a0336f11feb2c8a7721c5df9e288109c2d2d0b30c069b0f

Observation 99029762-b815-4202-903e-33770a0a3c76 · outbound

This paper cites The bias amplification paradox in text-to-image generation.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions The bias amplification paradox in text-to-image generation

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:47.536040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T23:42:43.854962Z digest=sha256:22b9146cba4ff426cda8d1aa21b738f8d06ac0b42133392de0422a1e31b49e22

Observation 716dcf68-43c4-4a15-97de-064c5d401727 · outbound

This paper cites Scaling autoregressive models for content-rich text-to-image generation.Transactions on Machine Learn- ing Research, 2022.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions Scaling autoregressive models for content-rich text-to-image generation.Transactions on Machine Learn- ing Research, 2022

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:47.312561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T23:42:43.941301Z digest=sha256:ae9d6d32dc02847b922a67335b7ad51d201426f8aec4bf2467421f0d90e44869

Observation 517e503f-be60-4ce8-938e-aaf294774c0c · outbound

This paper cites Sigmoid loss for language image pre-training.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions Sigmoid loss for language image pre-training

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:47.121589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T23:42:44.054838Z digest=sha256:fdabfc45a71ff1fbb123860199e242e3a762cf2fa0262fab3ac1e0847fc988e5

Observation c3150f97-83ce-48d9-837e-d5497f7ee3cc · outbound

This paper cites The unreasonable effectiveness of deep features as a perceptual metric.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions The unreasonable effectiveness of deep features as a perceptual metric

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:46.932291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T23:42:44.165791Z digest=sha256:80aa2b4eb1353a496638e2710a690e0ee988ef169ed92eb2d3e5102bed8ea255

Observation 0aef22c2-3279-4cba-a8ce-a278e4b1b84e · outbound

This paper cites SGLang: Efficient Execution of Structured Language Model Programs.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions SGLang: Efficient Execution of Structured Language Model Programs

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:44.284962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:44.284962Z digest=sha256:3e0577a93efa5f83c47128cdad07393891fbd9ebbbfd91abf1f8bee10b9359e4

Observation e9e566db-2c49-4d4b-bd6b-68a5572cd0ec · outbound

This paper cites an unresolved cited work.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:42:46.644753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T23:42:44.394752Z digest=sha256:7da2eac7834b89f7812149e6974cc3dca8f1154da81716624029b26230376bd1

Observation ee55de2c-6c89-4dc3-8284-a4fb32739b76 · outbound

This paper cites an unresolved cited work.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:42:46.457415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T23:42:44.514755Z digest=sha256:6de3d0886363d5558f0025167a5e51acd2d2bfc614406b35a934f144d9fb37a2

Observation 96d665c8-a8f2-4835-bba7-91f98df2851c · outbound

This paper cites an unresolved cited work.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:42:46.138421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T23:42:44.616284Z digest=sha256:41439592190ca556d43b0bb0a8601e324b992d90dd8cf53a73d75c759a68dd87

Observation 9fa2384b-0944-42cf-99c9-b5e9f07da3c9 · outbound

This paper cites psychologist.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions psychologist

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:45.864900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T23:42:44.693479Z digest=sha256:642cb1c3a799bdd53717c63173ddd560fb42ca3600420d5ddc72f78315835bba

Pith citing papers

No inbound Pith citation observations are available.