Pith. sign in

Paper Citation Record · LEDGER

Structured Captions Improve Prompt Adherence in Text-to-Image Models (Re-LAION-Caption 19M)

As of 7 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 1 inbound Pith citation observation for arXiv:2507.05300.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.05300 v1

Coverage vector

measured 16 of 16 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:50:21.910265Z

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T16:11:10.598413Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T09:11:02.341076Z

Reference resolution

16 of 16 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f885d8ca-7074-4cd9-913c-7f22f0cab520 · outbound

This paper cites Improving CLIP Training with Language Rewrites.

Structured Captions Improve Prompt Adherence in Text-to-Image Models (Re-LAION-Caption 19M) Improving CLIP Training with Language Rewrites

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:20.447559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:20.447559Z digest=sha256:7cd433c2b9f25f5d8611ed504193ea020616f4f7a9060f8d578b44999921f97b

Observation 88973b45-3c6a-4501-b3b5-3226af531b7e · outbound

This paper cites Mistral 7B.

Structured Captions Improve Prompt Adherence in Text-to-Image Models (Re-LAION-Caption 19M) Mistral 7B

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:20.713255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:20.713255Z digest=sha256:b54868f502366ed68c363bdfb121db052ef6a8f7295a6e7088c9ce24419db0cc

Observation 94b48528-2d7f-4432-b992-52d9ef528a87 · outbound

This paper cites VeCLIP: Improving CLIP Training via Visual-enriched Captions.

Structured Captions Improve Prompt Adherence in Text-to-Image Models (Re-LAION-Caption 19M) VeCLIP: Improving CLIP Training via Visual-enriched Captions

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:20.826610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:20.826610Z digest=sha256:435f2ee3bd12ecfceff9f96031aeefd811939a56f4d027e6356d8367ab994033

Observation 1e30366c-0674-43ef-9b1e-88fd609466fe · outbound

This paper cites What If We Recaption Billions of Web Images with LLaMA-3?.

Structured Captions Improve Prompt Adherence in Text-to-Image Models (Re-LAION-Caption 19M) What If We Recaption Billions of Web Images with LLaMA-3?

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:21.065213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:21.065213Z digest=sha256:1d0409de579d0aa09c8b5a30bff624af9929d5e5e317f51dc6e9b5e2277c2973

Observation 566015e5-ed23-4701-8800-34684a5e14bf · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

Structured Captions Improve Prompt Adherence in Text-to-Image Models (Re-LAION-Caption 19M) Learning Transferable Visual Models From Natural Language Supervision

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:21.319649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:21.319649Z digest=sha256:1511831b852b15428b31ab79dc98293126be091b983bb7f1f30fb62a176df207

Observation 340f21b8-295a-4eb3-abc5-e0e6b53459fc · outbound

This paper cites Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer.

Structured Captions Improve Prompt Adherence in Text-to-Image Models (Re-LAION-Caption 19M) Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:21.432622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:21.432622Z digest=sha256:a3d017e6272ab92878f1e9f1d42d52e6f4e84f6f7d6f4117ddbcd6c4f18eb03c

Observation 3b8c973b-ba5c-4166-9c49-159b4e9bd0fa · outbound

This paper cites Zero-Shot Text-to-Image Generation.

Structured Captions Improve Prompt Adherence in Text-to-Image Models (Re-LAION-Caption 19M) Zero-Shot Text-to-Image Generation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:21.587329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:21.587329Z digest=sha256:c322ecd909f3324fe0612c65f0f6cc5eb8796c672d5cee8ea327adee25f5bdc4

Observation 684798eb-30d9-4e21-aa95-74aebaaf8576 · outbound

This paper cites High-Resolution Image Synthesis with Latent Diffusion Models.

Structured Captions Improve Prompt Adherence in Text-to-Image Models (Re-LAION-Caption 19M) High-Resolution Image Synthesis with Latent Diffusion Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:21.722260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:21.722260Z digest=sha256:e29420c921904327501dd1ad4b613fdf80df4b783997aec69e0443a10ac063f9

Observation 3f51bf79-9eea-43ce-8bfd-cd8cdc3a2d0a · outbound

This paper cites LAION-5B: An open large-scale dataset for training next generation image-text models.

Structured Captions Improve Prompt Adherence in Text-to-Image Models (Re-LAION-Caption 19M) LAION-5B: An open large-scale dataset for training next generation image-text models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:21.801084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:21.801084Z digest=sha256:4ef113e1b73e1a2678041dc266f491c61a210cb304143fd12bb507b2542740d8

Observation af231c98-e5aa-41da-ad64-83e123ae1795 · outbound

This paper cites On the De-duplication of LAION-2B.

Structured Captions Improve Prompt Adherence in Text-to-Image Models (Re-LAION-Caption 19M) On the De-duplication of LAION-2B

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:21.910265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:21.910265Z digest=sha256:56e173bc4ff4b7b5e53f6ecfa4863201ed94a57b3f00e1d2b82ebc481579ebd3

Observation fc42e81c-c4a3-41ff-8b16-3565f1968b45 · outbound

This paper cites Real-time Scene Text Detection with Differentiable Binarization.

Structured Captions Improve Prompt Adherence in Text-to-Image Models (Re-LAION-Caption 19M) Real-time Scene Text Detection with Differentiable Binarization

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:21.181128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:21.181128Z digest=sha256:e17f9a40774114ac15ca053c7d288fb07a465c3f233b74023d2ae16d6db120d9

Observation b6a03b2e-96aa-4486-9572-0259eed5175a · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Structured Captions Improve Prompt Adherence in Text-to-Image Models (Re-LAION-Caption 19M) LoRA: Low-Rank Adaptation of Large Language Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:20.584536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:20.584536Z digest=sha256:a4a8700908f3a49ca4c436f66797a3f960417f9ea6194c78228ca8a913f739f2

Observation 3a50c2c2-4980-4c10-90f7-7c0efca2df78 · outbound

This paper cites Scaling Instruction-Finetuned Language Models.

Structured Captions Improve Prompt Adherence in Text-to-Image Models (Re-LAION-Caption 19M) Scaling Instruction-Finetuned Language Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:20.256972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:20.256972Z digest=sha256:d73f1052347044a3ef63466e4ebb515f310d1b5a373187bad1689d199b736186

Observation 09560a62-a540-4de7-8c60-a833ec3d9a4d · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

Structured Captions Improve Prompt Adherence in Text-to-Image Models (Re-LAION-Caption 19M) InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:20.352206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:20.352206Z digest=sha256:9fffee355c2ffaf31e3ae0aa5a5406b5ddd007a586a7bcc3100e7f01d8b6eb12

Observation 7a5cdab3-5c44-4c2d-857f-bdbbecbb83dc · outbound

This paper cites PixArt-\Sigma: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation.

Structured Captions Improve Prompt Adherence in Text-to-Image Models (Re-LAION-Caption 19M) PixArt-\Sigma: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:20.126209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:20.126209Z digest=sha256:fe2df56c7f3c764c067e3dd333f3358d6decc3d2d646384f3ea5ed337fafe134

Observation b5091e9e-6cf9-443f-b523-f03b45dbf721 · outbound

This paper cites OpenHumanVid: A Large-Scale High-Quality Dataset for Enhancing Human-Centric Video Generation.

Structured Captions Improve Prompt Adherence in Text-to-Image Models (Re-LAION-Caption 19M) OpenHumanVid: A Large-Scale High-Quality Dataset for Enhancing Human-Centric Video Generation

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:20.935002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:20.935002Z digest=sha256:2f28f13422e583648d8177d0c5beab093745c1e9d53eabc7f4b1f637a5e0770a

Pith citing papers

Observation d7a5f4ea-e055-4ebe-8ff9-309d4390ee0a · inbound

ReflectCAP: Detailed Image Captioning with Reflective Memory cites this paper.

ReflectCAP: Detailed Image Captioning with Reflective Memory Structured Captions Improve Prompt Adherence in Text-to-Image Models (Re-LAION-Caption 19M)

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:11:02.344304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T16:11:10.598413Z digest=sha256:5ce35447cbe0ac6589141e5a19bbf692f5620f7299e81c95b001a90645d8307d