Pith. sign in

Paper Citation Record · LEDGER

Structured Captions Improve Prompt Adherence in Text-to-Image Models (Re-LAION-Caption 19M)

As of 18 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 1 inbound Pith citation observation for arXiv:2507.05300.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.05300 v1

Coverage vector

measured 16 of 16 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:50:21.910265Z

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T16:11:10.598413Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T09:11:02.341076Z

Reference resolution

16 of 16 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f885d8ca-7074-4cd9-913c-7f22f0cab520 · outbound

This paper cites Improving CLIP Training with Language Rewrites.

Structured Captions Improve Prompt Adherence in Text-to-Image Models (Re-LAION-Caption 19M) Improving CLIP Training with Language Rewrites

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:20.447559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:20.447559Z digest=sha256:c38bc8c96ea76f1c15e728f5f33b60125515811e40b622f55b4f7a404c8b1233

Observation 88973b45-3c6a-4501-b3b5-3226af531b7e · outbound

This paper cites Mistral 7B.

Structured Captions Improve Prompt Adherence in Text-to-Image Models (Re-LAION-Caption 19M) Mistral 7B

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:20.713255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:20.713255Z digest=sha256:4830859004f3412e81f42fe7cabc6f84d023780885764c432377d9d37c2da21e

Observation 94b48528-2d7f-4432-b992-52d9ef528a87 · outbound

This paper cites VeCLIP: Improving CLIP Training via Visual-enriched Captions.

Structured Captions Improve Prompt Adherence in Text-to-Image Models (Re-LAION-Caption 19M) VeCLIP: Improving CLIP Training via Visual-enriched Captions

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:20.826610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:20.826610Z digest=sha256:ac5f42d9dca4d5e0ba3c36606a3a8c3ed3a9706fab81856c36bddf61e9eb6ef6

Observation 1e30366c-0674-43ef-9b1e-88fd609466fe · outbound

This paper cites What If We Recaption Billions of Web Images with LLaMA-3?.

Structured Captions Improve Prompt Adherence in Text-to-Image Models (Re-LAION-Caption 19M) What If We Recaption Billions of Web Images with LLaMA-3?

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:21.065213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:21.065213Z digest=sha256:f76fe9c7f9631059cffaff8bc15c721a5152d44cd2a69479cc6c6e12ba464791

Observation 566015e5-ed23-4701-8800-34684a5e14bf · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

Structured Captions Improve Prompt Adherence in Text-to-Image Models (Re-LAION-Caption 19M) Learning Transferable Visual Models From Natural Language Supervision

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:21.319649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:21.319649Z digest=sha256:9cd41e5d13f5a99f67b83ef7fee18df7ac7bb7fbf9aaa631f5fbda8551d4668b

Observation 340f21b8-295a-4eb3-abc5-e0e6b53459fc · outbound

This paper cites Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer.

Structured Captions Improve Prompt Adherence in Text-to-Image Models (Re-LAION-Caption 19M) Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:21.432622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:21.432622Z digest=sha256:ffbbff88436a3064465cfcfc8205558aece38b885c30fe751a4e4a8750f7e4e5

Observation 3b8c973b-ba5c-4166-9c49-159b4e9bd0fa · outbound

This paper cites Zero-Shot Text-to-Image Generation.

Structured Captions Improve Prompt Adherence in Text-to-Image Models (Re-LAION-Caption 19M) Zero-Shot Text-to-Image Generation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:21.587329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:21.587329Z digest=sha256:fcd8e13d2cc81cebd4ef21b75673a60a4cf8e96c3d51a5a8b5a0e27632b2432e

Observation 684798eb-30d9-4e21-aa95-74aebaaf8576 · outbound

This paper cites High-Resolution Image Synthesis with Latent Diffusion Models.

Structured Captions Improve Prompt Adherence in Text-to-Image Models (Re-LAION-Caption 19M) High-Resolution Image Synthesis with Latent Diffusion Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:21.722260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:21.722260Z digest=sha256:e34bbf64e45c1b5eeb73d94848629e312983ef4d1deab14737daa3ca93014bd3

Observation 3f51bf79-9eea-43ce-8bfd-cd8cdc3a2d0a · outbound

This paper cites LAION-5B: An open large-scale dataset for training next generation image-text models.

Structured Captions Improve Prompt Adherence in Text-to-Image Models (Re-LAION-Caption 19M) LAION-5B: An open large-scale dataset for training next generation image-text models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:21.801084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:21.801084Z digest=sha256:2736c8c1d67c7df5e60af049bebf75f7b0087eba291c7d0aadea78dcba075df4

Observation af231c98-e5aa-41da-ad64-83e123ae1795 · outbound

This paper cites On the De-duplication of LAION-2B.

Structured Captions Improve Prompt Adherence in Text-to-Image Models (Re-LAION-Caption 19M) On the De-duplication of LAION-2B

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:21.910265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:21.910265Z digest=sha256:1b8a1e2164d3639f54566d58a7d1bfad0b85bd2856ed9818661300af6607ca3e

Observation fc42e81c-c4a3-41ff-8b16-3565f1968b45 · outbound

This paper cites Real-time Scene Text Detection with Differentiable Binarization.

Structured Captions Improve Prompt Adherence in Text-to-Image Models (Re-LAION-Caption 19M) Real-time Scene Text Detection with Differentiable Binarization

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:21.181128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:21.181128Z digest=sha256:a9a57cba45b7fa5be5c2ae8178789d5484fa6e3592b8423abfe6cffc629dba98

Observation b6a03b2e-96aa-4486-9572-0259eed5175a · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Structured Captions Improve Prompt Adherence in Text-to-Image Models (Re-LAION-Caption 19M) LoRA: Low-Rank Adaptation of Large Language Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:20.584536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:20.584536Z digest=sha256:d34c93acda38976f3b1fcca7a517fb6185bc7773075fa5fddb1c6c01ec11c696

Observation 3a50c2c2-4980-4c10-90f7-7c0efca2df78 · outbound

This paper cites Scaling Instruction-Finetuned Language Models.

Structured Captions Improve Prompt Adherence in Text-to-Image Models (Re-LAION-Caption 19M) Scaling Instruction-Finetuned Language Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:20.256972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:20.256972Z digest=sha256:aff264c9a06bcdc68c49439e1eb7c3f2ff793022bfe283175860048d3b24b0e1

Observation 09560a62-a540-4de7-8c60-a833ec3d9a4d · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

Structured Captions Improve Prompt Adherence in Text-to-Image Models (Re-LAION-Caption 19M) InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:20.352206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:20.352206Z digest=sha256:7067cc26b5652dab62cb8efc79501a0f0b747a4e4c6afc0dc98064694de6d5b0

Observation 7a5cdab3-5c44-4c2d-857f-bdbbecbb83dc · outbound

This paper cites PixArt-\Sigma: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation.

Structured Captions Improve Prompt Adherence in Text-to-Image Models (Re-LAION-Caption 19M) PixArt-\Sigma: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:20.126209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:20.126209Z digest=sha256:1a7831000f3c611d8ad6e4436ebff14860ea5e2a749eda369f72633107046997

Observation b5091e9e-6cf9-443f-b523-f03b45dbf721 · outbound

This paper cites OpenHumanVid: A Large-Scale High-Quality Dataset for Enhancing Human-Centric Video Generation.

Structured Captions Improve Prompt Adherence in Text-to-Image Models (Re-LAION-Caption 19M) OpenHumanVid: A Large-Scale High-Quality Dataset for Enhancing Human-Centric Video Generation

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:20.935002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:20.935002Z digest=sha256:09ccb3bf8c4746e737c3b8aad68f9d3d3ec0f449e86145c5c65cbf12c87f6576

Pith citing papers

Observation d7a5f4ea-e055-4ebe-8ff9-309d4390ee0a · inbound

ReflectCAP: Detailed Image Captioning with Reflective Memory cites this paper.

ReflectCAP: Detailed Image Captioning with Reflective Memory Structured Captions Improve Prompt Adherence in Text-to-Image Models (Re-LAION-Caption 19M)

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:11:02.344304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T16:11:10.598413Z digest=sha256:02c1147a2e6566043d84e24128bfe5165ccef08f71e0c937539da594fbe5e492