Pith. sign in

Paper Citation Record · LEDGER

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models

As of 17 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 0 inbound Pith citation observations for arXiv:2501.00917.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.00917 v1

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:42:47.923687Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

28 of 28 outbound references displayed

  • verified exact2
  • verified fuzzy9
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9d5dcc33-ebaa-4faa-a274-ae49644a8e58 · outbound

This paper cites Learning transferable visual models from na tural language supervision,.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Learning transferable visual models from na tural language supervision,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T22:42:47.800949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:42:47.800949Z digest=sha256:5f5977ab66062c0b3b28f336f18f4ad82138013159cb06df5d000080784a318e

Observation 3e59bc6d-8dad-4484-b063-ceca84274ef6 · outbound

This paper cites Fla mingo: a visual language model for few-shot learning,.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Fla mingo: a visual language model for few-shot learning,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:42:48.429136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T22:42:47.806179Z digest=sha256:f72418532f9736aa1b91f9c52b94e956f97d2a50c57e6bcc72cfc21b35f07afa

Observation 6f6057e7-6a52-4df8-ae77-6a10ae1c2804 · outbound

This paper cites Rethinking Visual Dependency in Long-Context Reasoning for Large Vision-Language Models.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Rethinking Visual Dependency in Long-Context Reasoning for Large Vision-Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T22:42:47.810775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:42:47.810775Z digest=sha256:671d5a4660529c4247618f5d2358610f05659a61c10fad57a33e1515fa6e46e8

Observation c70eaa7e-504a-4f3f-8c2a-a913f2faba0b · outbound

This paper cites ARTIST: Improving the Generation of Text-rich Images with Disentangled Diffusion Models and Large Language Models.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models ARTIST: Improving the Generation of Text-rich Images with Disentangled Diffusion Models and Large Language Models

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-10T22:42:48.232622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T22:42:47.816050Z digest=sha256:92dd64176a72a5c3d75ea06a41f2c4c234f3cf1b588db2cb1f18bf75cfdcb7b4

Observation 839626a2-47b2-4584-8e5f-aee86d62c9c1 · outbound

This paper cites Textd iffuser: Diffusion models as text painters,.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Textd iffuser: Diffusion models as text painters,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:42:48.413065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T22:42:47.820974Z digest=sha256:bb7bae92c6da8ed790892a2467e105afcb770c4d703c3acf1cb6cac9d2de82f2

Observation 2842a8d5-f723-4814-a821-3dd074fbe122 · outbound

This paper cites Muse: Text-To-Image Generation via Masked Generative Transformers.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T22:42:47.825552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:42:47.825552Z digest=sha256:5cd1e57104b26cfa63a4cce69a87988fa350cd548aae3f61481d8c60ef6168a9

Observation 7d478ed2-71ab-4f85-8e57-07ab9d33de9c · outbound

This paper cites STAR: Scale-wise Text-conditioned AutoRegressive image generation.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models STAR: Scale-wise Text-conditioned AutoRegressive image generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T22:42:47.830666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:42:47.830666Z digest=sha256:ab484a75f8cf5ad23ace293a423b56951cead2c90cf8ffa71a5c11203e0665ca

Observation 3aad78ac-da31-432d-9ade-f8e89ecaf145 · outbound

This paper cites Region-Aware Text-to-Image Generation via Hard Binding and Soft Refinement.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Region-Aware Text-to-Image Generation via Hard Binding and Soft Refinement

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T22:42:47.835262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:42:47.835262Z digest=sha256:2c568a3e050faf0bb1d5a753791c69521826f6d9e6096f90fbfae56b048a9dda

Observation f9de7c66-9b72-4071-95e3-3028d3157f0b · outbound

This paper cites An analysis of the ingredients for learning interpretable symbolic regression models with human-in-the-loop and genetic prog ramming,.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models An analysis of the ingredients for learning interpretable symbolic regression models with human-in-the-loop and genetic prog ramming,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:42:48.397720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T22:42:47.839951Z digest=sha256:9fcd09967f13dd95088702371c1dd7c34264fb6ad7d08a88285a253e8d66e348

Observation d0ff0900-95db-490d-bbb6-8994ba843467 · outbound

This paper cites Improving cross-modal alignment f or text- guided image inpainting,.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Improving cross-modal alignment f or text- guided image inpainting,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T22:42:47.843949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:42:47.843949Z digest=sha256:746dace41fd2d1f730cf9057c323192402ed3f957b8730630da2590066f285c2

Observation ee5d1221-7ff3-4ab9-939b-13f0fb837801 · outbound

This paper cites Zero-shot text-to-image generation,.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Zero-shot text-to-image generation,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:42:48.373659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T22:42:47.847927Z digest=sha256:b683c18ee47c4f8be1ec8b5130f7c67259ce7e75a330e692823d74c443fb64f1

Observation cd633819-6bf3-46c3-a9c4-4a84c31ee2ed · outbound

This paper cites Towards language-driven video inpainting via multimoda l large language models,.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Towards language-driven video inpainting via multimoda l large language models,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:42:48.359278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T22:42:47.852061Z digest=sha256:48abab7e9577716fcf615ca84d63a1fbea7f9a6249e836437685e394b0f80e8e

Observation 7df5d13b-9bd0-4922-82e1-58ada63d4361 · outbound

This paper cites Prompt Expansion for Adaptive Text-to-Image Generation.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Prompt Expansion for Adaptive Text-to-Image Generation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T22:42:47.856445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:42:47.856445Z digest=sha256:476d1776293e8e21214a1366dd37a9d786f6e01992ca5aeb7377eee95fee7e96

Observation 5bd31eed-2ce6-4aca-b6db-8362b87569cd · outbound

This paper cites Training-free consistent text-to-image gener ation,.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Training-free consistent text-to-image gener ation,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:42:48.344419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T22:42:47.861291Z digest=sha256:5fb95f058e34e7840da341801b9fd38de9ed43fec312836c57679fe1c6142eea

Observation b92ebe12-0373-493b-baff-f96f1d5f8d6a · outbound

This paper cites Style-aware contrastive learning for multi-style image captioning,.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Style-aware contrastive learning for multi-style image captioning,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T22:42:47.865685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:42:47.865685Z digest=sha256:ea82fe0768a8ef0efe4c98e66dae67ff305417397850ba273b58cbe1a57f0af4

Observation eedea3d6-b472-4ad6-86aa-c758ef8bb01f · outbound

This paper cites Multimodal event transformer for image-guided st ory ending generation,.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Multimodal event transformer for image-guided st ory ending generation,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:42:48.320620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T22:42:47.869872Z digest=sha256:dffcc3c7c7e347283004a26c7bbbdd6028b34772601041966d226262e67b7910

Observation 4f700a97-5b4f-44d2-919d-8013fd81e193 · outbound

This paper cites Triple sequence generati ve adversarial nets for unsupervised image captioning,.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Triple sequence generati ve adversarial nets for unsupervised image captioning,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T22:42:47.874406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:42:47.874406Z digest=sha256:43f20af14d2f0f1b374464e35377d4c59d1be4c67781a585d4c30b04c3baeed2

Observation ccda3c91-5b5b-4ba9-b9f4-0efd3f4651b2 · outbound

This paper cites Sketch storytelling,.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Sketch storytelling,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T22:42:47.878734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:42:47.878734Z digest=sha256:b96176f32ef856662304ccfc4d6b314340f2e449eeddc3c7e7a6788cb7ade191

Observation 70acb0cc-a645-4cc5-9b12-9da7a0068e37 · outbound

This paper cites Generalist Multimodal AI: A Review of Architectures, Challenges and Opportunities.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Generalist Multimodal AI: A Review of Architectures, Challenges and Opportunities

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-10T22:42:48.155107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T22:42:47.882788Z digest=sha256:29541caf5e25fa619169880a11c4283909c74f6f2a43c8de9481f0db631d64cb

Observation 84865179-7f60-4189-a0b0-b616a27b54b6 · outbound

This paper cites Advancing Fine-Grained Visual Understanding with Multi-Scale Alignment in Multi-Modal Models.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Advancing Fine-Grained Visual Understanding with Multi-Scale Alignment in Multi-Modal Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T22:42:47.887679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:42:47.887679Z digest=sha256:dd2228a3880f00109c45aa00d56a6891bc983c599921ff5f5d83c76634174836

Observation 4647b3ae-dcd7-4ef3-89b5-b507a7c7b778 · outbound

This paper cites Visual in-context l earning for large vision-language models,.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Visual in-context l earning for large vision-language models,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T22:42:47.892120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:42:47.892120Z digest=sha256:fc52b04c1d6494a239496433bce69fd5821347097ecebe1d46354d074c906f16

Observation 050b7709-0670-4258-b108-1e6a93eb91af · outbound

This paper cites InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T22:42:47.896528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:42:47.896528Z digest=sha256:7ea9e81c489f86196b78fcd038f2b0154eb28b35ba6a671a5b7ea9fc58e58000

Observation 31b6ca13-b271-4ac6-8df8-52467eeb0a85 · outbound

This paper cites MoE-LLaVA: Mixture of Experts for Large Vision-Language Models.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models MoE-LLaVA: Mixture of Experts for Large Vision-Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T22:42:47.901252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:42:47.901252Z digest=sha256:a4401bc421a8267f431366c8a56e5f0f31ea142047add8aa24857b04c7ed639a

Observation 6672227c-d77f-4391-9712-fa755c0eb995 · outbound

This paper cites Understanding Reinforcement Learning-Based Fine-Tuning of Diffusion Models: A Tutorial and Review.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Understanding Reinforcement Learning-Based Fine-Tuning of Diffusion Models: A Tutorial and Review

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T22:42:47.905883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:42:47.905883Z digest=sha256:a6be1b10e3be4335385109aee875ed495fb3cda4ad8c8b4d235c6545c88e8d83

Observation 14e2375b-3758-4254-9be4-f3b03b6a1106 · outbound

This paper cites A Survey on Benchmarks of Multimodal Large Language Models.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models A Survey on Benchmarks of Multimodal Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T22:42:47.910468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:42:47.910468Z digest=sha256:cbe0b7aaf539082dfea62f3fe95a3c89bb9ac9d5d76efbc1439f1fa38a44d03a

Observation 4261da68-09ad-43cd-84e4-f7848eafff97 · outbound

This paper cites Ex ploring the frontier of vision-language models: A survey of current met hodologies and future directions,.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Ex ploring the frontier of vision-language models: A survey of current met hodologies and future directions,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T22:42:47.914817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:42:47.914817Z digest=sha256:1586009282ef49aa55b25f0ef756282d05435930e87a0d74f8ae42ec321db18f

Observation 4bc47a9f-7762-4e47-aefd-26d990754358 · outbound

This paper cites Visionllm: Large language model is also an open-ended decoder for vision-centric tasks,.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Visionllm: Large language model is also an open-ended decoder for vision-centric tasks,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:42:48.276806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T22:42:47.918957Z digest=sha256:bab440cdb4d26b047eba53425aa46ba16116b9abac2cec9b4a73a39e1d5cdf67

Observation b2e9a579-3d1b-43c4-9baf-103d1b168642 · outbound

This paper cites Less is more: Vision representation compression for efficient video gene ration with large language models,.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Less is more: Vision representation compression for efficient video gene ration with large language models,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:42:48.262399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T22:42:47.923687Z digest=sha256:cceb8eb75972be41019ba423880921523503df20e88dbb30da6ec64d0fd4e84f

Pith citing papers

No inbound Pith citation observations are available.