Pith. sign in

Paper Citation Record · LEDGER

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models

As of 17 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 0 inbound Pith citation observations for arXiv:2501.00917.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.00917 v1

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:42:47.923687Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

28 of 28 outbound references displayed

  • verified exact2
  • verified fuzzy9
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9d5dcc33-ebaa-4faa-a274-ae49644a8e58 · outbound

This paper cites Learning transferable visual models from na tural language supervision,.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Learning transferable visual models from na tural language supervision,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T22:42:47.800949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:42:47.800949Z digest=sha256:5f5977ab66062c0b3b28f336f18f4ad82138013159cb06df5d000080784a318e

Observation 3e59bc6d-8dad-4484-b063-ceca84274ef6 · outbound

This paper cites Fla mingo: a visual language model for few-shot learning,.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Fla mingo: a visual language model for few-shot learning,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:42:48.429136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:42:47.806179Z digest=sha256:346913a01e648b8fc66b29026d6c3255d67183475c9f3e397a0d851dadd04343

Observation 6f6057e7-6a52-4df8-ae77-6a10ae1c2804 · outbound

This paper cites Rethinking Visual Dependency in Long-Context Reasoning for Large Vision-Language Models.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Rethinking Visual Dependency in Long-Context Reasoning for Large Vision-Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T22:42:47.810775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:42:47.810775Z digest=sha256:671d5a4660529c4247618f5d2358610f05659a61c10fad57a33e1515fa6e46e8

Observation c70eaa7e-504a-4f3f-8c2a-a913f2faba0b · outbound

This paper cites ARTIST: Improving the Generation of Text-rich Images with Disentangled Diffusion Models and Large Language Models.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models ARTIST: Improving the Generation of Text-rich Images with Disentangled Diffusion Models and Large Language Models

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-10T22:42:48.232622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:42:47.816050Z digest=sha256:aeca3d532b327b217eb2367701619ae36a553b50f11ff8cb1fd53079a8bb8d9d

Observation 839626a2-47b2-4584-8e5f-aee86d62c9c1 · outbound

This paper cites Textd iffuser: Diffusion models as text painters,.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Textd iffuser: Diffusion models as text painters,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:42:48.413065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:42:47.820974Z digest=sha256:7de6f982de4efb191804ce928b2d9131f70c79f8ea225f6755a9862b72dd313a

Observation 2842a8d5-f723-4814-a821-3dd074fbe122 · outbound

This paper cites Muse: Text-To-Image Generation via Masked Generative Transformers.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T22:42:47.825552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:42:47.825552Z digest=sha256:5cd1e57104b26cfa63a4cce69a87988fa350cd548aae3f61481d8c60ef6168a9

Observation 7d478ed2-71ab-4f85-8e57-07ab9d33de9c · outbound

This paper cites STAR: Scale-wise Text-conditioned AutoRegressive image generation.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models STAR: Scale-wise Text-conditioned AutoRegressive image generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T22:42:47.830666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:42:47.830666Z digest=sha256:ab484a75f8cf5ad23ace293a423b56951cead2c90cf8ffa71a5c11203e0665ca

Observation 3aad78ac-da31-432d-9ade-f8e89ecaf145 · outbound

This paper cites Region-Aware Text-to-Image Generation via Hard Binding and Soft Refinement.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Region-Aware Text-to-Image Generation via Hard Binding and Soft Refinement

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T22:42:47.835262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:42:47.835262Z digest=sha256:2c568a3e050faf0bb1d5a753791c69521826f6d9e6096f90fbfae56b048a9dda

Observation f9de7c66-9b72-4071-95e3-3028d3157f0b · outbound

This paper cites An analysis of the ingredients for learning interpretable symbolic regression models with human-in-the-loop and genetic prog ramming,.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models An analysis of the ingredients for learning interpretable symbolic regression models with human-in-the-loop and genetic prog ramming,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:42:48.397720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:42:47.839951Z digest=sha256:67f7e6e5f850d8b070120a0fbd4a7f253490dfb1376f19ea7e34715afc2d1792

Observation d0ff0900-95db-490d-bbb6-8994ba843467 · outbound

This paper cites Improving cross-modal alignment f or text- guided image inpainting,.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Improving cross-modal alignment f or text- guided image inpainting,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T22:42:47.843949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:42:47.843949Z digest=sha256:746dace41fd2d1f730cf9057c323192402ed3f957b8730630da2590066f285c2

Observation ee5d1221-7ff3-4ab9-939b-13f0fb837801 · outbound

This paper cites Zero-shot text-to-image generation,.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Zero-shot text-to-image generation,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:42:48.373659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:42:47.847927Z digest=sha256:b7c984c5ed88f83e3261cf50c98415ee768176fa2e65719abf68a33388cffe54

Observation cd633819-6bf3-46c3-a9c4-4a84c31ee2ed · outbound

This paper cites Towards language-driven video inpainting via multimoda l large language models,.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Towards language-driven video inpainting via multimoda l large language models,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:42:48.359278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:42:47.852061Z digest=sha256:b6ca3efa73bd08af6d6d6458c84092e52f63ddef75d0a48cd682e8416e6ce908

Observation 7df5d13b-9bd0-4922-82e1-58ada63d4361 · outbound

This paper cites Prompt Expansion for Adaptive Text-to-Image Generation.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Prompt Expansion for Adaptive Text-to-Image Generation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T22:42:47.856445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:42:47.856445Z digest=sha256:476d1776293e8e21214a1366dd37a9d786f6e01992ca5aeb7377eee95fee7e96

Observation 5bd31eed-2ce6-4aca-b6db-8362b87569cd · outbound

This paper cites Training-free consistent text-to-image gener ation,.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Training-free consistent text-to-image gener ation,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:42:48.344419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:42:47.861291Z digest=sha256:f4f40bf259342128c7cd6c258adc33a08b6b0956fa99796c5f60d1521332ae6c

Observation b92ebe12-0373-493b-baff-f96f1d5f8d6a · outbound

This paper cites Style-aware contrastive learning for multi-style image captioning,.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Style-aware contrastive learning for multi-style image captioning,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T22:42:47.865685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:42:47.865685Z digest=sha256:ea82fe0768a8ef0efe4c98e66dae67ff305417397850ba273b58cbe1a57f0af4

Observation eedea3d6-b472-4ad6-86aa-c758ef8bb01f · outbound

This paper cites Multimodal event transformer for image-guided st ory ending generation,.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Multimodal event transformer for image-guided st ory ending generation,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:42:48.320620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:42:47.869872Z digest=sha256:13eeabd3075e110b8acb373f42f34be556fb72a3ea071adfd357d1cd6aa806ff

Observation 4f700a97-5b4f-44d2-919d-8013fd81e193 · outbound

This paper cites Triple sequence generati ve adversarial nets for unsupervised image captioning,.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Triple sequence generati ve adversarial nets for unsupervised image captioning,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T22:42:47.874406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:42:47.874406Z digest=sha256:43f20af14d2f0f1b374464e35377d4c59d1be4c67781a585d4c30b04c3baeed2

Observation ccda3c91-5b5b-4ba9-b9f4-0efd3f4651b2 · outbound

This paper cites Sketch storytelling,.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Sketch storytelling,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T22:42:47.878734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:42:47.878734Z digest=sha256:b96176f32ef856662304ccfc4d6b314340f2e449eeddc3c7e7a6788cb7ade191

Observation 70acb0cc-a645-4cc5-9b12-9da7a0068e37 · outbound

This paper cites Generalist Multimodal AI: A Review of Architectures, Challenges and Opportunities.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Generalist Multimodal AI: A Review of Architectures, Challenges and Opportunities

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-10T22:42:48.155107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:42:47.882788Z digest=sha256:5ede19f2d129f0166da49efcec913929487fc07a77637b934890e077cb7f4592

Observation 84865179-7f60-4189-a0b0-b616a27b54b6 · outbound

This paper cites Advancing Fine-Grained Visual Understanding with Multi-Scale Alignment in Multi-Modal Models.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Advancing Fine-Grained Visual Understanding with Multi-Scale Alignment in Multi-Modal Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T22:42:47.887679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:42:47.887679Z digest=sha256:dd2228a3880f00109c45aa00d56a6891bc983c599921ff5f5d83c76634174836

Observation 4647b3ae-dcd7-4ef3-89b5-b507a7c7b778 · outbound

This paper cites Visual in-context l earning for large vision-language models,.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Visual in-context l earning for large vision-language models,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T22:42:47.892120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:42:47.892120Z digest=sha256:fc52b04c1d6494a239496433bce69fd5821347097ecebe1d46354d074c906f16

Observation 050b7709-0670-4258-b108-1e6a93eb91af · outbound

This paper cites InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T22:42:47.896528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:42:47.896528Z digest=sha256:7ea9e81c489f86196b78fcd038f2b0154eb28b35ba6a671a5b7ea9fc58e58000

Observation 31b6ca13-b271-4ac6-8df8-52467eeb0a85 · outbound

This paper cites MoE-LLaVA: Mixture of Experts for Large Vision-Language Models.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models MoE-LLaVA: Mixture of Experts for Large Vision-Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T22:42:47.901252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:42:47.901252Z digest=sha256:a4401bc421a8267f431366c8a56e5f0f31ea142047add8aa24857b04c7ed639a

Observation 6672227c-d77f-4391-9712-fa755c0eb995 · outbound

This paper cites Understanding Reinforcement Learning-Based Fine-Tuning of Diffusion Models: A Tutorial and Review.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Understanding Reinforcement Learning-Based Fine-Tuning of Diffusion Models: A Tutorial and Review

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T22:42:47.905883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:42:47.905883Z digest=sha256:a6be1b10e3be4335385109aee875ed495fb3cda4ad8c8b4d235c6545c88e8d83

Observation 14e2375b-3758-4254-9be4-f3b03b6a1106 · outbound

This paper cites A Survey on Benchmarks of Multimodal Large Language Models.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models A Survey on Benchmarks of Multimodal Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T22:42:47.910468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:42:47.910468Z digest=sha256:cbe0b7aaf539082dfea62f3fe95a3c89bb9ac9d5d76efbc1439f1fa38a44d03a

Observation 4261da68-09ad-43cd-84e4-f7848eafff97 · outbound

This paper cites Ex ploring the frontier of vision-language models: A survey of current met hodologies and future directions,.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Ex ploring the frontier of vision-language models: A survey of current met hodologies and future directions,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T22:42:47.914817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:42:47.914817Z digest=sha256:1586009282ef49aa55b25f0ef756282d05435930e87a0d74f8ae42ec321db18f

Observation 4bc47a9f-7762-4e47-aefd-26d990754358 · outbound

This paper cites Visionllm: Large language model is also an open-ended decoder for vision-centric tasks,.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Visionllm: Large language model is also an open-ended decoder for vision-centric tasks,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:42:48.276806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:42:47.918957Z digest=sha256:5359de315daac5017d3afbf51e058a9a02e0c71976c453bc8d638793006bb478

Observation b2e9a579-3d1b-43c4-9baf-103d1b168642 · outbound

This paper cites Less is more: Vision representation compression for efficient video gene ration with large language models,.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Less is more: Vision representation compression for efficient video gene ration with large language models,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:42:48.262399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:42:47.923687Z digest=sha256:c5df75fa3d105a993411cb898fdc68008dc966688819f158fdda75c7c53b3f2a

Pith citing papers

No inbound Pith citation observations are available.