Pith. sign in

Paper Citation Record · LEDGER

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models

As of 11 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2501.07086.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.07086 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T20:53:27.913006Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact1
  • verified fuzzy20
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5c09b0f4-655d-4712-b47e-ace5d8872c24 · outbound

This paper cites GLIDE: towards photorealistic image generation and editing with text-guided diffusion models,.

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models GLIDE: towards photorealistic image generation and editing with text-guided diffusion models,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:28.561574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:53:27.669201Z digest=sha256:8b90d4f89b0fed719a9d813d05d67443ca296ad25f694470b4f4ecd5c36e8858

Observation bc7188f2-a1ad-4184-9c6b-93f933c60f96 · outbound

This paper cites Zero-shot text-to-image generation,.

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models Zero-shot text-to-image generation,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:28.546171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:53:27.676905Z digest=sha256:33fafd2870512cd9b6c7be9737678c434f67ada784feffd2be21b1d2f6f61720

Observation ff87c1b7-825f-46b8-9328-962e2d6ade18 · outbound

This paper cites Photorealistic text-to-image diffusion models with deep language understanding,.

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models Photorealistic text-to-image diffusion models with deep language understanding,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:28.528242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:53:27.706976Z digest=sha256:81806a433d9839f03c3a65cdaed39327004ddfeee980e7c26a0a05c532850aec

Observation e843a7ca-0d43-46e0-83f7-cdae742cdf4c · outbound

This paper cites The revolution of multimodal large language models: A survey,.

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models The revolution of multimodal large language models: A survey,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:28.512292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:53:27.757052Z digest=sha256:1052723b6386058669d6e5e819c6e5b2741a347d04b9fa2a6926f977b9e2bf3c

Observation 3b34aee4-f94c-43ab-aeb4-6d37ff67b7a5 · outbound

This paper cites Design guidelines for prompt engineering text- to-image generative models,.

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models Design guidelines for prompt engineering text- to-image generative models,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:28.496749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:53:27.789964Z digest=sha256:cb8e89143b56fe39b4b852aaca39616dea3b325333c108a13469065d92399c1c

Observation 64dee35c-00a8-4223-9320-0abecbadc0b8 · outbound

This paper cites Generative Multimodal Models are In-Context Learners.

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models Generative Multimodal Models are In-Context Learners

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:27.795138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:27.795138Z digest=sha256:e023b0c7f938976db8eb07672aa7cf424e19d522c7a398ee79b5005d870e4301

Observation 6c2327d7-e453-4c93-b25e-6cd6ba745609 · outbound

This paper cites Promptcot: Align prompt distribution via adapted chain- of-thought,.

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models Promptcot: Align prompt distribution via adapted chain- of-thought,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:28.480015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:53:27.804754Z digest=sha256:44977b7f9e344053d4cb72439f3f7a83367b859c0591588d1ceb01096b64a578

Observation d45e8b67-f011-4d74-b167-4f73e178c5f5 · outbound

This paper cites Not all languages are created equal in llms: Improving multilingual capability by cross-lingual-thought prompting,.

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models Not all languages are created equal in llms: Improving multilingual capability by cross-lingual-thought prompting,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:28.464282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:53:27.816471Z digest=sha256:c801310108bff6f67a62418d82a22cfb17a25f29acd5bb0714cc66e26bcb1d45

Observation 5d79981d-5e90-4e23-8910-71bee16d0996 · outbound

This paper cites Cross-lingual prompt- ing: Improving zero-shot chain-of-thought reasoning across languages,.

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models Cross-lingual prompt- ing: Improving zero-shot chain-of-thought reasoning across languages,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:28.432734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:53:27.825982Z digest=sha256:1b8af48824247fc41ecf498ea3366dad755f814eebac7cae2917b35cef0146d1

Observation da7a3b5b-2483-43b3-a1c7-525a0fb0cc1a · outbound

This paper cites PLUG: leveraging pivot language in cross-lingual instruction tuning,.

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models PLUG: leveraging pivot language in cross-lingual instruction tuning,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:28.413772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:53:27.830870Z digest=sha256:0931790ca4ff8e607e7a64aa8e56e2fa3bb0989bb4e63df767408992cfdc5a3a

Observation cf8ad285-c9ec-4bbf-b056-67851978658c · outbound

This paper cites Revealing the Parallel Multilingual Learning within Large Language Models.

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models Revealing the Parallel Multilingual Learning within Large Language Models

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-10T20:53:28.055231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:53:27.835654Z digest=sha256:99875576eec9add5b3645e28840d15faf16b758d21cccbeaf50d1af8fcfbca74

Observation ae032c7f-a368-466f-bcff-c2fdbf153b55 · outbound

This paper cites LAION-5B: an open large-scale dataset for training next generation image-text models,.

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models LAION-5B: an open large-scale dataset for training next generation image-text models,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:28.397020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:53:27.840560Z digest=sha256:c93536b74c2ac17e5ebe0d18babacf1a754ad84014ab3ba5b8e51422ee9825c3

Observation d6bf1b94-9ff5-47d2-805e-a1f6d0ee3f85 · outbound

This paper cites Denoising diffusion probabilistic models,.

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models Denoising diffusion probabilistic models,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:28.378402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:53:27.845353Z digest=sha256:63f9b0cb132db8137d3a7e93437e3e985b2afdc7341555de7a74c57bb4c6b414

Observation 245c3e8f-6ab6-461c-88c1-83c2602f244a · outbound

This paper cites Taming transformers for high- resolution image synthesis,.

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models Taming transformers for high- resolution image synthesis,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:28.362226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:53:27.850132Z digest=sha256:52cb954ecdacdccaaa8497a11377af4fee5fc80b8e67a4320d70af56fe9f7350

Observation b85d43a4-c851-4411-bc88-d6b2da11086b · outbound

This paper cites Clipscore: A reference-free evaluation metric for image captioning,.

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models Clipscore: A reference-free evaluation metric for image captioning,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:28.344341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:53:27.855077Z digest=sha256:ae11bb1f973c65fecad7f4ee57f720fe540708f2fb248c0c5a4478bc0ec20ef3

Observation eeed5c60-77dd-40a2-ac61-3415df90a96f · outbound

This paper cites Microsoft COCO: common objects in context,.

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models Microsoft COCO: common objects in context,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:27.859826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:27.859826Z digest=sha256:cbaab36592d846e3bab8e5593ffda5bc24c056367fa16b04f350724b5c0d85da

Observation 8e30d65f-3b00-49c4-bca0-7f73ba2739b0 · outbound

This paper cites Improving image generation with better captions,.

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models Improving image generation with better captions,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:27.864672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:27.864672Z digest=sha256:e985f082b875589c115e11f1fc80450d7be4e25612da81b641b226754f398219

Observation 2cc3403d-8551-4c7a-b4ff-b67900ccc57b · outbound

This paper cites T2i-compbench: A comprehensive benchmark for open-world compositional text-to-image generation,.

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models T2i-compbench: A comprehensive benchmark for open-world compositional text-to-image generation,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:28.302428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:53:27.874475Z digest=sha256:039353ffe1dcc5de9d73ad2161a5b74e3846d9fbc52759853f57d3587b192680

Observation aeb945c6-ee58-432a-b11f-023d886ea8e1 · outbound

This paper cites Magicbrush: A manually annotated dataset for instruction-guided image editing,.

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models Magicbrush: A manually annotated dataset for instruction-guided image editing,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:28.281489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:53:27.879790Z digest=sha256:df87ab747715d5ce9a992dfd4482eb87a1e0fc7ae72629dd197dd8fa372f8654

Observation 76e9cbcf-bcb4-4703-b250-ecfdefdfaa1c · outbound

This paper cites BLIP: bootstrapping language- image pre-training for unified vision-language understanding and genera- tion,.

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models BLIP: bootstrapping language- image pre-training for unified vision-language understanding and genera- tion,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:28.262795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:53:27.884483Z digest=sha256:b2459aabc6621d0964c942d88eed50a131bd94787fc0d04b01a07338c7509045

Observation 49515561-7c1d-4a3c-9bde-93083d2c8c80 · outbound

This paper cites Imagereward: Learning and evaluating human preferences for text- to-image generation,.

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models Imagereward: Learning and evaluating human preferences for text- to-image generation,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:28.235684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:53:27.889016Z digest=sha256:fb0584fc4209a8096b013ff88488bdc4dcff038ed77e41cdafbf251fe8c737d0

Observation d4fd8d51-e318-42f3-adb0-a86d990013b0 · outbound

This paper cites Lumina-Next: Making Lumina-T2X Stronger and Faster with Next-DiT.

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models Lumina-Next: Making Lumina-T2X Stronger and Faster with Next-DiT

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:27.893502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:27.893502Z digest=sha256:b58054db22be0e954de53726dbcb88fcde5a7108bf44e1c7f779764c8fd61395

Observation d40e715a-7169-44ab-b756-e860f0a55c8e · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models Gemma: Open Models Based on Gemini Research and Technology

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:27.898679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:27.898679Z digest=sha256:9ea8f765225c87b8f8cc2d15d16db31769c86487097994d340beb7605c2556fc

Observation aa3b67f8-2e80-4e97-ac35-e0e4c0e4dc06 · outbound

This paper cites Optimizing prompts for text- to-image generation,.

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models Optimizing prompts for text- to-image generation,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:28.183564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:53:27.903719Z digest=sha256:323907d6e7a73406e6cb7c417897bdf8a65101ebcd686cb7747a804d5452657b

Observation f22d0cdc-9799-45d9-af56-b3e61064f226 · outbound

This paper cites Aligner: Efficient Alignment by Learning to Correct.

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models Aligner: Efficient Alignment by Learning to Correct

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:27.908141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:27.908141Z digest=sha256:7bde6a4d5bf1ea1c681ec500a7a55e53307afca229a987c34fb802b6076b0215

Observation 17196ef3-5d1f-4214-b4b3-cdf3c1cb54e7 · outbound

This paper cites Weak-to-strong generalization: Eliciting strong capabilities with weak supervision,.

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models Weak-to-strong generalization: Eliciting strong capabilities with weak supervision,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:28.114770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:53:27.913006Z digest=sha256:92e82c0debb8f0801865e7e0e1da21574c6ef2ab3bd4dba304741ca0500f8bd6

Observation 1a4156a8-c63e-4c90-ba54-adc4cab0af09 · outbound

This paper cites 12 365– 12 394.

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models 12 365– 12 394

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:28.449198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:53:27.821191Z digest=sha256:9f6170620bde3f1dc3f02e2bd788ce47264097322e5dfb7af6ab1c9abf67c769

Pith citing papers

No inbound Pith citation observations are available.