Pith. sign in

Paper Citation Record · LEDGER

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models

As of 11 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2501.07086.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.07086 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T20:53:27.913006Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact1
  • verified fuzzy20
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5c09b0f4-655d-4712-b47e-ace5d8872c24 · outbound

This paper cites GLIDE: towards photorealistic image generation and editing with text-guided diffusion models,.

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models GLIDE: towards photorealistic image generation and editing with text-guided diffusion models,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:28.561574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:53:27.669201Z digest=sha256:418c9ed20cfcdfdc6bd8049ef690ded1b86c11c7063f82ceb51095fe32dd1f09

Observation bc7188f2-a1ad-4184-9c6b-93f933c60f96 · outbound

This paper cites Zero-shot text-to-image generation,.

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models Zero-shot text-to-image generation,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:28.546171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:53:27.676905Z digest=sha256:f0033faa75c58b51312c910ff7243e14b821cbcd8657d0406475345ff289f4d6

Observation ff87c1b7-825f-46b8-9328-962e2d6ade18 · outbound

This paper cites Photorealistic text-to-image diffusion models with deep language understanding,.

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models Photorealistic text-to-image diffusion models with deep language understanding,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:28.528242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:53:27.706976Z digest=sha256:9d9dd51fbb4248efcd383c151ae049623e8a0f49b981cf290cf3b6455bef7fa7

Observation e843a7ca-0d43-46e0-83f7-cdae742cdf4c · outbound

This paper cites The revolution of multimodal large language models: A survey,.

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models The revolution of multimodal large language models: A survey,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:28.512292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:53:27.757052Z digest=sha256:1a2512bb807121ebe8201b3915ab6181ce530b549616d191b2d8381290b29570

Observation 3b34aee4-f94c-43ab-aeb4-6d37ff67b7a5 · outbound

This paper cites Design guidelines for prompt engineering text- to-image generative models,.

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models Design guidelines for prompt engineering text- to-image generative models,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:28.496749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:53:27.789964Z digest=sha256:ab0a49c87f1ef1ea19d0d9b50abdc4e3d933363abe1e74096e83f26d53ed8de3

Observation 64dee35c-00a8-4223-9320-0abecbadc0b8 · outbound

This paper cites Generative Multimodal Models are In-Context Learners.

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models Generative Multimodal Models are In-Context Learners

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:27.795138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:27.795138Z digest=sha256:e023b0c7f938976db8eb07672aa7cf424e19d522c7a398ee79b5005d870e4301

Observation 6c2327d7-e453-4c93-b25e-6cd6ba745609 · outbound

This paper cites Promptcot: Align prompt distribution via adapted chain- of-thought,.

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models Promptcot: Align prompt distribution via adapted chain- of-thought,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:28.480015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:53:27.804754Z digest=sha256:f55f9fe57d73dcb2818812150ae2836a1f1808bf92bf974c77e8b1865dbbf0d5

Observation d45e8b67-f011-4d74-b167-4f73e178c5f5 · outbound

This paper cites Not all languages are created equal in llms: Improving multilingual capability by cross-lingual-thought prompting,.

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models Not all languages are created equal in llms: Improving multilingual capability by cross-lingual-thought prompting,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:28.464282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:53:27.816471Z digest=sha256:dd45957051d9262c2117c2e06bbd67d65868fb3554f7ff19b246219853645db5

Observation 5d79981d-5e90-4e23-8910-71bee16d0996 · outbound

This paper cites Cross-lingual prompt- ing: Improving zero-shot chain-of-thought reasoning across languages,.

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models Cross-lingual prompt- ing: Improving zero-shot chain-of-thought reasoning across languages,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:28.432734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:53:27.825982Z digest=sha256:8ba69933c1bc3195d53055c95b5e1c516a5dbac15213fae8d4a6903f48f90901

Observation da7a3b5b-2483-43b3-a1c7-525a0fb0cc1a · outbound

This paper cites PLUG: leveraging pivot language in cross-lingual instruction tuning,.

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models PLUG: leveraging pivot language in cross-lingual instruction tuning,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:28.413772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:53:27.830870Z digest=sha256:71dae5c1cab68817ccc95192a74c00d1184457bd390586ecb4c61c7c2e53c8b4

Observation cf8ad285-c9ec-4bbf-b056-67851978658c · outbound

This paper cites Revealing the Parallel Multilingual Learning within Large Language Models.

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models Revealing the Parallel Multilingual Learning within Large Language Models

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-10T20:53:28.055231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:53:27.835654Z digest=sha256:ba6e35d4af76806739c52828ffcef0665d00f5465f260f63eb10fd984e4e497d

Observation ae032c7f-a368-466f-bcff-c2fdbf153b55 · outbound

This paper cites LAION-5B: an open large-scale dataset for training next generation image-text models,.

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models LAION-5B: an open large-scale dataset for training next generation image-text models,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:28.397020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:53:27.840560Z digest=sha256:ff86321294116bc519ac9fe32bcd5a916d9ce4c873e765806047c26256be8e4a

Observation d6bf1b94-9ff5-47d2-805e-a1f6d0ee3f85 · outbound

This paper cites Denoising diffusion probabilistic models,.

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models Denoising diffusion probabilistic models,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:28.378402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:53:27.845353Z digest=sha256:a19e7a7238b2082686144af15d034587f86a4c6629c2f3d0c86ef83a4a6c1444

Observation 245c3e8f-6ab6-461c-88c1-83c2602f244a · outbound

This paper cites Taming transformers for high- resolution image synthesis,.

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models Taming transformers for high- resolution image synthesis,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:28.362226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:53:27.850132Z digest=sha256:272ba3658ee546f8eff9c775bd0f80d6b19ea03aa1a84aa7254f1e20e4c4b37f

Observation b85d43a4-c851-4411-bc88-d6b2da11086b · outbound

This paper cites Clipscore: A reference-free evaluation metric for image captioning,.

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models Clipscore: A reference-free evaluation metric for image captioning,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:28.344341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:53:27.855077Z digest=sha256:39075c402c7c3c847fb4bc09b18b6a3ca2f0195cbe6f0475ce60e850365accb1

Observation eeed5c60-77dd-40a2-ac61-3415df90a96f · outbound

This paper cites Microsoft COCO: common objects in context,.

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models Microsoft COCO: common objects in context,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:27.859826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:27.859826Z digest=sha256:cbaab36592d846e3bab8e5593ffda5bc24c056367fa16b04f350724b5c0d85da

Observation 8e30d65f-3b00-49c4-bca0-7f73ba2739b0 · outbound

This paper cites Improving image generation with better captions,.

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models Improving image generation with better captions,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:27.864672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:27.864672Z digest=sha256:e985f082b875589c115e11f1fc80450d7be4e25612da81b641b226754f398219

Observation 2cc3403d-8551-4c7a-b4ff-b67900ccc57b · outbound

This paper cites T2i-compbench: A comprehensive benchmark for open-world compositional text-to-image generation,.

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models T2i-compbench: A comprehensive benchmark for open-world compositional text-to-image generation,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:28.302428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:53:27.874475Z digest=sha256:6624deeca4f6062ef78104117fd36936ee23ec191dc075a93d10297d1f1199a4

Observation aeb945c6-ee58-432a-b11f-023d886ea8e1 · outbound

This paper cites Magicbrush: A manually annotated dataset for instruction-guided image editing,.

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models Magicbrush: A manually annotated dataset for instruction-guided image editing,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:28.281489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:53:27.879790Z digest=sha256:9cf09077590e11512c912e8bbbd5e4c21f7401da10fc6d5f49b9d16759f56ce7

Observation 76e9cbcf-bcb4-4703-b250-ecfdefdfaa1c · outbound

This paper cites BLIP: bootstrapping language- image pre-training for unified vision-language understanding and genera- tion,.

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models BLIP: bootstrapping language- image pre-training for unified vision-language understanding and genera- tion,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:28.262795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:53:27.884483Z digest=sha256:88b227eeed402c62acb5995cb3cc15b89b44f5f0292b12eeb500c1a562e3bf56

Observation 49515561-7c1d-4a3c-9bde-93083d2c8c80 · outbound

This paper cites Imagereward: Learning and evaluating human preferences for text- to-image generation,.

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models Imagereward: Learning and evaluating human preferences for text- to-image generation,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:28.235684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:53:27.889016Z digest=sha256:af6623f902c39658d8a883eeae17e8cba6c739369c2ae414a65ebddd419a49e7

Observation d4fd8d51-e318-42f3-adb0-a86d990013b0 · outbound

This paper cites Lumina-Next: Making Lumina-T2X Stronger and Faster with Next-DiT.

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models Lumina-Next: Making Lumina-T2X Stronger and Faster with Next-DiT

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:27.893502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:27.893502Z digest=sha256:b58054db22be0e954de53726dbcb88fcde5a7108bf44e1c7f779764c8fd61395

Observation d40e715a-7169-44ab-b756-e860f0a55c8e · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models Gemma: Open Models Based on Gemini Research and Technology

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:27.898679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:27.898679Z digest=sha256:9ea8f765225c87b8f8cc2d15d16db31769c86487097994d340beb7605c2556fc

Observation aa3b67f8-2e80-4e97-ac35-e0e4c0e4dc06 · outbound

This paper cites Optimizing prompts for text- to-image generation,.

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models Optimizing prompts for text- to-image generation,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:28.183564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:53:27.903719Z digest=sha256:3636bb82577206c3914cfae22f75d2ee7d3eab8b64fb27d3a96d53fe84025db5

Observation f22d0cdc-9799-45d9-af56-b3e61064f226 · outbound

This paper cites Aligner: Efficient Alignment by Learning to Correct.

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models Aligner: Efficient Alignment by Learning to Correct

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:27.908141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:27.908141Z digest=sha256:7bde6a4d5bf1ea1c681ec500a7a55e53307afca229a987c34fb802b6076b0215

Observation 17196ef3-5d1f-4214-b4b3-cdf3c1cb54e7 · outbound

This paper cites Weak-to-strong generalization: Eliciting strong capabilities with weak supervision,.

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models Weak-to-strong generalization: Eliciting strong capabilities with weak supervision,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:28.114770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:53:27.913006Z digest=sha256:924750c15ef721f3cbec0f6cac7ef7f1417146e95cfe7148fc6e97f0291cc3cb

Observation 1a4156a8-c63e-4c90-ba54-adc4cab0af09 · outbound

This paper cites 12 365– 12 394.

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models 12 365– 12 394

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:28.449198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:53:27.821191Z digest=sha256:46f40319e6824d5862b51ca280717fb142404918d647f29466c75b0bd6dd085f

Pith citing papers

No inbound Pith citation observations are available.