Pith. sign in

Paper Citation Record · LEDGER

CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder

As of 13 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 4 inbound Pith citation observations for arXiv:2412.17225.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.17225 v1

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T05:45:54.168593Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T05:51:12.656627Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T21:28:58.246299Z

Reference resolution

39 of 39 outbound references displayed

  • verified exact1
  • verified fuzzy21
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 32dcee00-c070-47fe-a50a-c395288bfe3c · outbound

This paper cites Improving Image Gen- eration with Better Captions.

CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder Improving Image Gen- eration with Better Captions

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:45:54.632585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T05:45:54.027737Z digest=sha256:35f84d399ff1e303d75e569d7bd4b93343c88626d4798b92084f7d1044f230e7

Observation d644f7f7-218a-4241-8348-8e52a477c2d7 · outbound

This paper cites Freeman, Michael Rubinstein, Yuanzhen Li, and Dilip Krishnan.

CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder Freeman, Michael Rubinstein, Yuanzhen Li, and Dilip Krishnan

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:45:54.620509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T05:45:54.032151Z digest=sha256:8c56c0c72a36e33504b4a7f4a863b37248777c5b346a90fa7eef564b4725bda6

Observation aa59fca9-478b-42bf-ad77-66a533438e79 · outbound

This paper cites Diffute: Universal text editing diffusion model.

CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder Diffute: Universal text editing diffusion model

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:45:54.610514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T05:45:54.037798Z digest=sha256:3ffb9116527251a55ddd79d3a2db02e6d9424b67b798b66e206c9382273f0662

Observation af6c4dbe-363e-4cd0-a785-3fd33ceaf91e · outbound

This paper cites TextDiffuser-2: Unleashing the Power of Language Models for Text Rendering.

CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder TextDiffuser-2: Unleashing the Power of Language Models for Text Rendering

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T05:45:54.042101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:45:54.042101Z digest=sha256:1a97b9cb03735243bba5cb0af914b3431210427c8cda2d6c455b0bd8dc9201b7

Observation 5be04580-7bc3-41bf-a3b5-123cfd3eeea4 · outbound

This paper cites Textdiffuser: Diffusion models as text painters.

CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder Textdiffuser: Diffusion models as text painters

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:45:54.600198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T05:45:54.046269Z digest=sha256:68b2584f0aee41a3e531af106630365a82fdf520e214f549787ca475592f56b7

Observation 0f6a3fba-73d8-4b6f-9915-8b934c5132e1 · outbound

This paper cites Diffusion Models Beat GANs on Image Synthesis.

CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder Diffusion Models Beat GANs on Image Synthesis

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T05:45:54.049939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:45:54.049939Z digest=sha256:c939bfaefe06d68aeead155279aa5e2dae40bb0a389b7dac3c7928ebc8ac5a17

Observation a105cabf-43bc-426b-b935-3f12b69ce34c · outbound

This paper cites Odm: A text-image further alignment pre-training approach for scene text detection and spotting.

CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder Odm: A text-image further alignment pre-training approach for scene text detection and spotting

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:45:54.589550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T05:45:54.054305Z digest=sha256:031bb74f25761fe41eb78369b61f21f189e4b2e4735a3a836762888830158fcd

Observation a33befb2-80d5-47d5-98fd-700c80dbde74 · outbound

This paper cites Duguangocr.

CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder Duguangocr

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:45:54.578511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T05:45:54.057863Z digest=sha256:f79fa6617ed5f46fcb76517de53c3ada222c57402afcd93cac4b71aa476ca9e5

Observation 4fdb2fe9-d5c1-46df-a71d-66b3a84b97ba · outbound

This paper cites Scaling recti- fied flow transformers for high-resolution image synthesis.

CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder Scaling recti- fied flow transformers for high-resolution image synthesis

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T05:45:54.061352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:45:54.061352Z digest=sha256:52cb462da6eed60e6085f447b8cf76b15ae5f7fb99074de9bf7245218d60408b

Observation f6730221-9216-40c4-bccd-9301fa904acb · outbound

This paper cites Wukong: 100 million large-scale chinese cross-modal pre-training dataset and a foundation frame- work, 2022.

CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder Wukong: 100 million large-scale chinese cross-modal pre-training dataset and a foundation frame- work, 2022

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:45:54.558944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T05:45:54.064964Z digest=sha256:d30e2a11d9bc79dddb1f8fe9e28d45abbe5f963762a393b4de3b5251f42982a9

Observation 78774e74-34d5-4000-bd64-99259f3d4ee9 · outbound

This paper cites Denoising Dif- fusion Probabilistic Models, 2020.

CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder Denoising Dif- fusion Probabilistic Models, 2020

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:45:54.546182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T05:45:54.068410Z digest=sha256:4c03ec3c7bcf0c4ddafacdb1f392389b97831698a6dc4565aa05a350514f6652

Observation 87be2020-c36d-43de-ad5c-6df12c40f426 · outbound

This paper cites Auto-encoding varia- tional bayes, 2022.

CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder Auto-encoding varia- tional bayes, 2022

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:45:54.533173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T05:45:54.072712Z digest=sha256:fd1d1f64f40bed59494cb1654b66a8b777dbb377f05458ceccfe903d1b528bef

Observation 11d9325a-7517-4838-9095-db005ac98a92 · outbound

This paper cites Pp-ocrv3: More attempts for the improvement of ultra lightweight ocr system,.

CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder Pp-ocrv3: More attempts for the improvement of ultra lightweight ocr system,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:45:54.521099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T05:45:54.076401Z digest=sha256:6a6611cf91f3aa11a6350f9e1d6423a262c4f5b7abeda6f67fb8e76c32c28680

Observation 4a50ce3e-f731-474e-8afe-a05b6511c937 · outbound

This paper cites Character-aware models improve visual text rendering, 2023.

CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder Character-aware models improve visual text rendering, 2023

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:45:54.509116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T05:45:54.079905Z digest=sha256:5a9e0481eda40532ddd7534e672f424836ddecb142df54da36333be3a92b4179

Observation ba0c7ce2-e223-4f68-b8c1-413f868882c1 · outbound

This paper cites Character-Aware Models Improve Visual Text Rendering.

CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder Character-Aware Models Improve Visual Text Rendering

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T05:45:54.083336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:45:54.083336Z digest=sha256:eb3c2f963565b1bb819397be129b5d0440c2b2cbb0f6b7a1d3c27b21f7cee948

Observation 5708ea8d-5278-4cc3-8bc1-46d4e43f1593 · outbound

This paper cites Glyph-ByT5: A Cus- tomized Text Encoder for Accurate Visual Text Rendering,.

CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder Glyph-ByT5: A Cus- tomized Text Encoder for Accurate Visual Text Rendering,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:45:54.495477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T05:45:54.086895Z digest=sha256:dd2a8ed4a64ffb34ab1b9b091ffabd481fd4778af9a355fb64422f93dd9893d6

Observation 0c45c0d0-fc72-4ab9-b27b-a285b2cdd64f · outbound

This paper cites Glyph-ByT5-v2: A Strong Aesthetic Baseline for Accurate Multilingual Visual Text Rendering.

CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder Glyph-ByT5-v2: A Strong Aesthetic Baseline for Accurate Multilingual Visual Text Rendering

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T05:45:54.094448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:45:54.094448Z digest=sha256:cb400c17a6e4c371c6b95d3ba211eafa5f4b0083c92d524e66e4a1d0612104c5

Observation 2ecb5dfd-4802-4cc8-b0dc-60425446d0bd · outbound

This paper cites GlyphDraw: Seamlessly Rendering Text with Intricate Spatial Structures in Text-to-Image Generation.

CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder GlyphDraw: Seamlessly Rendering Text with Intricate Spatial Structures in Text-to-Image Generation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T05:45:54.097886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:45:54.097886Z digest=sha256:6cb519786e1239232086d448aafaa3b37d0db851ecd5119315f1ddede2f21861

Observation ad03fc6a-a382-4c15-9bc4-a7aba992e08e · outbound

This paper cites Glyphdraw2: Automatic generation of complex glyph posters with diffusion models and large language models,.

CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder Glyphdraw2: Automatic generation of complex glyph posters with diffusion models and large language models,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:45:54.482727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T05:45:54.101959Z digest=sha256:a17ba403783183b0edf849c924e4b324ad0c432a6f205c79ed93791a43aa2ed1

Observation 5e344d80-5f47-487b-9baf-fec339db194f · outbound

This paper cites Midjourney.

CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder Midjourney

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:45:54.468583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T05:45:54.105504Z digest=sha256:201e60497ba9e40c1c083059af267fa977235f3fe29d8a7ceb6962dda3c782f3

Observation 37c395c1-3718-43a0-91ce-d99a458aecaf · outbound

This paper cites Improved denoising dif- fusion probabilistic models, 2021.

CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder Improved denoising dif- fusion probabilistic models, 2021

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:45:54.455745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T05:45:54.108873Z digest=sha256:1abc7128d2a9927ed567df33b25ba07930dcf2c0292cd9b84b89791ebe9a8f50

Observation fe117036-b7ff-4752-85c2-25ea340eb4ee · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder Learning transferable visual models from natural language supervi- sion

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T05:45:54.112155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:45:54.112155Z digest=sha256:ba56212e4a1135421ec82ba5c846f62e20d36520f4f90d3479904ad4aacef90f

Observation 098b4e94-2f61-4228-a1b5-6d8025581694 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T05:45:54.115472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:45:54.115472Z digest=sha256:3b60bb6e03f5dd410e6f706d1dc1f465b98d8bbcd2d0a512d83466320c32304f

Observation 59b35edd-5e0c-46be-8542-5ed4add82c13 · outbound

This paper cites Zero-shot text-to-image generation, 2021.

CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder Zero-shot text-to-image generation, 2021

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T05:45:54.118609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:45:54.118609Z digest=sha256:5a24c91027833e6999701cf73dda6f8e05586c94c6ba5b9bf69f83276011905f

Observation f74ce017-246f-456e-85b6-1d9d3748c3ed · outbound

This paper cites Hierarchical text-conditional image gener- ation with clip latents, 2022.

CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder Hierarchical text-conditional image gener- ation with clip latents, 2022

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T05:45:54.122161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:45:54.122161Z digest=sha256:222f954baf230d40264358e26af46219944e71ed10d64a149d634116c1e846da

Observation 8f16315b-4474-4658-a26c-796c1c0801fe · outbound

This paper cites Ocr-vqgan: Taming text- within-image generation.

CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder Ocr-vqgan: Taming text- within-image generation

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:45:54.410844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T05:45:54.126126Z digest=sha256:5954e56c4a81894729d2a8d0133632922424c26a402272e62270f640f695e44c

Observation 18160161-c409-440f-88f1-a2b150a5e58c · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder High-resolution image synthesis with latent diffusion models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:45:54.400593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T05:45:54.129746Z digest=sha256:714a1a75fa6d837e5983ca55b9cfa7dac1d23c7f99c2ab9eb79893868f8aef65

Observation 2b65121d-7aa4-46a7-8d28-d0f51b1b9144 · outbound

This paper cites Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding.

CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T05:45:54.134047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:45:54.134047Z digest=sha256:05f6352c6cf9cd6d69814b31234f31818e64426b25c9f2c489d156f8d397c506

Observation a1522277-adbe-44d6-8b8e-e58ff6b58c12 · outbound

This paper cites LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs.

CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T05:45:54.137766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:45:54.137766Z digest=sha256:bd4bd1c97756f298f78b6415a4efef099506983f13d7a905c02a7629452f429d

Observation 3dfd3bf8-f8d6-4cdb-a245-7165c583f528 · outbound

This paper cites Kingma, Ab- hishek Kumar, Stefano Ermon, and Ben Poole.

CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder Kingma, Ab- hishek Kumar, Stefano Ermon, and Ben Poole

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T05:45:54.141208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:45:54.141208Z digest=sha256:e532190234d19e8d4d7f3a4f7e49c4c8bd408b5606b2c8b845c3551ed040144b

Observation 0d02f40a-f4d3-4215-9539-a03046be485f · outbound

This paper cites AnyText: Multilingual Visual Text Generation And Editing.

CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder AnyText: Multilingual Visual Text Generation And Editing

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T05:45:54.144468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:45:54.144468Z digest=sha256:544c9f7ead3165c710eeb8fef55afb0a334db502547ea3ad2a16d40aff6b47c4

Observation 5c17193d-3251-4244-af5a-c4db9ddaa9a3 · outbound

This paper cites Glyphcontrol: Glyph conditional control for visual text generation.

CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder Glyphcontrol: Glyph conditional control for visual text generation

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:45:54.379824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T05:45:54.147962Z digest=sha256:dbc69e6b32a15d51b7e0060d58e7713feef2e99dc872fdae9db61a15de049978

Observation 2d8c04b4-8d3e-46c3-a1c3-eee0ee9c063d · outbound

This paper cites Long-clip: Unlocking the long-text capability of clip, 2024.

CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder Long-clip: Unlocking the long-text capability of clip, 2024

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:45:54.367440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T05:45:54.151114Z digest=sha256:01b6f692c06a2d1b26644361c2f795ae3ec1ed25d6a8453cbab4112e91ddef37

Observation f03aad97-a15e-49a2-ad58-e27632021fe1 · outbound

This paper cites Adding conditional control to text-to-image diffusion models.

CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder Adding conditional control to text-to-image diffusion models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:45:54.354679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T05:45:54.154138Z digest=sha256:8359cf7930058ad1f833a30831a83b89e9eeeb1c5dfcfd7fb126cf7567889cee

Observation c02d6877-585f-4d82-ab71-d409ae8d0b2b · outbound

This paper cites Adding Conditional Control to Text-to-Image Diffusion Models,.

CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder Adding Conditional Control to Text-to-Image Diffusion Models,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T05:45:54.157738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:45:54.157738Z digest=sha256:77cc0b9d6f9c5a29832561cd28bc8d681dc068cf0fff3b20a6a874c70ac137ef

Observation 40ef395d-b24b-4269-8b3e-ad783d5a255e · outbound

This paper cites Brush your text: Synthesize any scene text on im- ages via diffusion model.

CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder Brush your text: Synthesize any scene text on im- ages via diffusion model

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:45:54.335900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T05:45:54.165257Z digest=sha256:60410c9cf0f7f48596af56899d4ca2a0785935e407376324f7205c14922297b6

Observation 13b981e2-4f9b-4bc2-9c69-9a6572dca6a3 · outbound

This paper cites UDiffText: A Unified Framework for High-quality Text Synthesis in Arbitrary Images via Character-aware Diffusion Models.

CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder UDiffText: A Unified Framework for High-quality Text Synthesis in Arbitrary Images via Character-aware Diffusion Models

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-08-11T05:45:54.207012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T05:45:54.168593Z digest=sha256:5d15ae948756149d750e541645a07863327e471687236534b1707f760ebaa7c7

Observation bfd22a16-6f0b-4425-9200-70da3b6ed929 · outbound

This paper cites Adding Conditional Control to Text-to-Image Diffusion Models.

CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder Adding Conditional Control to Text-to-Image Diffusion Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T05:45:54.161125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:45:54.161125Z digest=sha256:1f837bfeba5c50d39e78f03cb499c0447aec652e82831dc6dac7db4d90eec46b

Observation fa5c5579-92f7-4b53-9ea1-4408977906f0 · outbound

This paper cites Glyph-ByT5: A Customized Text Encoder for Accurate Visual Text Rendering.

CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder Glyph-ByT5: A Customized Text Encoder for Accurate Visual Text Rendering

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T05:45:54.090483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:45:54.090483Z digest=sha256:b38f35fbb0bdc5a8fc1fb6f810d957180492311948d137adc1a6582ffc492706

Pith citing papers

Observation 800ea31d-774f-4249-835e-68c4aa3970ec · inbound

Holding the FP8 Quality Ceiling at 8-Bit Weights and Activations: INT8 and GGUF Post-Training Quantization of Ideogram 4.0 for Consumer GPUs cites this paper.

Holding the FP8 Quality Ceiling at 8-Bit Weights and Activations: INT8 and GGUF Post-Training Quantization of Ideogram 4.0 for Consumer GPUs CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:17:58.047058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T10:05:40.235610Z digest=sha256:07fecb4aab1e74c803c8f99a22c26f6d5f52ea994b72c3142fb3e6cfff3041c1

Observation 8daa24de-0235-4b42-a24e-0ddd934bb10e · inbound

GMO-E$^2$DIT: Grounded Multi-Operation Editing for E-Commerce Images cites this paper.

GMO-E$^2$DIT: Grounded Multi-Operation Editing for E-Commerce Images CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:56:58.996534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-02T13:56:43.671622Z digest=sha256:5576f24541d7f651068c59a35bcd63e78269646383180e2e5e0a51768d7eb46c

Observation f9fc143e-5a29-4ec9-8905-b2e272ea7998 · inbound

GMO-E$^2$DIT: Grounded Multi-Operation Editing for E-Commerce Images cites this paper.

GMO-E$^2$DIT: Grounded Multi-Operation Editing for E-Commerce Images CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-03T21:28:58.247930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-03T21:23:41.271521Z digest=sha256:f8da0c976b5695b468d752e5b00917621b9bac6b42e0981d55a6f84916460fcd

Observation df547916-2f60-4f9e-a17d-0164114a8eb5 · inbound

InnoText: A Unified Model for Visual Text Generation and Editing cites this paper.

InnoText: A Unified Model for Visual Text Generation and Editing CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T05:51:12.656627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T05:51:12.656627Z digest=sha256:af806002dee5201d886c5ecdf12631bd766a3071a0f03e1a704d4f5b472e4d6c