Pith. sign in

Paper Citation Record · LEDGER

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution

As of 20 August 2026, this Paper Citation Record lists 100 of 101 outbound references and 5 inbound Pith citation observations for arXiv:2412.15213.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.15213 v2

Coverage vector

measured 100 of 101 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T11:38:24.857582Z

measured 105 of 105 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T20:40:31.655763Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T10:44:47.556754Z

Reference resolution

100 of 101 outbound references displayed

  • verified exact2
  • verified fuzzy37
  • unresolved61
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 435ba7b8-7efb-48a3-87c1-6047480ad334 · outbound

This paper cites Building nor- malizing flows with stochastic interpolants.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution Building nor- malizing flows with stochastic interpolants

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T11:38:24.334557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:38:24.334557Z digest=sha256:934f01b35a1902b7a9089a45cc0b8ec0436f685984c916b3dc2c997c475acea8

Observation 696b35ef-42b4-43c9-9ddd-65a06ee14ed0 · outbound

This paper cites Stochastic Interpolants: A Unifying Framework for Flows and Diffusions.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution Stochastic Interpolants: A Unifying Framework for Flows and Diffusions

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T11:38:24.340460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:38:24.340460Z digest=sha256:558ddaf2cf411b3250afabed516d364268a3a54f1daba9bd425f58e88db94059

Observation 2535447a-ffc3-4f1c-af88-af1aebdfd324 · outbound

This paper cites Stochastic interpolants with data-dependent couplings.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution Stochastic interpolants with data-dependent couplings

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T11:38:24.346362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:38:24.346362Z digest=sha256:92664e5ea375bcc93d307e6f0b3affe0f0dc011bdb9859ee75319038241ea498

Observation 5fdcfd56-3ddb-4526-b399-3d481cf01238 · outbound

This paper cites Spice: Semantic propositional image cap- tion evaluation.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution Spice: Semantic propositional image cap- tion evaluation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T11:38:24.353376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:38:24.353376Z digest=sha256:aa3f961db7ad5b5269d55afacfb939a3d4285a884539b93df535d0cfc14cc3ff

Observation 33da7a0d-bd5f-4dbd-b185-9069abac68ae · outbound

This paper cites Meteor: An auto- matic metric for mt evaluation with improved correlation with human judgments.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution Meteor: An auto- matic metric for mt evaluation with improved correlation with human judgments

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T11:38:24.359056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:38:24.359056Z digest=sha256:877a3f529a4dd7df639ac6d61ebcadd9453f305d18e49a9fdf0cb825ac43df60

Observation 65599890-d97c-4cce-a532-db4ef936da1e · outbound

This paper cites All are worth words: A vit backbone for diffusion models.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution All are worth words: A vit backbone for diffusion models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T11:38:24.364185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:38:24.364185Z digest=sha256:91711558a5d152daa216ec55acb523fc606b25bf166f5a483bed8b564a4ee0f4

Observation cdb82db4-c223-4e47-b5fd-8f453104466c · outbound

This paper cites Adabins: Depth estimation using adaptive bins.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution Adabins: Depth estimation using adaptive bins

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T11:38:24.369364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:38:24.369364Z digest=sha256:90c26d493e104a5b4902268217fee6d05a0f52c80b510565bf2c74f817019256

Observation 2ed16f7a-ccda-478b-8d2c-df0c86a3e86f · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T11:38:24.374520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:38:24.374520Z digest=sha256:8582da942296f28c45354d84996e5cdae62113eff136efd3f2267b2d3ba3fbac

Observation 26773ebe-9b77-45d7-b855-67048c8f51d4 · outbound

This paper cites Align your latents: High-resolution video synthesis with latent diffusion models.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution Align your latents: High-resolution video synthesis with latent diffusion models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T11:38:24.380521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:38:24.380521Z digest=sha256:6000f860362000f1b284ac5e1988b5972ae82d9f6b6e53678f183946362ed67d

Observation 9d1002c0-5fe6-4b7e-a249-fd6911a694ae · outbound

This paper cites Virtual KITTI 2.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution Virtual KITTI 2

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T11:38:24.385580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:38:24.385580Z digest=sha256:403105d56351cbcfbe1cd6bd307f23e4bb28e0320e4aeae4d8f5c480309831ed

Observation 66d9c406-3b10-48f5-acd4-d8e348de535a · outbound

This paper cites PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T11:38:24.390885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:38:24.390885Z digest=sha256:951e29d08c06405cf03800720d1d66a2bc9c6b033029d276406d9ba0af8d7ec8

Observation 6f09a469-9974-4ea3-aebb-cedbe3b45c51 · outbound

This paper cites A simple framework for contrastive learning of visual representations.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution A simple framework for contrastive learning of visual representations

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T11:38:24.395917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:38:24.395917Z digest=sha256:01ceed5383996832e941920e91b14fbbb4d39bfc90d68e78ad4c91e843c0a580

Observation 2f56a1d2-0998-4c99-9fb4-a4bf73db0424 · outbound

This paper cites Analog Bits: Generating Discrete Data using Diffusion Models with Self-Conditioning.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution Analog Bits: Generating Discrete Data using Diffusion Models with Self-Conditioning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T11:38:24.400914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:38:24.400914Z digest=sha256:0e86254b819fe550226a6b7c062bc65d3e01f3578f78d3d4ade1b695c54b5a40

Observation 38413d0b-8797-463c-a5e4-aba891fdc666 · outbound

This paper cites Scannet: Richly-annotated 3d reconstructions of indoor scenes.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution Scannet: Richly-annotated 3d reconstructions of indoor scenes

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T11:38:24.406134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:38:24.406134Z digest=sha256:0741fba10be521798cc55a8b71bd5c5af2e4336621f29a0fb31aacbf7264b5c0

Observation 2cac98d1-e65c-40e1-8fe4-c2d930b6f7f4 · outbound

This paper cites Emu: Enhancing Image Generation Models Using Photogenic Needles in a Haystack.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution Emu: Enhancing Image Generation Models Using Photogenic Needles in a Haystack

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T11:38:24.411056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:38:24.411056Z digest=sha256:78843ab595373b68ae9c3cee54350c59c01bbf3631449f71374cec45bf8d105e

Observation c011adf8-ee11-470e-96c0-f880d242f1d5 · outbound

This paper cites Augmented Bridge Matching.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution Augmented Bridge Matching

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T11:38:24.415999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:38:24.415999Z digest=sha256:24bd477d1c6e0b5055fa31f98df8769a70c4cbdd25448ed10369c61c4ee32f4b

Observation a20e5b31-94d9-47cd-b1b8-4037fd8be738 · outbound

This paper cites Imagenet: A large-scale hierarchical im- age database.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution Imagenet: A large-scale hierarchical im- age database

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T11:38:24.421143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:38:24.421143Z digest=sha256:6acba12a1535532de0db60af3d4e8971a8cb4631e36eae54d067d3f3b5e2f52c

Observation b3311588-09b5-49b9-a560-864323b2e177 · outbound

This paper cites Diffusion models beat gans on image synthesis.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution Diffusion models beat gans on image synthesis

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T11:38:24.426383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:38:24.426383Z digest=sha256:d432dfb8bbbd0850c816bbbcf12174ae307cffb0efc4e04675bc767c625cc7e8

Observation 599c39d4-76e6-48be-afcd-d793974593e8 · outbound

This paper cites DiffusionDepth: Diffusion Denoising Approach for Monocular Depth Estimation.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution DiffusionDepth: Diffusion Denoising Approach for Monocular Depth Estimation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T11:38:24.431295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:38:24.431295Z digest=sha256:538e01829f3e17b770aad860ca32446c7666df4ad68ae6338c7ce395acf835d9

Observation 51798627-2b5f-4313-bd1a-3e28b72d9095 · outbound

This paper cites The Llama 3 Herd of Models.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution The Llama 3 Herd of Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T11:38:24.436432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:38:24.436432Z digest=sha256:295754c9aa2d150b14dd863b7ef8bcd767ffbc220aeb41d61cbcdf652678e1cf

Observation 664d3fb6-577b-40c5-99f8-4c93b5312a22 · outbound

This paper cites Omnidata: A scalable pipeline for making multi- task mid-level vision datasets from 3d scans.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution Omnidata: A scalable pipeline for making multi- task mid-level vision datasets from 3d scans

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T11:38:24.441443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:38:24.441443Z digest=sha256:950305233e690bdad2ae5fe1e7da949743d8e9e50b767ce02e8e3d4f2917bb57

Observation 93975b8c-a80f-4db1-9955-a67f1c4dd374 · outbound

This paper cites Depth map prediction from a single image using a multi-scale deep network.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution Depth map prediction from a single image using a multi-scale deep network

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T11:38:24.446667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:38:24.446667Z digest=sha256:338c05b35f0b07894e73b2ddbaed5b51ae0e7803fc0394710eac57527b531570

Observation 9222fab0-c631-4e50-8195-f03cee0d9db6 · outbound

This paper cites Scaling rectified flow transformers for high-resolution image synthesis.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution Scaling rectified flow transformers for high-resolution image synthesis

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T11:38:24.451550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:38:24.451550Z digest=sha256:705f623480ac9ac7c551d08d8407cd94f871aa33a676e13c3b86c2cfd7006f87

Observation fec6e34e-85be-4738-a281-ca6731ddbc63 · outbound

This paper cites Boosting Latent Diffusion with Flow Matching.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution Boosting Latent Diffusion with Flow Matching

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T11:38:24.456477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:38:24.456477Z digest=sha256:3e26d059a7e63d8a9efe03c0a494caa6da91c2067412ab279fe8e343eb654bac

Observation e768f982-6014-4e79-8086-eb73b7620658 · outbound

This paper cites Masked Non-Autoregressive Image Captioning.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution Masked Non-Autoregressive Image Captioning

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-08-11T11:38:25.434491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T11:38:24.461686Z digest=sha256:7933de197b74c3579dfae23c3aac8767a78404789c26664128f575c600caf042

Observation b1870aad-a429-4814-8c34-97c52515f1e5 · outbound

This paper cites Vision meets robotics: The kitti dataset.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution Vision meets robotics: The kitti dataset

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T11:38:24.467095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:38:24.467095Z digest=sha256:b986c554114304c3b44c722ac2681d9a2f74e93d7408c679f398085bf846e0d3

Observation 27bda5f3-d62d-40bd-8eef-470a6ddc235c · outbound

This paper cites Geneval: An object-focused framework for evaluating text- to-image alignment.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution Geneval: An object-focused framework for evaluating text- to-image alignment

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T11:38:24.471843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:38:24.471843Z digest=sha256:fe931ab9f5515c916ddd3ec9b4e9114a843f93cc1ca931b0ab83288e50e46a7e

Observation d08671be-7965-4847-96e7-52a5dd099746 · outbound

This paper cites Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T11:38:24.477461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:38:24.477461Z digest=sha256:10c44b7afa282a7e6056fd4e53c020f7cb7d4ba389bebef3b1e7264e96c2ad6f

Observation 68188682-a165-472d-bd0f-8d05c149f92a · outbound

This paper cites Non-Autoregressive Image Captioning with Counterfactuals-Critical Multi-Agent Learning.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution Non-Autoregressive Image Captioning with Counterfactuals-Critical Multi-Agent Learning

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-08-11T11:38:25.363900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T11:38:24.482454Z digest=sha256:3eacfcd37fda59acaad7131dc37aa814a50c570710374ff7ee6175b9f14a9239

Observation 454665f3-e7b9-443a-9bfc-d2f3987d2a25 · outbound

This paper cites Itera- tive α-(de) blending: A minimalist deterministic diffusion model.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution Itera- tive α-(de) blending: A minimalist deterministic diffusion model

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T11:38:24.487644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:38:24.487644Z digest=sha256:5c48101d4f233def307ac6519fc317118d5842847e675f98ee1947cb60f081ae

Observation 7b8aa261-e2d8-4152-ac33-1e2a5d902e26 · outbound

This paper cites CLIPScore: A Reference-free Evaluation Metric for Image Captioning.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution CLIPScore: A Reference-free Evaluation Metric for Image Captioning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T11:38:24.492578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:38:24.492578Z digest=sha256:b6c33e9c229d53c248f27ef6682eaad2ffbb4dddce98bf1db478b12b11f1928a

Observation cf9ca671-cd4c-447a-8dc7-32d32354586d · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equi- librium.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution Gans trained by a two time-scale update rule converge to a local nash equi- librium

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T11:38:24.497935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:38:24.497935Z digest=sha256:2819c5bd9b8654350c05336d6e658465582969a4a4336ec8247c05d01535c476

Observation b1fe33d6-73e2-4358-a2ab-8f54d17c1886 · outbound

This paper cites Classifier-Free Diffusion Guidance.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution Classifier-Free Diffusion Guidance

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T11:38:24.505068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:38:24.505068Z digest=sha256:4e4a25dc6337f0e083e6bfa988f962035dae6fab34abefcf11598b287612d7ee

Observation 5269352c-03d7-4eeb-aae4-6bc9a3e05b4d · outbound

This paper cites Denoising dif- fusion probabilistic models.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution Denoising dif- fusion probabilistic models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T11:38:24.510215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:38:24.510215Z digest=sha256:e636144625235d8f1c47c6b9e8e3d901c080d2524a3f2afaa89011a69ddebd70

Observation 904461c0-79e8-48f5-9713-3a75af07872d · outbound

This paper cites Imagen Video: High Definition Video Generation with Diffusion Models.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution Imagen Video: High Definition Video Generation with Diffusion Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T11:38:24.515325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:38:24.515325Z digest=sha256:2a245085ffde7dee62f34e3856dfda918f5c247fc68ee42c969573201e87de11

Observation 40eacc0c-6359-46dc-9c27-bd4895ad68f8 · outbound

This paper cites Cascaded diffusion models for high fidelity image generation.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution Cascaded diffusion models for high fidelity image generation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T11:38:24.521095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:38:24.521095Z digest=sha256:55d529739c476e03d48f18371c6b342eb957faf9dc81556fe14171df8f414beb

Observation d1b26eaf-861b-4539-bb0c-18591ad63f33 · outbound

This paper cites Estimation of non- normalized statistical models by score matching.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution Estimation of non- normalized statistical models by score matching

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T11:38:24.526093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:38:24.526093Z digest=sha256:e61b219637dd2226ec3a970baed54499aba2e715d41768d16242471deafec9b3

Observation 35e176c5-8e95-4d2f-9e44-34eb062a2c6d · outbound

This paper cites Deep visual-semantic alignments for generating image descriptions.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution Deep visual-semantic alignments for generating image descriptions

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T11:38:24.530967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:38:24.530967Z digest=sha256:12669b810bf8fa66f9450acdcc60f6e7926421705127c11f22833f8192361e09

Observation ffe6d42d-86ae-49b0-a74a-f4504553b7b7 · outbound

This paper cites Guiding a Diffusion Model with a Bad Version of Itself.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution Guiding a Diffusion Model with a Bad Version of Itself

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T11:38:24.535892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:38:24.535892Z digest=sha256:da08b33662d2658f134a6e3ca720545634b0491958309eaf72e541bbe9448746

Observation d6ae7475-ca56-40cb-adba-cde0e9688f66 · outbound

This paper cites Re- purposing diffusion-based image generators for monocular depth estimation.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution Re- purposing diffusion-based image generators for monocular depth estimation

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:38:27.231712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T11:38:24.541093Z digest=sha256:409c86bbb2dfb302d77dddd3f090100ae82913a6e46eab5ef84cf306381091a8

Observation cff8e08b-29fd-4092-9584-9c197ce3e7c8 · outbound

This paper cites Auto-Encoding Variational Bayes.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution Auto-Encoding Variational Bayes

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T11:38:24.546193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:38:24.546193Z digest=sha256:1598543ff5c524a5c756ba8f2212e5c3eb6bb2a93769b1c288798828306a10bb

Observation dcc5cccf-ed3f-4410-bad7-e5e6d6ba620c · outbound

This paper cites DiffWave: A Versatile Diffusion Model for Audio Synthesis.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution DiffWave: A Versatile Diffusion Model for Audio Synthesis

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T11:38:24.551473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:38:24.551473Z digest=sha256:e302cbd97ff6a3458abf83e9fa974b719884c83d8ac5f2ba3f12e93f69c9fdcf

Observation 77be2ad6-6aaa-4fa4-a125-39ab24c7af53 · outbound

This paper cites On informa- tion and sufficiency.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution On informa- tion and sufficiency

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:38:27.181171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T11:38:24.556906Z digest=sha256:4c55920defa7c728977410eb534aa5697ae3a91da78d868ae967a7433a4c54c0

Observation de3a114e-eb56-4342-9519-00280ec9bc4e · outbound

This paper cites Deterministic Non-Autoregressive Neural Sequence Modeling by Iterative Refinement.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution Deterministic Non-Autoregressive Neural Sequence Modeling by Iterative Refinement

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T11:38:24.561912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:38:24.561912Z digest=sha256:547b4f664d90428e6893624819236262c0160b5f19afd6899c32a8bb809ccfab

Observation 623e7f79-4038-46cf-8521-4957b81c72fd · outbound

This paper cites Blip: Bootstrapping language-image pre-training for uni- fied vision-language understanding and generation.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution Blip: Bootstrapping language-image pre-training for uni- fied vision-language understanding and generation

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:38:27.082916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T11:38:24.567406Z digest=sha256:4d55ebd6ec7e0b7750b0e1d0ef1f84c085a5c60bae3191fb1383f17df48e1f82

Observation fc182948-8f9b-4600-b299-6c8fcdfa9152 · outbound

This paper cites Depthformer: Exploiting long-range correlation and local information for accurate monocular depth estimation.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution Depthformer: Exploiting long-range correlation and local information for accurate monocular depth estimation

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:38:27.019954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T11:38:24.572426Z digest=sha256:178996a99e0bba05d26f023c98d60214e493e92a1d734284256e4ceb227e9eb2

Observation a454539e-bff4-4412-a474-793676f8a822 · outbound

This paper cites Binsformer: Revisiting adaptive bins for monocular depth estimation.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution Binsformer: Revisiting adaptive bins for monocular depth estimation

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:38:27.000184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T11:38:24.577220Z digest=sha256:ac9a655b7f60788a2ca0c0b0476db5cb67874c52689f55039e4007c18ba1e00e

Observation 1a1aab9e-2292-4646-ad03-d4cf113c37d3 · outbound

This paper cites Magic3d: High- resolution text-to-3d content creation.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution Magic3d: High- resolution text-to-3d content creation

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:38:26.935863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T11:38:24.582082Z digest=sha256:5a38f00fe76ae6f865d94254a5aa89bb6769c3974d8111f399aac0140819945e

Observation d341cd6d-65a3-45ae-b290-0f949296b28b · outbound

This paper cites Rouge: A package for automatic evaluation of summaries.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution Rouge: A package for automatic evaluation of summaries

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:38:26.905647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T11:38:24.586806Z digest=sha256:f08cac7672138a973cf8df25d867c1e80f30b2be69acd6186811a54f9e2165d3

Observation aa121f1e-86cf-49c0-9aea-2cb517f43940 · outbound

This paper cites Microsoft coco: Common objects in context.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution Microsoft coco: Common objects in context

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:38:26.825712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T11:38:24.591896Z digest=sha256:f860c5e2b98381d31c29401872e979eb7fb1a7ffb8e1180e1ef9b92cbb7284a4

Observation b2af447a-88aa-438e-b5fc-d7376ee09c64 · outbound

This paper cites Flow matching for generative modeling.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution Flow matching for generative modeling

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:38:26.745891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T11:38:24.596818Z digest=sha256:3c1914054b9522904cfd204d0fd03972017722e6d596d51bb1d9f20d99871183

Observation 2923650c-0bdd-4637-8d46-3dcc0d7b7fde · outbound

This paper cites Generalized Schr\"odinger Bridge Matching.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution Generalized Schr\"odinger Bridge Matching

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T11:38:24.601431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:38:24.601431Z digest=sha256:c379336a8ead9c75b3a3aec3ac999c54d4e37f1b38e72c813150717cee0ef792

Observation a0b07357-c22d-4dd5-b3a5-f0946dd25bfd · outbound

This paper cites I$^2$SB: Image-to-Image Schr\"odinger Bridge.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution I$^2$SB: Image-to-Image Schr\"odinger Bridge

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T11:38:24.606458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:38:24.606458Z digest=sha256:7567d159f55919a0859d3d5aa6e9e52aa6e816726b3ebd0593c12f29f47e35c8

Observation 812e8610-048f-429a-abc7-6137930d26cf · outbound

This paper cites AudioLDM: Text-to-Audio Generation with Latent Diffusion Models.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T11:38:24.611756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:38:24.611756Z digest=sha256:83e428668da64a87a2b3d2e61b2305b3aeff33b3ba1edad33e8853333d86377b

Observation 835c6f35-20e4-4488-b8d4-dba03ca0c2b1 · outbound

This paper cites One-2-3-45: Any sin- gle image to 3d mesh in 45 seconds without per-shape opti- mization.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution One-2-3-45: Any sin- gle image to 3d mesh in 45 seconds without per-shape opti- mization

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:38:26.690883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T11:38:24.616712Z digest=sha256:85347b2497d420281b308d593965bb8a1fa6fd1e6c6129dd8773705774ac8440

Observation 321dffdd-c1d9-442f-8bfa-01fe94b20a95 · outbound

This paper cites Alleviating Distortion in Image Generation via Multi-Resolution Diffusion Models and Time-Dependent Layer Normalization.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution Alleviating Distortion in Image Generation via Multi-Resolution Diffusion Models and Time-Dependent Layer Normalization

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T11:38:24.621341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:38:24.621341Z digest=sha256:af754cd7b4d05cc5f0798c3d70c94dd48fe937e07b625a59d464a6a2d13cc5b2

Observation 4b617335-9ec3-423d-a786-878056005954 · outbound

This paper cites Direct-3d: Learning direct text-to-3d gen- eration on massive noisy 3d data.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution Direct-3d: Learning direct text-to-3d gen- eration on massive noisy 3d data

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:38:26.611347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T11:38:24.626235Z digest=sha256:dbb7ab6e48980bd4f32ebe758090d959b7f5e485f926f7602068bbb667394645

Observation ef6ec8cf-e611-4bc2-a642-ae739ccae802 · outbound

This paper cites Flow straight and fast: Learning to generate and transfer data with rectified flow.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution Flow straight and fast: Learning to generate and transfer data with rectified flow

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:38:26.575060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T11:38:24.631203Z digest=sha256:d71f7eb2216dd2c981c3c0406ba83135adcea5a836891750191a9b159b07c219

Observation 927889f1-d8b6-43e8-abd7-dbdc8027fac8 · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution Swin transformer: Hierarchical vision transformer using shifted windows

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:38:26.557227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T11:38:24.635717Z digest=sha256:9ac2af963bfd19e95a6f61cd677d1eeb28a834aadcfecd7ca725a5fc1c096ecb

Observation 04824177-26b4-4ec3-abc1-6dc2df82fb8c · outbound

This paper cites Decoupled Weight Decay Regularization.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution Decoupled Weight Decay Regularization

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T11:38:24.640231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:38:24.640231Z digest=sha256:606d1a07bc5b5bb9ed69222f78f16d76d28b04cc9563940744d21fa8b5d11d82

Observation 315faa71-373e-4f38-97bb-2cbf0888416b · outbound

This paper cites Semantic-conditional diffu- sion networks for image captioning.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution Semantic-conditional diffu- sion networks for image captioning

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:38:26.538739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T11:38:24.644961Z digest=sha256:41049d7f3db5a9a986de93878d6061e4fb8007f033f7acfc11f85c8baeb1e96b

Observation 0256a1f6-d609-4304-aa0b-5fa077f86c2f · outbound

This paper cites Efficient Estimation of Word Representations in Vector Space.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution Efficient Estimation of Word Representations in Vector Space

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T11:38:24.649569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:38:24.649569Z digest=sha256:2dadc6cad4b3274c03b5e418f5111091cdb70b56ef89d30272ccb7aea13f81ac

Observation 32b8b15c-ce59-4417-b938-a10624f0947a · outbound

This paper cites GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T11:38:24.654718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:38:24.654718Z digest=sha256:8b2ae0de9c28d9522a0061fca3e1c2997cb107b88400190d0f9468bbcd6319b0

Observation 89a5d589-c5f0-4b85-935d-55a4ae6af2e1 · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution Representation Learning with Contrastive Predictive Coding

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-11T11:38:24.668281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:38:24.668281Z digest=sha256:0ac6f7d604d1568d963adfc72ad1a331548b5dbd38585db59698a1029f3cbd98

Observation 3b155c6b-0c3b-4525-8ea3-439e4e7ccc09 · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution Bleu: a method for automatic evaluation of machine translation

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:38:26.446857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T11:38:24.673374Z digest=sha256:542e9fb3dbf1a19d9e4ce4c15a3d4d68909ddf41d3d6bd9ad269ac13d18d52ac

Observation a12bc03f-d324-4e60-8b72-a4c5dd7258a7 · outbound

This paper cites Scalable diffusion mod- els with transformers.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution Scalable diffusion mod- els with transformers

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:38:26.365716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T11:38:24.678757Z digest=sha256:79b321dc963ab74a104241230e823355a6e45500dc2cd93e981d5aa8aa5ff9a0

Observation 54a5c6c8-a1cc-487d-b990-124b2292f91a · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-11T11:38:24.683392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:38:24.683392Z digest=sha256:5ebe0edbe7897a60d254a5284a0582abd2e3c508cb0260e318bc2ed04e7f4a9f

Observation b18bdcb4-870a-44db-88d0-c39dc7846247 · outbound

This paper cites Movie Gen: A Cast of Media Foundation Models.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution Movie Gen: A Cast of Media Foundation Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-11T11:38:24.688775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:38:24.688775Z digest=sha256:369f79b2810c8c179058f999a1bba4b0953398aa0e4e2cd40986849224769c80

Observation c15a7177-4620-4928-8b9d-c0cf646aa486 · outbound

This paper cites DreamFusion: Text-to-3D using 2D Diffusion.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution DreamFusion: Text-to-3D using 2D Diffusion

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-11T11:38:24.694914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:38:24.694914Z digest=sha256:fadd9574d2c4a224cd45a9aa0243e0e71472492dbdd54976898f1486a0b6769f

Observation ab10e947-585d-42d9-8138-b23976e2e1fd · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution Learn- ing transferable visual models from natural language super- vision

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:38:26.294938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T11:38:24.699829Z digest=sha256:285c3282b06e77366035c383a5338e69c24055149389810dda3a2109a1d9f0bc

Observation a8d61850-790e-4699-9da7-b24908f97c9e · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:38:26.272982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T11:38:24.705465Z digest=sha256:21bac0d8cac2d93070484c4d4ef076ddf0ade0da8b32682e9de0b7e55581a83d

Observation 6e9e38e0-f72f-47a3-87a3-a24b544c5970 · outbound

This paper cites Zero-shot text-to-image generation.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution Zero-shot text-to-image generation

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-11T11:38:24.710430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:38:24.710430Z digest=sha256:0551b09b4affe39e22791c2c604e9508b64ea220dbf88cf2e78ef6ebc93c51d2

Observation 8533bd67-8aa2-4980-9a5a-248b28fc9ef3 · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-11T11:38:24.715900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:38:24.715900Z digest=sha256:7398620cc383ecc7db6643194db4fff1663994b85541ab8d53ec55a8d8713810

Observation df2b7584-42a1-4bfd-b062-a872340291b9 · outbound

This paper cites Towards robust monocu- lar depth estimation: Mixing datasets for zero-shot cross- dataset transfer.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution Towards robust monocu- lar depth estimation: Mixing datasets for zero-shot cross- dataset transfer

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:38:26.235394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T11:38:24.721340Z digest=sha256:e1075cb3b873995564fef68e42b1633844a548935f1e257e09d9562e084ea287

Observation 3e51569b-60dd-4686-a8bb-b51507e0acb5 · outbound

This paper cites Vi- sion transformers for dense prediction.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution Vi- sion transformers for dense prediction

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:38:26.219654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T11:38:24.726433Z digest=sha256:aa977ce3263a4b70594c785f0c825eb5b7c306aa0dce8c1327ba39ac5fdb1e62

Observation 15e99f82-e625-4053-a5eb-c6b19719311a · outbound

This paper cites Hypersim: A photorealistic synthetic dataset for holistic indoor scene understanding.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution Hypersim: A photorealistic synthetic dataset for holistic indoor scene understanding

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:38:26.202106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T11:38:24.731625Z digest=sha256:be374edaf3c1eb2b26503ff0402b0949c6f47d400770593707b3035ddec211ce

Observation dbf73d05-a5b0-4baf-8693-e4e90f0bf986 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution High-resolution image synthesis with latent diffusion models

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:38:26.183955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T11:38:24.736881Z digest=sha256:c801effdb28f7dc2c8cf50664d5deb2cf84d27f65e001e1f063e30af0859bd4d

Observation e23db68c-0e50-4439-b001-ff52cf29a743 · outbound

This paper cites Photorealistic text-to-image diffusion models with deep language understanding.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution Photorealistic text-to-image diffusion models with deep language understanding

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:38:26.146434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T11:38:24.741724Z digest=sha256:1307ed7e4ad8f1129675e7419b34f11d179a0b9a15747d1fb1adee203b6a0284

Observation 1d2a88ab-eec0-4525-a706-ab8f24d18c53 · outbound

This paper cites Image super-resolution via iterative refinement.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution Image super-resolution via iterative refinement

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:38:26.087306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T11:38:24.746492Z digest=sha256:0d0b1372192182762c5de91c41b40df252eac841cd5addc0a5a233a00dda37b0

Observation 95b0e846-7ab3-4040-8233-402d1e254d80 · outbound

This paper cites A multi-view stereo benchmark with high-resolution images and multi-camera videos.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution A multi-view stereo benchmark with high-resolution images and multi-camera videos

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:38:26.025917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T11:38:24.751269Z digest=sha256:0ec7e4ee51c1ee38134372240c02f07315144fe66b22a8c40fee6184d181fe11

Observation 823f2e7e-25df-4f3a-885e-7bb6fd708b3f · outbound

This paper cites Diffusion schr¨odinger bridge matching.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution Diffusion schr¨odinger bridge matching

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:38:26.007530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T11:38:24.756551Z digest=sha256:5c37c041dea79434242e49b5486912164461d1087f36f08228b95e1c858443d9

Observation 5df961b1-ee8e-48c8-a3ce-55e616ec9800 · outbound

This paper cites Indoor segmentation and support inference from rgbd images.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution Indoor segmentation and support inference from rgbd images

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:38:25.989103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T11:38:24.761855Z digest=sha256:27ca779576e1b17480bf59cd1e55f7b12ac4bfd58409ee5b0050cd4212055411

Observation 7f16c4ad-2073-4eee-a700-bd737eace2ac · outbound

This paper cites Make-A-Video: Text-to-Video Generation without Text-Video Data.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution Make-A-Video: Text-to-Video Generation without Text-Video Data

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-11T11:38:24.766536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:38:24.766536Z digest=sha256:0ee328451153c7cb8e27bfe037bd683606507966b1a66bda6890bfaeab062140

Observation 3642203a-2cbb-4b89-b89e-f7b3725a8ba3 · outbound

This paper cites Deep unsupervised learning using nonequilibrium thermodynamics.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution Deep unsupervised learning using nonequilibrium thermodynamics

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-11T11:38:24.771386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:38:24.771386Z digest=sha256:6fa541d388174b1785de7def6305f685949eab3504a9ba53d80f2f9b685750e4

Observation 88e18801-2dd1-4499-8043-8e27ff2bc0a5 · outbound

This paper cites Generative modeling by estimating gradients of the data distribution.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution Generative modeling by estimating gradients of the data distribution

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:38:25.958130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T11:38:24.776146Z digest=sha256:e7ef3233a2dc9fc8bfda73d1aa54bd8a60d34e774dc9904aeafb338d600617ab

Observation fb1436ba-f435-44a6-9a86-b88626b238c8 · outbound

This paper cites Score- based generative modeling through stochastic differential equations.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution Score- based generative modeling through stochastic differential equations

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:38:25.940540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T11:38:24.781086Z digest=sha256:08d070c3cdc43169a50eea1c502850ad01709b16466dc582aec6351cb36f337d

Observation 419b73e6-a504-42c6-b44d-b237061b9551 · outbound

This paper cites Simplified Diffusion Schr\"odinger Bridge.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution Simplified Diffusion Schr\"odinger Bridge

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-11T11:38:24.786272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:38:24.786272Z digest=sha256:5e9ca8951530f5cb8eb3cfcbe431c84c394f3f328151fc6e01a4dd3f38cb3082

Observation ea3ea140-c91c-44a3-8512-863096c0be7d · outbound

This paper cites Simulation-free Schr\"odinger bridges via score and flow matching.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution Simulation-free Schr\"odinger bridges via score and flow matching

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-11T11:38:24.791769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:38:24.791769Z digest=sha256:698b9a9e59ccef7bdf13740595a2c5c8403b121baad6048bbc7ec61f26266ab5

Observation 2e03979d-5c27-4186-ad13-194b758e158b · outbound

This paper cites DIODE: A Dense Indoor and Outdoor DEpth Dataset.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution DIODE: A Dense Indoor and Outdoor DEpth Dataset

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-11T11:38:24.797704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:38:24.797704Z digest=sha256:96e1d68f5d8d42d5bf7ca004d5a359bf4cf8dbe7778deb08984505089107f6b0

Observation 8cb2929d-3c2c-4642-8d6b-047cb1a5a115 · outbound

This paper cites Attention is all you need.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution Attention is all you need

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:38:25.924116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T11:38:24.803597Z digest=sha256:df483ce1bdd6ec1929233aa0e1c6f064bb54805921c4cf8f9445153fce5b79a5

Observation 69ea2cca-8813-47bb-85af-d9f925f62c10 · outbound

This paper cites Cider: Consensus-based image description evalu- ation.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution Cider: Consensus-based image description evalu- ation

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:38:25.907788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T11:38:24.808441Z digest=sha256:1337bd048f42f6e16d6d669e6f70fbb5b17ba1495a79d84424fcc898f81d3b88

Observation 279b5a4c-03d9-4926-9f7b-cb9cbc8241e3 · outbound

This paper cites Raphael: Text-to- image generation via large mixture of diffusion paths.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution Raphael: Text-to- image generation via large mixture of diffusion paths

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:38:25.859071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T11:38:24.813786Z digest=sha256:e0b63955f0f0fad1a8af585760acb78201f0c12548dee5b1d4e58abcfe0b82e8

Observation dc822a03-7289-4a90-97c0-9321bcaba714 · outbound

This paper cites Transformer-based attention networks for con- tinuous pixel-wise prediction.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution Transformer-based attention networks for con- tinuous pixel-wise prediction

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:38:25.805563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T11:38:24.819080Z digest=sha256:5c880dad8b003f6ab41287bb28b5cf0b1fa30c7f3eca3e2756517c0026bcf2bb

Observation 1fa78ac5-c240-4d55-bd14-dfafa557a640 · outbound

This paper cites Depth anything: Un- leashing the power of large-scale unlabeled data.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution Depth anything: Un- leashing the power of large-scale unlabeled data

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-11T11:38:24.824360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:38:24.824360Z digest=sha256:0b3af49af25633205517c8601b999783169fd4d992fe4d973c8095780810675f

Observation e34d1de4-3209-4315-b037-4ab941705781 · outbound

This paper cites DiverseDepth: Affine-invariant Depth Prediction Using Diverse Data.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution DiverseDepth: Affine-invariant Depth Prediction Using Diverse Data

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-11T11:38:24.829998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:38:24.829998Z digest=sha256:9043e2429fb92c49e9f665a6d1827ec934326f23f3a612dcdd78f22dadf37a5d

Observation abf79ddf-1464-4d4e-8691-5e96135421a3 · outbound

This paper cites Learning to re- cover 3d scene shape from a single image.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution Learning to re- cover 3d scene shape from a single image

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:38:25.767038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T11:38:24.835494Z digest=sha256:e2fe101b5519d115f011ee1dbe19a3acb2c73977590e445d7bff63d0455816d9

Observation f30f524f-6022-4d28-aa4c-961f161a5ba9 · outbound

This paper cites Image captioning with semantic attention.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution Image captioning with semantic attention

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:38:25.742308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T11:38:24.841065Z digest=sha256:2467387afa5083a4d1b3d9a2857036e125cfba49bc05ef8b32c9260980dcd719

Observation 2bc72218-1650-473b-b085-efa2c254e99f · outbound

This paper cites Hierarchical normalization for robust monocular depth estimation.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution Hierarchical normalization for robust monocular depth estimation

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:38:25.720329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T11:38:24.846646Z digest=sha256:5a0814f08ab2c61ab4045a41847d82481e3439d372541c6726d76f6039d14c92

Observation a7b66d6a-538b-4d94-9439-af4ff69d4080 · outbound

This paper cites Denoising Diffusion Bridge Models.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution Denoising Diffusion Bridge Models

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-11T11:38:24.852458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:38:24.852458Z digest=sha256:ea0c1c00ef2cf1879f223baa0420cfd7190ae9d09e78421dc087b9c74e9b3acd

Observation 4623e797-82c7-4cda-83b5-126e6adc3829 · outbound

This paper cites Semi-autoregressive transformer for image captioning.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution Semi-autoregressive transformer for image captioning

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:38:25.696150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T11:38:24.857582Z digest=sha256:877dfac11bb0afd82eeff625ca6ca1e917e7c4103397ee317f0800521282aade

Pith citing papers

Observation 141762af-5843-40db-b580-5ceca126fcb2 · inbound

CellFlux: Simulating Cellular Morphology Changes via Flow Matching cites this paper.

CellFlux: Simulating Cellular Morphology Changes via Flow Matching Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T20:40:31.655763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T20:40:31.655763Z digest=sha256:52893efa4171392ee76b1f14c8da03ccbc6dbc76445acdc815d151fe581a6dd2

Observation 04d0ef4c-bfd5-4a6e-9111-2a2da7e92e8e · inbound

CellFluxRL: Biologically-Constrained Virtual Cell Modeling via Reinforcement Learning cites this paper.

CellFluxRL: Biologically-Constrained Virtual Cell Modeling via Reinforcement Learning Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:08:25.851113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T01:05:45.781048Z digest=sha256:ea5408a4b0735cf09bb27d5ee916e0ece47eb36329858b1cfd16374460265894

Observation 340fdcd4-0ad8-409c-acbd-9929ad1ff7ee · inbound

CellFluxRL: Biologically-Constrained Virtual Cell Modeling via Reinforcement Learning cites this paper.

CellFluxRL: Biologically-Constrained Virtual Cell Modeling via Reinforcement Learning Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-22T10:44:47.559518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T10:41:37.140709Z digest=sha256:3120f3481b97e2282c86e26bafddf496f0fd20a2dab74e2beab9ad21f12e4003

Observation 37d13483-8c03-4dcb-b348-7aef11f2072f · inbound

Structured Diffusion Bridges: Inductive Bias for Denoising Diffusion Bridges cites this paper.

Structured Diffusion Bridges: Inductive Bias for Denoising Diffusion Bridges Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution

Reference 71

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T10:46:01.243621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-10T15:21:02.100545Z digest=sha256:924a3f40224e11166c066b0715724eaef142f61bbe0269c793247c6fd5e153bd

Observation ac5d18ce-5e2e-4d04-a4e3-bf6ece9f341f · inbound

Structured Diffusion Bridges: Inductive Bias for Denoising Diffusion Bridges cites this paper.

Structured Diffusion Bridges: Inductive Bias for Denoising Diffusion Bridges Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution

Reference 71

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T07:12:28.651404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-13T07:08:43.088643Z digest=sha256:9a113b43cf73ffa402ddaae2ce7087eacba88dec7956efb5962e124293d8a18f