Pith. sign in

Paper Citation Record · LEDGER

Inference-time Scaling for Diffusion-based Audio Super-resolution

As of 7 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 1 inbound Pith citation observation for arXiv:2508.02391.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.02391 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T05:05:29.628800Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T07:49:20.194990Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T05:56:40.912157Z

Reference resolution

47 of 47 outbound references displayed

  • verified exact0
  • verified fuzzy12
  • unresolved35
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 48db7d8e-d0e6-4105-95d2-2a4543c366b7 · outbound

This paper cites MusicLM: Generating Music From Text.

Inference-time Scaling for Diffusion-based Audio Super-resolution MusicLM: Generating Music From Text

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T05:05:25.719480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:05:25.719480Z digest=sha256:fa7dd1e383a3c0b46e33be219ca45db44e58227eb95c5f07b95562ff9a93312c

Observation c06cd0e7-530e-43a7-9125-70cbc86b3060 · outbound

This paper cites Large Language Monkeys: Scaling Inference Compute with Repeated Sampling.

Inference-time Scaling for Diffusion-based Audio Super-resolution Large Language Monkeys: Scaling Inference Compute with Repeated Sampling

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T05:05:25.801277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:05:25.801277Z digest=sha256:d4c837501571cc285806c8e8ad09f51e79ac60c436598080631ee5a548eeb776

Observation 3673279d-2d10-4603-b4dc-35e525b8aeb8 · outbound

This paper cites Musicldm: Enhancing novelty in text-to-music generation using beat-synchronous mixup strategies.

Inference-time Scaling for Diffusion-based Audio Super-resolution Musicldm: Enhancing novelty in text-to-music generation using beat-synchronous mixup strategies

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:05:30.548978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:05:25.899510Z digest=sha256:d6754e6d2a696259f3ce91f49cee96ee8b75b1e3a3c722dfb075d9d3f6bfe13e

Observation 5c9013d8-5612-40b4-bc5f-5524d676e8f3 · outbound

This paper cites Wavlm: Large-scale self-supervised pre- training for full stack speech processing.IEEE Journal of Selected Topics in Signal Processing, 16(6):1505–1518, 2022.

Inference-time Scaling for Diffusion-based Audio Super-resolution Wavlm: Large-scale self-supervised pre- training for full stack speech processing.IEEE Journal of Selected Topics in Signal Processing, 16(6):1505–1518, 2022

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T05:05:26.001366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:05:26.001366Z digest=sha256:afb6133251d33de06b05830a31ed6159522c6b242d6b46afe8025ffda9b1d82e

Observation 95d9772d-68ba-4190-8073-53d2ae2d242e · outbound

This paper cites Qwen2-Audio Technical Report.

Inference-time Scaling for Diffusion-based Audio Super-resolution Qwen2-Audio Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T05:05:26.071505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:05:26.071505Z digest=sha256:a388be580a8c28ec0c305f95e5f01be41307b97f40b5790a6c17e78de34bb9c7

Observation 518c59ab-ee3e-47fc-94d1-251728d6e871 · outbound

This paper cites Directly Fine-Tuning Diffusion Models on Differentiable Rewards.

Inference-time Scaling for Diffusion-based Audio Super-resolution Directly Fine-Tuning Diffusion Models on Differentiable Rewards

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T05:05:26.135354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:05:26.135354Z digest=sha256:ac901a597e4efead104b243b36b524c466a4c5bd185ba1c1d94f6d560e778bd3

Observation 2748de54-886e-422f-9f23-377135b90ad9 · outbound

This paper cites Clap learning audio concepts from natural language supervision.

Inference-time Scaling for Diffusion-based Audio Super-resolution Clap learning audio concepts from natural language supervision

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T05:05:26.243912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:05:26.243912Z digest=sha256:09e0f21f09a764ed5f361092b5c889430be2633ce1958bcc37ea8a06c066a41f

Observation af261e11-cd76-4626-a8df-d031500f5e34 · outbound

This paper cites FunASR: A Fundamental End-to-End Speech Recognition Toolkit.

Inference-time Scaling for Diffusion-based Audio Super-resolution FunASR: A Fundamental End-to-End Speech Recognition Toolkit

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T05:05:26.342534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:05:26.342534Z digest=sha256:af9f76b11899c46e4a0ccefb40c662938dccff1adbf13abc31bbcf665f64e47a

Observation 152c3e68-9ab9-4f09-9092-41f59eebed06 · outbound

This paper cites AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning.

Inference-time Scaling for Diffusion-based Audio Super-resolution AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T05:05:26.435388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:05:26.435388Z digest=sha256:8996070c9be9a0c43951f638ad38d78407ec7e96a0e20db81c2806f07ab42c00

Observation ffaedfe7-3c15-4547-98b4-6cd388f1a795 · outbound

This paper cites NU-Wave 2: A General Neural Audio Upsampling Model for Various Sampling Rates.

Inference-time Scaling for Diffusion-based Audio Super-resolution NU-Wave 2: A General Neural Audio Upsampling Model for Various Sampling Rates

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T05:05:26.516349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:05:26.516349Z digest=sha256:ff67904d224c24a792bb14fea18155a0f33eb84a5f1b23d383ab46251ad1a68d

Observation 0c62d301-5dc5-4ce7-b698-6ed0dfdbf324 · outbound

This paper cites Denoising diffusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020.

Inference-time Scaling for Diffusion-based Audio Super-resolution Denoising diffusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T05:05:26.599483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:05:26.599483Z digest=sha256:07f3e924c9c396a513a63d13c9fc8a1baf288f397a63e8a71a18ed6c0ff31636

Observation e50f301c-ad54-4c9a-bd2a-4f9ae10809a1 · outbound

This paper cites Classifier-Free Diffusion Guidance.

Inference-time Scaling for Diffusion-based Audio Super-resolution Classifier-Free Diffusion Guidance

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T05:05:26.697388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:05:26.697388Z digest=sha256:303b6ad2b8134609e0bdb4d7d10aea188fb1f984a6985d9981d9277d78146628

Observation 023b804b-cf22-47f1-9408-d224b0a9f3f5 · outbound

This paper cites Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis.Advances in neural information processing systems, 33:17022–17033, 2020.

Inference-time Scaling for Diffusion-based Audio Super-resolution Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis.Advances in neural information processing systems, 33:17022–17033, 2020

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T05:05:26.799798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:05:26.799798Z digest=sha256:88d69f773eef7af82f4e44194e444f34c8981743b5a404421875a79fc132d73c

Observation 014c09e7-7854-42f0-aa36-8d28a67ce2ca · outbound

This paper cites NU-Wave: A Diffusion Probabilistic Model for Neural Audio Upsampling.

Inference-time Scaling for Diffusion-based Audio Super-resolution NU-Wave: A Diffusion Probabilistic Model for Neural Audio Upsampling

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T05:05:26.893661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:05:26.893661Z digest=sha256:877fe072146916341ef975e6cfdacff657db3d9edc231db2d0aef6a019ccc6aa

Observation 827cfac1-42dc-47e1-b3ff-6d89c3876da1 · outbound

This paper cites Noise-free optimization in early training steps for image super- resolution.

Inference-time Scaling for Diffusion-based Audio Super-resolution Noise-free optimization in early training steps for image super- resolution

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:05:30.513791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:05:27.012398Z digest=sha256:c9ad8692d73cdee6656b9b4d93fe1b943c547a40e2273f967ff37ef4e6f3d698

Observation ab16529c-e383-45ff-aaf6-4716bb46018c · outbound

This paper cites Audiosr: Versatile audio super-resolution at scale.

Inference-time Scaling for Diffusion-based Audio Super-resolution Audiosr: Versatile audio super-resolution at scale

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:05:30.503268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:05:27.067541Z digest=sha256:71731a93b7a4166cd7350c3316697b1824464b1d1620da69c336012f7280393f

Observation e0fee5c5-370f-447c-94d1-65406a061189 · outbound

This paper cites AudioLDM: Text-to-Audio Generation with Latent Diffusion Models.

Inference-time Scaling for Diffusion-based Audio Super-resolution AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T05:05:27.134039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:05:27.134039Z digest=sha256:37aa0637dfb357878b3eefad7275b875dd715999c63834c45d70242a466cc992

Observation b2bc81cf-3b27-4e90-a3f2-6b936bfc3ae5 · outbound

This paper cites Neural Vocoder is All You Need for Speech Super-resolution.

Inference-time Scaling for Diffusion-based Audio Super-resolution Neural Vocoder is All You Need for Speech Super-resolution

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T05:05:27.200465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:05:27.200465Z digest=sha256:1a9dd22e3c5b8559cd4861d9abd88b7c3553f738fae45478dfa303b8d042586b

Observation 3e068ebb-6c8c-4317-93d5-a1dc6e83723c · outbound

This paper cites VoiceFixer: Toward General Speech Restoration with Neural Vocoder.

Inference-time Scaling for Diffusion-based Audio Super-resolution VoiceFixer: Toward General Speech Restoration with Neural Vocoder

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T05:05:27.287407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:05:27.287407Z digest=sha256:a5caa612c50faf7a90eab5530dd6b7f71fa46579af12633a43df67e3c459412c

Observation d29ae95f-9b46-41b3-9a01-5effb4bc06e2 · outbound

This paper cites Dpm-solver: A fast ode solver for diffusion probabilistic model sampling in around 10 steps.Advances in Neural Information Processing Systems, 35:5775–5787, 2022.

Inference-time Scaling for Diffusion-based Audio Super-resolution Dpm-solver: A fast ode solver for diffusion probabilistic model sampling in around 10 steps.Advances in Neural Information Processing Systems, 35:5775–5787, 2022

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T05:05:27.368077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:05:27.368077Z digest=sha256:2268b63a1d7b7eded03cb7bd74f34606b12e1b9d73b17fbd3f854ac0108f6401

Observation f11eca33-6ed1-4d3b-91e2-a8800b66484e · outbound

This paper cites Scaling inference time compute for diffusion models.

Inference-time Scaling for Diffusion-based Audio Super-resolution Scaling inference time compute for diffusion models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T05:05:27.442554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:05:27.442554Z digest=sha256:5edddbca4b87397315c1624dbd685e1a696b80f173eccae67da87d64354323b2

Observation d014c32a-f416-493b-9aee-f7fc60768fa7 · outbound

This paper cites Uncertainty-driven loss for single image super-resolution.Advances in Neural Information Processing Systems, 34:16398–16409, 2021.

Inference-time Scaling for Diffusion-based Audio Super-resolution Uncertainty-driven loss for single image super-resolution.Advances in Neural Information Processing Systems, 34:16398–16409, 2021

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:05:30.481703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:05:27.503553Z digest=sha256:f0d7f85145968b3bab73a7460e88078fc0e33f972f3fc0fc0ceff8bce21de838

Observation 44414bcc-ba4b-4cb4-8992-01c2a6a7a066 · outbound

This paper cites The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models.

Inference-time Scaling for Diffusion-based Audio Super-resolution The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T05:05:27.596743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:05:27.596743Z digest=sha256:96989b2a0ac2e6c58ff09588b37b5043ce21f3ce1f68dfad97519a481301a3d7

Observation 1dd6f9d9-f804-468b-800a-17064386a1f9 · outbound

This paper cites Esc: Dataset for environmental sound classification.

Inference-time Scaling for Diffusion-based Audio Super-resolution Esc: Dataset for environmental sound classification

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:05:30.471212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:05:27.686704Z digest=sha256:0dc5fda132827150aca7d73970c465f995ea2249632f3bf150631280a7c135dc

Observation 7474cf9f-2393-45cf-b483-29ddd7c48efa · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

Inference-time Scaling for Diffusion-based Audio Super-resolution Robust speech recognition via large-scale weak supervision

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T05:05:27.768071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:05:27.768071Z digest=sha256:ad0dc56385b10530ffe9cf520c97731ebbc62f38046d124963695fa82b86c59b

Observation 68004d0d-dddb-45cc-b8d5-53f7c625913f · outbound

This paper cites FastSpeech 2: Fast and High-Quality End-to-End Text to Speech.

Inference-time Scaling for Diffusion-based Audio Super-resolution FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T05:05:27.864867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:05:27.864867Z digest=sha256:cae00b1a7590b5b7ca2a0d13265f1000797fbc90f475e39611007417588a62b8

Observation fd7b94e0-ed22-4fe2-941e-7ac4b062b715 · outbound

This paper cites Fastspeech: Fast, robust and controllable text to speech.Advances in neural information processing systems, 32, 2019.

Inference-time Scaling for Diffusion-based Audio Super-resolution Fastspeech: Fast, robust and controllable text to speech.Advances in neural information processing systems, 32, 2019

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:05:30.453031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:05:27.978066Z digest=sha256:727a70d4179d8c55c9b6268701ab1525e5b26e3bf261af3a78ac0814f9b9d864

Observation 5bfd7618-59e3-4446-b30a-2a323305096f · outbound

This paper cites High- resolution image synthesis with latent diffusion models.

Inference-time Scaling for Diffusion-based Audio Super-resolution High- resolution image synthesis with latent diffusion models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T05:05:28.080199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:05:28.080199Z digest=sha256:529503de1a456424f3456cc2c6dfc1ebb181aa65cc947c918fa4b52ad3370492

Observation e8139e27-ff8e-4e4d-86e3-ba24bdd26ffb · outbound

This paper cites NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers.

Inference-time Scaling for Diffusion-based Audio Super-resolution NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T05:05:28.156213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:05:28.156213Z digest=sha256:373d27cddd0665dd370217dc792dacbee36400361a085490459ce006363b910f

Observation da5f8c3c-c005-4d6c-a313-048a8742480e · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Inference-time Scaling for Diffusion-based Audio Super-resolution Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T05:05:28.206753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:05:28.206753Z digest=sha256:47057aa4ddec696557151af56c0f6d5f8d88b2fd2647f13b32b25c411dae5874

Observation 1bad2a83-4fc0-408d-a093-cf9fda9f040e · outbound

This paper cites Denoising Diffusion Implicit Models.

Inference-time Scaling for Diffusion-based Audio Super-resolution Denoising Diffusion Implicit Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T05:05:28.271130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:05:28.271130Z digest=sha256:a210a16e5440e3c1b65ce6480b6d6f4adfc0f5a3bbb9eb0ef72fc752d3ecfea0

Observation 421a2ab1-1d8c-48ed-b174-a78fcbeacf4c · outbound

This paper cites Score-Based Generative Modeling through Stochastic Differential Equations.

Inference-time Scaling for Diffusion-based Audio Super-resolution Score-Based Generative Modeling through Stochastic Differential Equations

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T05:05:28.360497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:05:28.360497Z digest=sha256:c55f85b02aee02c3242925b9ae060ad3497b13811995552370b5ccdfd86c6812

Observation e073ebd2-2993-4a1a-ac33-63b1d0d1c62b · outbound

This paper cites AudioX: A Unified Framework for Anything-to-Audio Generation.

Inference-time Scaling for Diffusion-based Audio Super-resolution AudioX: A Unified Framework for Anything-to-Audio Generation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T05:05:28.435197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:05:28.435197Z digest=sha256:b1feb6fe834290aa9554643e1cc1f6747dfc2c4454538e828911c2b1470e8d20

Observation 5df9d974-cdcf-4a0d-a221-71806c1dde49 · outbound

This paper cites Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound.

Inference-time Scaling for Diffusion-based Audio Super-resolution Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T05:05:28.514824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:05:28.514824Z digest=sha256:4acfb4d5e95d8bc8f713765f4994d31999844523d12947601c6575dfe2250122

Observation c40a1b6a-db08-4293-b150-fb6fbc2bd88b · outbound

This paper cites Towards robust speech super-resolution.IEEE/ACM transactions on audio, speech, and language processing, 29:2058–2066, 2021.

Inference-time Scaling for Diffusion-based Audio Super-resolution Towards robust speech super-resolution.IEEE/ACM transactions on audio, speech, and language processing, 29:2058–2066, 2021

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:05:30.435338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:05:28.592826Z digest=sha256:aee842fdaef48bdea21be001149a9d0fb9d7a318a4726299d3552c4a0393bd10

Observation fee83341-ce77-442d-9a4f-0cef3b860c8f · outbound

This paper cites Esrgan: Enhanced super-resolution generative adversarial networks.

Inference-time Scaling for Diffusion-based Audio Super-resolution Esrgan: Enhanced super-resolution generative adversarial networks

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:05:30.423997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:05:28.662941Z digest=sha256:e3aeb96b29e2d6436ba4a80721af5cfaf0bca7feb760b34a0455d1a76ca27d9a

Observation f7a061b5-c1ea-4829-a27e-8209352f00f1 · outbound

This paper cites Difix3d+: Improving 3d reconstructions with single-step diffusion models.

Inference-time Scaling for Diffusion-based Audio Super-resolution Difix3d+: Improving 3d reconstructions with single-step diffusion models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T05:05:28.728175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:05:28.728175Z digest=sha256:24f4ac85c51e8689f41458628236d5867809bb315a25fbe0f88f52884bcbbc52

Observation ae902334-f46f-4de1-8968-28ec79d9aa6c · outbound

This paper cites SANA 1.5: Efficient Scaling of Training-Time and Inference-Time Compute in Linear Diffusion Transformer.

Inference-time Scaling for Diffusion-based Audio Super-resolution SANA 1.5: Efficient Scaling of Training-Time and Inference-Time Compute in Linear Diffusion Transformer

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T05:05:28.821341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:05:28.821341Z digest=sha256:f865a9601bfa5a8db30da74f11d12a20a980685f531546a91e14ba7c94d5bb28

Observation 47215184-baae-4547-bdcc-39ff7bc441d1 · outbound

This paper cites Flashspeech: Efficient zero-shot speech synthesis.

Inference-time Scaling for Diffusion-based Audio Super-resolution Flashspeech: Efficient zero-shot speech synthesis

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:05:30.406623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:05:28.882782Z digest=sha256:e73e6c1da56c0b92fd859786dd8ae8b6a73facb323a5567da646396807e06249

Observation f925a107-9820-4bfc-ba08-83c3990f4c71 · outbound

This paper cites Comospeech: One-step speech and singing voice synthesis via consistency model.

Inference-time Scaling for Diffusion-based Audio Super-resolution Comospeech: One-step speech and singing voice synthesis via consistency model

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:05:30.395990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:05:28.955893Z digest=sha256:0c5d734371e676fd293f595aee886dd2fbe49f3f9a46fbd2906cb66f156624c8

Observation 751e1231-605c-4b79-85a1-fd9ae8a5d8d5 · outbound

This paper cites Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis.

Inference-time Scaling for Diffusion-based Audio Super-resolution Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T05:05:29.031835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:05:29.031835Z digest=sha256:ea81101c729fcade352314a6ab02bf1372862414b856421cc97d17f955939f01

Observation 98348ee1-b86b-421a-afa6-e30ad42aff16 · outbound

This paper cites GAN Vocoder: Multi-Resolution Discriminator Is All You Need.

Inference-time Scaling for Diffusion-based Audio Super-resolution GAN Vocoder: Multi-Resolution Discriminator Is All You Need

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T05:05:29.114926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:05:29.114926Z digest=sha256:1fb81daecf0c22906de6a5662f297516505c8860519dabeec54236b358e39705

Observation f9876195-b761-4a36-9249-f6796d051134 · outbound

This paper cites Conditioning and sampling in variational diffusion models for speech super-resolution.

Inference-time Scaling for Diffusion-based Audio Super-resolution Conditioning and sampling in variational diffusion models for speech super-resolution

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:05:30.385784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:05:29.190217Z digest=sha256:9c7119b1b16cc63e76ab9977f4796c2094980b548a94b89176cd4da149eaf35b

Observation f1ca83c3-1f35-4bbc-a6ef-48de064b3670 · outbound

This paper cites Uncertainty-guided perturbation for image super-resolution diffusion model.

Inference-time Scaling for Diffusion-based Audio Super-resolution Uncertainty-guided perturbation for image super-resolution diffusion model

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:05:30.364653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:05:29.256313Z digest=sha256:dbe87e1c5c53b0ba615a10a465f55fa48e1d4256a7feab62dc816368cdec8ea4

Observation c1d3dd04-2823-48ce-95c4-3e9ca5587cd6 · outbound

This paper cites Flashvideo: Flowing fidelity to detail for efficient high-resolution video generation.arXiv preprint arXiv:2502.05179, 2025.

Inference-time Scaling for Diffusion-based Audio Super-resolution Flashvideo: Flowing fidelity to detail for efficient high-resolution video generation.arXiv preprint arXiv:2502.05179, 2025

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T05:05:29.324913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:05:29.324913Z digest=sha256:c3b2c75a320731da484482d2e9343cf433b0c1586ceb24360cb1edd981deb102

Observation e6725dd0-fa92-42f8-93ce-be202f53b31d · outbound

This paper cites Inference-time scaling of diffusion models through classical search.arXiv preprint arXiv:2505.23614, 2025.

Inference-time Scaling for Diffusion-based Audio Super-resolution Inference-time scaling of diffusion models through classical search.arXiv preprint arXiv:2505.23614, 2025

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T05:05:29.441344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:05:29.441344Z digest=sha256:612d787373aebcfa1ecf0475a266152ac4972f1df8409f7dbc097777d68c622a

Observation 8d5bd33c-6925-46c0-856b-49e492e51a65 · outbound

This paper cites In-Context Edit: Enabling Instructional Image Editing with In-Context Generation in Large Scale Diffusion Transformer.

Inference-time Scaling for Diffusion-based Audio Super-resolution In-Context Edit: Enabling Instructional Image Editing with In-Context Generation in Large Scale Diffusion Transformer

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T05:05:29.628800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:05:29.628800Z digest=sha256:ef2e42e4dd7f596349021e6adbc8e93fe369c9edb34de86b383db4acdf08b9a7

Pith citing papers

Observation 59270332-aa16-4238-8b3f-21ab1d15d682 · inbound

Inference-Time Scaling for Joint Audio-Video Generation cites this paper.

Inference-Time Scaling for Joint Audio-Video Generation Inference-time Scaling for Diffusion-based Audio Super-resolution

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:56:40.913783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T07:49:20.194990Z digest=sha256:937e9c7ddcf0b8fa9eb5e673a3986327a5165dd0bd9d7a61e6607ae9478d1588