Pith. sign in

Paper Citation Record · LEDGER

Inference-time Scaling for Diffusion-based Audio Super-resolution

As of 18 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 1 inbound Pith citation observation for arXiv:2508.02391.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.02391 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T05:05:29.628800Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T07:49:20.194990Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T05:56:40.912157Z

Reference resolution

47 of 47 outbound references displayed

  • verified exact0
  • verified fuzzy12
  • unresolved35
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 48db7d8e-d0e6-4105-95d2-2a4543c366b7 · outbound

This paper cites MusicLM: Generating Music From Text.

Inference-time Scaling for Diffusion-based Audio Super-resolution MusicLM: Generating Music From Text

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T05:05:25.719480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:05:25.719480Z digest=sha256:a425830f700656e7fcb14f4b6999e7299630a220f3f6bc138be3a1432d9e41aa

Observation c06cd0e7-530e-43a7-9125-70cbc86b3060 · outbound

This paper cites Large Language Monkeys: Scaling Inference Compute with Repeated Sampling.

Inference-time Scaling for Diffusion-based Audio Super-resolution Large Language Monkeys: Scaling Inference Compute with Repeated Sampling

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T05:05:25.801277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:05:25.801277Z digest=sha256:ce524eac803255876d9725757750060fd1e07cdbae2a47c7cc8f5b780ea4c5cc

Observation 3673279d-2d10-4603-b4dc-35e525b8aeb8 · outbound

This paper cites Musicldm: Enhancing novelty in text-to-music generation using beat-synchronous mixup strategies.

Inference-time Scaling for Diffusion-based Audio Super-resolution Musicldm: Enhancing novelty in text-to-music generation using beat-synchronous mixup strategies

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:05:30.548978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T05:05:25.899510Z digest=sha256:817f64417be2e242d03d4c778a74b4c20bf91a50487487c2144e1aef7366c86c

Observation 5c9013d8-5612-40b4-bc5f-5524d676e8f3 · outbound

This paper cites Wavlm: Large-scale self-supervised pre- training for full stack speech processing.IEEE Journal of Selected Topics in Signal Processing, 16(6):1505–1518, 2022.

Inference-time Scaling for Diffusion-based Audio Super-resolution Wavlm: Large-scale self-supervised pre- training for full stack speech processing.IEEE Journal of Selected Topics in Signal Processing, 16(6):1505–1518, 2022

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T05:05:26.001366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:05:26.001366Z digest=sha256:1a7aeb7b74cb1efdaa83f62f313ac41eaed4df749d2fbd9204404f657c24b493

Observation 95d9772d-68ba-4190-8073-53d2ae2d242e · outbound

This paper cites Qwen2-Audio Technical Report.

Inference-time Scaling for Diffusion-based Audio Super-resolution Qwen2-Audio Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T05:05:26.071505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:05:26.071505Z digest=sha256:907b546521350350387d1908deb88bedba1e0df3230bf38e8739c875fa3e9005

Observation 518c59ab-ee3e-47fc-94d1-251728d6e871 · outbound

This paper cites Directly Fine-Tuning Diffusion Models on Differentiable Rewards.

Inference-time Scaling for Diffusion-based Audio Super-resolution Directly Fine-Tuning Diffusion Models on Differentiable Rewards

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T05:05:26.135354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:05:26.135354Z digest=sha256:c9c5ff01d17c2d119a4ec7d53fd8754751215e53088d69080f0c7f9aa0462b7e

Observation 2748de54-886e-422f-9f23-377135b90ad9 · outbound

This paper cites Clap learning audio concepts from natural language supervision.

Inference-time Scaling for Diffusion-based Audio Super-resolution Clap learning audio concepts from natural language supervision

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T05:05:26.243912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:05:26.243912Z digest=sha256:6fb4d42d19e62d8589f9275327a5e1b0f9bbed959c4f6ea9f24184fc3e153d78

Observation af261e11-cd76-4626-a8df-d031500f5e34 · outbound

This paper cites FunASR: A Fundamental End-to-End Speech Recognition Toolkit.

Inference-time Scaling for Diffusion-based Audio Super-resolution FunASR: A Fundamental End-to-End Speech Recognition Toolkit

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T05:05:26.342534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:05:26.342534Z digest=sha256:d40e256c5e6e24b62a969e43a6127f03db6a59209fbea436524d08c7640baa9f

Observation 152c3e68-9ab9-4f09-9092-41f59eebed06 · outbound

This paper cites AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning.

Inference-time Scaling for Diffusion-based Audio Super-resolution AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T05:05:26.435388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:05:26.435388Z digest=sha256:e768dbe84386c47b88a5a27908af272328b3298d4600c0e59458ec17ef9763ac

Observation ffaedfe7-3c15-4547-98b4-6cd388f1a795 · outbound

This paper cites NU-Wave 2: A General Neural Audio Upsampling Model for Various Sampling Rates.

Inference-time Scaling for Diffusion-based Audio Super-resolution NU-Wave 2: A General Neural Audio Upsampling Model for Various Sampling Rates

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T05:05:26.516349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:05:26.516349Z digest=sha256:c8897d24e1fc13b2d45f1bbfbd57ca540e9f765a332384141dc1a50491e84904

Observation 0c62d301-5dc5-4ce7-b698-6ed0dfdbf324 · outbound

This paper cites Denoising diffusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020.

Inference-time Scaling for Diffusion-based Audio Super-resolution Denoising diffusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T05:05:26.599483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:05:26.599483Z digest=sha256:d01dfc067e9e5b07a74a3f47e1cd1e85bb09f270cb76bd9f94690b5e97a8a613

Observation e50f301c-ad54-4c9a-bd2a-4f9ae10809a1 · outbound

This paper cites Classifier-Free Diffusion Guidance.

Inference-time Scaling for Diffusion-based Audio Super-resolution Classifier-Free Diffusion Guidance

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T05:05:26.697388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:05:26.697388Z digest=sha256:e579dd5db86d42317ae7ec4075b6e2e530d036411344e4e7270955b57221b630

Observation 023b804b-cf22-47f1-9408-d224b0a9f3f5 · outbound

This paper cites Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis.Advances in neural information processing systems, 33:17022–17033, 2020.

Inference-time Scaling for Diffusion-based Audio Super-resolution Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis.Advances in neural information processing systems, 33:17022–17033, 2020

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T05:05:26.799798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:05:26.799798Z digest=sha256:3e431625c0ea99b7e773037bbab8183fc56783c91febb619fe352ebccc62ea25

Observation 014c09e7-7854-42f0-aa36-8d28a67ce2ca · outbound

This paper cites NU-Wave: A Diffusion Probabilistic Model for Neural Audio Upsampling.

Inference-time Scaling for Diffusion-based Audio Super-resolution NU-Wave: A Diffusion Probabilistic Model for Neural Audio Upsampling

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T05:05:26.893661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:05:26.893661Z digest=sha256:4924ea78b3455d14a4a301d450abac487f539ac19aac50a79ea8329252ef9049

Observation 827cfac1-42dc-47e1-b3ff-6d89c3876da1 · outbound

This paper cites Noise-free optimization in early training steps for image super- resolution.

Inference-time Scaling for Diffusion-based Audio Super-resolution Noise-free optimization in early training steps for image super- resolution

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:05:30.513791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T05:05:27.012398Z digest=sha256:7ba54c3886ddbc145587b319f9b3f946ee63e75ed57cfa03b53e40aa1853ef48

Observation ab16529c-e383-45ff-aaf6-4716bb46018c · outbound

This paper cites Audiosr: Versatile audio super-resolution at scale.

Inference-time Scaling for Diffusion-based Audio Super-resolution Audiosr: Versatile audio super-resolution at scale

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:05:30.503268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T05:05:27.067541Z digest=sha256:b42dcf3770640cc3977b0d6ff82db7e8c68e5dcd6ebb8d57e2c37ac0cfc81054

Observation e0fee5c5-370f-447c-94d1-65406a061189 · outbound

This paper cites AudioLDM: Text-to-Audio Generation with Latent Diffusion Models.

Inference-time Scaling for Diffusion-based Audio Super-resolution AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T05:05:27.134039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:05:27.134039Z digest=sha256:9f7d8b7043858e1188c372de273095e770d126618a33b3c090eea9f0b45f6b66

Observation b2bc81cf-3b27-4e90-a3f2-6b936bfc3ae5 · outbound

This paper cites Neural Vocoder is All You Need for Speech Super-resolution.

Inference-time Scaling for Diffusion-based Audio Super-resolution Neural Vocoder is All You Need for Speech Super-resolution

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T05:05:27.200465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:05:27.200465Z digest=sha256:a8cc34edc2eea69a7b91e39510bc66fad293ce5d81bcb8238be6838a05ad0590

Observation 3e068ebb-6c8c-4317-93d5-a1dc6e83723c · outbound

This paper cites VoiceFixer: Toward General Speech Restoration with Neural Vocoder.

Inference-time Scaling for Diffusion-based Audio Super-resolution VoiceFixer: Toward General Speech Restoration with Neural Vocoder

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T05:05:27.287407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:05:27.287407Z digest=sha256:d0a864da4a05a2620ad7c9e7d7051a4b1cbf8f12f8c87b4316f9d86a8754e9ab

Observation d29ae95f-9b46-41b3-9a01-5effb4bc06e2 · outbound

This paper cites Dpm-solver: A fast ode solver for diffusion probabilistic model sampling in around 10 steps.Advances in Neural Information Processing Systems, 35:5775–5787, 2022.

Inference-time Scaling for Diffusion-based Audio Super-resolution Dpm-solver: A fast ode solver for diffusion probabilistic model sampling in around 10 steps.Advances in Neural Information Processing Systems, 35:5775–5787, 2022

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T05:05:27.368077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:05:27.368077Z digest=sha256:0c8ad5f80930ca3cedaea43d9c13c802da63b246d7f0e9254ec58f0603246557

Observation f11eca33-6ed1-4d3b-91e2-a8800b66484e · outbound

This paper cites Scaling inference time compute for diffusion models.

Inference-time Scaling for Diffusion-based Audio Super-resolution Scaling inference time compute for diffusion models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T05:05:27.442554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:05:27.442554Z digest=sha256:524fcb3205afe0b045366af902bec2530e77419aadf0e3b0608f12e35aaec40f

Observation d014c32a-f416-493b-9aee-f7fc60768fa7 · outbound

This paper cites Uncertainty-driven loss for single image super-resolution.Advances in Neural Information Processing Systems, 34:16398–16409, 2021.

Inference-time Scaling for Diffusion-based Audio Super-resolution Uncertainty-driven loss for single image super-resolution.Advances in Neural Information Processing Systems, 34:16398–16409, 2021

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:05:30.481703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T05:05:27.503553Z digest=sha256:eb69383bf6fda7da7cc6b7f4d48e9325263202b8e07238821656fe9ef4bcf357

Observation 44414bcc-ba4b-4cb4-8992-01c2a6a7a066 · outbound

This paper cites The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models.

Inference-time Scaling for Diffusion-based Audio Super-resolution The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T05:05:27.596743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:05:27.596743Z digest=sha256:6d1883a58383c5204442fc31391a4aeff505c53f94391a66ef3e1320f396cf60

Observation 1dd6f9d9-f804-468b-800a-17064386a1f9 · outbound

This paper cites Esc: Dataset for environmental sound classification.

Inference-time Scaling for Diffusion-based Audio Super-resolution Esc: Dataset for environmental sound classification

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:05:30.471212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T05:05:27.686704Z digest=sha256:a889e2fef0f9e3e02da4ff4e8dc3b0cec9c5250f7edd31565f5f8c2e523ba6b5

Observation 7474cf9f-2393-45cf-b483-29ddd7c48efa · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

Inference-time Scaling for Diffusion-based Audio Super-resolution Robust speech recognition via large-scale weak supervision

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T05:05:27.768071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:05:27.768071Z digest=sha256:2ad9efab83a69a00a0ac903b9867fcdc6a999dcb8ba8779003a672e432476fde

Observation 68004d0d-dddb-45cc-b8d5-53f7c625913f · outbound

This paper cites FastSpeech 2: Fast and High-Quality End-to-End Text to Speech.

Inference-time Scaling for Diffusion-based Audio Super-resolution FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T05:05:27.864867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:05:27.864867Z digest=sha256:20d48733be0d3efe11e87fc9f6ee61af8ee14863b9e178ed5135cb15c2b4c2da

Observation fd7b94e0-ed22-4fe2-941e-7ac4b062b715 · outbound

This paper cites Fastspeech: Fast, robust and controllable text to speech.Advances in neural information processing systems, 32, 2019.

Inference-time Scaling for Diffusion-based Audio Super-resolution Fastspeech: Fast, robust and controllable text to speech.Advances in neural information processing systems, 32, 2019

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:05:30.453031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T05:05:27.978066Z digest=sha256:0ffb003a8c8f32b9a6312ef53c9d614b336dbbf849024363d5ecfc283457c26f

Observation 5bfd7618-59e3-4446-b30a-2a323305096f · outbound

This paper cites High- resolution image synthesis with latent diffusion models.

Inference-time Scaling for Diffusion-based Audio Super-resolution High- resolution image synthesis with latent diffusion models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T05:05:28.080199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:05:28.080199Z digest=sha256:2a9d15fa1978e1a43fc8788d0d826fba6de0173498c08f85de5c618ed8392fdb

Observation e8139e27-ff8e-4e4d-86e3-ba24bdd26ffb · outbound

This paper cites NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers.

Inference-time Scaling for Diffusion-based Audio Super-resolution NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T05:05:28.156213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:05:28.156213Z digest=sha256:e5e71aab88ed09979a0b806a5e882c0e671baa2664eadee1020dfc695d65e317

Observation da5f8c3c-c005-4d6c-a313-048a8742480e · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Inference-time Scaling for Diffusion-based Audio Super-resolution Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T05:05:28.206753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:05:28.206753Z digest=sha256:baecd25ae1d509825286d8a8b45be68c5226d9e9568c522390de3b3e91994a2e

Observation 1bad2a83-4fc0-408d-a093-cf9fda9f040e · outbound

This paper cites Denoising Diffusion Implicit Models.

Inference-time Scaling for Diffusion-based Audio Super-resolution Denoising Diffusion Implicit Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T05:05:28.271130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:05:28.271130Z digest=sha256:6c3fef7fe42b0d0e415da8c413b4e3c062b7f2fa948ffd6bf8d5c1d3b003485f

Observation 421a2ab1-1d8c-48ed-b174-a78fcbeacf4c · outbound

This paper cites Score-Based Generative Modeling through Stochastic Differential Equations.

Inference-time Scaling for Diffusion-based Audio Super-resolution Score-Based Generative Modeling through Stochastic Differential Equations

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T05:05:28.360497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:05:28.360497Z digest=sha256:4da2b9abc105901204c2ebaca304a332c4ee7ad96921906ab3c29387d8dc11f1

Observation e073ebd2-2993-4a1a-ac33-63b1d0d1c62b · outbound

This paper cites AudioX: A Unified Framework for Anything-to-Audio Generation.

Inference-time Scaling for Diffusion-based Audio Super-resolution AudioX: A Unified Framework for Anything-to-Audio Generation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T05:05:28.435197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:05:28.435197Z digest=sha256:e1380bfe926396ad8d27e5a9e607f0577aa516ef4bd0f6b35ac4771ee4fdf07d

Observation 5df9d974-cdcf-4a0d-a221-71806c1dde49 · outbound

This paper cites Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound.

Inference-time Scaling for Diffusion-based Audio Super-resolution Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T05:05:28.514824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:05:28.514824Z digest=sha256:6ca8b5fe5a9ae8d06d8573b9157ff5e12e76267b7d17a88b57d1bda8d4b9c4a4

Observation c40a1b6a-db08-4293-b150-fb6fbc2bd88b · outbound

This paper cites Towards robust speech super-resolution.IEEE/ACM transactions on audio, speech, and language processing, 29:2058–2066, 2021.

Inference-time Scaling for Diffusion-based Audio Super-resolution Towards robust speech super-resolution.IEEE/ACM transactions on audio, speech, and language processing, 29:2058–2066, 2021

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:05:30.435338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T05:05:28.592826Z digest=sha256:cf61bd2416df110965527c2a84a829bd2415468393f85f64eace7fc3a6bac555

Observation fee83341-ce77-442d-9a4f-0cef3b860c8f · outbound

This paper cites Esrgan: Enhanced super-resolution generative adversarial networks.

Inference-time Scaling for Diffusion-based Audio Super-resolution Esrgan: Enhanced super-resolution generative adversarial networks

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:05:30.423997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T05:05:28.662941Z digest=sha256:6e4c5cfad4876eee0312c0176a16bc6ebc88a50cd0a020e4a47ecf1cf3301349

Observation f7a061b5-c1ea-4829-a27e-8209352f00f1 · outbound

This paper cites Difix3d+: Improving 3d reconstructions with single-step diffusion models.

Inference-time Scaling for Diffusion-based Audio Super-resolution Difix3d+: Improving 3d reconstructions with single-step diffusion models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T05:05:28.728175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:05:28.728175Z digest=sha256:22a789162c787dfe52854ded363e0d34690a2ef922b4cac0ce12d0632259c45a

Observation ae902334-f46f-4de1-8968-28ec79d9aa6c · outbound

This paper cites SANA 1.5: Efficient Scaling of Training-Time and Inference-Time Compute in Linear Diffusion Transformer.

Inference-time Scaling for Diffusion-based Audio Super-resolution SANA 1.5: Efficient Scaling of Training-Time and Inference-Time Compute in Linear Diffusion Transformer

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T05:05:28.821341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:05:28.821341Z digest=sha256:b42abd030e9529f7d0364fa00ab6d61424c1db0b1f5958059ff554e4d3f39a10

Observation 47215184-baae-4547-bdcc-39ff7bc441d1 · outbound

This paper cites Flashspeech: Efficient zero-shot speech synthesis.

Inference-time Scaling for Diffusion-based Audio Super-resolution Flashspeech: Efficient zero-shot speech synthesis

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:05:30.406623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T05:05:28.882782Z digest=sha256:93f644d6e07e257ca7cf45c9317c0987b19ac987d53fa4699860757390d0bad2

Observation f925a107-9820-4bfc-ba08-83c3990f4c71 · outbound

This paper cites Comospeech: One-step speech and singing voice synthesis via consistency model.

Inference-time Scaling for Diffusion-based Audio Super-resolution Comospeech: One-step speech and singing voice synthesis via consistency model

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:05:30.395990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T05:05:28.955893Z digest=sha256:41ba0d8ef30373d392d97fba0d1191886763ffd4b90fe4f4846fabf484b9db93

Observation 751e1231-605c-4b79-85a1-fd9ae8a5d8d5 · outbound

This paper cites Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis.

Inference-time Scaling for Diffusion-based Audio Super-resolution Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T05:05:29.031835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:05:29.031835Z digest=sha256:761db00c198362d74270a3b67918e5f9141683daac7cedcb2c064aa4248de141

Observation 98348ee1-b86b-421a-afa6-e30ad42aff16 · outbound

This paper cites GAN Vocoder: Multi-Resolution Discriminator Is All You Need.

Inference-time Scaling for Diffusion-based Audio Super-resolution GAN Vocoder: Multi-Resolution Discriminator Is All You Need

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T05:05:29.114926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:05:29.114926Z digest=sha256:8dd4f18213a42c377eea7cd4bdda6bcba7d75a01531e1bc6c2f03f8560587727

Observation f9876195-b761-4a36-9249-f6796d051134 · outbound

This paper cites Conditioning and sampling in variational diffusion models for speech super-resolution.

Inference-time Scaling for Diffusion-based Audio Super-resolution Conditioning and sampling in variational diffusion models for speech super-resolution

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:05:30.385784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T05:05:29.190217Z digest=sha256:60ca4c9bbe6009e2f6733de4394bffbe6541464da950c92aa71dde207f158c7a

Observation f1ca83c3-1f35-4bbc-a6ef-48de064b3670 · outbound

This paper cites Uncertainty-guided perturbation for image super-resolution diffusion model.

Inference-time Scaling for Diffusion-based Audio Super-resolution Uncertainty-guided perturbation for image super-resolution diffusion model

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:05:30.364653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T05:05:29.256313Z digest=sha256:6e0ba864798db7a05b9338fa7938500a393cefd832ec6446d404ad2a200d1568

Observation c1d3dd04-2823-48ce-95c4-3e9ca5587cd6 · outbound

This paper cites Flashvideo: Flowing fidelity to detail for efficient high-resolution video generation.arXiv preprint arXiv:2502.05179, 2025.

Inference-time Scaling for Diffusion-based Audio Super-resolution Flashvideo: Flowing fidelity to detail for efficient high-resolution video generation.arXiv preprint arXiv:2502.05179, 2025

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T05:05:29.324913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:05:29.324913Z digest=sha256:b3e0c435796166fdb206b220c439f2313c387f76cc03f814b42784e2c8700c2b

Observation e6725dd0-fa92-42f8-93ce-be202f53b31d · outbound

This paper cites Inference-time scaling of diffusion models through classical search.arXiv preprint arXiv:2505.23614, 2025.

Inference-time Scaling for Diffusion-based Audio Super-resolution Inference-time scaling of diffusion models through classical search.arXiv preprint arXiv:2505.23614, 2025

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T05:05:29.441344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:05:29.441344Z digest=sha256:2f9b7c1aa1233e1cef97fad9b5a4704f27aea15844057efad21c09f94c0aea5b

Observation 8d5bd33c-6925-46c0-856b-49e492e51a65 · outbound

This paper cites In-Context Edit: Enabling Instructional Image Editing with In-Context Generation in Large Scale Diffusion Transformer.

Inference-time Scaling for Diffusion-based Audio Super-resolution In-Context Edit: Enabling Instructional Image Editing with In-Context Generation in Large Scale Diffusion Transformer

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T05:05:29.628800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:05:29.628800Z digest=sha256:66303ec3c4d27ccdb68c212b8f0ff16707bf41104172b826cc8ac10b9051c9b6

Pith citing papers

Observation 59270332-aa16-4238-8b3f-21ab1d15d682 · inbound

Inference-Time Scaling for Joint Audio-Video Generation cites this paper.

Inference-Time Scaling for Joint Audio-Video Generation Inference-time Scaling for Diffusion-based Audio Super-resolution

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:56:40.913783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-28T07:49:20.194990Z digest=sha256:2f6030d21bd86b527c67a00b3a78f3aa09cc2b33b62bfbfa94df095cd2eb280a