Pith. sign in

Paper Citation Record · LEDGER

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet

As of 20 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 1 inbound Pith citation observation for arXiv:2505.16195.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.16195 v2

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:09:58.262045Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-18T08:20:02.986562Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T08:21:06.814521Z

Reference resolution

48 of 48 outbound references displayed

  • verified exact1
  • verified fuzzy27
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation abc09efc-c13d-4be5-8260-8f0b94acdb03 · outbound

This paper cites AudioGen: Textually Guided Audio Generation.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet AudioGen: Textually Guided Audio Generation

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:54.545341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:54.545341Z digest=sha256:ab8feeab027ab8995b3cba9446b0d3ded9eba8a0948bd62c221b791acedb8076

Observation 7953642e-47fc-455d-a2b1-c9fc8739fcc3 · outbound

This paper cites Make-An-Audio 2: Temporal-Enhanced Text-to-Audio Generation.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Make-An-Audio 2: Temporal-Enhanced Text-to-Audio Generation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:54.619538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:54.619538Z digest=sha256:e03dc01deaa38e4dc3e99378c517f7c3a071aec176df820d96350da4dd36d6f0

Observation 10e5f49b-3c03-4f1f-b5df-9f047e5e0bfd · outbound

This paper cites SpecMaskGIT: Masked Generative Modeling of Audio Spectrograms for Efficient Audio Synthesis and Beyond.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet SpecMaskGIT: Masked Generative Modeling of Audio Spectrograms for Efficient Audio Synthesis and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:54.712485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:54.712485Z digest=sha256:fb3dd660bb376a982e29e2722fcc246f4182ffd6a2245614c6ee3c300499e71a

Observation 66c067d3-bdd2-4644-98ba-2b6fbf9c6141 · outbound

This paper cites SoundCTM: Unifying Score-based and Consistency Models for Full-band Text-to-Sound Generation.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet SoundCTM: Unifying Score-based and Consistency Models for Full-band Text-to-Sound Generation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:54.810995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:54.810995Z digest=sha256:dbaf8fda7ac2f581f2a974d3abf4fb2645e1bc0152287310496583a032693cd9

Observation fe976c38-6e02-4d78-a072-a320e56aaf45 · outbound

This paper cites Stable audio open,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Stable audio open,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:02.801403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:09:54.911275Z digest=sha256:6b633445231b2de21937a77c28867704a0aec24a356ac56e6c2720d88550c07d

Observation 4ba353cf-9ed8-4907-825a-6f6705aa76b4 · outbound

This paper cites MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:54.999597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:54.999597Z digest=sha256:244e742af17ab6009e2654b1a7f4a51e4dd722e18c0c2524c1cea95faad63884

Observation 9b22797a-76b4-46ae-886b-f4261d8e063c · outbound

This paper cites Mm-diffusion: Learning multi-modal diffusion models for joint audio and video generation,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Mm-diffusion: Learning multi-modal diffusion models for joint audio and video generation,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:02.647919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:09:55.139356Z digest=sha256:0a32835508a3a5d4f3b4a0c5949de9503ed613c9e6e44dad8614206edae38038

Observation 23d57b05-ab77-49f4-88f7-9f99a863cba2 · outbound

This paper cites Visual Echoes: A Simple Unified Transformer for Audio-Visual Generation.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Visual Echoes: A Simple Unified Transformer for Audio-Visual Generation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:55.278233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:55.278233Z digest=sha256:00a293d32283fd225daaf59474a3fda96396f76b8631c8f830a06f188cecbe45

Observation e8818c41-138b-4f53-ab6e-af61e8b28b03 · outbound

This paper cites V2a-mapper: A lightweight solution for vision-to-audio generation by connecting foundation models,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet V2a-mapper: A lightweight solution for vision-to-audio generation by connecting foundation models,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:02.496790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:09:55.387681Z digest=sha256:f439617752929ccda4720b59dd99d2c8139207832f000c74e5ea7fdbded94bd6

Observation cfd683ff-f0a3-4a25-baab-cead609353e7 · outbound

This paper cites Seeing and hearing: Open-domain visual-audio generation with diffusion latent aligners,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Seeing and hearing: Open-domain visual-audio generation with diffusion latent aligners,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:02.352055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:09:55.487341Z digest=sha256:2bd7e38deb8f4d979f8b78ca82a9a7062917d54933ef0d34a6c02d304712e06c

Observation 9fef9217-ca06-447b-8cd2-15c042f266c9 · outbound

This paper cites Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:55.547205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:55.547205Z digest=sha256:be8150caa6b1d23908843b69e7bd220c11bffe94122684093783e97678c6c2df

Observation 0ddf2160-429f-49b1-b0ad-bcdd0b423e1f · outbound

This paper cites Synchformer: Efficient synchronization from sparse cues,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Synchformer: Efficient synchronization from sparse cues,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:02.207841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:09:55.606504Z digest=sha256:cfa30fb05641493401c004a9228dce815f687a217c1f0b0344aafd3bb96872c6

Observation 8994e6ca-0abb-4e83-b1c4-652cdca8e598 · outbound

This paper cites Adding conditional control to text- to-image diffusion models,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Adding conditional control to text- to-image diffusion models,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:02.020964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:09:55.683396Z digest=sha256:8e6754d381efad43ec93d55ec91fef1a79c9509ec542d1df4178ae6b7a469e4c

Observation 3cc6f8c6-e56a-4b03-b059-a741ae33fc09 · outbound

This paper cites Uni-controlnet: All-in-one control to text-to-image diffusion models,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Uni-controlnet: All-in-one control to text-to-image diffusion models,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:01.857838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:09:55.753711Z digest=sha256:3b2b5c735c7a0e35d4afb5b4f31aa3594e82be0ef7e5423c252597d4f387d1fe

Observation 110c6b42-6d76-4411-a425-12ee6748e1f7 · outbound

This paper cites Read, Watch and Scream! Sound Generation from Text and Video.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Read, Watch and Scream! Sound Generation from Text and Video

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:55.831770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:55.831770Z digest=sha256:b8fff14ae42a2d24be88bd678038cc037a5b646cffe507c4a9bd0617e1e66d9d

Observation 72a83755-a8c4-4920-b509-daf337c91156 · outbound

This paper cites FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:55.910310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:55.910310Z digest=sha256:fa5eadf532bf5aeee075863f50d5acb333558af9de6b68c9f223ae17feddeaeb

Observation 0dd4b443-d5f0-4fa6-a51c-394681685a84 · outbound

This paper cites Tell What You Hear From What You See -- Video to Audio Generation Through Text.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Tell What You Hear From What You See -- Video to Audio Generation Through Text

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:55.988743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:55.988743Z digest=sha256:f37bbd97089caeaab55a61fd0c6fe71feed9484cc83f176e12fc40e0d38d2caf

Observation 6278bfb9-8f1b-4e1a-8085-cc24982c664e · outbound

This paper cites Temporally aligned audio for video with autoregression,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Temporally aligned audio for video with autoregression,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:01.708710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:09:56.045325Z digest=sha256:9f0a995a0e28e8fb1ded5ba1fff757d1858a6272126e455c2f92fd0f0cfc7374

Observation f2a7234f-eedb-4911-a3eb-28b36912076e · outbound

This paper cites Frieren: Efficient video-to-audio generation network with rectified flow matching,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Frieren: Efficient video-to-audio generation network with rectified flow matching,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:01.422859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:09:56.119327Z digest=sha256:53a544b9426e14d9e35a95837a0a2f446ada63ba183d26574a0f5f4a093d28ec

Observation e6501fb0-d017-4ff1-9d9d-438c83298075 · outbound

This paper cites Mavil: Masked audio-video learners,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Mavil: Masked audio-video learners,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:01.261287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:09:56.195962Z digest=sha256:a7f442b3f8307f11451d5ef06c5fdf942239258d7ecccd5d35076203cf8ad66b

Observation 3cc382ef-08bf-40c2-a2e7-7cbbe25e8cc4 · outbound

This paper cites Diff-foley: Synchronized video- to-audio synthesis with latent diffusion models,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Diff-foley: Synchronized video- to-audio synthesis with latent diffusion models,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:01.089076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:09:56.264683Z digest=sha256:6c5e1154aba8793ab2da47b08b3f6a42d463e66eb5107431bed04cc17e0d911b

Observation 6cb4d445-1e69-4e49-ba26-f9ebff2361a3 · outbound

This paper cites FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:56.308460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:56.308460Z digest=sha256:b117a05fcf8dd88ede380efaf615122f7a92a305852f89bd511abe015b9b33b5

Observation 1153f749-b0bd-4fd4-b5e1-da4c0392e7da · outbound

This paper cites Smooth-foley: Creating continuous sound for video-to-audio generation under semantic guidance,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Smooth-foley: Creating continuous sound for video-to-audio generation under semantic guidance,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:00.923321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:09:56.376868Z digest=sha256:0bbe913807c09c2dbd6bbc8c9153175f1f4f1ff1ceea9e8417d10c658c739a0d

Observation 18b0951f-4fe7-4acb-a6e8-fc7fc4972a5f · outbound

This paper cites Learning transferable visual models from natural language supervision,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Learning transferable visual models from natural language supervision,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:00.765043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:09:56.424504Z digest=sha256:3e17cd919d9417788fadd9f872c70841084902db8ef2171a653aee15bfd3afc9

Observation 34ad2f69-31e0-421e-aa52-a94edb43a1cc · outbound

This paper cites Maskgit: Masked generative image transformer,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Maskgit: Masked generative image transformer,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:00.596551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:09:56.497923Z digest=sha256:bbb9d73f696eff55ccf846a94f8a9a263f3842161b064fc2dde9727fe502ce9c

Observation 389e82fd-b576-42d8-b9dc-c7eb68d08266 · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:56.595415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:56.595415Z digest=sha256:a7164f58bd103425af0fbbce29626b1a9bc3c2a62a3eda4a589bd7ad255553da

Observation 1648a738-ef21-43fe-8484-e30a2be13b05 · outbound

This paper cites Music Foundation Model as Generic Booster for Music Downstream Tasks.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Music Foundation Model as Generic Booster for Music Downstream Tasks

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:56.680530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:56.680530Z digest=sha256:4f890a04a4343bce72bbe1b167bd144939e2760bc0e12e876cee563bac332386

Observation 4ba595d1-8f83-41b5-b393-85b15143268d · outbound

This paper cites High Fidelity Neural Audio Compression.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet High Fidelity Neural Audio Compression

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:56.761212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:56.761212Z digest=sha256:e53c0eb20964eaad17bbaf77a63711bb87d1fffb271052788dcaeb253aa1fcf5

Observation bc0b6b04-276c-4456-8638-737ada8b65fd · outbound

This paper cites High- fidelity audio compression with improved rvqgan,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet High- fidelity audio compression with improved rvqgan,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:00.473120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:09:56.840072Z digest=sha256:b8f084750d93c1b51307137b3d6f026dc32ff32ff679b3e7c38075f6badb6fb2

Observation 52abba22-8b04-45dd-aacc-3e6d083918f9 · outbound

This paper cites Taming Visually Guided Sound Generation.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Taming Visually Guided Sound Generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:56.889056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:56.889056Z digest=sha256:378f0ab92d9dad98fbd12d5e5aa2859450c3bd11e679df1d15077ffa6eb1440c

Observation ea1a425d-b872-4299-bf93-254e01388d49 · outbound

This paper cites Masked autoencoders are scalable vision learners,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Masked autoencoders are scalable vision learners,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:00.382535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:09:56.939962Z digest=sha256:98280fc499a4cbe1d3e06cc48f8cdf9d48d49f159a591a78975ec6e197ea5af5

Observation 3e56c4ab-42ed-407a-8092-60bd86077502 · outbound

This paper cites Extending audio masked autoencoders toward audio restoration,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Extending audio masked autoencoders toward audio restoration,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:00.237928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:09:57.006139Z digest=sha256:88429c88b41adf8778b6cff9d3672d9f0e796c7a84bf5c1b2cc926c0bea7449b

Observation 87d13178-79e3-43d7-b824-dede84d37966 · outbound

This paper cites Mage: Masked generative encoder to unify representation learning and image synthesis,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Mage: Masked generative encoder to unify representation learning and image synthesis,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:00.021600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:09:57.058562Z digest=sha256:31122d0605ec16e145934f1733f73b12a747dabf04be706a865e84281c163f0b

Observation ead1e612-6841-4a2b-817b-ab9faebc6d7d · outbound

This paper cites AudioLDM: Text-to-Audio Generation with Latent Diffusion Models.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:57.138047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:57.138047Z digest=sha256:e7b0f1066a90bdf2f77aa629b88b9f00358827564fefa43541e8630e3b82ca7d

Observation 158aff55-058b-43f9-a4a9-4fc26dd9b46b · outbound

This paper cites PIXART-{\delta}: Fast and Controllable Image Generation with Latent Consistency Models.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet PIXART-{\delta}: Fast and Controllable Image Generation with Latent Consistency Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:57.179030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:57.179030Z digest=sha256:b7f8da50ff8ea550da9a3fc2c34e19cf10efdf4b1abecdd7fe8574473da62fc0

Observation c4db9e3c-35ea-45c6-aa2f-f3f918488ffb · outbound

This paper cites COCOLA: Coherence-Oriented Contrastive Learning of Musical Audio Representations.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet COCOLA: Coherence-Oriented Contrastive Learning of Musical Audio Representations

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:09:58.450395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:09:57.248961Z digest=sha256:40b4fd417ea56b512bb8acfe129c60fe975a59779703ad9e74e58b25104606a2

Observation d64c8f29-6bdc-49fe-a011-3afb21821ef2 · outbound

This paper cites Editing music with melody and text: Using controlnet for diffusion transformer,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Editing music with melody and text: Using controlnet for diffusion transformer,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:09:59.778127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:09:57.354123Z digest=sha256:7db499bfe76dbfd19c693ff69908c53c3b787d82ab14ccaae1a0e82944c829fc

Observation 5e669fe6-ece2-4d8e-88cf-6ced1feec5be · outbound

This paper cites Classifier-Free Diffusion Guidance.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Classifier-Free Diffusion Guidance

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:57.420079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:57.420079Z digest=sha256:13b9b4638b108c53deb020a4dd0fdf8a8fbd1632a3af45b81d7a1dd372eeb794

Observation e77c68c5-45da-461b-aeba-416f1234010b · outbound

This paper cites Muse: Text-To-Image Generation via Masked Generative Transformers.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:57.486525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:57.486525Z digest=sha256:1fcc4d43bacab7b89f7d0830b34cfbb30e4ae1129ef4dce538e9dfd85af907e1

Observation b1ce6187-a2e0-41d4-842d-2858f15c4994 · outbound

This paper cites Stemgen: A music generation model that listens,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Stemgen: A music generation model that listens,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:09:59.611073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:09:57.567373Z digest=sha256:d592df1e256702ab35d88699f615f0fc18ee971580c32a87b6e32c42af4bda7e

Observation 1c375ab3-c629-4e01-8ec7-e5d4b76227e4 · outbound

This paper cites Imagebind: One embedding space to bind them all,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Imagebind: One embedding space to bind them all,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:09:59.447837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:09:57.668416Z digest=sha256:291a06e41db5b0f5f23d958f67771763eb4ac577a512dc4dcbb6e4738b487dfa

Observation 899ae236-eb3b-43c4-83cd-44ea45f0aa47 · outbound

This paper cites Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:09:59.318967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:09:57.772268Z digest=sha256:49944306301c20907e02533c5d8a60fcd44380728bce2c5d0c9982dd692bc654

Observation f6054d9e-adc8-4eb1-aca5-4fb921a2b459 · outbound

This paper cites Audio set: An ontology and human-labeled dataset for audio events,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Audio set: An ontology and human-labeled dataset for audio events,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:09:59.207305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:09:57.862288Z digest=sha256:21689d2173c663c82fcede48851d735a864217b255314dea8a56b091b4b636e1

Observation a066d024-477c-4e68-9145-818d1990d3b7 · outbound

This paper cites Efficient Training of Audio Transformers with Patchout.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Efficient Training of Audio Transformers with Patchout

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:57.929534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:57.929534Z digest=sha256:4fcde99cfe27b1a4683f0c4e8dcf7181ddfe7a32ef878eaa46703f328b5974dd

Observation 9d1e8033-5093-4836-8fc2-544e8816659a · outbound

This paper cites Vggsound: A large- scale audio-visual dataset,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Vggsound: A large- scale audio-visual dataset,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:09:59.139077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:09:58.028545Z digest=sha256:13a6124b2a824855c751ff272c80c4b757fc077c320661fba3ac0ad18c6f553b

Observation 330a7253-e81a-4013-98aa-49fd7a4ff704 · outbound

This paper cites Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:09:59.019857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:09:58.119664Z digest=sha256:1e7cb7483109f15a4f2bbdcaf2f4a167a5e09df75f757f04b1a68757d03f63e9

Observation ecda695c-59e4-4d15-b432-4efb0dcb829f · outbound

This paper cites Cnn architectures for large-scale audio classification,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Cnn architectures for large-scale audio classification,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:09:58.866089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:09:58.194238Z digest=sha256:8053e902cb5a7923298afd8039a2723b1611e0dafbad0e91c2e95e76df4dcb60

Observation 9efe0af6-4ce8-454b-8bf9-b24274ca4411 · outbound

This paper cites Panns: Large-scale pretrained audio neural networks for audio pattern recognition,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Panns: Large-scale pretrained audio neural networks for audio pattern recognition,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:09:58.751950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:09:58.262045Z digest=sha256:08a71b9d190bf93493d7bd3afe262644fbd672024596c8cf203f7b81426f76c4

Pith citing papers

Observation 2678bf73-aeaa-4275-b270-963fb3ea0938 · inbound

MMAudioSep: Taming Video-to-Audio Generative Model Towards Video/Text-Queried Sound Separation cites this paper.

MMAudioSep: Taming Video-to-Audio Generative Model Towards Video/Text-Queried Sound Separation SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-18T08:21:06.817478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T08:20:02.986562Z digest=sha256:d12f35947a9a7e93b1f2c84ce73a1f1d3978423f57cacdb1181895f8236c2751