Pith. sign in

Paper Citation Record · LEDGER

Visual Representation Matters: Exploiting Temporal Differences in Video-to-Audio Generation

As of 10 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 0 inbound Pith citation observations for arXiv:2608.04902.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.04902 v1

Coverage vector

measured 17 of 17 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T14:02:31.138343Z

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

17 of 17 outbound references displayed

  • verified exact1
  • verified fuzzy11
  • unresolved4
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 45471de0-2efc-470e-b0c4-ccd50d499613 · outbound

This paper cites The Million Song Dataset.

Visual Representation Matters: Exploiting Temporal Differences in Video-to-Audio Generation The Million Song Dataset

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:02:31.499333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:02:31.067776Z digest=sha256:1f7db8d30895e67adbf168bfdde22abe3b59ddb04c89e0be7c511f7b3ea48a9c

Observation 02dce2a7-0f15-4d4f-9322-65f13d136b38 · outbound

This paper cites InICCV, 4195–4205.

Visual Representation Matters: Exploiting Temporal Differences in Video-to-Audio Generation InICCV, 4195–4205

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:02:31.429678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:02:31.094932Z digest=sha256:fc763de89eecab173bba5c51e1c95b8861c29972a85f51f4db4355fe0903acda

Observation dfdb8c95-54b5-42cc-bb6f-9ceff5005d74 · outbound

This paper cites HunyuanVideo-Foley: Multimodal Diffusion with Representation Alignment for High-Fidelity Foley Audio Generation.

Visual Representation Matters: Exploiting Temporal Differences in Video-to-Audio Generation HunyuanVideo-Foley: Multimodal Diffusion with Representation Alignment for High-Fidelity Foley Audio Generation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T14:02:31.099756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:02:31.099756Z digest=sha256:695038d9620c8c7cffa695a8a47c828d08a661df2e0739015d93781e2b2dc615

Observation 2b52b8b3-99e7-4d01-add4-fafc8a8e1fea · outbound

This paper cites InICASSP, 1–5.

Visual Representation Matters: Exploiting Temporal Differences in Video-to-Audio Generation InICASSP, 1–5

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:02:31.414717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:02:31.104491Z digest=sha256:42f53a99a8b279f2bf55f916fbadb61e2c0bf60f7fb084d346f0805ca4b360d2

Observation 3945e800-7001-402d-8428-18eb7951b55e · outbound

This paper cites an unresolved cited work.

Visual Representation Matters: Exploiting Temporal Differences in Video-to-Audio Generation Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:02:31.386492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:02:31.112670Z digest=sha256:b64c832a2f95145fa58ec5608152cb07c45060880bd0e0b18a8e39cfe9fb6f86

Observation 6c0ff6be-6ef1-4bc6-9df1-5f6eb67ecdd7 · outbound

This paper cites InICML, 62596–62626.

Visual Representation Matters: Exploiting Temporal Differences in Video-to-Audio Generation InICML, 62596–62626

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:02:31.371486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:02:31.117017Z digest=sha256:80bf9d6f34a57fca4e80f06ee5b6a0bdc688c943ec2967491cf78d2cc9c06a06

Observation 66eed211-a3b8-45f9-88d3-94138eb509cb · outbound

This paper cites Video-to-Audio Generation with Hidden Alignment.

Visual Representation Matters: Exploiting Temporal Differences in Video-to-Audio Generation Video-to-Audio Generation with Hidden Alignment

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T14:02:31.121546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:02:31.121546Z digest=sha256:2a89ce164d6f34bcd886a5c46adc48ae879838d2e951cc35f2c4f947236a4f15

Observation 6a24dab8-44da-4224-8599-e384618994c1 · outbound

This paper cites baby crying.

Visual Representation Matters: Exploiting Temporal Differences in Video-to-Audio Generation baby crying

Reference 17

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T14:02:31.310949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:02:31.138343Z digest=sha256:a4da98b69a0e5cdb08139528b6796e0bbbfc6d9951482c4d6f55f855091cb137

Observation c311c4f5-a012-47b8-ba7a-5376eeac1565 · outbound

This paper cites Snake ac- tivation functions are applied throughout the network, and no final tanh activation is used in the decoder.

Visual Representation Matters: Exploiting Temporal Differences in Video-to-Audio Generation Snake ac- tivation functions are applied throughout the network, and no final tanh activation is used in the decoder

Reference 640

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:02:31.326121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:02:31.134405Z digest=sha256:e6adaabc9ec33bf12a8d79aec83a618ead712a332de77c318e1d227363e7744e

Observation dd7e5510-caa3-461c-95c7-a3634daa1938 · outbound

This paper cites K.; Ishii, M.; Hayakawa, A.; Shibuya, T.; Schwing, A.; andMitsufuji,Y.2025.MMAudio:TamingMultimodalJointTrain- ing for High-Quality Video-to-Audio Synthesis.

Visual Representation Matters: Exploiting Temporal Differences in Video-to-Audio Generation K.; Ishii, M.; Hayakawa, A.; Shibuya, T.; Schwing, A.; andMitsufuji,Y.2025.MMAudio:TamingMultimodalJointTrain- ing for High-Quality Video-to-Audio Synthesis

Reference 725

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:02:31.485720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:02:31.072369Z digest=sha256:f2e68171603d85a346602315c134d10b633a6d77e38992339b255caa93bdfd09

Observation 6f2360e0-6179-4b73-99c8-0f77d5b01876 · outbound

This paper cites Dataset Task Hours (h) Source AudioCaps T2A 109 (Kim et al.

Visual Representation Matters: Exploiting Temporal Differences in Video-to-Audio Generation Dataset Task Hours (h) Source AudioCaps T2A 109 (Kim et al

Reference 2013

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:02:31.341681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:02:31.130418Z digest=sha256:df43b03669a82ada2a306033c8e2c21ca1cd4026ab7c134fc495217114ea1eda

Observation e5e2d683-b662-418a-8a5c-cffad615da28 · outbound

This paper cites In ICASSP, 776–780.

Visual Representation Matters: Exploiting Temporal Differences in Video-to-Audio Generation In ICASSP, 776–780

Reference 2017

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:02:31.458005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:02:31.086327Z digest=sha256:2ac146aad2f4d15796d88396e2570a8e8512100585194b40c3a890c1ed9e19b1

Observation 2f9cbe26-a44d-4c5a-8ffd-a1aca9ac9a19 · outbound

This paper cites InECCV, 570–586.

Visual Representation Matters: Exploiting Temporal Differences in Video-to-Audio Generation InECCV, 570–586

Reference 2018

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:02:31.356595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:02:31.126304Z digest=sha256:3fb1095ba4b335c89635971825f5f412d85de6aaeed5235a2748baf761f8c7ef

Observation 301fe667-d513-4690-8b29-1bcc1ac6b95f · outbound

This paper cites InICML, 21450–21474.

Visual Representation Matters: Exploiting Temporal Differences in Video-to-Audio Generation InICML, 21450–21474

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:02:31.443727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:02:31.090647Z digest=sha256:c47da588c2c9dec3655afe36eff1b58444d87612b5f5e3cbe22a350a2f06323f

Observation 5de61743-68ed-4801-ae66-0c8a82a90731 · outbound

This paper cites InICML, 46804–46822.

Visual Representation Matters: Exploiting Temporal Differences in Video-to-Audio Generation InICML, 46804–46822

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:02:31.400600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:02:31.108705Z digest=sha256:7cb4243c2734f2a23efefae98f3a870c38278f1f70291053e98dce4404e60924

Observation 7e551071-0bbc-4330-9fb5-20bc92722f03 · outbound

This paper cites an unresolved cited work.

Visual Representation Matters: Exploiting Temporal Differences in Video-to-Audio Generation Unresolved cited work

Reference 2025

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:02:31.471592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:02:31.081777Z digest=sha256:06cf7e69bf20773203fb540b4fd53b47210ab514acf6b83517c2b3cd73ebddcb

Observation 3cdb3f24-129f-4066-a256-8e4cced7f829 · outbound

This paper cites arXiv:2606.15956.

Visual Representation Matters: Exploiting Temporal Differences in Video-to-Audio Generation arXiv:2606.15956

Reference 2026

Resolution
verified exact
raw_fallback, observed 2026-08-06T14:02:31.296477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:02:31.076800Z digest=sha256:3b3a33227121dc5715865e89203417582cd658638abe1ab8c9ca5e5baba38907

Pith citing papers

No inbound Pith citation observations are available.