Pith. sign in

Paper Citation Record · LEDGER

Towards Video to Piano Music Generation with Chain-of-Perform Support Benchmarks

As of 17 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 0 inbound Pith citation observations for arXiv:2505.20038.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.20038 v1

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:05:12.150659Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

25 of 25 outbound references displayed

  • verified exact3
  • verified fuzzy8
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5001ff69-54ec-4cfe-8d99-1f5e3c17a8b5 · outbound

This paper cites Video-Guided Foley Sound Generation with Multimodal Controls.

Towards Video to Piano Music Generation with Chain-of-Perform Support Benchmarks Video-Guided Foley Sound Generation with Multimodal Controls

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:05:09.819840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:05:09.819840Z digest=sha256:6b0e8a4a05d1a71839ec1be026fc73850639c848d3b3d6325855b2a5eca607cf

Observation bade64f8-cec2-4ab2-ab56-2e7f85d2b730 · outbound

This paper cites YingSound: Video-Guided Sound Effects Generation with Multi-modal Chain-of-Thought Controls.

Towards Video to Piano Music Generation with Chain-of-Perform Support Benchmarks YingSound: Video-Guided Sound Effects Generation with Multi-modal Chain-of-Thought Controls

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:05:12.786853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:05:09.874214Z digest=sha256:0d52f5c1936834d70b1ab6d5706664317f7ec937e4430496f8afa6f695e4579f

Observation 2309e642-f146-464a-a083-36fb2b903a66 · outbound

This paper cites MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis.

Towards Video to Piano Music Generation with Chain-of-Perform Support Benchmarks MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:05:09.943750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:05:09.943750Z digest=sha256:b4bcac1eff4b7f9a30f198efbd39aab8e3c6040ca0916a1998b5a9d0317b6937

Observation 0d41f6e1-c5b0-4aba-8b10-537ae6f9645d · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Towards Video to Piano Music Generation with Chain-of-Perform Support Benchmarks DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:05:10.009315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:05:10.009315Z digest=sha256:571c8068dcaf7f63be6cdaadef3bf1734ccce2d41dd9886a250a3ecb1bbfefd2

Observation 26bf0cc3-e0af-410d-8e83-5b219d1d5a88 · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilib- rium.

Towards Video to Piano Music Generation with Chain-of-Perform Support Benchmarks Gans trained by a two time-scale update rule converge to a local nash equilib- rium

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:05:10.132451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:05:10.132451Z digest=sha256:05a1d581dcd3e0ffd0523347f652f1c7ff482983fe7014206dd8ff88f6916442

Observation be61c9ec-bfb7-4bdc-9141-01856861c8e0 · outbound

This paper cites Taming visually guided sound generation.

Towards Video to Piano Music Generation with Chain-of-Perform Support Benchmarks Taming visually guided sound generation

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:05:14.505860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:05:10.225285Z digest=sha256:6fc4b63899bf831db78d9bc22ce99e9f10a6a158a9fbe2f032d5c92e28fb5846

Observation 0562a702-3591-42da-91a1-9f6291dbbe66 · outbound

This paper cites Sophia Koepke, Olivia Wiles, Yael Moses, and Andrew Zisserman.

Towards Video to Piano Music Generation with Chain-of-Perform Support Benchmarks Sophia Koepke, Olivia Wiles, Yael Moses, and Andrew Zisserman

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:05:14.327614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:05:10.374623Z digest=sha256:0c2138e76d5239af2bc4a4e307c757a8766807c41569b3abbc78efa80e17647a

Observation 84250236-4a2f-410b-8b7e-80c0110144ff · outbound

This paper cites Crandall, and Christopher Raphael.

Towards Video to Piano Music Generation with Chain-of-Perform Support Benchmarks Crandall, and Christopher Raphael

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:05:14.101676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:05:10.495877Z digest=sha256:e3398f3bb8d60cd539ab80c758526984305ab46d027e979bf887a6e763d947b0

Observation 8e7bcbcc-767a-40d2-ae68-369c5ecf4d75 · outbound

This paper cites Tri-Ergon: Fine-grained Video-to-Audio Generation with Multi-modal Conditions and LUFS Control.

Towards Video to Piano Music Generation with Chain-of-Perform Support Benchmarks Tri-Ergon: Fine-grained Video-to-Audio Generation with Multi-modal Conditions and LUFS Control

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:05:12.592926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:05:10.608172Z digest=sha256:4b8153a74f21ca568acf6749cc38cc75f9cd0ea07db1f3f09ef24004c7fa31fc

Observation 58c5112c-09e0-4fee-8b63-9ea593e6c88e · outbound

This paper cites Diff-foley: Synchronized video-to-audio synthesis with la- tent diffusion models, 2023.

Towards Video to Piano Music Generation with Chain-of-Perform Support Benchmarks Diff-foley: Synchronized video-to-audio synthesis with la- tent diffusion models, 2023

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:05:13.840027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:05:10.693496Z digest=sha256:ba401b7771186cab7a460fa4588803afe88d511943b3263bd14fce5636f70a63

Observation 142adb7a-b282-4e1f-b92e-0b5caa682562 · outbound

This paper cites Foleygen: Visually-guided audio generation.

Towards Video to Piano Music Generation with Chain-of-Perform Support Benchmarks Foleygen: Visually-guided audio generation

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:05:13.604068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:05:10.818067Z digest=sha256:534739e0914caff80f8b7d593dbac54e7d3ee7b294d2ccc904cd0c51cb5248fc

Observation 5726a169-6270-466d-b1fa-5fbad09bc8d6 · outbound

This paper cites Qwen2.5 technical report, 2025.

Towards Video to Piano Music Generation with Chain-of-Perform Support Benchmarks Qwen2.5 technical report, 2025

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:05:13.403506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:05:10.960063Z digest=sha256:071af4b075fb1a645202cd0586df541f38d8c4b840124325d744c5003b3251ba

Observation 571820b7-73f0-4787-ac9a-5534892b862e · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Towards Video to Piano Music Generation with Chain-of-Perform Support Benchmarks Direct preference optimization: Your language model is secretly a reward model

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:05:11.035226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:05:11.035226Z digest=sha256:995d681e8f82fb7942888d5767172ce360e6ca5dca2a53a249a9cb7d3734c1f6

Observation 0c40d3e8-45a1-4301-9083-8ca8d04573a2 · outbound

This paper cites Improved techniques for training gans.

Towards Video to Piano Music Generation with Chain-of-Perform Support Benchmarks Improved techniques for training gans

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:05:11.152150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:05:11.152150Z digest=sha256:151e2498b6b9c343e425f3dadd78efe56a52bd815d4dd8be4244ec52a009c435

Observation 2fde6f11-e80f-45b7-9201-2879022d1fe5 · outbound

This paper cites Audeo: Audio Generation for a Silent Performance Video.

Towards Video to Piano Music Generation with Chain-of-Perform Support Benchmarks Audeo: Audio Generation for a Silent Performance Video

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:05:12.402036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:05:11.255279Z digest=sha256:677c674490c15ec68fd136f28f280fbe652d9ae173264bb915c47d7dac1e7ab8

Observation bfde4dd2-da87-4953-a989-e1eb11484bd0 · outbound

This paper cites AudioX: A Unified Framework for Anything-to-Audio Generation.

Towards Video to Piano Music Generation with Chain-of-Perform Support Benchmarks AudioX: A Unified Framework for Anything-to-Audio Generation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:05:11.305127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:05:11.305127Z digest=sha256:73351c2db2154c7266fd96d046db1c5cb488e1cf94a1cf95d1cb2322cda9cc09

Observation 8aab714f-f38e-46bd-abbc-3c3d4a717b46 · outbound

This paper cites Temporally Aligned Audio for Video with Autoregression.

Towards Video to Piano Music Generation with Chain-of-Perform Support Benchmarks Temporally Aligned Audio for Video with Autoregression

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:05:11.397850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:05:11.397850Z digest=sha256:b39fcc5960ab1684ad007906f6ca2fdb616801ca757f3bfd7f1670cd1eccb08e

Observation 4163fa1e-cbdb-4479-8a77-037d6e2d1217 · outbound

This paper cites V2a-mapper: A lightweight solution for vision-to-audio generation by connecting foun- dation models, 2023.

Towards Video to Piano Music Generation with Chain-of-Perform Support Benchmarks V2a-mapper: A lightweight solution for vision-to-audio generation by connecting foun- dation models, 2023

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:05:13.199507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:05:11.478150Z digest=sha256:f5256b2e14b1970360a92b2e6f92246e8985cdad005da8d5d2aca4bc15c8e7c2

Observation c847c1af-9770-42ea-b268-32cd39ee6b86 · outbound

This paper cites Frieren: Efficient Video-to-Audio Generation Network with Rectified Flow Matching.

Towards Video to Piano Music Generation with Chain-of-Perform Support Benchmarks Frieren: Efficient Video-to-Audio Generation Network with Rectified Flow Matching

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:05:11.602718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:05:11.602718Z digest=sha256:bfd1a86d5a3917bb1007e410f7a9e02ffa1d66fc512d996b087b77f3dff6e5f7

Observation d01e4aff-f9f4-49a1-abdd-c810a9e3ff68 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large lan- guage models.

Towards Video to Piano Music Generation with Chain-of-Perform Support Benchmarks Chain-of-thought prompting elicits reasoning in large lan- guage models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:05:11.717744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:05:11.717744Z digest=sha256:f6541efbc7c65128ac8abb46f4ad02a8636db3ae426025544827d060b7fabd97

Observation 04a16f20-97f0-4ae9-ba51-3c4b66cfaf48 · outbound

This paper cites Large-scale con- trastive language-audio pretraining with feature fusion and keyword-to-caption augmentation.

Towards Video to Piano Music Generation with Chain-of-Perform Support Benchmarks Large-scale con- trastive language-audio pretraining with feature fusion and keyword-to-caption augmentation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:05:11.786319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:05:11.786319Z digest=sha256:ace87f0db13563493c0067af9af2acb31fdbe16b0326c92d28d0aea2588c8e0d

Observation c82a12f4-a32d-4160-8063-9d5fe817a8ea · outbound

This paper cites LLaVA-CoT: Let Vision Language Models Reason Step-by-Step.

Towards Video to Piano Music Generation with Chain-of-Perform Support Benchmarks LLaVA-CoT: Let Vision Language Models Reason Step-by-Step

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:05:11.869274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:05:11.869274Z digest=sha256:e5ee903ce5a9b02932c78d35ffe0ec22879737d0d6440e71df55cfe420bfd97c

Observation 408206cd-881b-4eb6-be8a-b9c0006f1583 · outbound

This paper cites Diverse and aligned audio-to-video genera- tion via text-to-video model adaptation.

Towards Video to Piano Music Generation with Chain-of-Perform Support Benchmarks Diverse and aligned audio-to-video genera- tion via text-to-video model adaptation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:05:12.973198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:05:11.940525Z digest=sha256:47ef3ce2f3e82b0c63c5d702fc24e8c5b15c35c8a3878cab66b7274329b170bd

Observation 10a5b411-b034-47a1-b8e3-b003ce232157 · outbound

This paper cites Improve Vision Language Model Chain-of-thought Reasoning.

Towards Video to Piano Music Generation with Chain-of-Perform Support Benchmarks Improve Vision Language Model Chain-of-thought Reasoning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:05:12.046711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:05:12.046711Z digest=sha256:c36434993c48b05f36a5142e8ce4bb0d79539b8661342915dadae0ebd6436f6f

Observation e5935e36-2c37-445e-81ae-9d1e08b05f4c · outbound

This paper cites FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds.

Towards Video to Piano Music Generation with Chain-of-Perform Support Benchmarks FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:05:12.150659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:05:12.150659Z digest=sha256:9fc12c012a3fe3ee4ccfef932a792ab7b1cec662e6f5565b74e214a800fa3233

Pith citing papers

No inbound Pith citation observations are available.