Pith. sign in

Paper Citation Record · LEDGER

AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences

As of 10 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2507.00475.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.00475 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:19:53.784375Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact3
  • verified fuzzy16
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bff84d31-ef32-4d78-84c4-69ce575a5ab6 · outbound

This paper cites Leveraging AI to Generate Audio for User-generated Content in Video Games.

AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences Leveraging AI to Generate Audio for User-generated Content in Video Games

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:19:54.404688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:19:52.218751Z digest=sha256:8a692bd313cd54f6e402806e153b59c6d0a84484b43e78d3c991f7f9fb84af52

Observation ced172d0-bccc-4d23-9345-70dfbb000bfd · outbound

This paper cites How should we evaluate synthesized environmental sounds,.

AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences How should we evaluate synthesized environmental sounds,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:57.881042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:19:52.286779Z digest=sha256:7571964ccead3fca575450c5715cf264c5ee58f1ac65cf36f8bf6062858a996e

Observation 82ffe15f-da9e-4bac-b782-0f8ef3f81e72 · outbound

This paper cites An effective quality evaluation protocol for speech enhancement algorithms,.

AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences An effective quality evaluation protocol for speech enhancement algorithms,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:57.622417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:19:52.327726Z digest=sha256:fffd7631c32c02fa07940694d0ed39a5d23a22fa51f5f2350fabfd285ca26bab

Observation 0201f1fa-03c6-481f-ae34-d95e75947a30 · outbound

This paper cites Mel-cepstral distance measure for objective speech quality assessment,.

AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences Mel-cepstral distance measure for objective speech quality assessment,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:57.378676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:19:52.405548Z digest=sha256:870c5ed72f141c7f8bc05f955e8c99ac70cca717870859313fdce4484e6ced04

Observation b14f7aa4-b61b-4b6c-a8ba-379dac7f31fb · outbound

This paper cites Human-CLAP: Human-perception-based contrastive language-audio pretraining,.

AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences Human-CLAP: Human-perception-based contrastive language-audio pretraining,

Reference 5

Resolution
verified exact
raw_fallback, observed 2026-08-06T21:19:54.276450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:19:52.469206Z digest=sha256:36836edc1882a92f45a75e47c05eebd81e3733e5b185e8551045a9739719433a

Observation 265ce1c3-27ec-45c7-9cf2-6d144038162f · outbound

This paper cites Self-supervised Audio Teacher-Student Transformer for Both Clip-level and Frame-level Tasks.

AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences Self-supervised Audio Teacher-Student Transformer for Both Clip-level and Frame-level Tasks

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:19:54.006373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:19:52.531504Z digest=sha256:d0f396f0f8a08c28be3bccf09dcfca2e1e61c6016a8e8aed0fadae3082981249

Observation 133a3232-6ffe-4d4a-94d8-936a231fffbd · outbound

This paper cites Foley Sound Synthesis at the DCASE 2023 Challenge.

AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences Foley Sound Synthesis at the DCASE 2023 Challenge

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:52.610462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:52.610462Z digest=sha256:366232c019751b93538f37195a555e1897183676cc9fc66af6df04093fd870ad

Observation 6d174fd1-5325-4c00-9d6a-46fcbfaa7a26 · outbound

This paper cites Using dynamic time warping to find patterns in time series,.

AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences Using dynamic time warping to find patterns in time series,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:57.177899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:19:52.646054Z digest=sha256:3a3375d34eed04e0b56c25c49c5e7d757784dbcfd0d039ec5609a5ebfe4e6cce

Observation a6d0491c-3747-4e14-bcc1-f031f6ba3cdb · outbound

This paper cites Dynamic time warping under subsequence,.

AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences Dynamic time warping under subsequence,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:56.920610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:19:52.719921Z digest=sha256:1f4e381d05a6812b6331d4e3626cf21176989f0fc5efe1d6fef3c169e0dcc621

Observation 80ddb99f-504a-460e-9362-ec6e1f1bcf00 · outbound

This paper cites Relate: Subjective evaluation dataset for automatic evaluation of relevance between text and audio,.

AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences Relate: Subjective evaluation dataset for automatic evaluation of relevance between text and audio,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:56.721799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:19:52.782560Z digest=sha256:a0504c50b69b1cd00ec6dc03e0a811f77c31e0e3a641b4cf9a43f76355eb8445

Observation d1a964c5-3e2a-4db2-b122-421b597c6a0d · outbound

This paper cites PAM: Prompting Audio-Language Models for Audio Quality Assessment.

AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences PAM: Prompting Audio-Language Models for Audio Quality Assessment

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:52.851246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:52.851246Z digest=sha256:a00b99b1f43c832163b3002e8072b96367c9bbb253776cea574bcc017f1e62c2

Observation 7feaf2e8-91c9-45f7-9929-e699670a212b · outbound

This paper cites Speech- BERTScore: Reference-aware automatic evaluation of speech generation leveraging nlp evaluation metrics,.

AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences Speech- BERTScore: Reference-aware automatic evaluation of speech generation leveraging nlp evaluation metrics,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:56.481697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:19:52.932662Z digest=sha256:aa145c13c39976fffdc39f36746643c0396acf03ee2614ecb0a490afdd2e5b37

Observation 4eba9c28-1718-4af8-b975-33ad5893b314 · outbound

This paper cites BERTScore: Evaluating text generation with bert,.

AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences BERTScore: Evaluating text generation with bert,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:56.295310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:19:53.023947Z digest=sha256:751fca58bf5d572aa3760c4707c8c56c58f2d3429ee21d3f1affbc23f5eed714

Observation 47102028-a538-4314-89c2-0e946891576c · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:53.067501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:53.067501Z digest=sha256:5ca2bd60bb26d022e25e533b3933f4ae159f97692031293eac0ead5f0a8eba1a

Observation b742a959-b4ce-4b31-9cbe-a9fde70627e1 · outbound

This paper cites HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,.

AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:56.105669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:19:53.127818Z digest=sha256:a252a2854ff367f8046337a8f181ec1a5cb7a3ef0f60675649115eb3f9b69cc1

Observation eee17444-cae4-44fc-8873-4ab47fe9dca4 · outbound

This paper cites AudioCaps: Generating captions for audios in the wild,.

AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences AudioCaps: Generating captions for audios in the wild,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:55.889041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:19:53.208905Z digest=sha256:0e8da9c197475ac64ca75ed02c0cca60bd69608491e0c7bbaf929de284c15959

Observation 17d59307-b38d-4257-8c92-a740fd511246 · outbound

This paper cites AudioLDM 2: Learning holistic audio generation with self-supervised pretraining,.

AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences AudioLDM 2: Learning holistic audio generation with self-supervised pretraining,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:53.251131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:53.251131Z digest=sha256:e4eeba4008401e819efb9c694651759e7dff3d8cae1120e0c83db2636fc70d32

Observation 0a6dbeca-be39-4df6-811f-c9d77ee23099 · outbound

This paper cites AudioLDM: Text-to-audio generation with latent diffusion models,.

AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences AudioLDM: Text-to-audio generation with latent diffusion models,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:55.652067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:19:53.316231Z digest=sha256:ac2f8c5230824aeed88b39817eebbf5728b076b7c90df163e1bf5f68135f8ac1

Observation 22a2f17d-6983-401a-9a7e-2026ed796308 · outbound

This paper cites AudioGen: Textually Guided Audio Generation.

AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences AudioGen: Textually Guided Audio Generation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:53.373430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:53.373430Z digest=sha256:668f9715b7267f34b49fa47bd841adedadc9676cce68eee37b6010ce6584ac0b

Observation 57db31d8-1b16-40b2-b399-f4ef562e7b20 · outbound

This paper cites BYOL for Audio: Exploring pre-trained general-purpose audio representations,.

AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences BYOL for Audio: Exploring pre-trained general-purpose audio representations,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:55.362877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:19:53.409957Z digest=sha256:4924fb4d5899cb973ab0351acc9084cb34c954f6acd83e7b2703050f8872b583

Observation 7be1632e-89a8-472e-ab05-4b491b7f8565 · outbound

This paper cites Backpropagation applied to handwritten zip code recognition,.

AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences Backpropagation applied to handwritten zip code recognition,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:53.471791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:53.471791Z digest=sha256:f5d8882a15d417ded64cb2d28794689880e0a475f693fa07ad03dbf23fdd061d

Observation cadd9621-dfc4-458e-aa49-1478b4d5916c · outbound

This paper cites Attention Is All You Need.

AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences Attention Is All You Need

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:53.511609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:53.511609Z digest=sha256:72594c33abf9d4d44cc6779ee4beca67e9fcf49dcd7c24f53666363c23430325

Observation 1ff852fa-0a3a-448d-a0ba-b0c7ac9a08ce · outbound

This paper cites AST: Audio spectrogram trans- former,.

AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences AST: Audio spectrogram trans- former,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:55.153487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:19:53.573314Z digest=sha256:3e998a79aba2d1e0d1a521269e98a06b1f31d1707e1ee817cc4fd0e8b0666bba

Observation c69a0d0a-f007-4243-84ca-d10ffcdb4952 · outbound

This paper cites W ARP-Q: Quality prediction for generative neural speech codecs,.

AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences W ARP-Q: Quality prediction for generative neural speech codecs,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:54.978164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:19:53.630560Z digest=sha256:5415ee0371e290650154a59e72352fb0ebeb0c590b8f81ee60a25ecf2be0d387

Observation 86d194db-b93a-432f-becf-fed237d3f5be · outbound

This paper cites Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms.

AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:53.679676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:53.679676Z digest=sha256:e37d5b999db0bdaa67c802ccb5ba3f888d36e4dd2e6ce5f8e2c94c44514b7969

Observation 69fde75b-1511-4bf7-9cb4-46accb36684e · outbound

This paper cites A reference-free metric for language-queried audio source separation using contrastive language-audio pretraining,.

AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences A reference-free metric for language-queried audio source separation using contrastive language-audio pretraining,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:54.738962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:19:53.743187Z digest=sha256:42ce450b6db35089018034170c24ba05d22f4c8227e5d495ab49245e65873ba0

Observation c1a49af6-baf8-462e-97bd-73da735584c0 · outbound

This paper cites Conformer: Local features coupling global representations for recognition and detection,.

AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences Conformer: Local features coupling global representations for recognition and detection,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:54.507500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:19:53.784375Z digest=sha256:feac4222a3df4de4f7ebe0287070c26ea67c1fac3dfcf00574b874741b4809ac

Pith citing papers

No inbound Pith citation observations are available.