Pith. sign in

Paper Citation Record · LEDGER

Improving Perceptual Audio Aesthetic Assessment via Triplet Loss and Self-Supervised Embeddings

As of 14 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 1 inbound Pith citation observation for arXiv:2509.03292.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.03292 v1

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T11:04:46.304755Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-26T03:16:39.283647Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T14:39:57.486114Z

Reference resolution

22 of 22 outbound references displayed

  • verified exact0
  • verified fuzzy21
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ef0f1f39-2be5-4036-8465-7526be8fa775 · outbound

This paper cites The V oiceMOS Challenge 2022,.

Improving Perceptual Audio Aesthetic Assessment via Triplet Loss and Self-Supervised Embeddings The V oiceMOS Challenge 2022,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T11:04:49.956454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T11:04:45.715077Z digest=sha256:7a6ae45a5c55b4848a66bb5d08e895af1adbb5b005db8e086c94a36eac297e40

Observation 9567268a-eef9-40fd-a8e9-2b82272ab0a7 · outbound

This paper cites The voicemos challenge 2023: Zero-shot subjective speech quality prediction for multiple domains,.

Improving Perceptual Audio Aesthetic Assessment via Triplet Loss and Self-Supervised Embeddings The voicemos challenge 2023: Zero-shot subjective speech quality prediction for multiple domains,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T11:04:49.804754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T11:04:45.764757Z digest=sha256:0d594ed65d74b47c78daa1de999eb2d7c0cc01c4897cc4abf5ee3c4a7b3ffeb9

Observation c2619186-770a-466a-a56c-a712488fe67f · outbound

This paper cites The voicemos challenge 2024: Beyond speech quality prediction,.

Improving Perceptual Audio Aesthetic Assessment via Triplet Loss and Self-Supervised Embeddings The voicemos challenge 2024: Beyond speech quality prediction,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T11:04:49.664750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T11:04:45.815103Z digest=sha256:8c4449523dca5d3c569f1a6a29fb25806a3e8bd1f58309d4b085457a580ac880

Observation 63b89fdb-bc2e-4edd-af12-75407fc573ce · outbound

This paper cites Fusion of self-supervised learned models for MOS prediction,.

Improving Perceptual Audio Aesthetic Assessment via Triplet Loss and Self-Supervised Embeddings Fusion of self-supervised learned models for MOS prediction,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T11:04:49.574796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T11:04:45.829892Z digest=sha256:bdb17d18753e2418902880452d3c1c14e8278006caf2f89957e4a12b43718150

Observation 2f397553-9393-4d7d-90a1-558ec1d3de1b · outbound

This paper cites UTMOS: UTokyo-SaruLab system for V oiceMOS Chal- lenge 2022,.

Improving Perceptual Audio Aesthetic Assessment via Triplet Loss and Self-Supervised Embeddings UTMOS: UTokyo-SaruLab system for V oiceMOS Chal- lenge 2022,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T11:04:49.287377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T11:04:45.861257Z digest=sha256:bfd0d85861e2448b4c8da70dc1b1e9054c3dd5b99f797387ae6a75b0ae828de5

Observation e24e072a-f79a-4620-ae78-c5ef5863ce3e · outbound

This paper cites A study on incorporating Whisper for robust speech assessment,.

Improving Perceptual Audio Aesthetic Assessment via Triplet Loss and Self-Supervised Embeddings A study on incorporating Whisper for robust speech assessment,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T11:04:49.044899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T11:04:45.877952Z digest=sha256:e295293177600f4647ac4202d338fc97c290094430a4cc6b0558d06cf9e3e4e3

Observation 525cc080-8a22-4e23-963e-6c0f47a24838 · outbound

This paper cites Corn: Co-trained full- and no-reference speech quality assessment,.

Improving Perceptual Audio Aesthetic Assessment via Triplet Loss and Self-Supervised Embeddings Corn: Co-trained full- and no-reference speech quality assessment,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T11:04:48.819873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T11:04:45.897598Z digest=sha256:9b43cb27406c37585f31eadff05e8f660c8b53ff4b05c4623bb0fe2eaba69fa2

Observation a585ad7b-8c1f-489e-b033-09fc09661743 · outbound

This paper cites Enabling auditory large language models for automatic speech quality evaluation,.

Improving Perceptual Audio Aesthetic Assessment via Triplet Loss and Self-Supervised Embeddings Enabling auditory large language models for automatic speech quality evaluation,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T11:04:48.642848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T11:04:45.924752Z digest=sha256:8992af24535e6f173197333c2d8fb67528042a03a465edf2e62ed4d8f2f529d2

Observation 5ff3ec75-e2b8-4988-9d3e-1c5cb503a623 · outbound

This paper cites Speech foundation models on intelligibility prediction for hearing-impaired listeners,.

Improving Perceptual Audio Aesthetic Assessment via Triplet Loss and Self-Supervised Embeddings Speech foundation models on intelligibility prediction for hearing-impaired listeners,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T11:04:48.358044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T11:04:45.965072Z digest=sha256:f7239e4110b16b092bfac22c279abe176bbb069790ae3bcccc8ddf1627b8d735

Observation 6b0422e6-2a41-4b30-aab6-f039c8a1485a · outbound

This paper cites Non-intrusive speech intelligibility prediction for hearing- impaired users using intermediate ASR features and human memory models,.

Improving Perceptual Audio Aesthetic Assessment via Triplet Loss and Self-Supervised Embeddings Non-intrusive speech intelligibility prediction for hearing- impaired users using intermediate ASR features and human memory models,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T11:04:48.234178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T11:04:45.984833Z digest=sha256:8c5e11e7651c7c696e195140515e8bc82700e2819644e46ce8c867bc5149c114

Observation e96fcfb8-51fc-4839-8e77-a5be341db9f6 · outbound

This paper cites A review on subjective and objective evaluation of synthetic speech,.

Improving Perceptual Audio Aesthetic Assessment via Triplet Loss and Self-Supervised Embeddings A review on subjective and objective evaluation of synthetic speech,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T11:04:48.090620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T11:04:46.034755Z digest=sha256:7d9cd5e9fdbd4cce0a7c8332b6be0aeafe0e8afb60f174ca6c84b39d0e665a5a

Observation 3d00225f-8c76-4cf9-81ef-fd826fae7a74 · outbound

This paper cites Self-supervised speech quality estimation and enhancement using only clean speech,.

Improving Perceptual Audio Aesthetic Assessment via Triplet Loss and Self-Supervised Embeddings Self-supervised speech quality estimation and enhancement using only clean speech,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T11:04:47.954912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T11:04:46.076073Z digest=sha256:bf8a2f5235b76aa37437ca2a0eea986ad7bcc829c181a1edad0d2a3088d2dea0

Observation 956ee022-43bc-4f78-bf2d-c9fa8c2f8ab6 · outbound

This paper cites Generalization ability of MOS prediction networks,.

Improving Perceptual Audio Aesthetic Assessment via Triplet Loss and Self-Supervised Embeddings Generalization ability of MOS prediction networks,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T11:04:47.807494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T11:04:46.083341Z digest=sha256:fb88d87b7f90edfa7e54b25bf6f389a358c5670a5386b583ee773129f9161986

Observation 1c7fed5d-9104-4af0-8c3c-5fec5e99ef6f · outbound

This paper cites Deep learning-based non-intrusive multi-objective speech assessment model with cross-domain features,.

Improving Perceptual Audio Aesthetic Assessment via Triplet Loss and Self-Supervised Embeddings Deep learning-based non-intrusive multi-objective speech assessment model with cross-domain features,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T11:04:47.617912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T11:04:46.097720Z digest=sha256:9f550e431f8198faf95e9deaa3e91aedcd5f8108efad1f5e4a2e95b92b50788f

Observation edefb18b-f081-48bb-96c6-4d6210cf10a9 · outbound

This paper cites Speechlmscore: Evaluat- ing speech generation using speech language model,.

Improving Perceptual Audio Aesthetic Assessment via Triplet Loss and Self-Supervised Embeddings Speechlmscore: Evaluat- ing speech generation using speech language model,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T11:04:47.416887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T11:04:46.162639Z digest=sha256:c104b7be5236ecdb927c4cd026ce1a9b3adfbe402e55b4da9a81ad2a36493c19

Observation f77c807d-9c99-4570-89d1-b8a4cef47988 · outbound

This paper cites BEATs: Audio pre-training with acoustic tokenizers,.

Improving Perceptual Audio Aesthetic Assessment via Triplet Loss and Self-Supervised Embeddings BEATs: Audio pre-training with acoustic tokenizers,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T11:04:47.215679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T11:04:46.168984Z digest=sha256:f1bb02df853ee30af6506ef3f1f4d1c890670a59a6d058c5383c93d4e12b2d20

Observation 83ef93fd-94b9-42be-9f3d-71b756dafc1a · outbound

This paper cites MBNet: MOS prediction for synthesized speech with mean-bias network,.

Improving Perceptual Audio Aesthetic Assessment via Triplet Loss and Self-Supervised Embeddings MBNet: MOS prediction for synthesized speech with mean-bias network,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T11:04:46.999369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T11:04:46.188563Z digest=sha256:7fe9ac5bc7315e0d23679a573e4db0c8557ccaf5c1a5006e7ff8393e9837c03d

Observation 35cfd7a4-db72-4a07-a230-6b365b7cfbfa · outbound

This paper cites LDNet: Unified listener dependent modeling in MOS prediction for synthetic speech,.

Improving Perceptual Audio Aesthetic Assessment via Triplet Loss and Self-Supervised Embeddings LDNet: Unified listener dependent modeling in MOS prediction for synthetic speech,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T11:04:46.797786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T11:04:46.207045Z digest=sha256:ec3f59cea3ae4f8f39b964deb7d8e78e0044151da60514f44041702554b565c8

Observation 31fcf490-1523-45ff-a782-7b7d19f25e0d · outbound

This paper cites HAAQI- Net: A non-intrusive neural music audio quality assessment model for hearing aids,.

Improving Perceptual Audio Aesthetic Assessment via Triplet Loss and Self-Supervised Embeddings HAAQI- Net: A non-intrusive neural music audio quality assessment model for hearing aids,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T11:04:46.737263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T11:04:46.211742Z digest=sha256:45506e42a52160482cb49ec98121335151c9a05b41b352391cc592175b39c61e

Observation 384d0956-531d-40ee-8827-e20475294701 · outbound

This paper cites FaceNet: A unified embed- ding for face recognition and clustering,.

Improving Perceptual Audio Aesthetic Assessment via Triplet Loss and Self-Supervised Embeddings FaceNet: A unified embed- ding for face recognition and clustering,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T11:04:46.659417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T11:04:46.225739Z digest=sha256:0c5090a994164b5cce53ddfcd58960a93cad38df9bba769b7de8ff260ee06376

Observation 0b919c50-5eae-4dad-8989-9c092755494a · outbound

This paper cites Unsupervised feature learning via non-parametric instance discrimination,.

Improving Perceptual Audio Aesthetic Assessment via Triplet Loss and Self-Supervised Embeddings Unsupervised feature learning via non-parametric instance discrimination,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T11:04:46.558959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T11:04:46.274755Z digest=sha256:f97fcde42581f78f1f29f2bc5dbdb9e44149660bf2e4ddfa322d7577302d3e40

Observation a7c5feaa-4371-4fa3-acac-4a755db43373 · outbound

This paper cites Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound.

Improving Perceptual Audio Aesthetic Assessment via Triplet Loss and Self-Supervised Embeddings Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T11:04:46.304755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:04:46.304755Z digest=sha256:59329c852c2935751444b2c01d60720ff44b7b3dff71834f8191b79e13705ad6

Pith citing papers

Observation 5ed07b03-18c9-470d-8a5c-1519df88cac1 · inbound

DNSMOS-C: Improving End-to-end Speech Quality Models via Contrastive Learning cites this paper.

DNSMOS-C: Improving End-to-end Speech Quality Models via Contrastive Learning Improving Perceptual Audio Aesthetic Assessment via Triplet Loss and Self-Supervised Embeddings

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-04T14:39:57.487688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T03:16:39.283647Z digest=sha256:75caf43c847f346e4113f8818b333bf92738956e818da7269fb3fb89311c62fc