Pith. sign in

Paper Citation Record · LEDGER

Improving Perceptual Audio Aesthetic Assessment via Triplet Loss and Self-Supervised Embeddings

As of 20 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 1 inbound Pith citation observation for arXiv:2509.03292.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.03292 v1

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T11:04:46.304755Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-26T03:16:39.283647Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T14:39:57.486114Z

Reference resolution

22 of 22 outbound references displayed

  • verified exact0
  • verified fuzzy21
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ef0f1f39-2be5-4036-8465-7526be8fa775 · outbound

This paper cites The V oiceMOS Challenge 2022,.

Improving Perceptual Audio Aesthetic Assessment via Triplet Loss and Self-Supervised Embeddings The V oiceMOS Challenge 2022,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T11:04:49.956454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T11:04:45.715077Z digest=sha256:fdbd5fa846718c5223d3fe8a9ff001929281aefd74ed6cd093de461943931a14

Observation 9567268a-eef9-40fd-a8e9-2b82272ab0a7 · outbound

This paper cites The voicemos challenge 2023: Zero-shot subjective speech quality prediction for multiple domains,.

Improving Perceptual Audio Aesthetic Assessment via Triplet Loss and Self-Supervised Embeddings The voicemos challenge 2023: Zero-shot subjective speech quality prediction for multiple domains,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T11:04:49.804754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T11:04:45.764757Z digest=sha256:caa1e2e5ea548a8cd06e5d7df35fdbad48c6da55c2a1195c937237a29685d79a

Observation c2619186-770a-466a-a56c-a712488fe67f · outbound

This paper cites The voicemos challenge 2024: Beyond speech quality prediction,.

Improving Perceptual Audio Aesthetic Assessment via Triplet Loss and Self-Supervised Embeddings The voicemos challenge 2024: Beyond speech quality prediction,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T11:04:49.664750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T11:04:45.815103Z digest=sha256:8aeea7a6acb8c176e5b09c9e00f6c0a4b0e9d23b7792abaf46844de49aeabc22

Observation 63b89fdb-bc2e-4edd-af12-75407fc573ce · outbound

This paper cites Fusion of self-supervised learned models for MOS prediction,.

Improving Perceptual Audio Aesthetic Assessment via Triplet Loss and Self-Supervised Embeddings Fusion of self-supervised learned models for MOS prediction,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T11:04:49.574796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T11:04:45.829892Z digest=sha256:e55a1f04f541cd7e0ee7d51591c7ea7190cb280d802b59ab5a2f0cd5bccb283b

Observation 2f397553-9393-4d7d-90a1-558ec1d3de1b · outbound

This paper cites UTMOS: UTokyo-SaruLab system for V oiceMOS Chal- lenge 2022,.

Improving Perceptual Audio Aesthetic Assessment via Triplet Loss and Self-Supervised Embeddings UTMOS: UTokyo-SaruLab system for V oiceMOS Chal- lenge 2022,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T11:04:49.287377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T11:04:45.861257Z digest=sha256:42dd96dbaac2c996a755c6662a3cc66d6eccbb1d24606a9522bfd8ea0cc2a757

Observation e24e072a-f79a-4620-ae78-c5ef5863ce3e · outbound

This paper cites A study on incorporating Whisper for robust speech assessment,.

Improving Perceptual Audio Aesthetic Assessment via Triplet Loss and Self-Supervised Embeddings A study on incorporating Whisper for robust speech assessment,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T11:04:49.044899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T11:04:45.877952Z digest=sha256:0c696b2d663a41d84cc3f25581d47767045a0605a5f65bd243a391a326027efe

Observation 525cc080-8a22-4e23-963e-6c0f47a24838 · outbound

This paper cites Corn: Co-trained full- and no-reference speech quality assessment,.

Improving Perceptual Audio Aesthetic Assessment via Triplet Loss and Self-Supervised Embeddings Corn: Co-trained full- and no-reference speech quality assessment,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T11:04:48.819873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T11:04:45.897598Z digest=sha256:7ca007c2926b2c5116614c45a9fb651a625313823c144029d27df32dfa974e21

Observation a585ad7b-8c1f-489e-b033-09fc09661743 · outbound

This paper cites Enabling auditory large language models for automatic speech quality evaluation,.

Improving Perceptual Audio Aesthetic Assessment via Triplet Loss and Self-Supervised Embeddings Enabling auditory large language models for automatic speech quality evaluation,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T11:04:48.642848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T11:04:45.924752Z digest=sha256:1907f3bb83ebee1115f755f09a9ee7b6ff809cf09929e0a5fd885185653e2889

Observation 5ff3ec75-e2b8-4988-9d3e-1c5cb503a623 · outbound

This paper cites Speech foundation models on intelligibility prediction for hearing-impaired listeners,.

Improving Perceptual Audio Aesthetic Assessment via Triplet Loss and Self-Supervised Embeddings Speech foundation models on intelligibility prediction for hearing-impaired listeners,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T11:04:48.358044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T11:04:45.965072Z digest=sha256:8bcb6f5b0b13aea83c43bf68891dc92567a541e2f3867992d2180b7ce03cb33b

Observation 6b0422e6-2a41-4b30-aab6-f039c8a1485a · outbound

This paper cites Non-intrusive speech intelligibility prediction for hearing- impaired users using intermediate ASR features and human memory models,.

Improving Perceptual Audio Aesthetic Assessment via Triplet Loss and Self-Supervised Embeddings Non-intrusive speech intelligibility prediction for hearing- impaired users using intermediate ASR features and human memory models,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T11:04:48.234178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T11:04:45.984833Z digest=sha256:676fd930339db1041f2c5d32dc0ee52a1f142b31c3095b36e5853835e3a62cfb

Observation e96fcfb8-51fc-4839-8e77-a5be341db9f6 · outbound

This paper cites A review on subjective and objective evaluation of synthetic speech,.

Improving Perceptual Audio Aesthetic Assessment via Triplet Loss and Self-Supervised Embeddings A review on subjective and objective evaluation of synthetic speech,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T11:04:48.090620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T11:04:46.034755Z digest=sha256:d4e9a21b485f54d0f677df192ae82b40afbfad784f850eac15addb13efd54906

Observation 3d00225f-8c76-4cf9-81ef-fd826fae7a74 · outbound

This paper cites Self-supervised speech quality estimation and enhancement using only clean speech,.

Improving Perceptual Audio Aesthetic Assessment via Triplet Loss and Self-Supervised Embeddings Self-supervised speech quality estimation and enhancement using only clean speech,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T11:04:47.954912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T11:04:46.076073Z digest=sha256:2c9ddfb3de87234de4111360c30c5c31e73180474b8bf579d76259526290db49

Observation 956ee022-43bc-4f78-bf2d-c9fa8c2f8ab6 · outbound

This paper cites Generalization ability of MOS prediction networks,.

Improving Perceptual Audio Aesthetic Assessment via Triplet Loss and Self-Supervised Embeddings Generalization ability of MOS prediction networks,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T11:04:47.807494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T11:04:46.083341Z digest=sha256:68b8a7f6c3691efd8943b72e87f6d368a36343eeb9f7e78c1e2ff24dd430582b

Observation 1c7fed5d-9104-4af0-8c3c-5fec5e99ef6f · outbound

This paper cites Deep learning-based non-intrusive multi-objective speech assessment model with cross-domain features,.

Improving Perceptual Audio Aesthetic Assessment via Triplet Loss and Self-Supervised Embeddings Deep learning-based non-intrusive multi-objective speech assessment model with cross-domain features,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T11:04:47.617912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T11:04:46.097720Z digest=sha256:ac98681fa2afea7c43b0d2441dcf29066110038997e942d1cb78c0c5f8b2d7f4

Observation edefb18b-f081-48bb-96c6-4d6210cf10a9 · outbound

This paper cites Speechlmscore: Evaluat- ing speech generation using speech language model,.

Improving Perceptual Audio Aesthetic Assessment via Triplet Loss and Self-Supervised Embeddings Speechlmscore: Evaluat- ing speech generation using speech language model,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T11:04:47.416887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T11:04:46.162639Z digest=sha256:c4182961d5820ca92dd9fa3a3f051ed69931d228e647202550b45901c9403faa

Observation f77c807d-9c99-4570-89d1-b8a4cef47988 · outbound

This paper cites BEATs: Audio pre-training with acoustic tokenizers,.

Improving Perceptual Audio Aesthetic Assessment via Triplet Loss and Self-Supervised Embeddings BEATs: Audio pre-training with acoustic tokenizers,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T11:04:47.215679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T11:04:46.168984Z digest=sha256:e301d22a14c7b22265531eb3f053862acd750a465ad6f636fcab2d7bbdb81a54

Observation 83ef93fd-94b9-42be-9f3d-71b756dafc1a · outbound

This paper cites MBNet: MOS prediction for synthesized speech with mean-bias network,.

Improving Perceptual Audio Aesthetic Assessment via Triplet Loss and Self-Supervised Embeddings MBNet: MOS prediction for synthesized speech with mean-bias network,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T11:04:46.999369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T11:04:46.188563Z digest=sha256:e97d5a78b7ad3c282498ed0df48daef97321bd66aa1d1b241b74b4352d785c8c

Observation 35cfd7a4-db72-4a07-a230-6b365b7cfbfa · outbound

This paper cites LDNet: Unified listener dependent modeling in MOS prediction for synthetic speech,.

Improving Perceptual Audio Aesthetic Assessment via Triplet Loss and Self-Supervised Embeddings LDNet: Unified listener dependent modeling in MOS prediction for synthetic speech,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T11:04:46.797786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T11:04:46.207045Z digest=sha256:87a655bb35fdc3d2f2a4eaed1e87b608b05c6d1983ee50df01e7a83a9c20e77f

Observation 31fcf490-1523-45ff-a782-7b7d19f25e0d · outbound

This paper cites HAAQI- Net: A non-intrusive neural music audio quality assessment model for hearing aids,.

Improving Perceptual Audio Aesthetic Assessment via Triplet Loss and Self-Supervised Embeddings HAAQI- Net: A non-intrusive neural music audio quality assessment model for hearing aids,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T11:04:46.737263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T11:04:46.211742Z digest=sha256:41e5d50a6e939ab764f8774569c852b931d86ca2a82fb2b03d391fc93fb38b08

Observation 384d0956-531d-40ee-8827-e20475294701 · outbound

This paper cites FaceNet: A unified embed- ding for face recognition and clustering,.

Improving Perceptual Audio Aesthetic Assessment via Triplet Loss and Self-Supervised Embeddings FaceNet: A unified embed- ding for face recognition and clustering,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T11:04:46.659417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T11:04:46.225739Z digest=sha256:d5d8d7c5220d157dd28a8fd0f50eea9d33bf9c9ea6dc60d36e6e3cac3a89a772

Observation 0b919c50-5eae-4dad-8989-9c092755494a · outbound

This paper cites Unsupervised feature learning via non-parametric instance discrimination,.

Improving Perceptual Audio Aesthetic Assessment via Triplet Loss and Self-Supervised Embeddings Unsupervised feature learning via non-parametric instance discrimination,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T11:04:46.558959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T11:04:46.274755Z digest=sha256:c9f1cf4afbabe6a7eb03c134f960240c1f3c668bce4c0bb6b438536703dafa6c

Observation a7c5feaa-4371-4fa3-acac-4a755db43373 · outbound

This paper cites Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound.

Improving Perceptual Audio Aesthetic Assessment via Triplet Loss and Self-Supervised Embeddings Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T11:04:46.304755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:04:46.304755Z digest=sha256:9d23bc214859bf101115ded6d2b2c0572cbdb2bd2e33c324291f46be5bc5a4bc

Pith citing papers

Observation 5ed07b03-18c9-470d-8a5c-1519df88cac1 · inbound

DNSMOS-C: Improving End-to-end Speech Quality Models via Contrastive Learning cites this paper.

DNSMOS-C: Improving End-to-end Speech Quality Models via Contrastive Learning Improving Perceptual Audio Aesthetic Assessment via Triplet Loss and Self-Supervised Embeddings

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-04T14:39:57.487688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-26T03:16:39.283647Z digest=sha256:fa862da70f66c357a675adfe8255a42aaf4ab203bf794e23612c0a86f74c6e36