Pith. sign in

Paper Citation Record · LEDGER

SAME: A Semantically-Aligned Music Autoencoder

As of 19 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 7 inbound Pith citation observations for arXiv:2605.18613.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.18613 v1

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-20T07:59:11.802632Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:14:47.082431Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T06:06:50.507426Z

Reference resolution

41 of 41 outbound references displayed

  • verified exact8
  • verified fuzzy33
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation adb8357d-4e68-4eff-9fb8-b18c758cf6bf · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

SAME: A Semantically-Aligned Music Autoencoder High-resolution image synthesis with latent diffusion models

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T08:28:09.271236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T07:59:11.802632Z digest=sha256:a02d893b0075eae326b7f7118e190be44b3fb7a29030dd02303bd2e251a2d1bd

Observation 38f9a095-b68e-4b0e-9d2c-1677b0afe847 · outbound

This paper cites SoundStream: An end-to-end neural audio codec.

SAME: A Semantically-Aligned Music Autoencoder SoundStream: An end-to-end neural audio codec

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T08:28:09.279086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T07:59:11.802632Z digest=sha256:403c923f1d4412f1ee3292027263ca025d1c91ae51f3f7a662203d9ec74390bd

Observation 754e839a-16cb-4065-adec-e86f9599b9cf · outbound

This paper cites High fidelity neural audio compression.

SAME: A Semantically-Aligned Music Autoencoder High fidelity neural audio compression

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T08:28:09.276890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T07:59:11.802632Z digest=sha256:8c4717d73935756c0f35d9732c8ed05904fd11d34249e50b089ed1b36b679c92

Observation 25dc52c9-261a-45c6-bba1-574a4d983594 · outbound

This paper cites High-fidelity audio compression with improved RVQGAN.

SAME: A Semantically-Aligned Music Autoencoder High-fidelity audio compression with improved RVQGAN

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T08:28:09.272951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T07:59:11.802632Z digest=sha256:2190e986632ecfbe47ec632403de9f092d015419129d2f47d21a6984ac717db7

Observation 395ae8f6-115f-4ed6-aa1a-4731a64f4f86 · outbound

This paper cites Neural discrete representation learning.

SAME: A Semantically-Aligned Music Autoencoder Neural discrete representation learning

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T08:28:09.269401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T07:59:11.802632Z digest=sha256:d8862a6542d6de2c0d5529e27248d122cf2a3ac3ec1db904f72f847dd64ed2a5

Observation 19d374bf-1b08-4c10-a091-0ebd5d7a4dbe · outbound

This paper cites AudioLM: A language modeling approach to audio generation.

SAME: A Semantically-Aligned Music Autoencoder AudioLM: A language modeling approach to audio generation

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T08:28:09.274942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T07:59:11.802632Z digest=sha256:117d50c8f7663719f0267bd79c40336fdf700c445f8b31b97a27569699f65bd6

Observation 61fbd16a-cfb2-4794-8673-07753333fbf9 · outbound

This paper cites Simple and controllable music generation.

SAME: A Semantically-Aligned Music Autoencoder Simple and controllable music generation

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T08:28:09.255696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T07:59:11.802632Z digest=sha256:0f2309f0cbc0b14d43bd5be03967710caaa6af05506710c0220f484ebd8203a7

Observation 7ec80fcf-3b6a-4722-afe5-5ca518746800 · outbound

This paper cites Stable Audio Open.

SAME: A Semantically-Aligned Music Autoencoder Stable Audio Open

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T08:28:09.242129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T07:59:11.802632Z digest=sha256:af5e208aa43ad32820196b9efcc74eabd2c2de68a340ee984c10ec154fc9c876

Observation 1f6ea139-343f-47f5-b763-26282fc6d72b · outbound

This paper cites Back to ear: Perceptually driven high fidelity music reconstruction.

SAME: A Semantically-Aligned Music Autoencoder Back to ear: Perceptually driven high fidelity music reconstruction

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-20T08:03:09.028422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T07:59:11.802632Z digest=sha256:0ff013e15fc6a45ea726a84019e0dc3282f4109f2e6aaf3dc7a0f17f0d0b4373

Observation 8d4006ea-7bd0-47c6-9968-39da1c0c5482 · outbound

This paper cites HILCodec: High-fidelity and lightweight neural audio codec.

SAME: A Semantically-Aligned Music Autoencoder HILCodec: High-fidelity and lightweight neural audio codec

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T08:28:09.238439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T07:59:11.802632Z digest=sha256:311eb27c6399c970026b2cb837c78aaef8541ed1a9b672c1c8f236d4d6549a8b

Observation df03b5b0-66eb-46fa-80f3-568a0a4772a7 · outbound

This paper cites Music2Latent: Consistency autoencoders for latent audio compression.

SAME: A Semantically-Aligned Music Autoencoder Music2Latent: Consistency autoencoders for latent audio compression

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T08:28:09.259419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T07:59:11.802632Z digest=sha256:421eedc2e32aaeba08d2eba4f8783a85b6922dde1e79f998c5f0d17ed0cea8f7

Observation ed81b00c-62af-4b44-b00b-9b8512495ede · outbound

This paper cites Music2Latent2: Audio compression with summary embeddings and autoregressive decoding.

SAME: A Semantically-Aligned Music Autoencoder Music2Latent2: Audio compression with summary embeddings and autoregressive decoding

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T08:28:09.228821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T07:59:11.802632Z digest=sha256:f18e71dd62730cb2f3b33ae65a0a19bc4054186a00b8c483edda80c33291e9b7

Observation e20e19ec-f288-4f18-ba60-d71069b872c7 · outbound

This paper cites CoDiCodec: Unifying continuous and discrete compressed representations of audio.

SAME: A Semantically-Aligned Music Autoencoder CoDiCodec: Unifying continuous and discrete compressed representations of audio

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T08:28:09.267142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T07:59:11.802632Z digest=sha256:44bd371699c7474d104dd133ea41fa24d0236cf1fe366f4c4211f453c31cc0a0

Observation d5d10548-ff6b-412a-86cc-f4709e763016 · outbound

This paper cites Scaling transformers for low-bitrate high-quality speech coding.

SAME: A Semantically-Aligned Music Autoencoder Scaling transformers for low-bitrate high-quality speech coding

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T08:28:09.249916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T07:59:11.802632Z digest=sha256:bec97683aa92c13e982432d5fc7fce14b6469f56c36699a377dcdab7ce24c175

Observation da66e041-6e04-417a-8450-c86b4032862b · outbound

This paper cites TS3-Codec: Transformer-based simple streaming single codec.

SAME: A Semantically-Aligned Music Autoencoder TS3-Codec: Transformer-based simple streaming single codec

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T08:28:09.240324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T07:59:11.802632Z digest=sha256:a09e821ddfc9b563a9274c5dfb9477671050316ed64ad5fba50b6b33569de6ab

Observation f84b2963-6f99-42c7-97d7-3eb79e46cc37 · outbound

This paper cites ALMTokenizer: A low-bitrate and semantic-rich audio codec tokenizer for audio language modeling.

SAME: A Semantically-Aligned Music Autoencoder ALMTokenizer: A low-bitrate and semantic-rich audio codec tokenizer for audio language modeling

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T08:28:09.226692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T07:59:11.802632Z digest=sha256:babd5e43bb5244e9af3090606c9c8ce8ca0e17b5b729624dae5281746c66748e

Observation c860cb7c-ce36-40cc-a634-29b2ff486b78 · outbound

This paper cites SpeechTokenizer: Unified speech tokenizer for speech language models.

SAME: A Semantically-Aligned Music Autoencoder SpeechTokenizer: Unified speech tokenizer for speech language models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T08:28:09.234627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T07:59:11.802632Z digest=sha256:ad1f8cfceae5d6f096b6d668e0f1b213c9c18bcc0589a0b3dc5277bc1f3d5d52

Observation b9e36cce-9356-466b-904a-5be3a68d981d · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

SAME: A Semantically-Aligned Music Autoencoder Moshi: a speech-text foundation model for real-time dialogue

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-20T08:03:09.024785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T07:59:11.802632Z digest=sha256:6589687ddbbbd36268f1d58d6d72dc9bf10baa975b5c2dc757beb5396f293ac2

Observation 56aa87db-235f-4331-8757-27235bb81219 · outbound

This paper cites FunCodec: A fundamental, reproducible and integrable open-source toolkit for neural speech codec.

SAME: A Semantically-Aligned Music Autoencoder FunCodec: A fundamental, reproducible and integrable open-source toolkit for neural speech codec

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T08:28:09.224626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T07:59:11.802632Z digest=sha256:821bcc24d862fe267de9625e4332fbe60d378673fa547ff5653a3f4060b9005d

Observation 047b2720-b6dc-41db-9725-374e1e076988 · outbound

This paper cites An image is worth 32 tokens for reconstruction and generation.

SAME: A Semantically-Aligned Music Autoencoder An image is worth 32 tokens for reconstruction and generation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T08:28:09.218578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T07:59:11.802632Z digest=sha256:1d9dfdd426d9c06bba9923d3de8a6731f0d504d1e8828735f9dd7c98221a5d73

Observation d124999f-cfb1-4dfc-878d-2b66ea87c73c · outbound

This paper cites Perceiver: General perception with iterative attention.

SAME: A Semantically-Aligned Music Autoencoder Perceiver: General perception with iterative attention

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T08:28:09.251777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T07:59:11.802632Z digest=sha256:3fb1da4290a87b11bf7592841efd7cf927c88c4a4e206feefd06c484ff853491

Observation 91179220-522f-4873-bb6e-e749510338a2 · outbound

This paper cites Differential Transformer.

SAME: A Semantically-Aligned Music Autoencoder Differential Transformer

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T08:28:09.261137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T07:59:11.802632Z digest=sha256:bb4852a582108aa7ad83cb61424c395a68e72b7b547ffe9d786e8021b83c6562

Observation 0b3b56eb-1bc4-40e2-9534-4f2d7760c3b8 · outbound

This paper cites RoFormer: Enhanced transformer with Rotary Position Embedding.

SAME: A Semantically-Aligned Music Autoencoder RoFormer: Enhanced transformer with Rotary Position Embedding

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T08:28:09.265281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T07:59:11.802632Z digest=sha256:f31294d36039dc2f487188f8cbcb8b430dcc45b462d4f48e20608847b7531818

Observation 05207c42-c685-4519-8b62-30d5a4122fa5 · outbound

This paper cites Transformers without normalization.

SAME: A Semantically-Aligned Music Autoencoder Transformers without normalization

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T08:28:09.230589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T07:59:11.802632Z digest=sha256:3f12baeaa0a4f954adc7fe41e09222856ebcded3c04d828fdf90b18af0861778

Observation d64d97cd-8c4a-4467-a0bf-f79a916ce770 · outbound

This paper cites LiteRT: On-device runtime for cross-platform machine learning inference.

SAME: A Semantically-Aligned Music Autoencoder LiteRT: On-device runtime for cross-platform machine learning inference

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T08:28:09.257565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T07:59:11.802632Z digest=sha256:128a11634563a721dfc6ceaf48cd6518329d47243c30b686b7ade54a403d6c60

Observation f30f953c-9e4d-4d8b-b5c0-5045efcdcd51 · outbound

This paper cites Diffusion Transformers with Representation Autoencoders.

SAME: A Semantically-Aligned Music Autoencoder Diffusion Transformers with Representation Autoencoders

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-20T08:03:09.020437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T07:59:11.802632Z digest=sha256:e2d13dc119f88f0d746dfdddb7230a0aa8bd6695dc93688c074fa5733a187f1c

Observation 860db250-3eca-4a2a-b5fd-0690b60f3ee5 · outbound

This paper cites Unified latents (ul): How to train your latents.

SAME: A Semantically-Aligned Music Autoencoder Unified latents (ul): How to train your latents

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-20T08:03:09.032895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T07:59:11.802632Z digest=sha256:08938be7ff04693a07f319a4ab9a3d387c94791ab305031f9e56ebf1c3587e16

Observation eedc3626-4891-4aa6-9664-0792d6ea6d97 · outbound

This paper cites Parallel WaveGAN: A fast waveform generation model based on Generative Adversarial Networks with multi-resolution spectrogram.

SAME: A Semantically-Aligned Music Autoencoder Parallel WaveGAN: A fast waveform generation model based on Generative Adversarial Networks with multi-resolution spectrogram

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T08:28:09.263077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T07:59:11.802632Z digest=sha256:b1e8c7dcdf6c744f4429c10165410cc8519cd66ea513e78576f6a75ffabda677

Observation b814ddc0-74e4-419c-a5c5-c1cf9881b8b0 · outbound

This paper cites Multi-scale spectral loss revisited.

SAME: A Semantically-Aligned Music Autoencoder Multi-scale spectral loss revisited

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T08:28:09.246017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T07:59:11.802632Z digest=sha256:d303af605155ff0236a60ced7f8b6bbbf26b96175c19b8d5d49bb825558ddebb

Observation 0314a4b3-539c-4380-ba65-07a18e8ad0b6 · outbound

This paper cites The relativistic discriminator: A key element missing from standard GAN.

SAME: A Semantically-Aligned Music Autoencoder The relativistic discriminator: A key element missing from standard GAN

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T08:28:09.232460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T07:59:11.802632Z digest=sha256:9c0adfceb4e305e170ce7ca21b96798140b0c66da7c600c88967603cdca8342e

Observation 45fb1f45-d492-4f3a-ba4f-eb57e47f68f9 · outbound

This paper cites Near-perfect-reconstruction pseudo-QMF banks.

SAME: A Semantically-Aligned Music Autoencoder Near-perfect-reconstruction pseudo-QMF banks

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T08:28:09.253805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T07:59:11.802632Z digest=sha256:9ce600d50e3682a7457e83cadcbf079a931327063da54471214dc2b7fab68268

Observation bf927c57-c9ab-47f4-801b-be8dd709aaea · outbound

This paper cites LARP: Tokenizing videos with a learned autoregressive generative prior.

SAME: A Semantically-Aligned Music Autoencoder LARP: Tokenizing videos with a learned autoregressive generative prior

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T08:28:09.236593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T07:59:11.802632Z digest=sha256:0a24e6642aedc3d344c07f0e5c750457495da12c93be4a1d8169349d6c407fe3

Observation 9275c599-08d7-40ae-b519-10060e6293f2 · outbound

This paper cites Flow Matching for generative modeling.

SAME: A Semantically-Aligned Music Autoencoder Flow Matching for generative modeling

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T08:28:09.220494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T07:59:11.802632Z digest=sha256:6321db4b9592cc8ef08a144cf19a48b929191338d95c5c29a613a3e45d5bb093

Observation 1573ce80-2a56-477e-8d13-1ba159c9f785 · outbound

This paper cites Biorthogonal bases of compactly supported wavelets.

SAME: A Semantically-Aligned Music Autoencoder Biorthogonal bases of compactly supported wavelets

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T08:28:09.222271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T07:59:11.802632Z digest=sha256:f986ca6b3587992b92d536cf75a710312ade774ca2f62484d60a999a741c7551

Observation b10ee692-b4e4-44c4-8a12-1e155bdc15dc · outbound

This paper cites Encoder-Decoder Gemma: Improving the Quality-Efficiency Trade-Off via Adaptation.

SAME: A Semantically-Aligned Music Autoencoder Encoder-Decoder Gemma: Improving the Quality-Efficiency Trade-Off via Adaptation

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-20T08:03:09.037981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T07:59:11.802632Z digest=sha256:6979c3dd522e38bc11a9eacb19c89cdef033ad0de747608d79c22d668718940b

Observation 70a4a902-5082-4f81-be6d-6f7154172ba3 · outbound

This paper cites Cautious optimizers: Improving training with one line of code.arXiv preprint arXiv:2411.16085.

SAME: A Semantically-Aligned Music Autoencoder Cautious optimizers: Improving training with one line of code.arXiv preprint arXiv:2411.16085

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-20T08:03:09.016799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T07:59:11.802632Z digest=sha256:b3da985989980baab61ee056ecc5951cd6ae642c3e145763824cd9f57eb77fa9

Observation 1da41848-9534-4da7-9215-00d3cf41146c · outbound

This paper cites Long-form music generation with latent diffusion.

SAME: A Semantically-Aligned Music Autoencoder Long-form music generation with latent diffusion

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T08:28:09.244129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T07:59:11.802632Z digest=sha256:d5be56e9f39701aab38b553ea68d00bbec021e2238897f5587a7bb96c0258b70

Observation 4746feac-d84f-4f90-8583-83520adc999e · outbound

This paper cites The Song Describer Dataset: a corpus of audio captions for music-and-language evaluation.

SAME: A Semantically-Aligned Music Autoencoder The Song Describer Dataset: a corpus of audio captions for music-and-language evaluation

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T08:28:09.248002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T07:59:11.802632Z digest=sha256:c8048362066485dc8033574f3e830ec04ffb2e4e20411ea76ca7df548f549eb9

Observation d2046df9-936a-40ca-9172-59b902dac0a8 · outbound

This paper cites Adapting Fréchet Audio Distance for generative music evaluation.

SAME: A Semantically-Aligned Music Autoencoder Adapting Fréchet Audio Distance for generative music evaluation

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T08:28:09.216452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T07:59:11.802632Z digest=sha256:347aabbac7f376d8f39eb71901c3e1e408e43e10d54b5d67a92a4b769f4f6e34

Observation a383438d-e2ab-4c17-9c99-5ef644ceb757 · outbound

This paper cites MuQ-Eval: An open-source per-sample quality metric for AI music generation evaluation.

SAME: A Semantically-Aligned Music Autoencoder MuQ-Eval: An open-source per-sample quality metric for AI music generation evaluation

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-20T08:03:09.012344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T07:59:11.802632Z digest=sha256:5572debd19e15716a5ea65bb7998ecdd0bdb45cad6b7ce834ffe28140d1d2a32

Observation 571cd417-4f01-4901-85ba-0475ffed1fc0 · outbound

This paper cites ACE-Step 1.5: Pushing the boundaries of open-source music generation.

SAME: A Semantically-Aligned Music Autoencoder ACE-Step 1.5: Pushing the boundaries of open-source music generation

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-20T08:03:09.008463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T07:59:11.802632Z digest=sha256:587944f2594af1be69ca3a097bb9e6a51db334a43ff2a5ad341f5596e213246d

Pith citing papers

Observation 57720e0b-3974-4a4a-8738-682cef2bcc82 · inbound

A Quantized Native Runtime for On-Device Semantic Audio Generation cites this paper.

A Quantized Native Runtime for On-Device Semantic Audio Generation SAME: A Semantically-Aligned Music Autoencoder

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-10T06:06:50.509304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-10T05:58:17.100557Z digest=sha256:f6d0f3fee59589eb11349f94df70f0174ff7f3a9e574a839b209f7be1beec077

Observation 741b6961-9843-4368-945a-67c1e2663b33 · inbound

Qwen-Music Technical Report cites this paper.

Qwen-Music Technical Report SAME: A Semantically-Aligned Music Autoencoder

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-14T03:47:22.936776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T03:47:22.936776Z digest=sha256:9732c3c13817f6a4b69095c2a771dc48303d1342331cfbe45e372156717731fc

Observation ec0df292-a9c6-4338-8012-b284af50115e · inbound

Qwen-Music Technical Report cites this paper.

Qwen-Music Technical Report SAME: A Semantically-Aligned Music Autoencoder

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-02T06:52:28.298503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:52:28.298503Z digest=sha256:b29dd2a939c52f254e06d3fa410b9145f96476ba548a41d8c004e4a88a3540f8

Observation c2728c85-db52-44b2-a69b-d7f618dea03c · inbound

Qwen-Audio-3.0-Gen-Preview Technical Report cites this paper.

Qwen-Audio-3.0-Gen-Preview Technical Report SAME: A Semantically-Aligned Music Autoencoder

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-30T14:08:14.347392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T14:08:14.347392Z digest=sha256:c5294b0a744ccd0fe691614d6d915ffd8951a3e028cb6737d86ee14ba2aae434

Observation b4620f63-e242-4900-acfc-50dc690bca6b · inbound

Qwen-Audio-3.0-Gen-Preview Technical Report cites this paper.

Qwen-Audio-3.0-Gen-Preview Technical Report SAME: A Semantically-Aligned Music Autoencoder

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T10:17:25.066467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:17:25.066467Z digest=sha256:ded5d6e087befc18c006bf600adbd10b856c0115f491c74736a2cd1b1fe52f3a

Observation 95ebb923-c51c-437d-80c3-c2bb4a451048 · inbound

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks cites this paper.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks SAME: A Semantically-Aligned Music Autoencoder

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-04T16:29:26.723412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:29:26.723412Z digest=sha256:760a5baa9c194662c6f01874e21d48d143a02d953506609c8f933193e60cfc08

Observation 7cecd235-87b3-40c3-8f52-32a819df9aab · inbound

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks cites this paper.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks SAME: A Semantically-Aligned Music Autoencoder

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:47.082431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:47.082431Z digest=sha256:d2ac819f1e86d97da825671aa80f0231b2e620609db5d729b01e6fe6dceab48f