Pith. sign in

Paper Citation Record · LEDGER

Discrete Audio Representations for Automated Audio Captioning

As of 15 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 1 inbound Pith citation observation for arXiv:2505.14989.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.14989 v1

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:29:41.089739Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:29:37.637530Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T15:29:41.315494Z

Reference resolution

36 of 36 outbound references displayed

  • verified exact1
  • verified fuzzy17
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 99f40907-dc3f-4326-98e4-a199b836bdaa · outbound

This paper cites Discrete Audio Representations for Automated Audio Captioning.

Discrete Audio Representations for Automated Audio Captioning Discrete Audio Representations for Automated Audio Captioning

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:29:41.407940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T15:29:37.637530Z digest=sha256:156cde9039ceec6fe32cf1ecca07269b26e1173471653c098724583dafc72944

Observation 33ad1828-4fd5-4d8e-bb50-0923888c62ad · outbound

This paper cites We construct BART-based and GPT-2 XL-based AAC systems, both utilizing semantic and acoustic tokens.

Discrete Audio Representations for Automated Audio Captioning We construct BART-based and GPT-2 XL-based AAC systems, both utilizing semantic and acoustic tokens

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:29:44.012926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T15:29:37.685143Z digest=sha256:fcaa5b5ae63635d23dcd26819d6654040e4eb248e04c7fb9e3debeafe9e59694

Observation cefd1b20-81c4-4ea7-a9c9-c6e2de374566 · outbound

This paper cites an unresolved cited work.

Discrete Audio Representations for Automated Audio Captioning Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:29:43.877697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T15:29:37.771872Z digest=sha256:eb1fae26a25986d323d5de3a16cda3b628c1e624d1c08dfe758dfbda9fadd42b

Observation 72d3ae5e-6dc2-4d14-8e49-086a70a44d25 · outbound

This paper cites Our findings indicate that semantic tokens significantly outperform acoustic tokens in this context.

Discrete Audio Representations for Automated Audio Captioning Our findings indicate that semantic tokens significantly outperform acoustic tokens in this context

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:29:43.729200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T15:29:37.837897Z digest=sha256:d9627aaef091745389e689a1f3dd1927b3ca8f2bdf4d68338d1e320af022627f

Observation 603fc8e4-6430-4c5f-b5f4-a5dffa6819bf · outbound

This paper cites Automated audio captioning: An overview of recent progress and new challenges,.

Discrete Audio Representations for Automated Audio Captioning Automated audio captioning: An overview of recent progress and new challenges,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:29:43.606953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T15:29:37.933826Z digest=sha256:18f2824fabc4f712e326c33b4dee30521cf4649b7aebdf0f2efd882caa348887

Observation c8825f6b-0b1c-4c98-b2d7-44a980073a55 · outbound

This paper cites Audio captioning based on transformer and pre-trained cnn,.

Discrete Audio Representations for Automated Audio Captioning Audio captioning based on transformer and pre-trained cnn,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:29:43.460031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T15:29:38.002613Z digest=sha256:ea1f205d68cfcc8a1ae66a41b5104081a7b169563088c9d1a9cdf018d6fee259

Observation 59b5a379-4928-4a17-ad1d-929c748ca906 · outbound

This paper cites Automated audio caption- ing by fine-tuning bart with audioset tags,.

Discrete Audio Representations for Automated Audio Captioning Automated audio caption- ing by fine-tuning bart with audioset tags,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:29:43.305054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T15:29:38.078918Z digest=sha256:6b49f7181b8746d214562fc8c2b104e46f95099b28adc9509064c586a0af58a8

Observation f8b710d6-2282-4147-9d68-64b5259cbd75 · outbound

This paper cites Leveraging pre-trained bert for audio captioning,.

Discrete Audio Representations for Automated Audio Captioning Leveraging pre-trained bert for audio captioning,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:29:43.170929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T15:29:38.173354Z digest=sha256:c517e905aa441ba4f6b60e21d9479e7a8dc705c674552e9393d984fab0a0fbb8

Observation 83fc5ec3-ab4e-4553-bf30-3d2aac66599b · outbound

This paper cites Conette: An efficient au- dio captioning system leveraging multiple datasets with task em- bedding,.

Discrete Audio Representations for Automated Audio Captioning Conette: An efficient au- dio captioning system leveraging multiple datasets with task em- bedding,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:29:43.032026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T15:29:38.274473Z digest=sha256:ee542be91f708f0b2310136636358ef36e195ade17e6578c1ea77583f23510c3

Observation e6243da2-8ba2-4ee1-ac41-5b5c5bb07e3f · outbound

This paper cites Improving audio captioning mod- els with fine-grained audio features, text embedding supervision, and llm mix-up augmentation,.

Discrete Audio Representations for Automated Audio Captioning Improving audio captioning mod- els with fine-grained audio features, text embedding supervision, and llm mix-up augmentation,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:29:42.889413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T15:29:38.375285Z digest=sha256:1252fa05283e8bbff0cfcc6b612238cb8b91f293a3f9f0509ca61335dc77048f

Observation 35d13838-760a-4507-a964-a3d39b0c42e4 · outbound

This paper cites Panns: Large-scale pretrained audio neural networks for audio pattern recognition,.

Discrete Audio Representations for Automated Audio Captioning Panns: Large-scale pretrained audio neural networks for audio pattern recognition,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:38.478425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:29:38.478425Z digest=sha256:1dd6c0a858a4d4b9f99584acb94e9a4c7806ad138a9a07deda3f1f244f6e0670

Observation bb3ef9d7-fdaa-49fb-a8eb-c94b7d2dcecf · outbound

This paper cites Adapting a convnext model to audio classification on au- dioset,.

Discrete Audio Representations for Automated Audio Captioning Adapting a convnext model to audio classification on au- dioset,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:29:42.746482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T15:29:38.560986Z digest=sha256:3c30e35efac8a4b10d20907f7b2a6057ccf4b4416aca13a4eb9212f30b9d8902

Observation 46b92cf6-81f9-48cc-933a-4c058438af93 · outbound

This paper cites Beats: audio pre-training with acoustic tok- enizers,.

Discrete Audio Representations for Automated Audio Captioning Beats: audio pre-training with acoustic tok- enizers,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:29:42.551062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T15:29:38.654264Z digest=sha256:a52fad8ed8b8eb6b36175bc5f4d90f367e26dd9bf7bf8fc48c967007eb381f54

Observation 6b9661e5-d984-425b-8ea9-a1c8e8aff9c9 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Discrete Audio Representations for Automated Audio Captioning BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:38.730373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:29:38.730373Z digest=sha256:ada87aa0e06018d6d97efeceea9ca031cae0fd5133ec1844c7d31b0f8971407c

Observation f97881d5-1890-40a5-8d5a-44f7c8b214a5 · outbound

This paper cites BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension.

Discrete Audio Representations for Automated Audio Captioning BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:38.816322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:29:38.816322Z digest=sha256:6732b8985951aa4f9aa7479a6281f94dd8ebdc24074ede29c1c775a057737152

Observation faeb6c81-bccb-4221-a163-554f51667a31 · outbound

This paper cites Language models are unsupervised multitask learners,.

Discrete Audio Representations for Automated Audio Captioning Language models are unsupervised multitask learners,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:38.907144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:29:38.907144Z digest=sha256:1ad74f2f6458040d177e7dc2fed4ef93704e794536dd5bbd6050d4e6b04f1db0

Observation 6c960552-04af-4b3a-ba53-1ebb82abe3a5 · outbound

This paper cites STAB: Speech Tokenizer Assessment Benchmark.

Discrete Audio Representations for Automated Audio Captioning STAB: Speech Tokenizer Assessment Benchmark

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:39.010252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:29:39.010252Z digest=sha256:efdd0f789f4a02e577c649f520aed7aee371bd8045c74511365cc95eccb319a5

Observation 1bf8414e-bbe7-412c-926d-3c551e889850 · outbound

This paper cites Soundstream: An end-to-end neural audio codec,.

Discrete Audio Representations for Automated Audio Captioning Soundstream: An end-to-end neural audio codec,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:39.090136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:29:39.090136Z digest=sha256:5668819ab619e28f6f3d3c4fcc8b6844b2b7c840d97c10aba2b1ba7768fea9ec

Observation 604426d6-34a9-45e0-8b98-d9a7518c4e3a · outbound

This paper cites High Fidelity Neural Audio Compression.

Discrete Audio Representations for Automated Audio Captioning High Fidelity Neural Audio Compression

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:39.236427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:29:39.236427Z digest=sha256:2ef60036f648d9b023c5526e7772dec7e2f8d46cb2548b334f2312afb373642a

Observation a635e2e1-ee8c-4774-a573-a3dd8fd5abe8 · outbound

This paper cites High-fidelity audio compression with improved rvqgan,.

Discrete Audio Representations for Automated Audio Captioning High-fidelity audio compression with improved rvqgan,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:39.371761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:29:39.371761Z digest=sha256:d32b827c16dbdb6e507e8405e37614c395069e4af27cb8e97afc58dade69d5a7

Observation db345d99-1cc7-44c3-b2e9-cff7d83fdb29 · outbound

This paper cites Hubert: Self-supervised speech represen- tation learning by masked prediction of hidden units,.

Discrete Audio Representations for Automated Audio Captioning Hubert: Self-supervised speech represen- tation learning by masked prediction of hidden units,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:39.476168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:29:39.476168Z digest=sha256:07fde3c3d35a1094ede4a332b6008df71a50d5223a459496a92490532dfc38b9

Observation d440acd7-51cc-45cf-99e7-d54ad8af7e6c · outbound

This paper cites Wavlm: Large-scale self- supervised pre-training for full stack speech processing,.

Discrete Audio Representations for Automated Audio Captioning Wavlm: Large-scale self- supervised pre-training for full stack speech processing,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:39.544115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:29:39.544115Z digest=sha256:14ee454f0dbc4d08942134ccfd04e9b8ae26a1e1822f30428d80f9a121ade8c6

Observation 820f0797-cf0f-4bd5-aad6-9be4af9109d9 · outbound

This paper cites RepCodec: A speech represen- tation codec for speech tokenization,.

Discrete Audio Representations for Automated Audio Captioning RepCodec: A speech represen- tation codec for speech tokenization,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:29:42.358530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T15:29:39.614586Z digest=sha256:d48ed8543b93f6b47c0845ed9e00edc022c3454fdceecbdb42babdcf5af1a8c2

Observation 2365d3a4-9a06-480b-a9e9-738cec03d94e · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

Discrete Audio Representations for Automated Audio Captioning Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:39.691319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:29:39.691319Z digest=sha256:1840ee2ca761fca8e3020200f5e7653337727085e90d19d118b4ad72324679b9

Observation 7e1079c0-39a2-4f82-bd04-ba21a7d587b4 · outbound

This paper cites BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data.

Discrete Audio Representations for Automated Audio Captioning BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:39.917468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:29:39.917468Z digest=sha256:d2242fe8f0da4669b6108e05ef5425a24348a276d943d4011aeb981762fe7e80

Observation d6434d90-b764-4b0a-b87d-a5606774bcdf · outbound

This paper cites Exploration of efficient end-to-end asr using discretized input from self-supervised learning,.

Discrete Audio Representations for Automated Audio Captioning Exploration of efficient end-to-end asr using discretized input from self-supervised learning,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:29:42.206416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T15:29:40.021098Z digest=sha256:7a05161e3390f86c8f250f18ed5a383952475bd7db6e71a789d635740e9e8d41

Observation 9031b10f-1abd-4b2c-9429-6d768514a162 · outbound

This paper cites How should we extract discrete audio tokens from self-supervised models?.

Discrete Audio Representations for Automated Audio Captioning How should we extract discrete audio tokens from self-supervised models?

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:40.136058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:29:40.136058Z digest=sha256:89d0f762895109a6db63ea5e496ce0a75881b42806338d6b8d2d0df18bda8bde

Observation f071cc86-68db-4404-ba0a-947998171708 · outbound

This paper cites Discrete audio representation as an alternative to mel-spectrograms for speaker and speech recognition,.

Discrete Audio Representations for Automated Audio Captioning Discrete audio representation as an alternative to mel-spectrograms for speaker and speech recognition,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:29:42.081468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T15:29:40.274350Z digest=sha256:21ee29b1a8039527bb751c0e0657696b5eba88c4b51c3f012aff28734f54905d

Observation e716f219-6b97-4506-b963-09661381826f · outbound

This paper cites Enclap: Combining neu- ral audio codec and audio-text joint embedding for automated au- dio captioning,.

Discrete Audio Representations for Automated Audio Captioning Enclap: Combining neu- ral audio codec and audio-text joint embedding for automated au- dio captioning,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:29:41.951056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T15:29:40.400959Z digest=sha256:e11cb52c44c7fd02923971424d67f825afb415f23a78e013e8e2ce566465f673

Observation 559d8f9d-1118-412f-8dc3-f2ee8e0538a6 · outbound

This paper cites Neural discrete represen- tation learning,.

Discrete Audio Representations for Automated Audio Captioning Neural discrete represen- tation learning,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:40.513005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:29:40.513005Z digest=sha256:3417d7e7fbcb3caa4ac90b2a11488adc4b18ee3f5ea09d456e35b613b166fb7b

Observation 5ca75e92-a442-4ac8-9b6b-6902dbdd3bb9 · outbound

This paper cites Investigating lo- cal and global information for automated audio captioning with transfer learning,.

Discrete Audio Representations for Automated Audio Captioning Investigating lo- cal and global information for automated audio captioning with transfer learning,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:29:41.823902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T15:29:40.581171Z digest=sha256:20d9df290904928aefe584bec1e79fc4495c619d6e0fb4115418655662a60166

Observation a11eb655-7160-47d5-9313-fe082c58d128 · outbound

This paper cites Prefix tuning for auto- mated audio captioning,.

Discrete Audio Representations for Automated Audio Captioning Prefix tuning for auto- mated audio captioning,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:29:41.713383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T15:29:40.658035Z digest=sha256:1deafdfc9898128ca90466e692aa588dba563b3744573fb74b0c67c4c4404155

Observation fa09ffca-3051-4b66-b7ac-c9a49cbceec3 · outbound

This paper cites Audio set: An ontology and human-labeled dataset for audio events,.

Discrete Audio Representations for Automated Audio Captioning Audio set: An ontology and human-labeled dataset for audio events,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:40.769821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:29:40.769821Z digest=sha256:fdaccb20ca4dbad23dd342587983a2a74e31bdfa96cd87d965978c91a4c5d1f1

Observation 6c7a4798-49e0-4a63-8e6a-5c18735ef4c8 · outbound

This paper cites Clotho: An audio cap- tioning dataset,.

Discrete Audio Representations for Automated Audio Captioning Clotho: An audio cap- tioning dataset,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:40.890406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:29:40.890406Z digest=sha256:29d3a7a76f7934f213508e21d41940ca8b6163688a15c7528d7d68d30e6da3b8

Observation ca949e47-509f-4b35-9b34-a44501a28463 · outbound

This paper cites Improved image captioning via policy gradient optimization of spider,.

Discrete Audio Representations for Automated Audio Captioning Improved image captioning via policy gradient optimization of spider,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:40.981070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:29:40.981070Z digest=sha256:3d17c7d3f1a24fdf75659152b242eb9c132e9f43bb0fcb802bb75d0277ef7fee

Observation 5e43b8e5-c476-4d03-a339-29dcf1715581 · outbound

This paper cites Can audio captions be evaluated with image caption metrics?.

Discrete Audio Representations for Automated Audio Captioning Can audio captions be evaluated with image caption metrics?

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:29:41.528936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T15:29:41.089739Z digest=sha256:e32b4cb6ffbe82bf38d548e060caa427a6cfa45c7b17b3df74d80cfc08acaab6

Pith citing papers

Observation 99f40907-dc3f-4326-98e4-a199b836bdaa · inbound

Discrete Audio Representations for Automated Audio Captioning cites this paper.

Discrete Audio Representations for Automated Audio Captioning Discrete Audio Representations for Automated Audio Captioning

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:29:41.407940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T15:29:37.637530Z digest=sha256:156cde9039ceec6fe32cf1ecca07269b26e1173471653c098724583dafc72944