Pith. sign in

Paper Citation Record · LEDGER

Discrete Audio Representations for Automated Audio Captioning

As of 9 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 1 inbound Pith citation observation for arXiv:2505.14989.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.14989 v1

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:29:41.089739Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:29:37.637530Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T15:29:41.315494Z

Reference resolution

36 of 36 outbound references displayed

  • verified exact1
  • verified fuzzy17
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 99f40907-dc3f-4326-98e4-a199b836bdaa · outbound

This paper cites Discrete Audio Representations for Automated Audio Captioning.

Discrete Audio Representations for Automated Audio Captioning Discrete Audio Representations for Automated Audio Captioning

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:29:41.407940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:29:37.637530Z digest=sha256:10d0bfd280e3fa41b5f8bdef8de9a60dfc2a19ad58a4ecc7b98e42e22fa88b68

Observation 33ad1828-4fd5-4d8e-bb50-0923888c62ad · outbound

This paper cites We construct BART-based and GPT-2 XL-based AAC systems, both utilizing semantic and acoustic tokens.

Discrete Audio Representations for Automated Audio Captioning We construct BART-based and GPT-2 XL-based AAC systems, both utilizing semantic and acoustic tokens

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:29:44.012926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:29:37.685143Z digest=sha256:8d2f2a9855597a321b6958ad6e64156e0dee84f2a415fa9029c4afe077c587cf

Observation cefd1b20-81c4-4ea7-a9c9-c6e2de374566 · outbound

This paper cites an unresolved cited work.

Discrete Audio Representations for Automated Audio Captioning Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:29:43.877697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:29:37.771872Z digest=sha256:27a5108b77d2d55144a9f8e3cae6fc909f4d0bf2782c62ff23878b38e68804ca

Observation 72d3ae5e-6dc2-4d14-8e49-086a70a44d25 · outbound

This paper cites Our findings indicate that semantic tokens significantly outperform acoustic tokens in this context.

Discrete Audio Representations for Automated Audio Captioning Our findings indicate that semantic tokens significantly outperform acoustic tokens in this context

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:29:43.729200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:29:37.837897Z digest=sha256:74ff7ed4f47201e58165c11886db584ba1822a96ef363bba54e6a475aeddee1e

Observation 603fc8e4-6430-4c5f-b5f4-a5dffa6819bf · outbound

This paper cites Automated audio captioning: An overview of recent progress and new challenges,.

Discrete Audio Representations for Automated Audio Captioning Automated audio captioning: An overview of recent progress and new challenges,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:29:43.606953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:29:37.933826Z digest=sha256:8cc5fa49c7375ae4fd1cbb2b38b55ffe8578ea526e497835208dfe7b7a2033ab

Observation c8825f6b-0b1c-4c98-b2d7-44a980073a55 · outbound

This paper cites Audio captioning based on transformer and pre-trained cnn,.

Discrete Audio Representations for Automated Audio Captioning Audio captioning based on transformer and pre-trained cnn,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:29:43.460031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:29:38.002613Z digest=sha256:95ec2d59ef33b4c974b80971d15df44f2480bddc32f2eae7b10ebc9858264e1e

Observation 59b5a379-4928-4a17-ad1d-929c748ca906 · outbound

This paper cites Automated audio caption- ing by fine-tuning bart with audioset tags,.

Discrete Audio Representations for Automated Audio Captioning Automated audio caption- ing by fine-tuning bart with audioset tags,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:29:43.305054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:29:38.078918Z digest=sha256:f6e98ac59c7c577746d20260c02a12a0409482117c1b37bc83cebe12c7a03c23

Observation f8b710d6-2282-4147-9d68-64b5259cbd75 · outbound

This paper cites Leveraging pre-trained bert for audio captioning,.

Discrete Audio Representations for Automated Audio Captioning Leveraging pre-trained bert for audio captioning,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:29:43.170929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:29:38.173354Z digest=sha256:a17d4273ab2974824a3b1979210b4d91936d82b4fbf1c46e12e6eab5c1c7d324

Observation 83fc5ec3-ab4e-4553-bf30-3d2aac66599b · outbound

This paper cites Conette: An efficient au- dio captioning system leveraging multiple datasets with task em- bedding,.

Discrete Audio Representations for Automated Audio Captioning Conette: An efficient au- dio captioning system leveraging multiple datasets with task em- bedding,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:29:43.032026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:29:38.274473Z digest=sha256:b29dab9e0da088ee9735596a63cb9c60edde4e248e91800b11f25ad2cea79a1a

Observation e6243da2-8ba2-4ee1-ac41-5b5c5bb07e3f · outbound

This paper cites Improving audio captioning mod- els with fine-grained audio features, text embedding supervision, and llm mix-up augmentation,.

Discrete Audio Representations for Automated Audio Captioning Improving audio captioning mod- els with fine-grained audio features, text embedding supervision, and llm mix-up augmentation,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:29:42.889413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:29:38.375285Z digest=sha256:18bc25011874cd808ae9a1dffb217065b58bb7c3bfd5ed2c6ddcccdfea34a0d2

Observation 35d13838-760a-4507-a964-a3d39b0c42e4 · outbound

This paper cites Panns: Large-scale pretrained audio neural networks for audio pattern recognition,.

Discrete Audio Representations for Automated Audio Captioning Panns: Large-scale pretrained audio neural networks for audio pattern recognition,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:38.478425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:29:38.478425Z digest=sha256:9953a74e490db4a037699629815bb81cd85061461ca60fdfad9e1239fdd98b5e

Observation bb3ef9d7-fdaa-49fb-a8eb-c94b7d2dcecf · outbound

This paper cites Adapting a convnext model to audio classification on au- dioset,.

Discrete Audio Representations for Automated Audio Captioning Adapting a convnext model to audio classification on au- dioset,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:29:42.746482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:29:38.560986Z digest=sha256:7a014094b4f289bf9d2be884e5f236529b12948256e4b40cf9ffea061445504e

Observation 46b92cf6-81f9-48cc-933a-4c058438af93 · outbound

This paper cites Beats: audio pre-training with acoustic tok- enizers,.

Discrete Audio Representations for Automated Audio Captioning Beats: audio pre-training with acoustic tok- enizers,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:29:42.551062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:29:38.654264Z digest=sha256:e5fbebb5f3e04d128d99b1a4330c8f6de71881a4ab96c08a67f65ec056814071

Observation 6b9661e5-d984-425b-8ea9-a1c8e8aff9c9 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Discrete Audio Representations for Automated Audio Captioning BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:38.730373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:29:38.730373Z digest=sha256:6a432f9031eb3836b21a802c79694d2f4accefe549cfd312289f158e3c8f71a4

Observation f97881d5-1890-40a5-8d5a-44f7c8b214a5 · outbound

This paper cites BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension.

Discrete Audio Representations for Automated Audio Captioning BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:38.816322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:29:38.816322Z digest=sha256:8292df8155c7c94645534307f046de65111147569fb8e82b204d44f600fe106e

Observation faeb6c81-bccb-4221-a163-554f51667a31 · outbound

This paper cites Language models are unsupervised multitask learners,.

Discrete Audio Representations for Automated Audio Captioning Language models are unsupervised multitask learners,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:38.907144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:29:38.907144Z digest=sha256:b65f6e35699d7f48fbc0127c351dc4f97aa3be195470bab69c7ef93c991c31e4

Observation 6c960552-04af-4b3a-ba53-1ebb82abe3a5 · outbound

This paper cites STAB: Speech Tokenizer Assessment Benchmark.

Discrete Audio Representations for Automated Audio Captioning STAB: Speech Tokenizer Assessment Benchmark

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:39.010252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:29:39.010252Z digest=sha256:0d18920cc7203cb90394693c6c0981fa1135cf0147e5513ba28b39d8c6166c10

Observation 1bf8414e-bbe7-412c-926d-3c551e889850 · outbound

This paper cites Soundstream: An end-to-end neural audio codec,.

Discrete Audio Representations for Automated Audio Captioning Soundstream: An end-to-end neural audio codec,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:39.090136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:29:39.090136Z digest=sha256:95d8957879f36b752c4b509ead9a81bea7f11b65f4d9c0eec06e861c5a8ed8d4

Observation 604426d6-34a9-45e0-8b98-d9a7518c4e3a · outbound

This paper cites High Fidelity Neural Audio Compression.

Discrete Audio Representations for Automated Audio Captioning High Fidelity Neural Audio Compression

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:39.236427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:29:39.236427Z digest=sha256:74db96f75231ccc2c2b0510a9594b6974f5da7c1eabf30dcdaa84a5909cb53db

Observation a635e2e1-ee8c-4774-a573-a3dd8fd5abe8 · outbound

This paper cites High-fidelity audio compression with improved rvqgan,.

Discrete Audio Representations for Automated Audio Captioning High-fidelity audio compression with improved rvqgan,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:39.371761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:29:39.371761Z digest=sha256:76f66d2e74ff957406089f261574cf79ab7a1fbdbb3f706180afb5383f263344

Observation db345d99-1cc7-44c3-b2e9-cff7d83fdb29 · outbound

This paper cites Hubert: Self-supervised speech represen- tation learning by masked prediction of hidden units,.

Discrete Audio Representations for Automated Audio Captioning Hubert: Self-supervised speech represen- tation learning by masked prediction of hidden units,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:39.476168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:29:39.476168Z digest=sha256:f00528a576e55ec93fb75afd4752e592c6b04cd79244a4dcbb37c430c302038a

Observation d440acd7-51cc-45cf-99e7-d54ad8af7e6c · outbound

This paper cites Wavlm: Large-scale self- supervised pre-training for full stack speech processing,.

Discrete Audio Representations for Automated Audio Captioning Wavlm: Large-scale self- supervised pre-training for full stack speech processing,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:39.544115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:29:39.544115Z digest=sha256:125889072d3c07988dbb3d889c375b04b2ddbae1435120d734b497da39e5374c

Observation 820f0797-cf0f-4bd5-aad6-9be4af9109d9 · outbound

This paper cites RepCodec: A speech represen- tation codec for speech tokenization,.

Discrete Audio Representations for Automated Audio Captioning RepCodec: A speech represen- tation codec for speech tokenization,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:29:42.358530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:29:39.614586Z digest=sha256:3f90bd8d56e38d59a623ac505f6d1fe8324385d9fc9e9fa0fe514e71c98c11c6

Observation 2365d3a4-9a06-480b-a9e9-738cec03d94e · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

Discrete Audio Representations for Automated Audio Captioning Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:39.691319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:29:39.691319Z digest=sha256:17d76adf1895e67b53cbbe0f7ef04c1710783f19861ee460977cf58e27aa3fd4

Observation 7e1079c0-39a2-4f82-bd04-ba21a7d587b4 · outbound

This paper cites BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data.

Discrete Audio Representations for Automated Audio Captioning BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:39.917468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:29:39.917468Z digest=sha256:3e28f95a396f05b81fc7c9ef152d947b5fc1d136a12f73dcbb22c47a61263047

Observation d6434d90-b764-4b0a-b87d-a5606774bcdf · outbound

This paper cites Exploration of efficient end-to-end asr using discretized input from self-supervised learning,.

Discrete Audio Representations for Automated Audio Captioning Exploration of efficient end-to-end asr using discretized input from self-supervised learning,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:29:42.206416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:29:40.021098Z digest=sha256:7f5c300efd568709563fbcbb3f32478cc4861fa5466b8265f79efc327cdaa57b

Observation 9031b10f-1abd-4b2c-9429-6d768514a162 · outbound

This paper cites How should we extract discrete audio tokens from self-supervised models?.

Discrete Audio Representations for Automated Audio Captioning How should we extract discrete audio tokens from self-supervised models?

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:40.136058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:29:40.136058Z digest=sha256:f24199c0d8a9d45f59e741848370d7e9d8084d9b8e293cf31a925b3aa13450fc

Observation f071cc86-68db-4404-ba0a-947998171708 · outbound

This paper cites Discrete audio representation as an alternative to mel-spectrograms for speaker and speech recognition,.

Discrete Audio Representations for Automated Audio Captioning Discrete audio representation as an alternative to mel-spectrograms for speaker and speech recognition,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:29:42.081468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:29:40.274350Z digest=sha256:443f876c1e78f3d8daefab8c09fd33f4c85dbdcf4b23576cfab961b753b4c0f9

Observation e716f219-6b97-4506-b963-09661381826f · outbound

This paper cites Enclap: Combining neu- ral audio codec and audio-text joint embedding for automated au- dio captioning,.

Discrete Audio Representations for Automated Audio Captioning Enclap: Combining neu- ral audio codec and audio-text joint embedding for automated au- dio captioning,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:29:41.951056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:29:40.400959Z digest=sha256:f5e75a16ca5da26c0d02583c6c2ac4f3f57a47459c41b013754ba0fae25639c3

Observation 559d8f9d-1118-412f-8dc3-f2ee8e0538a6 · outbound

This paper cites Neural discrete represen- tation learning,.

Discrete Audio Representations for Automated Audio Captioning Neural discrete represen- tation learning,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:40.513005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:29:40.513005Z digest=sha256:86a94649eb7698d26793e486ebc9eda5557e5eb506e1ec2cf866d235635cc9c4

Observation 5ca75e92-a442-4ac8-9b6b-6902dbdd3bb9 · outbound

This paper cites Investigating lo- cal and global information for automated audio captioning with transfer learning,.

Discrete Audio Representations for Automated Audio Captioning Investigating lo- cal and global information for automated audio captioning with transfer learning,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:29:41.823902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:29:40.581171Z digest=sha256:afb8f5a5f5542b3d6673394f974d63f90eb778cac4abb110e40b73d2804dc1b8

Observation a11eb655-7160-47d5-9313-fe082c58d128 · outbound

This paper cites Prefix tuning for auto- mated audio captioning,.

Discrete Audio Representations for Automated Audio Captioning Prefix tuning for auto- mated audio captioning,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:29:41.713383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:29:40.658035Z digest=sha256:daef2f8904d74c7a45a85f0eed19d2cae4d34acf3e9d4f4df6e56a812e343910

Observation fa09ffca-3051-4b66-b7ac-c9a49cbceec3 · outbound

This paper cites Audio set: An ontology and human-labeled dataset for audio events,.

Discrete Audio Representations for Automated Audio Captioning Audio set: An ontology and human-labeled dataset for audio events,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:40.769821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:29:40.769821Z digest=sha256:92964160f478db263c2a5ac63aba2909e807456f4e09fa864ea8814ec914aa72

Observation 6c7a4798-49e0-4a63-8e6a-5c18735ef4c8 · outbound

This paper cites Clotho: An audio cap- tioning dataset,.

Discrete Audio Representations for Automated Audio Captioning Clotho: An audio cap- tioning dataset,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:40.890406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:29:40.890406Z digest=sha256:4270eb828cb53f7995effb6e0fca2fc48e4d57d1c8197f651877870bef1a275d

Observation ca949e47-509f-4b35-9b34-a44501a28463 · outbound

This paper cites Improved image captioning via policy gradient optimization of spider,.

Discrete Audio Representations for Automated Audio Captioning Improved image captioning via policy gradient optimization of spider,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:40.981070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:29:40.981070Z digest=sha256:d9a5e544e21f1cf309eee968786fc380f4d7681e391546df7ca2fe309599acde

Observation 5e43b8e5-c476-4d03-a339-29dcf1715581 · outbound

This paper cites Can audio captions be evaluated with image caption metrics?.

Discrete Audio Representations for Automated Audio Captioning Can audio captions be evaluated with image caption metrics?

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:29:41.528936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:29:41.089739Z digest=sha256:168843c64ddf69e7cc1a8cdd1ae7567dc574c26f969b62d8223b60f1d40f3428

Pith citing papers

Observation 99f40907-dc3f-4326-98e4-a199b836bdaa · inbound

Discrete Audio Representations for Automated Audio Captioning cites this paper.

Discrete Audio Representations for Automated Audio Captioning Discrete Audio Representations for Automated Audio Captioning

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:29:41.407940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:29:37.637530Z digest=sha256:10d0bfd280e3fa41b5f8bdef8de9a60dfc2a19ad58a4ecc7b98e42e22fa88b68