Pith. sign in

Paper Citation Record · LEDGER

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer

As of 18 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 1 inbound Pith citation observation for arXiv:2506.00800.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.00800 v1

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:01:35.993878Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:01:33.023609Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T12:01:36.150100Z

Reference resolution

39 of 39 outbound references displayed

  • verified exact0
  • verified fuzzy32
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d9fec535-2a2f-45d1-8589-33a4efdc3814 · outbound

This paper cites CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T12:01:36.208170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T12:01:33.023609Z digest=sha256:19467e9b50ce1d3df90f1eb0f561218f83a80453ce3f73f4ab55e068e95938a2

Observation c26a34a6-401c-444b-b56b-9d4426b7316c · outbound

This paper cites Automated Audio Captioning Methods using Pre- trained Language Model Pre-trained language models have been utilized in some stud- ies to improve AAC performance [4–11].

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Automated Audio Captioning Methods using Pre- trained Language Model Pre-trained language models have been utilized in some stud- ies to improve AAC performance [4–11]

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:40.671760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T12:01:33.069426Z digest=sha256:f536bfd06c4e5ec4fda7d8726b3efb65c29cea5d893170a5100fe2880fc0ea9f

Observation 22ad47f7-076f-4c97-94c5-2ddeb94a5f13 · outbound

This paper cites First, En- CLAP transforms an input audio waveform into a CLAP audio embedding and EnCodec discrete tokens.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer First, En- CLAP transforms an input audio waveform into a CLAP audio embedding and EnCodec discrete tokens

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:40.662089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T12:01:33.141557Z digest=sha256:1095c2d88902cfc2f55e3468b3beb43e375e8949af5e8861d7a6d25556ca72cb

Observation 2b899848-a6e9-43d5-a3ab-5cb4e692de14 · outbound

This paper cites an unresolved cited work.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:01:40.652837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T12:01:33.268808Z digest=sha256:464ac1796f053a673aafe101b08e649c6f52c0bd54faae4c4bd329ee4a4f8e94

Observation d9178af3-636b-4d95-b1ab-9fbb316c1c21 · outbound

This paper cites Experimental Setup We conducted experiments on two AAC datasets: Audio- Caps [23] and Clotho [24].

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Experimental Setup We conducted experiments on two AAC datasets: Audio- Caps [23] and Clotho [24]

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:40.643800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T12:01:33.325366Z digest=sha256:ef96464064e1794e44f644398e8672ee64d9ed29d7a22a75825490afccf6a290

Observation c9cbd217-355f-43af-8165-8c4a49217fbf · outbound

This paper cites semantic-rich and discrete.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer semantic-rich and discrete

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:40.633679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T12:01:33.360355Z digest=sha256:aa8938701fdda2ceb4b34b6f80ac2de419cfe042d71eda581e2e94859fd5e71e

Observation e5b382c7-cdf9-45c2-ae30-53407d95080f · outbound

This paper cites an unresolved cited work.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:01:40.624033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T12:01:33.427951Z digest=sha256:a16e5b2f1af42c1a91b70c9095dc05bd98794a9dd4af69191c9677431cfdb5bb

Observation 2b8053d3-a6a5-4e22-bdb8-2f33fd944383 · outbound

This paper cites Automated audio captioning with recurrent neural networks,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Automated audio captioning with recurrent neural networks,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:40.615126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T12:01:33.508524Z digest=sha256:f6d5771d54f2cdfde7f0be99cd41189a5d98201ea1e0ffff0d55f7cfa208e58c

Observation 21aaba24-c546-4ab3-a3fd-5ecc41bf560d · outbound

This paper cites Automated audio captioning: An overview of recent progress and new challenges,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Automated audio captioning: An overview of recent progress and new challenges,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:40.606743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T12:01:33.600862Z digest=sha256:97159de971a1bbb77ed7ad64b39b55ec8379d450eb3b34feb5561c5607c97504

Observation 327d26fd-73b7-45c0-9581-f4d75e449ba1 · outbound

This paper cites Beyond the status quo: A contemporary survey of advances and challenges in audio cap- tioning,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Beyond the status quo: A contemporary survey of advances and challenges in audio cap- tioning,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:40.597069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T12:01:33.656600Z digest=sha256:aca9de36c0fb086a1659a586cf0b2811777144bbcc7a26b2021e1f999faa6a4c

Observation c32a6fbe-9650-4e07-b6a3-7baa2371cee5 · outbound

This paper cites Audio Captioning using Pre-Trained Large-Scale Language Model Guided by Audio-based Similar Caption Retrieval.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Audio Captioning using Pre-Trained Large-Scale Language Model Guided by Audio-based Similar Caption Retrieval

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:33.727705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:33.727705Z digest=sha256:2133191f91fe1b7140143cbcb678525ff60ba9d365b68201ac26922eed1c3b71

Observation b6605451-fc0a-494c-abe4-d485ddffa251 · outbound

This paper cites Automated audio cap- tioning by fine-tuning bart with audioset tags,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Automated audio cap- tioning by fine-tuning bart with audioset tags,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:40.586909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T12:01:33.801814Z digest=sha256:9a4442f10a5db3bfa5a44a848a5be2e26ea091b8acd2273e127ada3a76460005

Observation 405f177c-82a0-4e4f-b9b0-fbffb039535e · outbound

This paper cites Enclap: Combining neural audio codec and audio-text joint embedding for automated audio captioning,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Enclap: Combining neural audio codec and audio-text joint embedding for automated audio captioning,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:40.572547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T12:01:33.866374Z digest=sha256:0f56703c1871037c59688856bed28c5f5744c509b154a4b6e71219869c700ebd

Observation b04eacd2-2fe4-4413-a3e5-caf3b85c35fb · outbound

This paper cites Wavcaps: A chatgpt-assisted weakly- labelled audio captioning dataset for audio-language multimodal research,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Wavcaps: A chatgpt-assisted weakly- labelled audio captioning dataset for audio-language multimodal research,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:40.432143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T12:01:33.905971Z digest=sha256:38adb73558eeaa7e1d59cc5f4e38f27c361771e51de835dff183a063ff465840

Observation fd3a5ab3-c46d-4b97-8b81-c063f3e49c89 · outbound

This paper cites Improving audio captioning models with fine- grained audio features, text embedding supervision, and llm mix- up augmentation,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Improving audio captioning models with fine- grained audio features, text embedding supervision, and llm mix- up augmentation,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:40.220996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T12:01:33.966127Z digest=sha256:fefe2f06c0dba4665435de8969f115f76e4a69cf05b29c105351605f0eda6fae

Observation 0199daa7-7f7b-48a3-8f11-3e7babeef12a · outbound

This paper cites Recap: Retrieval-augmented audio captioning,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Recap: Retrieval-augmented audio captioning,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:40.143102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T12:01:34.051316Z digest=sha256:aec4b4f0b579498025afe8f075d857b9bd31da214b88f2846faaa31cc9e2706f

Observation ab9921c6-7f27-4813-8fdb-9715031c8ae5 · outbound

This paper cites Taming Data and Transformers for Audio Generation.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Taming Data and Transformers for Audio Generation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:34.136572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:34.136572Z digest=sha256:dcd56a344710e2a61638b4161bf29b4f6d92883b95247b6bd6875bf2c8ba8e39

Observation 95fd2992-8ce8-463c-ad70-4910f955bce7 · outbound

This paper cites Enhancing automated audio captioning via large language models with optimized audio encoding,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Enhancing automated audio captioning via large language models with optimized audio encoding,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:39.973182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T12:01:34.232063Z digest=sha256:9f8cb545bf72f540a8dfd17c3e1ea6d1a20e4b2685a4cba4622648d7c939951a

Observation e5fb0d28-9ed9-4735-b5bf-fa0b242bd93f · outbound

This paper cites BEATs: Audio pre-training with acoustic tokenizers,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer BEATs: Audio pre-training with acoustic tokenizers,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:39.792153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T12:01:34.312586Z digest=sha256:7e929f4a90cdf93fa56e6f84e08e0c8eabe1ad973bb37b6f781b9b49f4ecdd24

Observation 1c71dbdf-29d1-4713-b39f-5cc7e576fecd · outbound

This paper cites Audio Set: An ontology and human-labeled dataset for audio events,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Audio Set: An ontology and human-labeled dataset for audio events,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:39.645364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T12:01:34.374746Z digest=sha256:3c3df42347906822f0f40d4323238d0eb468dcb36f8c539b8bbbb84d6fddd8c5

Observation 0e361f16-f008-4518-b797-c9ae4fd2cf9d · outbound

This paper cites Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:39.413088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T12:01:34.445594Z digest=sha256:1d4f445445e6ba166c41f14107dfd42bb4bee54a17a074fa1271ab2385921ee3

Observation 849d4c47-fa88-4064-a327-d07ebe4e913d · outbound

This paper cites High fidelity neural audio compression,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer High fidelity neural audio compression,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:39.203206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T12:01:34.506940Z digest=sha256:1a9e4996904502d97454272f9592194000e4c39797bd9608fa64e6fd7c1e1281

Observation 39e2a501-5016-464b-ac77-b9882674171b · outbound

This paper cites BART: Denoising sequence-to-sequence pre-training for natural language genera- tion, translation, and comprehension,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer BART: Denoising sequence-to-sequence pre-training for natural language genera- tion, translation, and comprehension,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:38.968506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T12:01:34.589090Z digest=sha256:06f4fd34f0be73d84eae4fdfda353ec402e1d1ea592add29a177303b4974bd39

Observation 7244c56a-c1e5-4a44-bd02-48dcc4a1123e · outbound

This paper cites SemantiCodec: An Ultra Low Bitrate Semantic Audio Codec for General Sound.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer SemantiCodec: An Ultra Low Bitrate Semantic Audio Codec for General Sound

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:34.694185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:34.694185Z digest=sha256:3928e17cfe126a22021e85c0983ac2c4000bc8913650d41e46bc57537fdb152a

Observation 1daff668-6f25-4c57-859f-d70c503c2e27 · outbound

This paper cites Speechtok- enizer: Unified speech tokenizer for speech language models,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Speechtok- enizer: Unified speech tokenizer for speech language models,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:38.710089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T12:01:34.783763Z digest=sha256:7c11a30550c30f26853bb2e9df55f33005a47e691f975def28cb6cbaad4fd562

Observation 825707c1-2a9f-4374-806e-0fcdaa537847 · outbound

This paper cites Hubert: Self-supervised speech representation learning by masked prediction of hidden units,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Hubert: Self-supervised speech representation learning by masked prediction of hidden units,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:38.416120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T12:01:34.886016Z digest=sha256:7175f9b2aaaed45c6bfb1cd6f11ee8ec6b26793f52c22d4f8014497681a34350

Observation dde80387-82ca-484d-9e41-15118b8e9377 · outbound

This paper cites SoundStream: An end-to-end neural audio codec,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer SoundStream: An end-to-end neural audio codec,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:38.234586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T12:01:34.967305Z digest=sha256:db28561fd6bd4d9541575b1be12236b4f811e77883d985f15fa6061577854d49

Observation cad56588-70cf-4805-afe2-cb9b317a6268 · outbound

This paper cites vq- wav2vec: Self-supervised learning of discrete speech representa- tions,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer vq- wav2vec: Self-supervised learning of discrete speech representa- tions,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:38.051717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T12:01:35.045177Z digest=sha256:d1e2b2f1401fb885cc284ee92f510841655e4e12cb02ea06f74f3d05eadf27be

Observation dacc8b28-076a-42f1-999d-7607c3e29081 · outbound

This paper cites How should we extract discrete audio tokens from self-supervised models?,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer How should we extract discrete audio tokens from self-supervised models?,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:37.876465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T12:01:35.127376Z digest=sha256:275bafedc571996046bf3c8c897b7bf6488ec96cdaa77c89c1d899e82524b119

Observation a310985e-a820-492c-96f7-f08ca33af77f · outbound

This paper cites Audiocaps: Generating captions for audios in the wild,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Audiocaps: Generating captions for audios in the wild,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:37.725909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T12:01:35.207642Z digest=sha256:1593eb3b20ad49902e81666a231b86a33231124aa215eb74bf2c8ec6f32e2618

Observation 6b35c6c1-310e-4ee6-8a4e-c1c384c400d2 · outbound

This paper cites Clotho: An audio cap- tioning dataset,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Clotho: An audio cap- tioning dataset,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:37.529375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T12:01:35.295830Z digest=sha256:09ba1283ac97e6c7696adbf9b8894a8d858a2244b69cb7666818e2bb22c26d48

Observation 78e4e4d8-f205-4695-a62a-f2b7e767aa50 · outbound

This paper cites Decoupled Weight Decay Regularization.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Decoupled Weight Decay Regularization

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:35.359333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:35.359333Z digest=sha256:d0bafe9311e0f9f1ff711776fae9bf375f61de3d1873444ba18958d83e92288a

Observation e32cd370-503b-4519-ad3a-18545a2f37e6 · outbound

This paper cites Meteor universal: Language spe- cific translation evaluation for any target language,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Meteor universal: Language spe- cific translation evaluation for any target language,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:37.362079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T12:01:35.442458Z digest=sha256:c5f345f4a22b852caf60cd0a38fdd2df9704f1a40a526f9269d0921c03b94291

Observation 61ddb5b5-50de-4892-bba4-d68103ef3633 · outbound

This paper cites CIDEr: Consensus-based image description evaluation,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer CIDEr: Consensus-based image description evaluation,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:37.197754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T12:01:35.563960Z digest=sha256:ed1a6ff8f7d166bd2853496860623285d119c4f0cb874d9ec5ec9ce9d46e7a54

Observation 89e7d13c-cac2-43b4-b452-0bba65e4a275 · outbound

This paper cites SPICE: Semantic propositional image caption evaluation,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer SPICE: Semantic propositional image caption evaluation,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:37.030840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T12:01:35.643613Z digest=sha256:021dc2dd10044120507ae4b69cadea37b2f0868eb899ffeddd1be305b1955f47

Observation a78f4136-0a21-428d-a152-04d6e1cace00 · outbound

This paper cites Improved image captioning via policy gradient optimization of spider,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Improved image captioning via policy gradient optimization of spider,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:36.844380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T12:01:35.725864Z digest=sha256:405891702b1e16094e3c11517884ff9d94e57d8f8b5575eb9a7959bf01fe4222

Observation ffd1cc54-6e51-4411-8c84-438a46ef7573 · outbound

This paper cites Can audio captions be evaluated with image caption metrics?,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Can audio captions be evaluated with image caption metrics?,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:36.701552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T12:01:35.799846Z digest=sha256:ccff39fca2ecbc067e99138d8896561385514a686cf92e066cfb3330ea26767f

Observation 0e90ee26-cb6f-4015-84ee-ea6ba9f10b91 · outbound

This paper cites Enclap++: Ana- lyzing the enclap framework for optimizing automated audio cap- tioning performance,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Enclap++: Ana- lyzing the enclap framework for optimizing automated audio cap- tioning performance,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:36.552520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T12:01:35.906298Z digest=sha256:d10e5f9609f4837c4d95aebcd39a4a091e10cca0546ab799aa1a92d610143094

Observation 0a1131d4-b1e1-4a29-91fa-1082570c3239 · outbound

This paper cites Slam-aac: Enhancing audio captioning with paraphras- ing augmentation and clap-refine through llms,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Slam-aac: Enhancing audio captioning with paraphras- ing augmentation and clap-refine through llms,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:36.366008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T12:01:35.993878Z digest=sha256:27bd6b248ac667e3a0d4db2a81e21e4d4b8e3c59b90f3471f390d1646ac6738c

Pith citing papers

Observation d9fec535-2a2f-45d1-8589-33a4efdc3814 · inbound

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer cites this paper.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T12:01:36.208170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T12:01:33.023609Z digest=sha256:19467e9b50ce1d3df90f1eb0f561218f83a80453ce3f73f4ab55e068e95938a2