Pith. sign in

Paper Citation Record · LEDGER

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer

As of 10 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 1 inbound Pith citation observation for arXiv:2506.00800.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.00800 v1

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:01:35.993878Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:01:33.023609Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T12:01:36.150100Z

Reference resolution

39 of 39 outbound references displayed

  • verified exact0
  • verified fuzzy32
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d9fec535-2a2f-45d1-8589-33a4efdc3814 · outbound

This paper cites CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T12:01:36.208170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:01:33.023609Z digest=sha256:f332a502754a983b6f9498225af16675911166d7dbef57408df4b9eeeb8aa6a5

Observation c26a34a6-401c-444b-b56b-9d4426b7316c · outbound

This paper cites Automated Audio Captioning Methods using Pre- trained Language Model Pre-trained language models have been utilized in some stud- ies to improve AAC performance [4–11].

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Automated Audio Captioning Methods using Pre- trained Language Model Pre-trained language models have been utilized in some stud- ies to improve AAC performance [4–11]

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:40.671760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:01:33.069426Z digest=sha256:4bb56c26bf095b5d68a30f15d03b510b7d539bf6d10b039ecffd66e0d54312fa

Observation 22ad47f7-076f-4c97-94c5-2ddeb94a5f13 · outbound

This paper cites First, En- CLAP transforms an input audio waveform into a CLAP audio embedding and EnCodec discrete tokens.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer First, En- CLAP transforms an input audio waveform into a CLAP audio embedding and EnCodec discrete tokens

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:40.662089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:01:33.141557Z digest=sha256:24340f05578eb64c33b588644fd92d4dcb1eeecd0e83b50d9f7f60dea1d308f1

Observation 2b899848-a6e9-43d5-a3ab-5cb4e692de14 · outbound

This paper cites an unresolved cited work.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:01:40.652837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:01:33.268808Z digest=sha256:12927290e2c592746a396d8cbd0d8e9fed617d5d108e3a36e26b9f3ddec4e29d

Observation d9178af3-636b-4d95-b1ab-9fbb316c1c21 · outbound

This paper cites Experimental Setup We conducted experiments on two AAC datasets: Audio- Caps [23] and Clotho [24].

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Experimental Setup We conducted experiments on two AAC datasets: Audio- Caps [23] and Clotho [24]

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:40.643800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:01:33.325366Z digest=sha256:bdfc067948fef107f3fdcb7c2aaf59fb5fdbe0e9d099f2a6a449ba916b7673b2

Observation c9cbd217-355f-43af-8165-8c4a49217fbf · outbound

This paper cites semantic-rich and discrete.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer semantic-rich and discrete

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:40.633679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:01:33.360355Z digest=sha256:d78849cca9243afc71227af69d4cd7829f22d96842a5c00920c4acf46e33a7c0

Observation e5b382c7-cdf9-45c2-ae30-53407d95080f · outbound

This paper cites an unresolved cited work.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:01:40.624033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:01:33.427951Z digest=sha256:0e20bf5223d084c359b51930d52fa16b7a5b6f94bc95f215ddac34dfa75406d0

Observation 2b8053d3-a6a5-4e22-bdb8-2f33fd944383 · outbound

This paper cites Automated audio captioning with recurrent neural networks,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Automated audio captioning with recurrent neural networks,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:40.615126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:01:33.508524Z digest=sha256:a7b39b64808863a42b827e8ff613be94289b84238f796a078f0cc049ec1d8ad0

Observation 21aaba24-c546-4ab3-a3fd-5ecc41bf560d · outbound

This paper cites Automated audio captioning: An overview of recent progress and new challenges,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Automated audio captioning: An overview of recent progress and new challenges,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:40.606743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:01:33.600862Z digest=sha256:a778ef083df9db770301c1d064e57dec92eed0523505de370b3bf17249ea4fab

Observation 327d26fd-73b7-45c0-9581-f4d75e449ba1 · outbound

This paper cites Beyond the status quo: A contemporary survey of advances and challenges in audio cap- tioning,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Beyond the status quo: A contemporary survey of advances and challenges in audio cap- tioning,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:40.597069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:01:33.656600Z digest=sha256:3fbf0418b8a0c3d7b96c76382d7bd3c88a81b61f50bbe5db821a9a7d6cf16714

Observation c32a6fbe-9650-4e07-b6a3-7baa2371cee5 · outbound

This paper cites Audio Captioning using Pre-Trained Large-Scale Language Model Guided by Audio-based Similar Caption Retrieval.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Audio Captioning using Pre-Trained Large-Scale Language Model Guided by Audio-based Similar Caption Retrieval

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:33.727705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:33.727705Z digest=sha256:93082ca65220b09ed0de4300a6b96d5bca87f7581a0d4701a1bfa98264c08240

Observation b6605451-fc0a-494c-abe4-d485ddffa251 · outbound

This paper cites Automated audio cap- tioning by fine-tuning bart with audioset tags,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Automated audio cap- tioning by fine-tuning bart with audioset tags,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:40.586909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:01:33.801814Z digest=sha256:16db9a6398d37dd832be7169025a14c452729880c8cc609088b43ba1f20c1944

Observation 405f177c-82a0-4e4f-b9b0-fbffb039535e · outbound

This paper cites Enclap: Combining neural audio codec and audio-text joint embedding for automated audio captioning,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Enclap: Combining neural audio codec and audio-text joint embedding for automated audio captioning,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:40.572547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:01:33.866374Z digest=sha256:05311437d367ec8d2138cb1a4313a3d3cb1ec2c9413a3cff707a0badcfca1fa8

Observation b04eacd2-2fe4-4413-a3e5-caf3b85c35fb · outbound

This paper cites Wavcaps: A chatgpt-assisted weakly- labelled audio captioning dataset for audio-language multimodal research,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Wavcaps: A chatgpt-assisted weakly- labelled audio captioning dataset for audio-language multimodal research,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:40.432143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:01:33.905971Z digest=sha256:df973f1c29499a32a504f2225d25d83bfa3c05ef4e8343e0e9ba2924467f59f1

Observation fd3a5ab3-c46d-4b97-8b81-c063f3e49c89 · outbound

This paper cites Improving audio captioning models with fine- grained audio features, text embedding supervision, and llm mix- up augmentation,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Improving audio captioning models with fine- grained audio features, text embedding supervision, and llm mix- up augmentation,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:40.220996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:01:33.966127Z digest=sha256:6c8e04f34322efaeb03e1b80a51027ec460dca0430668cb0a535e0b291b0f75e

Observation 0199daa7-7f7b-48a3-8f11-3e7babeef12a · outbound

This paper cites Recap: Retrieval-augmented audio captioning,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Recap: Retrieval-augmented audio captioning,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:40.143102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:01:34.051316Z digest=sha256:c93aca89e5fe3724fab237e39bf5217df057472c83ae91c0aa8a7dfd2dc48c9a

Observation ab9921c6-7f27-4813-8fdb-9715031c8ae5 · outbound

This paper cites Taming Data and Transformers for Audio Generation.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Taming Data and Transformers for Audio Generation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:34.136572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:34.136572Z digest=sha256:2db7d5dccf89c271e883188d660e6e82d7dece7635b732655ab8f7ed690d70a6

Observation 95fd2992-8ce8-463c-ad70-4910f955bce7 · outbound

This paper cites Enhancing automated audio captioning via large language models with optimized audio encoding,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Enhancing automated audio captioning via large language models with optimized audio encoding,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:39.973182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:01:34.232063Z digest=sha256:267699166f8d13b03dcab1e5e824117e0427e836926076e9356e6031c6e9bee9

Observation e5fb0d28-9ed9-4735-b5bf-fa0b242bd93f · outbound

This paper cites BEATs: Audio pre-training with acoustic tokenizers,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer BEATs: Audio pre-training with acoustic tokenizers,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:39.792153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:01:34.312586Z digest=sha256:f5656a2b0249aeccb88a019fd07fca9d5ac5a2fb26c31170f0875cfe6cadf920

Observation 1c71dbdf-29d1-4713-b39f-5cc7e576fecd · outbound

This paper cites Audio Set: An ontology and human-labeled dataset for audio events,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Audio Set: An ontology and human-labeled dataset for audio events,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:39.645364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:01:34.374746Z digest=sha256:8e67ee937eb9578bcc636f36bac7874894a1b96cd7eb3f65726af986598ec690

Observation 0e361f16-f008-4518-b797-c9ae4fd2cf9d · outbound

This paper cites Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:39.413088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:01:34.445594Z digest=sha256:8fc2279ba62526c05bfef0e6bcc7c9c2b123ea0319f0217612c8b4ae259cfd66

Observation 849d4c47-fa88-4064-a327-d07ebe4e913d · outbound

This paper cites High fidelity neural audio compression,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer High fidelity neural audio compression,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:39.203206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:01:34.506940Z digest=sha256:799715a2d72b4f7097767196edf812fa9ee7dc61b7b64735b30f464d50a7b14f

Observation 39e2a501-5016-464b-ac77-b9882674171b · outbound

This paper cites BART: Denoising sequence-to-sequence pre-training for natural language genera- tion, translation, and comprehension,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer BART: Denoising sequence-to-sequence pre-training for natural language genera- tion, translation, and comprehension,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:38.968506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:01:34.589090Z digest=sha256:14575d84761173e8654f404565a175319af9252cc7a63bae2f52dc50bd5f8136

Observation 7244c56a-c1e5-4a44-bd02-48dcc4a1123e · outbound

This paper cites SemantiCodec: An Ultra Low Bitrate Semantic Audio Codec for General Sound.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer SemantiCodec: An Ultra Low Bitrate Semantic Audio Codec for General Sound

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:34.694185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:34.694185Z digest=sha256:92e3470b8dcfcdb844dd130a5c4886a9de833731ec33ba25f54aa461141d7fd0

Observation 1daff668-6f25-4c57-859f-d70c503c2e27 · outbound

This paper cites Speechtok- enizer: Unified speech tokenizer for speech language models,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Speechtok- enizer: Unified speech tokenizer for speech language models,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:38.710089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:01:34.783763Z digest=sha256:75b54d585ee16f1ecd8d0effad4151ba8f0efed0eab9f6ecf386a53cc1e74b31

Observation 825707c1-2a9f-4374-806e-0fcdaa537847 · outbound

This paper cites Hubert: Self-supervised speech representation learning by masked prediction of hidden units,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Hubert: Self-supervised speech representation learning by masked prediction of hidden units,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:38.416120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:01:34.886016Z digest=sha256:47d37711ed7a7b0093315fac76a48ea962f8fce01f81e4b448bc2eaba5b7645b

Observation dde80387-82ca-484d-9e41-15118b8e9377 · outbound

This paper cites SoundStream: An end-to-end neural audio codec,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer SoundStream: An end-to-end neural audio codec,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:38.234586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:01:34.967305Z digest=sha256:39ec9e6006ba037bf3780f2ee2ed518363e4d720057ec72a84c6b9a4eefa2139

Observation cad56588-70cf-4805-afe2-cb9b317a6268 · outbound

This paper cites vq- wav2vec: Self-supervised learning of discrete speech representa- tions,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer vq- wav2vec: Self-supervised learning of discrete speech representa- tions,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:38.051717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:01:35.045177Z digest=sha256:9d5d4bd161821382525dce6cabaa32aa44bd726a71700dd579a3360945aff284

Observation dacc8b28-076a-42f1-999d-7607c3e29081 · outbound

This paper cites How should we extract discrete audio tokens from self-supervised models?,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer How should we extract discrete audio tokens from self-supervised models?,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:37.876465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:01:35.127376Z digest=sha256:9be66431367df362a535d310abca963396e71bd0f22f4c61e53790b57377a144

Observation a310985e-a820-492c-96f7-f08ca33af77f · outbound

This paper cites Audiocaps: Generating captions for audios in the wild,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Audiocaps: Generating captions for audios in the wild,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:37.725909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:01:35.207642Z digest=sha256:b6c53dbddf01bb3300daedebad6f1e6d2ad28b062252901bec441d9dba906df9

Observation 6b35c6c1-310e-4ee6-8a4e-c1c384c400d2 · outbound

This paper cites Clotho: An audio cap- tioning dataset,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Clotho: An audio cap- tioning dataset,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:37.529375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:01:35.295830Z digest=sha256:200ff85b08588e0321b0d433655014c8d9e13fb05ba7b889a3fc2e5f0ba5d8ae

Observation 78e4e4d8-f205-4695-a62a-f2b7e767aa50 · outbound

This paper cites Decoupled Weight Decay Regularization.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Decoupled Weight Decay Regularization

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:35.359333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:35.359333Z digest=sha256:d59b6a48e287cff278f274b14f5dda7548e2c43be06ea2c5d878f9e3156aa066

Observation e32cd370-503b-4519-ad3a-18545a2f37e6 · outbound

This paper cites Meteor universal: Language spe- cific translation evaluation for any target language,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Meteor universal: Language spe- cific translation evaluation for any target language,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:37.362079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:01:35.442458Z digest=sha256:fa25f8adbd5c2b9c3c25676f32f536cdf4b660e131581dc2f53a7f22ad94fcb2

Observation 61ddb5b5-50de-4892-bba4-d68103ef3633 · outbound

This paper cites CIDEr: Consensus-based image description evaluation,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer CIDEr: Consensus-based image description evaluation,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:37.197754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:01:35.563960Z digest=sha256:e8e68fc9449e21dbe73d1d98e7aeebb3c830a08717667538331db88923d7978c

Observation 89e7d13c-cac2-43b4-b452-0bba65e4a275 · outbound

This paper cites SPICE: Semantic propositional image caption evaluation,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer SPICE: Semantic propositional image caption evaluation,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:37.030840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:01:35.643613Z digest=sha256:7b949a8913351d6b63ca4669dd64856ccb8cc4e9107b9cc83aad14e6cf0b0d7d

Observation a78f4136-0a21-428d-a152-04d6e1cace00 · outbound

This paper cites Improved image captioning via policy gradient optimization of spider,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Improved image captioning via policy gradient optimization of spider,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:36.844380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:01:35.725864Z digest=sha256:5f021b8f9ea328f2abc60938ae3743c673d8c6abf5a2043e14d402bba811ded7

Observation ffd1cc54-6e51-4411-8c84-438a46ef7573 · outbound

This paper cites Can audio captions be evaluated with image caption metrics?,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Can audio captions be evaluated with image caption metrics?,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:36.701552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:01:35.799846Z digest=sha256:8604b8b77c5d89769272dca5f91601bce7f096a1842cc1d1c518a763e8d4800f

Observation 0e90ee26-cb6f-4015-84ee-ea6ba9f10b91 · outbound

This paper cites Enclap++: Ana- lyzing the enclap framework for optimizing automated audio cap- tioning performance,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Enclap++: Ana- lyzing the enclap framework for optimizing automated audio cap- tioning performance,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:36.552520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:01:35.906298Z digest=sha256:513fc7974cddd4e6e49087a43031d230d2c931c17ed5ee2e49f321549aacefbe

Observation 0a1131d4-b1e1-4a29-91fa-1082570c3239 · outbound

This paper cites Slam-aac: Enhancing audio captioning with paraphras- ing augmentation and clap-refine through llms,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Slam-aac: Enhancing audio captioning with paraphras- ing augmentation and clap-refine through llms,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:36.366008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:01:35.993878Z digest=sha256:ff1860614f8428e72cfdc7aeefc2fc5e326624138135dd6b7d300cc6f3623567

Pith citing papers

Observation d9fec535-2a2f-45d1-8589-33a4efdc3814 · inbound

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer cites this paper.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T12:01:36.208170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:01:33.023609Z digest=sha256:f332a502754a983b6f9498225af16675911166d7dbef57408df4b9eeeb8aa6a5