Pith. sign in

Paper Citation Record · LEDGER

AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models

As of 12 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 4 inbound Pith citation observations for arXiv:2411.18953.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.18953 v1

Coverage vector

measured 62 of 62 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T10:46:23.625101Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T14:36:19.810821Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T00:23:19.246908Z

Reference resolution

62 of 62 outbound references displayed

  • verified exact1
  • verified fuzzy42
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5412e09b-4750-44a0-81b6-866377de6ce9 · outbound

This paper cites WavCaps: A chatgpt-assisted weakly-labelled audio captioning dataset for audio-language multimodal research,.

AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models WavCaps: A chatgpt-assisted weakly-labelled audio captioning dataset for audio-language multimodal research,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:46:24.046844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:46:23.472444Z digest=sha256:9ff8d488d662c0ca4432f661f618928f8b6bf7054d2ab858640bec34ba1fbaaa

Observation fda6a710-5d53-4d4f-8cb7-a5448786865f · outbound

This paper cites Pengi: An audio language model for audio tasks,.

AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models Pengi: An audio language model for audio tasks,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:46:24.039573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:46:23.475396Z digest=sha256:1f5d616a78478e04628a407768bb10d6a355077de3a51254127bfdf8baa80880

Observation 1cce0362-de32-4d59-aae4-00b756525d12 · outbound

This paper cites A Survey of Large Language Models.

AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models A Survey of Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T10:46:23.477868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:46:23.477868Z digest=sha256:dd5a787e15c413f9adc59c677a5d4538c678d7bcb5deacc16d861f85bcdbe6aa

Observation 5f309895-9776-45e8-bca3-ad3d57ba8667 · outbound

This paper cites Prompting large language models with speech recognition abilities,.

AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models Prompting large language models with speech recognition abilities,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:46:24.031909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:46:23.481022Z digest=sha256:57dd4867d486674a7959015aa24eb7529e0a8631b841e9c2644a4af87a882a38

Observation 1c6c905b-d0bd-4504-827b-2fdac303ae4e · outbound

This paper cites AudioLDM: Text-to-audio generation with latent diffusion models,.

AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models AudioLDM: Text-to-audio generation with latent diffusion models,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:46:24.024571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:46:23.483625Z digest=sha256:ff9022b3a7f181bb18307e67de111543f0226b6e52ebbf26c338d1bea0dd6fb6

Observation 6111ccf1-dcdc-4d26-bb53-b98f14742a81 · outbound

This paper cites AudioLDM 2: Learning holistic audio generation with self-supervised pretraining,.

AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models AudioLDM 2: Learning holistic audio generation with self-supervised pretraining,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T10:46:23.486192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:46:23.486192Z digest=sha256:da776ee5d8b8395dd6d4f7d9bbe7311bc819f6e723196a0e8f484c590a0e6c98

Observation be342558-9bb6-4582-9953-923f192bc141 · outbound

This paper cites Audio retrieval with natural language queries: A benchmark study,.

AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models Audio retrieval with natural language queries: A benchmark study,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:46:24.013342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:46:23.488951Z digest=sha256:ad2c6b356633272585d6a8766507e4c2eb76add762e33783e26d41825346f872

Observation 153faeac-14dd-41d7-907e-b1f231193a76 · outbound

This paper cites Audiolog: LLMs-powered long audio logging with hybrid token- semantic contrastive learning,.

AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models Audiolog: LLMs-powered long audio logging with hybrid token- semantic contrastive learning,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:46:24.006758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:46:23.491426Z digest=sha256:9ecad3f0464d93c7c9785d487cdeb601e27f3027a6831f8c872589dc1661b892

Observation bb9a2217-7c62-492d-87ad-1945a3870248 · outbound

This paper cites AIR-Bench: Benchmarking Large Audio-Language Models via Generative Comprehension.

AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models AIR-Bench: Benchmarking Large Audio-Language Models via Generative Comprehension

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T10:46:23.493650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:46:23.493650Z digest=sha256:bda0a87e452995b68eddee9b259118547cc3b75efb102498cb36271dd12981c8

Observation a4c3c30a-c2b8-40b7-8c22-efc1038db471 · outbound

This paper cites Sparks of Large Audio Models: A Survey and Outlook.

AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models Sparks of Large Audio Models: A Survey and Outlook

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T10:46:23.496448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:46:23.496448Z digest=sha256:0515832dda5fc26233fc846f705625c0598fa5f8a95cf10e31d9a18c7b19c738

Observation 70798fde-685b-4eb8-8f9d-94e74084282b · outbound

This paper cites Audio-Language Datasets of Scenes and Events: A Survey.

AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models Audio-Language Datasets of Scenes and Events: A Survey

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T10:46:23.499141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:46:23.499141Z digest=sha256:484463aef8ee177e1fda52b0fc6b3121995779a551ef4fada20c1b67c387b7ac

Observation 4a4c5665-ff4c-4ec6-a60a-1029d8409f5d · outbound

This paper cites Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,.

AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:46:23.999997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:46:23.501922Z digest=sha256:68097160c76ef009220b55593b379f839d164bf8cbdca7f17cc72a706afa0aaa

Observation 4d9f7c2f-491f-4c55-9202-38437eb3e36b · outbound

This paper cites CLAP learning audio concepts from natural language supervision,.

AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models CLAP learning audio concepts from natural language supervision,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:46:23.993308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:46:23.504346Z digest=sha256:8f9eeccda30d36aa12bde850e0e6946367e607de8cb6eb4f3d67accbb0a89fd7

Observation c452f5cb-cd64-4616-8bbc-94f3eefa1848 · outbound

This paper cites EmotionCaps: Enhancing Audio Captioning Through Emotion-Augmented Data Generation.

AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models EmotionCaps: Enhancing Audio Captioning Through Emotion-Augmented Data Generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T10:46:23.506748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:46:23.506748Z digest=sha256:829aa9de874b3e0400d0b196566dee28840efb06297e428199fb2d7cb1bc8c2c

Observation 8187fd51-6258-45fb-b0c5-6d2453337a10 · outbound

This paper cites Auto-ACD: A large-scale dataset for audio-language representation learning,.

AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models Auto-ACD: A large-scale dataset for audio-language representation learning,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:46:23.986628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:46:23.509510Z digest=sha256:1a78779e0e120f5868c3f8c63133df386d79c4db418c646fe8a1b0fdbe23c2ff

Observation 9ee79af6-7f1e-4d93-86f8-67acf816466b · outbound

This paper cites Sound-VECaps: Improving Audio Generation with Visual Enhanced Captions.

AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models Sound-VECaps: Improving Audio Generation with Visual Enhanced Captions

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T10:46:23.511855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:46:23.511855Z digest=sha256:f740721ac57bbb1fa598f247aea0f163641f0c005e2e709ab7590294681fc30b

Observation 34df367f-df9c-4851-9380-ba9f08c25cf3 · outbound

This paper cites Listen, think, and understand,.

AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models Listen, think, and understand,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:46:23.980087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:46:23.514902Z digest=sha256:434d0d0e1fecdef595a07ced855f29dd6194c4d465c02ca52766f20fd71e9612

Observation 37b41303-ac18-4d8b-9192-6c210e4aa9a3 · outbound

This paper cites Joint audio and speech understanding,.

AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models Joint audio and speech understanding,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:46:23.973414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:46:23.517262Z digest=sha256:36b230fa49111fa8b46a503f694ce067e2f2084b524c913da6182edb62ccd2fa

Observation 19b40a60-6b0d-4cc3-af48-8d55b32759de · outbound

This paper cites Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models.

AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T10:46:23.519965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:46:23.519965Z digest=sha256:555b40bfc9d62ceaf305ec1b29f85cd5b03eb392c62f44e13f1dd2ad9e5b53e8

Observation 7598409f-d473-4ab0-a335-7a5eab74f37c · outbound

This paper cites Audio Set: An ontology and human-labeled dataset for audio events,.

AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models Audio Set: An ontology and human-labeled dataset for audio events,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:46:23.966618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:46:23.522806Z digest=sha256:db5854515ce94e4a8cde0711a390c4859bb24f2fd54010166d9223dd7c4b35c2

Observation e2cd386a-2a6b-4180-ae51-059a378cbc13 · outbound

This paper cites AudioCaps: Generating captions for audios in the wild,.

AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models AudioCaps: Generating captions for audios in the wild,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:46:23.959760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:46:23.525231Z digest=sha256:2cf100fea2e5a50911dcd0c09be6093bea128f73603d81ef89e5170a9a6bed87

Observation 512e81bb-4ea8-4c8b-a1af-67228eb499cb · outbound

This paper cites Clotho: an audio captioning dataset,.

AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models Clotho: an audio captioning dataset,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:46:23.952599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:46:23.527709Z digest=sha256:0b1627192944ea33c2d50243ba0fa5bf91ed46cf7d583c791f82d18d985a6548

Observation 8b82374e-64cf-48ec-ab1b-32d3124a7b24 · outbound

This paper cites Diversity and bias in audio cap- tioning datasets,.

AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models Diversity and bias in audio cap- tioning datasets,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:46:23.945778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:46:23.530093Z digest=sha256:061c1accbf0f9a321ed3c053f2036ffde7d10bafba9e8e82e28a03488476d43c

Observation 6087bed5-8891-48d3-83c8-9abcdd51d92e · outbound

This paper cites Mistral 7B.

AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models Mistral 7B

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T10:46:23.532458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:46:23.532458Z digest=sha256:58009aaaa4e23128d95860d5b83b7b0a922aa459bc79eac69a877214f8e78400

Observation 815b8c94-99bb-43bc-9be0-a4e1b8331bb4 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T10:46:23.535022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:46:23.535022Z digest=sha256:d1affb95f7afdccd87e7cbedd5911adc37df765cd0fb38c1906c85a11eaf2f87

Observation edb17653-531d-4f7c-8172-406fa979f9af · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models LLaMA: Open and Efficient Foundation Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T10:46:23.537522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:46:23.537522Z digest=sha256:ea013aa4d0baa72212b542ac34affc77194082d7a31a1e2a71158e1ad97984ce

Observation bd9fd254-456a-48c5-b36e-da9558d15077 · outbound

This paper cites A Reference-free Metric for Language-Queried Audio Source Separation using Contrastive Language-Audio Pretraining.

AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models A Reference-free Metric for Language-Queried Audio Source Separation using Contrastive Language-Audio Pretraining

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-12T10:46:23.677547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:46:23.539990Z digest=sha256:e138495328a58ff5c491da34ee9c805170c6898aa5335ee2458c2305eea1d905

Observation e2cf8811-6424-4476-8026-61e30ce46a5c · outbound

This paper cites AI Chains: Transparent and controllable human-ai interaction by chaining large language model prompts,.

AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models AI Chains: Transparent and controllable human-ai interaction by chaining large language model prompts,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:46:23.938524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:46:23.542900Z digest=sha256:34cf988637d52312501b8d6a8301e3966681aa8a6f4412e9b7e186785a6a78ce

Observation 748f18a3-2bbb-45e2-b8c8-d796dbf03383 · outbound

This paper cites Classify first, and then extract: Prompt chaining technique for information extraction,.

AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models Classify first, and then extract: Prompt chaining technique for information extraction,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:46:23.931928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:46:23.545262Z digest=sha256:e108d90dd19cb3eb456a8afba4aee467ebf9a225e97cd56dfa07c1c2ed2ab3da

Observation e668f465-a80e-4cc5-982e-254274855cce · outbound

This paper cites Prompt Chaining or Stepwise Prompt? Refinement in Text Summarization.

AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models Prompt Chaining or Stepwise Prompt? Refinement in Text Summarization

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T10:46:23.547684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:46:23.547684Z digest=sha256:0b3ef4fddc1e9731357e9156f9e7ce680a28a7504306abaa9d523f80236c8343

Observation 6a600e20-b37a-4f8e-aefe-129ce631f31a · outbound

This paper cites Automated audio captioning with recurrent neural networks,.

AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models Automated audio captioning with recurrent neural networks,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:46:23.925196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:46:23.550512Z digest=sha256:df69ffd540532a6554b56e3621cab373fd9a7286143bda95ca04c28449b061cd

Observation e55bcf0f-efbe-4541-b911-c6493ac10e93 · outbound

This paper cites Automated audio cap- tioning: An overview of recent progress and new challenges,.

AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models Automated audio cap- tioning: An overview of recent progress and new challenges,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:46:23.918341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:46:23.552716Z digest=sha256:a584b4a24e7c5a84c9e86ffc9058bfc6631181fb74571204c93adc92e3f56f97

Observation 5cdef9eb-96ed-4ca9-b7e7-bacbc9f3636d · outbound

This paper cites Beyond the Status Quo: A contemporary survey of advances and challenges in audio captioning,.

AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models Beyond the Status Quo: A contemporary survey of advances and challenges in audio captioning,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:46:23.911555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:46:23.555098Z digest=sha256:21c2b81db85f170220b54be6ac669b3a9f5231b17c4269141d93cdb434e9aa9f

Observation b6f33fba-db54-44e9-ae5e-222ca397ce3b · outbound

This paper cites HTS-AT: A hierarchical token-semantic audio transformer for sound classification and detection,.

AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models HTS-AT: A hierarchical token-semantic audio transformer for sound classification and detection,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:46:23.904939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:46:23.557535Z digest=sha256:a9732e8d68b5500bade3473879eb0fdfc195758e637fb8b4b8d907957659a49a

Observation aa36887d-76ef-4160-944d-771faaf6b155 · outbound

This paper cites BART: Denoising sequence-to- sequence pre-training for natural language generation, translation, and comprehension,.

AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models BART: Denoising sequence-to- sequence pre-training for natural language generation, translation, and comprehension,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:46:23.898276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:46:23.559850Z digest=sha256:5ba50577b5ca06103bb3b63659406227ce61a46b4566887ff2018bda7638ee70

Observation 85acfd53-8061-4388-8457-9f498f86c114 · outbound

This paper cites BLEU: a method for automatic evaluation of machine translation,.

AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models BLEU: a method for automatic evaluation of machine translation,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:46:23.891478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:46:23.562395Z digest=sha256:cd07d7b649caf2f1e0c817dd7894b7e711c8b7dff9e8f6510087904cf740eeb7

Observation 3e43268d-f253-4f4b-8beb-59223e87714c · outbound

This paper cites ROUGE: A package for automatic evaluation of summaries,.

AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models ROUGE: A package for automatic evaluation of summaries,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:46:23.884067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:46:23.564750Z digest=sha256:2f4a0c762a01f2aa309d4e35a4b58abaa198165ed4ca323889bd410b63d39533

Observation ae0e2482-55a8-46d8-9cf0-f0a94091ee2b · outbound

This paper cites METEOR: An automatic metric for mt evalu- ation with improved correlation with human judgments,.

AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models METEOR: An automatic metric for mt evalu- ation with improved correlation with human judgments,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:46:23.877431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:46:23.567063Z digest=sha256:5eff789377e872ec7ca7ea45bbfe19069b8fb5f9c3cc90692f100392b0b4bc4a

Observation 99f8ee5e-7073-4463-a19d-8dd8ec3bdf39 · outbound

This paper cites CIDEr: Consensus- based image description evaluation,.

AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models CIDEr: Consensus- based image description evaluation,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:46:23.869698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:46:23.569233Z digest=sha256:f6a50d00a26021013401987d10ec1ea1d37f3cf41e27165e183d1628684e9e26

Observation 7effa66e-3450-4e30-a92b-6431f4de1500 · outbound

This paper cites SPICE: Semantic propositional image caption evaluation,.

AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models SPICE: Semantic propositional image caption evaluation,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:46:23.862108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:46:23.571594Z digest=sha256:42df88cbb46ca75511a217d8312714d7c41a5261276e4e05dfdc40f7820cee32

Observation 0defaa85-deb6-4f3b-94db-e971433ea7c4 · outbound

This paper cites Improved image captioning via policy gradient optimization of SPIDEr,.

AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models Improved image captioning via policy gradient optimization of SPIDEr,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:46:23.854949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:46:23.573955Z digest=sha256:962d96e597eed40359849530df486efb7b81fb7e8e811d36e50ff6d0a8945fd1

Observation b17b16cf-a201-4da7-9d4a-8e4c5e4a9e3b · outbound

This paper cites EnCLAP: Combining neural audio codec and audio-text joint embedding for automated audio cap- tioning,.

AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models EnCLAP: Combining neural audio codec and audio-text joint embedding for automated audio cap- tioning,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:46:23.847422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:46:23.576341Z digest=sha256:84a2b03d7f2a26c030e286de244704fa4528d819d0259d5273d2870bf35c0d2a

Observation ea6a31d0-8fb8-451d-babe-cc4b4977a19c · outbound

This paper cites CoNeTTE: An efficient audio captioning system leveraging multiple datasets with task embedding,.

AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models CoNeTTE: An efficient audio captioning system leveraging multiple datasets with task embedding,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:46:23.840158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:46:23.578773Z digest=sha256:93f669ee107a283b95a35768e5e0ef2833f1804c61e299fd4fc5621bd46bb187

Observation ca402966-e9f1-484e-8561-a58021840176 · outbound

This paper cites Taming Data and Transformers for Audio Generation.

AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models Taming Data and Transformers for Audio Generation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T10:46:23.581165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:46:23.581165Z digest=sha256:0576e950941091da0097d90a0f61e5eea2247f2dbc485270da3465986459bb58

Observation 7b14bb3e-303a-4526-85d9-a59f16d7b0c3 · outbound

This paper cites Bridging Language Gaps in Audio-Text Retrieval.

AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models Bridging Language Gaps in Audio-Text Retrieval

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T10:46:23.583660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:46:23.583660Z digest=sha256:b42e9bd7f6e6a43ded537389ac219412d82201ead1728f98c6dca7d667377850

Observation 47cbc6da-e85c-41e2-8fe3-0814f149b2bf · outbound

This paper cites Language-based Audio Retrieval Task in DCASE 2022 Challenge,.

AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models Language-based Audio Retrieval Task in DCASE 2022 Challenge,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:46:23.832888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:46:23.586339Z digest=sha256:359ec6654769d6b5a406959e982986a7b987fdb29be208cb1f154638a3845dd5

Observation 716f0ad4-8b3a-4022-81cd-22d38ca135eb · outbound

This paper cites On metric learning for audio-text cross-modal retrieval,.

AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models On metric learning for audio-text cross-modal retrieval,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:46:23.825638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:46:23.588770Z digest=sha256:de85dba0684e8109e507e4ace83c7a9aa5df5596b8e9de8c4f4af7362e341289

Observation c239b82e-ec42-407c-a247-6a915d972e0a · outbound

This paper cites Audio-text retrieval in context,.

AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models Audio-text retrieval in context,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:46:23.818473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:46:23.591035Z digest=sha256:196cf38a27fb25d4d6f85ff129accd6e158096637a602d77c02f9ef88810be11

Observation b09eaca3-bdf7-4705-bced-ae5a5300e589 · outbound

This paper cites Momentum contrast for unsupervised visual representation learning,.

AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models Momentum contrast for unsupervised visual representation learning,

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T10:46:23.593657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:46:23.593657Z digest=sha256:84ba99a3c43bcfbb27893f4b18e8c052524cd4ec99aa31e6056e81439bbe68c4

Observation cee41d20-b763-4d4b-b797-5cdf188e1bfe · outbound

This paper cites BERT: Pre- training of deep bidirectional transformers for language understanding,.

AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models BERT: Pre- training of deep bidirectional transformers for language understanding,

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T10:46:23.595993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:46:23.595993Z digest=sha256:2448f1d49d6d5bb86eb6be8fdd7a8a7652af28ab85e8af52bc53c30020929149

Observation 1dc61dd2-70c1-4e9a-b933-87385846af75 · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T10:46:23.598481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:46:23.598481Z digest=sha256:3f93cc23da7bdf23d01bdd6e6d7e89f9cc480e439b8e04396d6543467ad1e2fe

Observation c987fde5-e302-4034-b548-a4eb27d8b031 · outbound

This paper cites CED: Consistent ensemble distillation for audio tagging,.

AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models CED: Consistent ensemble distillation for audio tagging,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:46:23.803377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:46:23.601051Z digest=sha256:8500e9f9f9246dc1b97288f64c20c1617a256fee7ca722716e4a4a3822ff7ffa

Observation 8ed14750-33cc-4f36-a9a6-d3f5d8ebf5de · outbound

This paper cites SONAR: Sentence-Level Multimodal and Language-Agnostic Representations.

AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models SONAR: Sentence-Level Multimodal and Language-Agnostic Representations

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T10:46:23.603547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:46:23.603547Z digest=sha256:613e9564ef810b743a4be692ab1cdbff64e932e908eda62a0b7528d1bb2ee184

Observation ee550096-6068-4f51-96ee-bfd8a58a5077 · outbound

This paper cites Zero-shot audio classification based on class label embeddings,.

AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models Zero-shot audio classification based on class label embeddings,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:46:23.795741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:46:23.606297Z digest=sha256:5af5f172a9240cfab6ac202d084d0c0ac97e006377a1622ee505f8114ca01fcb

Observation ec1a6be6-1f34-41c2-9881-b14f3e405b75 · outbound

This paper cites Zero-shot audio classification via semantic embeddings,.

AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models Zero-shot audio classification via semantic embeddings,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:46:23.788393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:46:23.608703Z digest=sha256:8cf3058edd3b3d7c54c7aae833e74a625ec2b7a8cd38801de756151d0f80ba4a

Observation c63467ed-dff0-4033-91c6-9eb6d43fd350 · outbound

This paper cites A dataset and taxonomy for urban sound research,.

AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models A dataset and taxonomy for urban sound research,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:46:23.781137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:46:23.610973Z digest=sha256:9173cb59404d462a0cde82b2039589ccb649c4ed417a7dd2625262d05893ee63

Observation 1a70be6d-78e9-42d0-829f-28064a60410f · outbound

This paper cites ESC: Dataset for environmental sound classification,.

AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models ESC: Dataset for environmental sound classification,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:46:23.773913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:46:23.613339Z digest=sha256:7d4885e74bc9d4d255a2cc40814118af5b62d8199d6a9da0607a06bb57d20c7b

Observation 2fac8381-5f92-4251-bfee-977d69e70d3d · outbound

This paper cites Common V oice: A massively-multilingual speech corpus,.

AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models Common V oice: A massively-multilingual speech corpus,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:46:23.766729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:46:23.615535Z digest=sha256:1da64fb85bdce23dc2ac3b24f777e973e1704e30485250b63e3bdbaac0e547d4

Observation d4f10cf4-21ac-40f6-b90d-aa60b18d2462 · outbound

This paper cites CREMA-D: Crowd-sourced emotional multimodal actors dataset,.

AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models CREMA-D: Crowd-sourced emotional multimodal actors dataset,

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-12T10:46:23.617990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:46:23.617990Z digest=sha256:66f2ea2d061f5fca96a31cd00ef80e64ac12b04b19e9ec59284a8a060099b477

Observation 2eff404b-cbdf-4265-84e2-48aa01e9330f · outbound

This paper cites The ryerson audio-visual database of emotional speech and song (RA VDESS): A dynamic, multimodal set of facial and vocal expressions in north american english,.

AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models The ryerson audio-visual database of emotional speech and song (RA VDESS): A dynamic, multimodal set of facial and vocal expressions in north american english,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:46:23.755928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:46:23.620375Z digest=sha256:e2655c77ce0e9364cbff16ab4b751d131d3e843c36cea464d7024b13545eabad

Observation c6807c65-e7c8-47e2-9b09-65f31a822756 · outbound

This paper cites Automatic musical genre clas- sification of audio signals,.

AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models Automatic musical genre clas- sification of audio signals,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:46:23.748565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:46:23.622753Z digest=sha256:dd96a99f615ff6570a1e3f3e50d38367ab4dc263ffd2f28e5f1351e4014a01a1

Observation b8237517-42ab-46ae-8f75-d837defccb9f · outbound

This paper cites OpenMIC-2018: An open data-set for multiple instrument recognition,.

AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models OpenMIC-2018: An open data-set for multiple instrument recognition,

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:46:23.740943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:46:23.625101Z digest=sha256:d244732d01758d8faa3dc10835561a75a91d58c8bc63874f3b7cc2c0d317d6a5

Pith citing papers

Observation e0dc5f4c-f6ea-4aff-9abf-fa28969b7b5a · inbound

Audio-Language Models for Audio-Centric Tasks: A Systematic Survey cites this paper.

Audio-Language Models for Audio-Centric Tasks: A Systematic Survey AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models

Reference 137

Resolution
unresolved
no resolver link, observed 2026-08-10T14:36:19.810821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:36:19.810821Z digest=sha256:b8b3aab25a4da55bfa8871e0dbedb00406bbbe1812449c342feb2b3a8856497a

Observation 11c2d147-4f94-4224-a045-a88b69bd12b0 · inbound

CLaMP 3: Universal Music Information Retrieval Across Unaligned Modalities and Unseen Languages cites this paper.

CLaMP 3: Universal Music Information Retrieval Across Unaligned Modalities and Unseen Languages AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:13.836268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:13.836268Z digest=sha256:fbf73f14357dcd74cce3437a4bc9cbdd11b86bf3b8fa0f12485d1ec5971409c0

Observation d4795b6c-b14e-4649-ba61-c559c05a85d5 · inbound

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval cites this paper.

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:23:19.253506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T00:23:16.841565Z digest=sha256:71964ccf87372df63ac37160b434c6fe7685408b9dceb4d82d0826043991f10e

Observation 03c50d2d-29da-42ee-90ed-d6ba7bb606b0 · inbound

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio cites this paper.

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T14:46:34.886995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:46:34.886995Z digest=sha256:bc2661f5b59c91d2d39aec0bcc0c43d7ce1127c76ffeb9b9cd69ed04efdae6b8