Pith. sign in

Paper Citation Record · LEDGER

FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion

As of 20 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 4 inbound Pith citation observations for arXiv:2506.01111.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.01111 v1

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:54:26.533962Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T15:41:34.152062Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T07:39:38.511905Z

Reference resolution

48 of 48 outbound references displayed

  • verified exact1
  • verified fuzzy23
  • unresolved24
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cc370d3f-ef9f-49e8-87e9-f98a4d298af9 · outbound

This paper cites Large-scale contrastive language-audio pretraining with feature fusion and keyword- to-caption augmentation.

FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion Large-scale contrastive language-audio pretraining with feature fusion and keyword- to-caption augmentation

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:54:21.897043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:54:21.897043Z digest=sha256:ba056091455845572b06e6f218d38c6ba832d1b59438f92c4b1dcf8173441e64

Observation 38d7254e-d4ca-4bdf-80a0-79f7f82527d3 · outbound

This paper cites GAMA: A large audio-language model with advanced audio understanding and complex reasoning abilities.

FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion GAMA: A large audio-language model with advanced audio understanding and complex reasoning abilities

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:33.855577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:54:21.982508Z digest=sha256:8b984c1fe2ad0e19c2fbbc60f65349d9add4f7fc71d035989ad5300febcf3b6f

Observation d820f451-eaf9-45fe-9417-c12eb40bc215 · outbound

This paper cites Qwen2-Audio Technical Report.

FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion Qwen2-Audio Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:54:22.155841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:54:22.155841Z digest=sha256:6fee79ba07ef8fe65a452efbd7a3170586dc1e5c39304f5f0613d161131bd239

Observation b78b85a1-16ac-4e6c-a2d5-bd3926cfae93 · outbound

This paper cites Clotho: An audio captioning dataset, 2019.

FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion Clotho: An audio captioning dataset, 2019

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:33.667717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:54:22.269803Z digest=sha256:6a108ff3fc00dfc9477c112ff730fd3bb0504aaf812d1c45f77917bb76ca8f92

Observation 5eecb123-44b0-4203-b7d2-969117562a76 · outbound

This paper cites AudioCaps: Gen- erating captions for audios in the wild.

FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion AudioCaps: Gen- erating captions for audios in the wild

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:33.459637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:54:22.389758Z digest=sha256:7a8374e226273b8fd61a8764cf03ed7974e8e2b1d423b520915ea908830ac876

Observation b35f10ac-e745-4208-a2a2-9fcf797c9f94 · outbound

This paper cites Laion-audio-630k dataset, 2023.

FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion Laion-audio-630k dataset, 2023

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:33.252807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:54:22.579658Z digest=sha256:de8f1e1c64b8fb899f74e833434ad856133bcc158fb84c87b3f36d04f1b59482

Observation 2451f32d-b313-47aa-b839-3bddeac183c0 · outbound

This paper cites Plumbley, Yuexian Zou, and Wenwu Wang.

FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion Plumbley, Yuexian Zou, and Wenwu Wang

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:54:22.732392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:54:22.732392Z digest=sha256:6cb789f2d581cd45bfee2cd10de58c6726c93f34159045e7700fb74d586c4325

Observation 5a50eb9d-dd42-4b28-b624-f47d9fb5e24e · outbound

This paper cites Plumbley, Woon- Seng Gan, and Jianfeng Chen.

FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion Plumbley, Woon- Seng Gan, and Jianfeng Chen

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:33.064023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:54:22.838399Z digest=sha256:79fe09b417d9572fcdae3f2344d601018c7c9a7642fe0b5cf1113f06c2792c5c

Observation fc93b794-3e01-472c-876e-f7598d659ac8 · outbound

This paper cites Auto-acd: A large-scale dataset for audio-language representation learning, 2024.

FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion Auto-acd: A large-scale dataset for audio-language representation learning, 2024

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:32.823625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:54:22.937909Z digest=sha256:324b8e068101096c5b1c849b0eea536f6f4a704f14d85d389bbecf80b52dcd36

Observation 6644e575-145e-4f33-a8c6-998eed237ab5 · outbound

This paper cites Plumbley, and Wenwu Wang.

FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion Plumbley, and Wenwu Wang

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:32.603896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:54:23.016763Z digest=sha256:e82dcde8742ff540dd5cb2d5b04c5b0ec2b0006c4cf634ae9c4b2f88163de49e

Observation 2595efec-ca1b-4177-baf7-4db62f5eb713 · outbound

This paper cites Air-bench: Benchmarking large audio-language models via generative comprehension, 2024.

FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion Air-bench: Benchmarking large audio-language models via generative comprehension, 2024

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:54:23.123758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:54:23.123758Z digest=sha256:9adb6ccccf5163efb82a25896b07f2e3aaa6862536f5568d63fd9f9d5a3b1dc2

Observation 7526a4a6-4242-4a4a-92b1-275753e75560 · outbound

This paper cites an unresolved cited work.

FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:54:32.366303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:54:23.284501Z digest=sha256:7250f3a51cc1584e6b73f855089f9de08dd51247dc758830fcaf6f6780f6faad

Observation 623cae2a-1a2d-484e-bd27-603b159ac3dc · outbound

This paper cites Logothetis, and Stefano Panzeri.

FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion Logothetis, and Stefano Panzeri

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:32.130971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:54:23.369380Z digest=sha256:9adaa906448f2435b24a373dba1a13a5535fa3982993fdff986b5b14c780659d

Observation 66fa02db-07a5-46a8-b7b0-df63937ee244 · outbound

This paper cites Ernst and Heinrich H.

FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion Ernst and Heinrich H

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:31.933877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:54:23.514952Z digest=sha256:eeec5209f30916d4930639e1fc92c4e6a2e8b643a693b1ee12cd627f4e106d4a

Observation f5478d8a-f51b-41a3-9ce4-9edf31aa7945 · outbound

This paper cites Bregman.Auditory scene analysis: The perceptual organization of sound.

FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion Bregman.Auditory scene analysis: The perceptual organization of sound

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:31.717289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:54:23.606715Z digest=sha256:7e2924e94ad3a8fca71a90d5353a374381fadc1035b9c873d7f7ba0bee4b1dae

Observation db545bcc-c8e5-44d6-b9d6-a62097f4f62a · outbound

This paper cites an unresolved cited work.

FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:54:31.475796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:54:23.696649Z digest=sha256:6bfc7aba59742f2cbc8ba0fe1e53fc51cdbcbb97b8a1b06357ef320d284aae8f

Observation fcc38876-3201-4953-8df3-a34d43f5d720 · outbound

This paper cites Robust speech recognition via large-scale weak supervision, 2022.

FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion Robust speech recognition via large-scale weak supervision, 2022

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:54:23.791108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:54:23.791108Z digest=sha256:6b55a71bfee50f6745646db03adbc8e7b763405e249d0361f40d956618010461

Observation 1a7114b2-1ba5-47c6-a4ce-75b965456850 · outbound

This paper cites OpenMU: Your Swiss Army Knife for Music Understanding.

FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion OpenMU: Your Swiss Army Knife for Music Understanding

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:54:23.874122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:54:23.874122Z digest=sha256:3bc2f8f075cab1835690450fdf31b300fa3a327ebe3283e348e918d60270b924

Observation c05f132e-b094-48e9-842c-ed8375b1003c · outbound

This paper cites Qwen2.5-VL Technical Report.

FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion Qwen2.5-VL Technical Report

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:54:23.957188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:54:23.957188Z digest=sha256:ccd301561907214242453639c75374acd161bd05904e8a3a9a6be9eb148e717a

Observation 246d2c6d-8a8a-4502-adc1-9c8eb70c8476 · outbound

This paper cites Qwq-32b: Embracing the power of reinforcement learning, March 2025.

FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion Qwq-32b: Embracing the power of reinforcement learning, March 2025

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T11:54:24.010409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:54:24.010409Z digest=sha256:cf4415ea85c21942011a76bf5cb8f9878b9bc6f54fc39c3fe15bc3f13c0c70b0

Observation 5b530e6b-57c4-448c-a732-c3f507502ff5 · outbound

This paper cites Elizalde, S.

FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion Elizalde, S

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:31.317776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:54:24.064260Z digest=sha256:42de91e9e441042d245ce105778bbc42337aa5681769ecaae030fb3a83d1672d

Observation 96a26cd4-5438-442b-8501-83c7f5e3ccbf · outbound

This paper cites CTAL: Pre-training Cross-modal Transformer for Audio-and-Language Representations.

FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion CTAL: Pre-training Cross-modal Transformer for Audio-and-Language Representations

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:54:26.810143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:54:24.139798Z digest=sha256:f9919ab2acfbfc407f9814f3e08c63f8be710b290a6e26dc9d056ce29e2b2cad

Observation 21d36ae2-4189-491f-a8d3-e1e5c3946cab · outbound

This paper cites Yeh, P.-Y.

FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion Yeh, P.-Y

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:31.096452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:54:24.192860Z digest=sha256:5133977af5b71332cf59a1a9c34acd8c58c16eaadf94c00f197ff409429edbd2

Observation 93b38794-9dc5-4c3f-b4e9-219efef4dac9 · outbound

This paper cites UniAudio: An Audio Foundation Model Toward Universal Audio Generation.

FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T11:54:24.284181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:54:24.284181Z digest=sha256:cdac68d8cde7a4bf819b42a194594ab3b1f335884074fe4184786275ff5acb80

Observation 934e5a04-3a41-46a3-a014-3186c6fd475d · outbound

This paper cites an unresolved cited work.

FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:54:30.908490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:54:24.399481Z digest=sha256:457d1e6cb346a36e08782281ca913d12f2186ff8ac142e0ca557fa795d8fbb7f

Observation 8ee85056-7646-4160-aa5d-1a7aa678299e · outbound

This paper cites Pengi: An audio language model for audio tasks.

FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion Pengi: An audio language model for audio tasks

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:30.733035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:54:24.533014Z digest=sha256:d9b5311bf3af247dde0e737638e6bc983ce8363d114b56b906913a9f9b279e0c

Observation 49ce8f76-5915-466e-9e36-442b9caf9ff3 · outbound

This paper cites Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models.

FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:54:24.590613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:54:24.590613Z digest=sha256:83ef94d170a5aaf3893fae0f248d31acbdc7e421e508cd16f8cdd0ea1a4a6b49

Observation ddf48656-4821-4cf4-a23e-f436705725c2 · outbound

This paper cites Audio Flamingo: A Novel Audio Language Model with Few-Shot Learning and Dialogue Abilities.

FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion Audio Flamingo: A Novel Audio Language Model with Few-Shot Learning and Dialogue Abilities

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:54:24.703966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:54:24.703966Z digest=sha256:82008f628e724d9c12cbb03f13993d1339f3ef9f054b04f82bd6f69951cf7cc7

Observation 72eafdcd-927d-4425-8bee-60ce0ed98965 · outbound

This paper cites Hybrid transformers for music source separation.

FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion Hybrid transformers for music source separation

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:30.538543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:54:24.807745Z digest=sha256:e626c1104077a181732cb1e4e244276c1f393dd4f96abd16ab9b78511398da97

Observation 02d9edf9-e02d-4217-9fc5-4aff04ff41d0 · outbound

This paper cites Yamnet: Audio event classification.

FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion Yamnet: Audio event classification

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:30.293154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:54:24.919854Z digest=sha256:d2341b63c19a1025eb040a541b468ca8d5bb7f93ba2d2566f3605b7e540f8212

Observation 30d3e3b0-1c19-4c46-96f8-6d74b5b4b9b1 · outbound

This paper cites Gemmeke, Daniel P.

FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion Gemmeke, Daniel P

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:30.078863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:54:25.029097Z digest=sha256:fef175fc64fadea12e15ac8869de394f023a71da51f3e9f909aef913ea172934

Observation c2a87ce3-ce16-4470-ab22-7d48c573fdd7 · outbound

This paper cites Visualizing data using t-sne.Journal of Machine Learning Research, 9(86):2579–2605, 2008.

FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion Visualizing data using t-sne.Journal of Machine Learning Research, 9(86):2579–2605, 2008

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T11:54:25.099327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:54:25.099327Z digest=sha256:1de3dc16ab8c1f1aa68ff12df4700b740aef53c63f0cf14775e59ee17b0ec84e

Observation 1cdb47c3-095c-49ba-9dea-11f5b8c8a35a · outbound

This paper cites Hts- at: A hierarchical token-semantic audio transformer for sound classification and detection.

FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion Hts- at: A hierarchical token-semantic audio transformer for sound classification and detection

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:29.832360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:54:25.166218Z digest=sha256:4908c49b42c75257231a45eff285691452dec9c0de248c0bbc0821577b52b556

Observation 5a1aa497-cdd8-43e2-af63-140883b60f41 · outbound

This paper cites Bert: Pre-training of deep bidirectional transformers for language understanding.

FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion Bert: Pre-training of deep bidirectional transformers for language understanding

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:29.622647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:54:25.272799Z digest=sha256:d42c6cf6c0e223c82443d385c72755a2aeabb6ec569894bb96ba33b31d4eb93b

Observation 5f3d0f4a-b6a0-48ba-a7a0-2de3f300d16d · outbound

This paper cites • If the intensity or emotional context of the sound is conveyed (e.g., the dog barking intensely or the doorbell ringing in a rapid succession).

FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion • If the intensity or emotional context of the sound is conveyed (e.g., the dog barking intensely or the doorbell ringing in a rapid succession)

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:29.400615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:54:25.370191Z digest=sha256:0cf0a6da18770190aa767dbae73f9022fa1b6af5501d6a383c3d09bd53e6d1bb

Observation f76d6af2-fcf9-4a94-948d-b022ba748b5e · outbound

This paper cites trumpets,.

FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion trumpets,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:29.201077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:54:25.449949Z digest=sha256:d85e4a877f6d22c453513b6e90109dd1a80cd446d042ed2c009bb7af79053c73

Observation 942d7622-30eb-4367-ad85-108b4391203d · outbound

This paper cites an unresolved cited work.

FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:54:29.016598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:54:25.515475Z digest=sha256:f01ea5ce361ef0359ecb183d4523c252b3fedd7784dd50c9385b0a941bf2ca56

Observation 887a1896-de18-4dae-81fe-055598b02891 · outbound

This paper cites an unresolved cited work.

FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:54:28.819671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:54:25.614622Z digest=sha256:02e201a9eaa5f231bfb5355a85a90871ddf496a3c72feaabb0b39b4334daa30a

Observation 2f4dff6b-52f8-4c4b-be0b-301752e4994c · outbound

This paper cites an unresolved cited work.

FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:54:28.618007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:54:25.722899Z digest=sha256:ead3531408f417799a706d14f96580b962d58b6ccdbdc592986f5b3083d56c25

Observation 647bcb52-c431-4d05-bcc4-bf9144e0205b · outbound

This paper cites an unresolved cited work.

FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:54:28.468326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:54:25.834488Z digest=sha256:baeba279d9ca0ab23b3371670637b65d7c02c00d369db5fb4093e8738d656eaf

Observation 84987096-14f7-4a6a-93c5-df9dea7ea3c5 · outbound

This paper cites an unresolved cited work.

FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:54:28.257431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:54:25.909754Z digest=sha256:92c05c1adb78dd5bc88d496399902cd4afc491b7eea8da0b4fc22cc7de7d2047

Observation a2bcbc8d-1215-46ca-a725-a5dbb990fb6d · outbound

This paper cites an unresolved cited work.

FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:54:28.031006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:54:26.039626Z digest=sha256:3549b51536d422a8bed4b945074cbd15331684918f1ce85c53769df849df00f8

Observation 7fe5bfa9-5d6a-423e-bd72-555cdececd27 · outbound

This paper cites instrument.

FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion instrument

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:27.849529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:54:26.116448Z digest=sha256:d102bd44ba621745a9465d7ba3444c531ae2747eddf2f58b6d00cfc26e11d067

Observation ff0f6336-9620-4f16-b60c-9a53810b553d · outbound

This paper cites an unresolved cited work.

FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:54:27.616673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:54:26.220834Z digest=sha256:4b7005d34608045ea59e6410620c59e93e58a7eec74d6ba994ff0d6c6ce33ea5

Observation 98da82d0-c3fe-4687-b349-4a6ed02d6fe9 · outbound

This paper cites an unresolved cited work.

FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:54:27.452235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:54:26.317122Z digest=sha256:45d42d090aa2edade5589112a18d03498bd0f42b7bd9130de2fd3fdecf4c15da

Observation b7ad94d4-13eb-415c-8f32-d6641bdda7f4 · outbound

This paper cites A car’s engine roars as it accelerates.

FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion A car’s engine roars as it accelerates

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:27.249922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:54:26.419983Z digest=sha256:60d9d73e09a3992dabb1a1b76fe1f8d598c8415cd0747fb9bfb383e5d0a96f10

Observation fdc345c5-9d2c-405e-9690-b471e65f7f27 · outbound

This paper cites Audio Description.

FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion Audio Description

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:27.030893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:54:26.533962Z digest=sha256:6a15e27f60c4741b4b36d94c28791816c94494f25ac4e6cfbaa4fddfaf3a2f14

Observation afbe7848-d998-486f-aed2-0171e0a7c2ae · outbound

This paper cites an unresolved cited work.

FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion Unresolved cited work

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T11:54:22.082802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:54:22.082802Z digest=sha256:fcc863ff788b0e53f48b5eea08157934aa99c720e151d00570d35463f8015bec

Pith citing papers

Observation 2dca371f-5f41-4521-8e81-5b79e200dd2f · inbound

Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation cites this paper.

Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T21:06:08.068024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:06:08.068024Z digest=sha256:aeecc4d86d156e87c738cc773f94248667e8b3f295c3f54581cfb6c952154a8c

Observation 27bde799-b3b5-4866-86e4-ecf1e2a11ab2 · inbound

EvA: An Evidence-First Audio Understanding Paradigm for LALMs cites this paper.

EvA: An Evidence-First Audio Understanding Paradigm for LALMs FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-13T16:50:28.996341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:50:28.996341Z digest=sha256:96af2c6e08115423de400ea2284b79953a4f01a9d4bf882925399ee5840ad962

Observation 1e3d27f5-dfab-44d1-80d7-b2c6cfffbafe · inbound

Bagpiper-Edit: Zero-Shot Open-Ended Audio Editing via Rich-Caption cites this paper.

Bagpiper-Edit: Zero-Shot Open-Ended Audio Editing via Rich-Caption FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:39:38.513487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-26T13:17:40.718000Z digest=sha256:01ef2f2b1079d2e33ded9ea0921e20107b76c2adb5803f1b3db4e96621d952ca

Observation 81f7bfc2-eef1-4023-80a4-976ee2762f86 · inbound

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework cites this paper.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.152062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.152062Z digest=sha256:af241160137d6e53353ff5c55c0ab4522f4c32ac4f4f808ba120a6b93adbe2ae