Pith. sign in

Paper Citation Record · LEDGER

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts

As of 16 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 0 inbound Pith citation observations for arXiv:2607.06611.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.06611 v1

Coverage vector

measured 53 of 53 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-11T01:53:00.127636Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

53 of 53 outbound references displayed

  • verified exact7
  • verified fuzzy45
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7f693040-bd69-46b8-ae57-62bf2b178b88 · outbound

This paper cites The evolution of sentiment analysis and conversational AI: Techniques applications and future research directions.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts The evolution of sentiment analysis and conversational AI: Techniques applications and future research directions

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:48.942291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:d95e06723a1154024c5aa79b17f468972b5a44f2bd4aaffa1f53c137a74abedc

Observation c797137c-7de2-4edd-ae15-66f2ec2735aa · outbound

This paper cites Sentiment analysis and emotion recognition from speech using universal speech representations.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts Sentiment analysis and emotion recognition from speech using universal speech representations

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:48.881121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:1dd822a7698dd77ea6930dfc7e48c51ee090e824cee7f1caacc48b729004634f

Observation f07b8f8e-8391-4d93-8498-6176ad613e87 · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech representations, in: Proceedings of NeurIPS, pp.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts wav2vec 2.0: A framework for self-supervised learning of speech representations, in: Proceedings of NeurIPS, pp

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:48.703203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:60783f3dcbb7df9f29afa341d1d718932bb9333f56a0b5d7e50b37f5db151bc7

Observation d21193c9-b6e7-431c-bbec-bdcb507027ae · outbound

This paper cites TweetEval: Unified benchmark and comparative evaluation for tweet classification, in: Findings of EMNLP, pp.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts TweetEval: Unified benchmark and comparative evaluation for tweet classification, in: Findings of EMNLP, pp

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:48.341409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:8fe4b4d624a2640f87bb0955db09032d5e9449523822f06426341ce76348220a

Observation 1e7a7188-cd28-45ff-9ddc-32c0a3f60164 · outbound

This paper cites Sentiment Analysis of Customer Feedback and Reviews in E-Commerce Systems, in: Proceedings of ICTCS, pp.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts Sentiment Analysis of Customer Feedback and Reviews in E-Commerce Systems, in: Proceedings of ICTCS, pp

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:48.623963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:5e05f021468a4f7bf3fc64120b2f3d01739d3dbf410d26fa290a47c3e14586f6

Observation 9f5f687a-49a2-4fcc-beff-cb3be8a79c5b · outbound

This paper cites IEMOCAP: Interactive emotional dyadic motion capture database.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts IEMOCAP: Interactive emotional dyadic motion capture database

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:49.008905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:7f4bdebe3bdc470b4b6aa7b6963828e7f003b7d26bd18625a1226e8a3a6a581e

Observation 6bf36bf0-23ef-4bd0-a977-ea0adbef5d99 · outbound

This paper cites The MSP-Podcast Corpus.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts The MSP-Podcast Corpus

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-07-11T01:57:52.125343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:c364b907fa4d0c8ba639c6adcf5b0e662ddc1a1ba075011bdaa8bd3cf2510c99

Observation 5581cd06-56bc-48e7-a1d8-cc58f3a6080f · outbound

This paper cites German’s next language model, in: Proceedings of COLING, pp.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts German’s next language model, in: Proceedings of COLING, pp

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:48.561625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:c4b52ddd2b3b3a983b34ccd9ee8695efd881b7242d5eb8f36e194bd04886f2d5

Observation 639c29ce-3311-40e5-ac9e-dc5278c8b5b3 · outbound

This paper cites WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:49.538198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:dff65c451ae738a4dc4c51f9cfbc3f608fe6190d50275511768845eae62d67a4

Observation b7b2eaf5-d8ad-4360-9fbb-d7fe89dfd425 · outbound

This paper cites BERT: Pre-training of deep bidirectional transformers for language understanding, in: Proceedings of NAACT-HLT, pp.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts BERT: Pre-training of deep bidirectional transformers for language understanding, in: Proceedings of NAACT-HLT, pp

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:49.571023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:92036ab7f75b1c230d274eb719aad14da95115b76f5c6a57e73fa85cc21abcdf

Observation 0f7eca62-d1b6-4a2c-ab23-17fd0297fd19 · outbound

This paper cites Optimized Sentiment Analysis in Tagalog Speech Using PCA and BRNN on Prosodic Suprasegmental and MFCC Features, in: Proceedings of ICTC, pp.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts Optimized Sentiment Analysis in Tagalog Speech Using PCA and BRNN on Prosodic Suprasegmental and MFCC Features, in: Proceedings of ICTC, pp

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:49.601183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:17e995aad71662bd190ea6dbd2233dd37bb963eb801116a75bed8ddb63f842e3

Observation 0f0b3717-bb94-4ffc-aecc-b9e4d329bfec · outbound

This paper cites A review on speech emotion recognition: A survey, recent advances, challenges, and the influence of noise.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts A review on speech emotion recognition: A survey, recent advances, challenges, and the influence of noise

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:48.404930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:53abe453623705410d086702fedf5e5759807c47e8cd0d1b4de657c8fc5e6bf6

Observation cdeb7a1d-0ee1-46fa-9d83-315bb24b73f7 · outbound

This paper cites Teacher-Student Training and Triplet Loss for Facial Expression Recognition under Occlusion, in: Proceedings of ICPR, pp.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts Teacher-Student Training and Triplet Loss for Facial Expression Recognition under Occlusion, in: Proceedings of ICPR, pp

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:49.726874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:403f0eaf96c54ef36675049db48ac848ee2147d7e328d09e6bf3f80d783e4c1e

Observation 7e0b00e8-8c0e-40c6-acd1-5426bd551698 · outbound

This paper cites AST: Audio Spectrogram Transformer, in: Proceedings of INTERSPEECH, pp.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts AST: Audio Spectrogram Transformer, in: Proceedings of INTERSPEECH, pp

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:48.944100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:29162ce5b3d031e315896a2ca6c57311d8cb5911f084dffec7de2cef9aedef7b

Observation 173827fc-beb8-4ce7-b67f-01a05ee034c6 · outbound

This paper cites Distilling the Knowledge in a Neural Network.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts Distilling the Knowledge in a Neural Network

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-11T01:57:52.151963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:79f3ef924ffc332e1fe2d8b55de210e6741b8f4c9a9133133f21e8e71cad1e6e

Observation fc45812c-5ef5-406c-bc89-7e90096498fc · outbound

This paper cites HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:49.352399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:79980a0441979637400c04489571305b697734c8179f639eee8fdc3718106a80

Observation c9ae9372-df68-4ce1-ac83-bdbdeda3aba4 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models, in: Proceedings of ICLR.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts LoRA: Low-Rank Adaptation of Large Language Models, in: Proceedings of ICLR

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:49.383973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:2b8f937fc99f423b9876c85c76c87aad9d63ef179b499e8c43da4569b7508fba

Observation 44ae0fd6-1df5-4da5-99f2-d10fdf2153b6 · outbound

This paper cites Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech, in: Pro- ceedings of ICML, pp.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech, in: Pro- ceedings of ICML, pp

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:49.474514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:453dd106e14963d643bbf10f4c1a315307c699bd98ad210794e9b80113b40f7d

Observation 2daecc1d-f41f-4b82-973e-41c02b5f2660 · outbound

This paper cites faster-whisper: Faster Whisper transcription with CTranslate2.https://github.com/SYSTRAN/faster-whisper.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts faster-whisper: Faster Whisper transcription with CTranslate2.https://github.com/SYSTRAN/faster-whisper

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:49.381658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:fb05f587e6b1b7fb3770722083bb6da3a38ae55ef511cc77a32648eed5f18067

Observation b98508a2-01cb-49c5-bd1f-3374212a27c6 · outbound

This paper cites Incorporating end-to-end speech recognition models for sentiment analysis, in: Proceedings of ICRA, pp.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts Incorporating end-to-end speech recognition models for sentiment analysis, in: Proceedings of ICRA, pp

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:48.562121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:eb07ce9060b40feb6088d38ad51596f29f57e092148c9269c432f021f993eb99

Observation 7a028ca4-f932-4d96-92e0-96691cde939e · outbound

This paper cites Unimodal-driven distillation in multimodal emotion recognition with dynamic fusion, in: Proceedings of ICME, pp.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts Unimodal-driven distillation in multimodal emotion recognition with dynamic fusion, in: Proceedings of ICME, pp

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:49.410591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:cc07c56a318ad19610d2ebb959f9022f37bf3ebbdb8342f062fafa90be70a51e

Observation 5bd5a4c2-13a7-4ecd-ac3f-a003d5b00c58 · outbound

This paper cites Speech Emotion Recognition With ASR Transcripts: a Comprehensive Study on Word Error Rate and Fusion Techniques, in: Proceedings of SLT, pp.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts Speech Emotion Recognition With ASR Transcripts: a Comprehensive Study on Word Error Rate and Fusion Techniques, in: Proceedings of SLT, pp

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:49.261230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:6d0e69b134a99af450f1635c8a851381e8d9137f198aa8b05437b942d43fe751

Observation 8ac18c4f-f752-4748-87f7-d42d163a878e · outbound

This paper cites Decoupled multimodal distilling for emotion recognition, in: Proceedings of CVPR, pp.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts Decoupled multimodal distilling for emotion recognition, in: Proceedings of CVPR, pp

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:49.322770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:b04ce12d719fb7adca4bbf95e8662f0b1830335f58f24c103b7b1521865409ce

Observation 3aea66e4-3850-4ebb-9bcd-e5852175771b · outbound

This paper cites A survey of deep learning-based multimodal emotion recognition: Speech, text, and face.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts A survey of deep learning-based multimodal emotion recognition: Speech, text, and face

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:49.507364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:38ee4d3686413d7bb03a7f6f078d3fd0a0e466671020b10a8b6a044952800686

Observation 2ebb9d24-e0ba-4c28-9175-b8ad767de9a0 · outbound

This paper cites Development of interactive English e-learning video entertainment teaching environment based on virtual reality and game teaching emotion analysis.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts Development of interactive English e-learning video entertainment teaching environment based on virtual reality and game teaching emotion analysis

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:49.435195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:6b78f657311c752dfdf5c237d726ad62df1395411c2733bdf263d6a143279166

Observation 470f6a7e-e1b9-4267-a186-e558dc4a4df9 · outbound

This paper cites Unifying distillation and privileged information, in: Proceedings of ICLR.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts Unifying distillation and privileged information, in: Proceedings of ICLR

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:48.727337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:1e52243a52e0dffcc698cd6849eede8961e425aacef1f16094b15a14348ad0e0

Observation ffdce282-5df0-4fcb-bfb4-1c1f1d883b29 · outbound

This paper cites ScaleVLAD: Improving Multimodal Sentiment Analysis via Multi-Scale Fusion of Locally Descriptors.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts ScaleVLAD: Improving Multimodal Sentiment Analysis via Multi-Scale Fusion of Locally Descriptors

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-07-11T01:57:52.113290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:c1402d08277a4c85bdcd38cd7ffc562e6e87679178e7a878f1c46bf5e9de840c

Observation becdda28-330f-4f03-83d9-c406ffb7824f · outbound

This paper cites Audio sentiment analysis by heterogeneous signal features learned from utterance-based parallel neural network, in: Proceedings of AffCon@AAAI, pp.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts Audio sentiment analysis by heterogeneous signal features learned from utterance-based parallel neural network, in: Proceedings of AffCon@AAAI, pp

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:49.096140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:f55ce888d71a1293c568c86739b5d77a25f4f05304f92aef651141fd996c0d47

Observation ffef5c22-058d-4bc6-b934-a8361559a09d · outbound

This paper cites CamemBERT: a tasty French language model, in: Proceedings of ACL, pp.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts CamemBERT: a tasty French language model, in: Proceedings of ACL, pp

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:49.356234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:a05f556e50b1fdcfdf6d731434e396cfd0be1ce99232088b59d20bec32ed657e

Observation b74604ae-6a73-4ab8-b296-85bad7aa1bed · outbound

This paper cites Learning using generated privileged information by text-to-image diffusion models, in: Proceedings of ICPR, pp.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts Learning using generated privileged information by text-to-image diffusion models, in: Proceedings of ICPR, pp

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:49.328815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:88c92459594829fb8117ca15a257543b629d660456e32cb7f9c9f66f85322840

Observation 9d77120e-2752-4f43-a185-3bdf66a764c7 · outbound

This paper cites Verbal sentiment analysis and detection using recurrent neural network, in: Advanced Data Mining Tools and Methods for Social Computing.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts Verbal sentiment analysis and detection using recurrent neural network, in: Advanced Data Mining Tools and Methods for Social Computing

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:48.882398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:e8a09510e48a75a84f88770dc5fbe0087224d12df463b7a321b11b915fc0efda

Observation 5c34b6f5-2a3b-4f4d-af1c-0af8fd4e45ba · outbound

This paper cites Bridging Modalities: Knowledge Distillation and Masked Training for Translating Multi-Modal Emotion Recognition to Uni-Modal, Speech-Only Emotion Recognition.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts Bridging Modalities: Knowledge Distillation and Masked Training for Translating Multi-Modal Emotion Recognition to Uni-Modal, Speech-Only Emotion Recognition

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-07-11T01:57:52.031561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:4208ea872c2dd1542ca4204c51b63058f28f0d4879b7fa0d8043162c7caa1960

Observation bfaf5578-4fc6-4d6f-8a12-8647fb491780 · outbound

This paper cites an unresolved cited work.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-07-11T02:07:49.174874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:f0083098034d7e83ca85195f1c7727c4ff1c55b7df5e3ab8f5ff337601d6d433

Observation 7fcf8860-9b95-48d0-8dcb-d715a2158b6e · outbound

This paper cites 4668–4672.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts 4668–4672

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:48.433904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:b96b941770d8b7a7446574b06ce36449a8ac6d266d0b91c70e502c9f697f1cff

Observation 8adb56cc-3895-4df3-8298-bf67eef50be9 · outbound

This paper cites No Language Left Behind: Scaling Human-Centered Machine Translation.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts No Language Left Behind: Scaling Human-Centered Machine Translation

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-07-11T01:57:52.159386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:b27a049565c50a588bd0b0f3e4b00edf72f5846b535209bc386c40154d24b966

Observation e3810431-9f63-4532-bbbb-7601d50ccc77 · outbound

This paper cites RoBERTuito: a pre-trained language model for social media text in Spanish, in: Proceedings of LREC, pp.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts RoBERTuito: a pre-trained language model for social media text in Spanish, in: Proceedings of LREC, pp

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:49.034169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:35b57af1a0cd0365fe5651db9608e749d8fd3a093de349bdda5d1abdecc2be3b

Observation f4e197aa-3d68-4436-b2cc-caa9b2cea59c · outbound

This paper cites MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in Conversations, in: Proceedings of ACL, pp.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in Conversations, in: Proceedings of ACL, pp

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:49.133028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:97049f40de725d9227bf9cc65fec824dc2075ad1f379038efd0724e5135502ab

Observation 31401229-8535-41aa-abe5-79ab315dbe2a · outbound

This paper cites Scaling speech technology to 1,000+languages.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts Scaling speech technology to 1,000+languages

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:49.201643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:06217f63e1226a3355ad44aac77e92db8a45e628f3b7d1e159e08a702cc05d8a

Observation 2b27edf8-5f60-414a-839d-285db828943c · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision, in: Proceedings of ICML, pp.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts Robust Speech Recognition via Large-Scale Weak Supervision, in: Proceedings of ICML, pp

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:49.443273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:bfb6d1fe071ccf75cdc13f9f643bdc8e489b906356dff912ca8748a7ec280003

Observation 7687e83f-2f1a-478f-8084-fbcba3c8d3b2 · outbound

This paper cites Cascaded cross-modal transformer for audio-textual classification.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts Cascaded cross-modal transformer for audio-textual classification

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:49.696578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:28512e30e191172b1b373cd2a5509a9178d4e58f2c145d39228fa0f0ed3757b3

Observation 17fe5ee7-83eb-411a-814d-bb49c472a88e · outbound

This paper cites Cascaded cross-modal transformer for request and complaint detection, in: Proceedings of ACMMM, pp.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts Cascaded cross-modal transformer for request and complaint detection, in: Proceedings of ACMMM, pp

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:48.792233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:61700f25846eb63479bb222c07d068375bf98b407aaae6944ed95a879cfe212a

Observation f3cdb596-e0bf-4499-b398-13ab84f207f4 · outbound

This paper cites SepTr: Separable Transformer for Audio Spectrogram Processing, in: Proceedings of INTER- SPEECH, pp.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts SepTr: Separable Transformer for Audio Spectrogram Processing, in: Proceedings of INTER- SPEECH, pp

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:49.758610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:d9dc5d930e8d526600c53aaadc26fa288fcf7dfc6f66c54e35130e8f5f65cc6c

Observation 050246b2-a0f1-436a-8777-1513cb80f3a5 · outbound

This paper cites An integrated approach for mental health assessment using emotion analysis and scales.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts An integrated approach for mental health assessment using emotion analysis and scales

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:48.665945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:15df14f24f2b86b97356362aed6aedeef0fc0f00e0fc18790383e63e109e2996

Observation 13152bb5-235f-490d-a9b3-8ff72df03477 · outbound

This paper cites Leveraging pre-trained language model for speech sentiment analysis, in: Proceedings of INTERSPEECH, pp.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts Leveraging pre-trained language model for speech sentiment analysis, in: Proceedings of INTERSPEECH, pp

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:49.276886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:fb9b89685df92e457d2c4a5ac016c58ed89a7784ead46e1a7b658eaea6bb8340

Observation 9a154783-b248-45ab-9fcc-c64df45744a8 · outbound

This paper cites A comparative study on Bengali speech sentiment analysis based on audio data, in: Proceedings of BigComp, pp.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts A comparative study on Bengali speech sentiment analysis based on audio data, in: Proceedings of BigComp, pp

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:49.142487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:87ef3c26c87a153356fa45ec73aedd034c6b6ec70beba066c8a59e79f4189e8c

Observation 3ad7cd62-36fd-4bd7-8197-18dbb4cbf17d · outbound

This paper cites Multimodal transformer for unaligned multimodal language sequences, in: Proceedings of ACL, pp.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts Multimodal transformer for unaligned multimodal language sequences, in: Proceedings of ACL, pp

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:49.106555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:09bca59a6a8d1983afc8811d0d1085f43ed508e4ee77a6e65fc1d7e7e82d25b6

Observation d4331dbd-6328-40ab-8fa6-2c0ad0ab1bc1 · outbound

This paper cites Hierarchical cross-modal attention and dual audio pathways for enhanced multimodal sentiment analysis.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts Hierarchical cross-modal attention and dual audio pathways for enhanced multimodal sentiment analysis

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:48.912555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:7c9ff42dd1ac48ac0a3b36678e55bfe5602507c43935996fe59d26b6f092028b

Observation d41049fb-4d5a-4183-8be6-3675fee23ca4 · outbound

This paper cites Learning using privileged information: Similarity control and knowledge transfer.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts Learning using privileged information: Similarity control and knowledge transfer

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:49.245218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:8e2354d524f47efd5c89b49281821db5e04b245dc0247b1c73b3f417ef138efc

Observation 2b030a15-2b87-45b2-8590-d35b6fbd9979 · outbound

This paper cites A Fine-tuned Wav2vec 2.0/HuBERT Benchmark For Speech Emotion Recognition, Speaker Verification and Spoken Language Understanding.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts A Fine-tuned Wav2vec 2.0/HuBERT Benchmark For Speech Emotion Recognition, Speaker Verification and Spoken Language Understanding

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-07-11T01:57:52.061776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:fd5c795e83baaab743868823f373a461be840083896aff88bc44fd8775a3a61a

Observation 55ee1f27-4db3-4afe-ac07-226c743311bf · outbound

This paper cites Enriching multimodal sentiment analysis through textual emotional descriptions of visual-audio content, in: Proceedings of AAAI, pp.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts Enriching multimodal sentiment analysis through textual emotional descriptions of visual-audio content, in: Proceedings of AAAI, pp

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:49.664775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:cae298940a4869f8c2c1ec76ba427291745ff08ebe1cf1f0f11a90178c6667a6

Observation 53ca2bb2-41f4-4adb-9b3f-f91edeec52d4 · outbound

This paper cites A Self-Adjusting Fusion Representation Learning Model for Unaligned Text-Audio Sequences.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts A Self-Adjusting Fusion Representation Learning Model for Unaligned Text-Audio Sequences

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-07-11T01:57:52.092558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:762c362907d873540685edb0e82970924ae2ae3de46128f34bfbd04d340ccc2a

Observation d1bbf595-ef6c-43fa-a12b-4ec23e26d35b · outbound

This paper cites Multimodal speech emotion recognition using audio and text, in: Proceedings of SLT, pp.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts Multimodal speech emotion recognition using audio and text, in: Proceedings of SLT, pp

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:49.302756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:2c0d987bb4cc26304c7ea5b1bce71581745e36df5e94d3b653cb2418200a1c9e

Observation 7c0dee6f-1b7b-4c3b-971c-47b38a7c9102 · outbound

This paper cites Personality-aware multimodal driver emotion recognition towards intelligent connected vehicles.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts Personality-aware multimodal driver emotion recognition towards intelligent connected vehicles

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T02:07:49.632327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:1b521ce08dde37f2082fc4fe9c6b78843a1b590a148f87728ead7c99d088fb0f

Pith citing papers

No inbound Pith citation observations are available.