Pith. sign in

Paper Citation Record · LEDGER

Are Mamba-based Audio Foundation Models the Best Fit for Non-Verbal Emotion Recognition?

As of 13 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 0 inbound Pith citation observations for arXiv:2506.02258.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.02258 v1

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:29:54.084398Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

23 of 23 outbound references displayed

  • verified exact0
  • verified fuzzy15
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fc71a14e-b9b5-4e86-b62e-1eccbc0621ee · outbound

This paper cites Speech emotion recognition based on hmm and svm,.

Are Mamba-based Audio Foundation Models the Best Fit for Non-Verbal Emotion Recognition? Speech emotion recognition based on hmm and svm,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:55.216794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T11:29:52.141074Z digest=sha256:7bf3965c37eca6f1172b2fd61846fa6e54fc09694379126e4d13b6d9f2a17aaa

Observation c9b6b090-75e2-4ef8-97e9-15441a19ba65 · outbound

This paper cites Emotion recognition in speech using mfcc and wavelet features,.

Are Mamba-based Audio Foundation Models the Best Fit for Non-Verbal Emotion Recognition? Emotion recognition in speech using mfcc and wavelet features,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:55.203011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T11:29:52.180343Z digest=sha256:639e9e71967d0440272ef82975101ee25a0bd931797705ed7b2d9c093a64f262

Observation 1b912b5f-c264-4fae-b6e5-828260b43db7 · outbound

This paper cites Speech emotion recognition based on feature selection and extreme learning machine decision tree,.

Are Mamba-based Audio Foundation Models the Best Fit for Non-Verbal Emotion Recognition? Speech emotion recognition based on feature selection and extreme learning machine decision tree,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:55.190003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T11:29:52.236085Z digest=sha256:3aac22f6d330121c28b4730db4a801406c4cb1c7f5220ba7b74ae66082580a86

Observation 05cb934a-48bf-4892-b403-97ca842b7208 · outbound

This paper cites Speech emotion recognition with dual-sequence lstm architecture,.

Are Mamba-based Audio Foundation Models the Best Fit for Non-Verbal Emotion Recognition? Speech emotion recognition with dual-sequence lstm architecture,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:55.173025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T11:29:52.283010Z digest=sha256:eefe0357fec06a8557d028adcf709d65d123eb2e27901c492c9dd13df26df508

Observation 1f062ee6-db11-4b59-8db9-c7498806e1c2 · outbound

This paper cites Convolution neural network based automatic speech emotion recognition using mel-frequency cepstrum coefficients,.

Are Mamba-based Audio Foundation Models the Best Fit for Non-Verbal Emotion Recognition? Convolution neural network based automatic speech emotion recognition using mel-frequency cepstrum coefficients,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:55.158322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T11:29:52.344930Z digest=sha256:31ea74cc3db78b29ce657381cb5517e4af1d3bc5fe44e7e7b4de6376c60896c6

Observation 8b4d47a9-4339-41fe-9960-235e3d22a499 · outbound

This paper cites Ctnet: Conversational transformer network for emotion recognition,.

Are Mamba-based Audio Foundation Models the Best Fit for Non-Verbal Emotion Recognition? Ctnet: Conversational transformer network for emotion recognition,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:55.146277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T11:29:52.398994Z digest=sha256:42c8f0ff8497fb104177c0626bc62cabd011695b92eac939ebadedb76329a6b0

Observation f15749ac-0c24-496d-aae4-a5f7766a748f · outbound

This paper cites Emotion Recognition from Speech Using Wav2vec 2.0 Embeddings.

Are Mamba-based Audio Foundation Models the Best Fit for Non-Verbal Emotion Recognition? Emotion Recognition from Speech Using Wav2vec 2.0 Embeddings

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:52.459509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:29:52.459509Z digest=sha256:841edeb19fb7f1b9f7ed2e2136b029a0a697411892bc3657ceb9f9adf6bbe9e5

Observation 1228f4ed-326d-4082-b7cf-063425fb9bbd · outbound

This paper cites Transforming the Embeddings: A Lightweight Technique for Speech Emotion Recognition Tasks.

Are Mamba-based Audio Foundation Models the Best Fit for Non-Verbal Emotion Recognition? Transforming the Embeddings: A Lightweight Technique for Speech Emotion Recognition Tasks

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:52.540223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:29:52.540223Z digest=sha256:96e1b4d810dadcd669c517a06a44736480374ffc0475d07d34c1caa916ee5c3a

Observation c2055c74-ac05-47b6-b767-cc3b7bec6d69 · outbound

This paper cites Adapting wavlm for speech emotion recognition,.

Are Mamba-based Audio Foundation Models the Best Fit for Non-Verbal Emotion Recognition? Adapting wavlm for speech emotion recognition,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:55.135147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T11:29:52.604361Z digest=sha256:ca357f42b9e68e945bc9d837d0ad21ad72f1f4c4a562fd0c7f06bec7594de578

Observation dcf272e4-ab17-47b4-bd5d-66874f4c8801 · outbound

This paper cites Audio mamba: Selective state spaces for self-supervised audio representations,.

Are Mamba-based Audio Foundation Models the Best Fit for Non-Verbal Emotion Recognition? Audio mamba: Selective state spaces for self-supervised audio representations,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:55.104324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T11:29:52.691205Z digest=sha256:d628be29988bf3e90fe0c3a1b646f47a687987873b2958b1966a1b665afb8729

Observation a954a949-eb11-4c71-91e0-39eab3549663 · outbound

This paper cites Speech emotion recognition considering nonverbal vocalization in affective conver- sations,.

Are Mamba-based Audio Foundation Models the Best Fit for Non-Verbal Emotion Recognition? Speech emotion recognition considering nonverbal vocalization in affective conver- sations,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:54.977226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T11:29:52.700086Z digest=sha256:81233d446413092d7b10224b9f54f75a889cc427f238c89b1bd273302fb05f7b

Observation 3d6631ce-bc9b-4570-868c-032edd60a33a · outbound

This paper cites Jvnv: A corpus of japanese emotional speech with verbal content and nonverbal expressions,.

Are Mamba-based Audio Foundation Models the Best Fit for Non-Verbal Emotion Recognition? Jvnv: A corpus of japanese emotional speech with verbal content and nonverbal expressions,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:52.780585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:29:52.780585Z digest=sha256:ff03627913bdf4c364358ddce5f7ac76cee2df130e9bb403a0c2cc08ac914dbe

Observation 16b4b1d8-020d-4e38-8091-32195789b2fc · outbound

This paper cites Large-scale nonverbal vocalization detection using transformers,.

Are Mamba-based Audio Foundation Models the Best Fit for Non-Verbal Emotion Recognition? Large-scale nonverbal vocalization detection using transformers,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:54.797553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T11:29:52.917984Z digest=sha256:a0c796e526081c3de4e1e7209386dbcbbfff1f6f1b1473dd5b0451746600678c

Observation f3947231-5db5-4fc1-a000-36366dc4d0f7 · outbound

This paper cites Investigation of ensemble of self-supervised models for speech emotion recognition,.

Are Mamba-based Audio Foundation Models the Best Fit for Non-Verbal Emotion Recognition? Investigation of ensemble of self-supervised models for speech emotion recognition,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:54.641425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T11:29:53.110719Z digest=sha256:c3f0aed510b7d718644fb1c30149d67d293140c202947fa3d7bb993f50ab70f1

Observation ca4873d6-0c0d-4e6b-a43a-e65f13184559 · outbound

This paper cites Heterogeneity over homogeneity: Investigating multilingual speech pre-trained models for detecting audio deepfake,.

Are Mamba-based Audio Foundation Models the Best Fit for Non-Verbal Emotion Recognition? Heterogeneity over homogeneity: Investigating multilingual speech pre-trained models for detecting audio deepfake,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:54.588877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T11:29:53.269332Z digest=sha256:0373b897302f42fd1f050f480bb22eb140023281ee5e94df55342dd390d920a6

Observation acfaa700-c80b-48a7-8cf7-14352455a022 · outbound

This paper cites Wavlm: Large-scale self-supervised pre- training for full stack speech processing,.

Are Mamba-based Audio Foundation Models the Best Fit for Non-Verbal Emotion Recognition? Wavlm: Large-scale self-supervised pre- training for full stack speech processing,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:53.455978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:29:53.455978Z digest=sha256:b7c45890dd1a13e836fa06b742c745cf41306c9b77e618b77a53933b27da4bf0

Observation cb1e393d-8591-4161-be9a-ea9244a82828 · outbound

This paper cites Unispeech-sat: Universal speech representation learning with speaker aware pre-training,.

Are Mamba-based Audio Foundation Models the Best Fit for Non-Verbal Emotion Recognition? Unispeech-sat: Universal speech representation learning with speaker aware pre-training,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:54.455841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T11:29:53.619048Z digest=sha256:463a6c86fa313c539c13af037b4eaa1f6b3621fd314717c92f774325a2d117f2

Observation f5cf2178-ceb3-4c9e-9b77-b447c2c7174a · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech representations,.

Are Mamba-based Audio Foundation Models the Best Fit for Non-Verbal Emotion Recognition? wav2vec 2.0: A framework for self-supervised learning of speech representations,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:53.770157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:29:53.770157Z digest=sha256:b295508b2ceb85ad1b67954cf2a74ca5f375ecd9e6dfa7bf076315b8f9aeef26

Observation a3d5bfc7-3246-4ca6-9c49-d8b336f159fc · outbound

This paper cites Hubert: Self-supervised speech representation learning by masked prediction of hidden units,.

Are Mamba-based Audio Foundation Models the Best Fit for Non-Verbal Emotion Recognition? Hubert: Self-supervised speech representation learning by masked prediction of hidden units,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:53.913426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:29:53.913426Z digest=sha256:a7bad4bfef910a9e872abc4945ea139494a572beb87315cbb8c0ab81573d6512

Observation 36458d9a-1854-4b5c-8ab7-553d023f5c4e · outbound

This paper cites R ´enyi divergence and kullback-leibler divergence,.

Are Mamba-based Audio Foundation Models the Best Fit for Non-Verbal Emotion Recognition? R ´enyi divergence and kullback-leibler divergence,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:54.057247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:29:54.057247Z digest=sha256:542f67e0d5ae405efb2f993a36211b1b38486fc05c07e6f454dbb6983228636d

Observation 5ff0ecd1-5f3b-40b7-9a2d-3b9b7d77c9a7 · outbound

This paper cites Asvp-esd: A dataset and its benchmark for emotion recognition using both speech and non-speech utterances,.

Are Mamba-based Audio Foundation Models the Best Fit for Non-Verbal Emotion Recognition? Asvp-esd: A dataset and its benchmark for emotion recognition using both speech and non-speech utterances,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:54.077393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:29:54.077393Z digest=sha256:b7102ef330fb54a4d0b78ce0929c9d72f1d672efc4bf75818d2232739446a0cb

Observation 82c93a67-9c89-4e6a-995f-fb5e3a48f7f6 · outbound

This paper cites Jnv corpus: A corpus of japanese nonverbal vocalizations with diverse phrases and emotions,.

Are Mamba-based Audio Foundation Models the Best Fit for Non-Verbal Emotion Recognition? Jnv corpus: A corpus of japanese nonverbal vocalizations with diverse phrases and emotions,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:54.319362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T11:29:54.080965Z digest=sha256:c1224d6131c065a18c814601a028e005a97c85dcf3f90a001bc100b1db69933b

Observation 1b27c255-fff4-428e-a2a8-86c7ea9f4a45 · outbound

This paper cites The variably intense vocalizations of affect and emotion (vivae) corpus prompts new per- spective on nonspeech perception.

Are Mamba-based Audio Foundation Models the Best Fit for Non-Verbal Emotion Recognition? The variably intense vocalizations of affect and emotion (vivae) corpus prompts new per- spective on nonspeech perception

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:54.187393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T11:29:54.084398Z digest=sha256:5b06d350e6fbcdf74b4f304f07decc806cef70a21c16af5304d76b6bb8c4dff4

Pith citing papers

No inbound Pith citation observations are available.