Pith. sign in

Paper Citation Record · LEDGER

Speaker-Aware Temporal Aggregation Strategies on Segment Representations for Depression Detection in Dyadic Interaction: A Benchmark Study

As of 8 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 0 inbound Pith citation observations for arXiv:2607.02904.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.02904 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-12T06:14:42.321767Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

29 of 29 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation de41dc07-afb8-48ad-9bf8-2c0f5c33ba3e · outbound

This paper cites Depressive disorder (depression),.

Speaker-Aware Temporal Aggregation Strategies on Segment Representations for Depression Detection in Dyadic Interaction: A Benchmark Study Depressive disorder (depression),

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-12T06:14:42.321767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:14:42.321767Z digest=sha256:ef42e3375cb7b71249fc245cd2bd8f8765300fa233ee3c795c0de02c7556219e

Observation 81be1974-b3f1-4ddc-9123-4449b0720bbd · outbound

This paper cites A review of depression and suicide risk assessment using speech analysis,.

Speaker-Aware Temporal Aggregation Strategies on Segment Representations for Depression Detection in Dyadic Interaction: A Benchmark Study A review of depression and suicide risk assessment using speech analysis,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-12T06:14:42.321767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:14:42.321767Z digest=sha256:dc153b6670bf483767da7fc5cfe635c501bd51f3f38a68a23867ec2202b6fb5c

Observation 3e73c4da-4309-4495-a98a-5509ad875ca2 · outbound

This paper cites Automated assessment of psychiatric disorders using speech: A systematic review,.

Speaker-Aware Temporal Aggregation Strategies on Segment Representations for Depression Detection in Dyadic Interaction: A Benchmark Study Automated assessment of psychiatric disorders using speech: A systematic review,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-12T06:14:42.321767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:14:42.321767Z digest=sha256:a7f07f9416830dc1966201121517a46c68d1c3b80af970373f05a33c00f10c94

Observation 73b5a47f-5e83-4829-a16a-a55e1aca617b · outbound

This paper cites WavLM: Large-scale self-supervised pre-training for full stack speech processing,.

Speaker-Aware Temporal Aggregation Strategies on Segment Representations for Depression Detection in Dyadic Interaction: A Benchmark Study WavLM: Large-scale self-supervised pre-training for full stack speech processing,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-12T06:14:42.321767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:14:42.321767Z digest=sha256:4fe799ee41776cdea12192d8da2adf5e9d6b0a9faf2bee969b608bb910f081b3

Observation 8511b00e-59bb-4ae5-a6e3-df843c6306be · outbound

This paper cites The distress analysis interview corpus of human and computer interviews,.

Speaker-Aware Temporal Aggregation Strategies on Segment Representations for Depression Detection in Dyadic Interaction: A Benchmark Study The distress analysis interview corpus of human and computer interviews,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-12T06:14:42.321767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:14:42.321767Z digest=sha256:f2d0d5660bdf77fac1fc80843fbeb04df107908405150ba78cbbe7d3392bbda4

Observation 4f1af0eb-1202-49b0-9160-a424fcea8c81 · outbound

This paper cites MODMA dataset: a Multi-modal Open Dataset for Mental-disorder Analysis.

Speaker-Aware Temporal Aggregation Strategies on Segment Representations for Depression Detection in Dyadic Interaction: A Benchmark Study MODMA dataset: a Multi-modal Open Dataset for Mental-disorder Analysis

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-12T06:14:42.321767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:14:42.321767Z digest=sha256:b0a98ad39233024daa7570be12765ff58c348ae39c7f2ba2ada622fa2858ead0

Observation 28de5576-2bb0-49ca-9446-7276fb253d47 · outbound

This paper cites HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units.

Speaker-Aware Temporal Aggregation Strategies on Segment Representations for Depression Detection in Dyadic Interaction: A Benchmark Study HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-12T06:14:42.321767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:14:42.321767Z digest=sha256:09a6e3413495aeed000dd58b01c13ea35eb6f0210c7492b89a022a8520a939d6

Observation 65531f6e-da62-45c0-b5b8-2a255405c98c · outbound

This paper cites Robust wav2vec 2.0: Analyzing domain shift in self-supervised pre-training,.

Speaker-Aware Temporal Aggregation Strategies on Segment Representations for Depression Detection in Dyadic Interaction: A Benchmark Study Robust wav2vec 2.0: Analyzing domain shift in self-supervised pre-training,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-12T06:14:42.321767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:14:42.321767Z digest=sha256:96c16b47216c7a7ecf5c6fd3eac4853efa241bb420e499e86b2b77b97145ea1e

Observation 905556f7-761f-4022-846c-18c036a4ba48 · outbound

This paper cites data2vec: A general framework for self-supervised learning in speech, vision and language,.

Speaker-Aware Temporal Aggregation Strategies on Segment Representations for Depression Detection in Dyadic Interaction: A Benchmark Study data2vec: A general framework for self-supervised learning in speech, vision and language,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-12T06:14:42.321767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:14:42.321767Z digest=sha256:46b6a3e7c0f03945b831ce490bfa157278cf9895dd3da19b4cdd2468db7379e6

Observation 5a231095-8343-43cd-9b28-7b84916812e4 · outbound

This paper cites XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale.

Speaker-Aware Temporal Aggregation Strategies on Segment Representations for Depression Detection in Dyadic Interaction: A Benchmark Study XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-12T06:14:42.321767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:14:42.321767Z digest=sha256:37511ec332d2a6f1b4892c1ba5f97dc4f518b6a9e28169d087d0866b57f90f6d

Observation d3e2531e-8dc2-4a71-9c8f-b5a86be9239b · outbound

This paper cites Inves- tigation of layer-wise speech representations in self-supervised learning models: A cross-lingual study in detecting depression,.

Speaker-Aware Temporal Aggregation Strategies on Segment Representations for Depression Detection in Dyadic Interaction: A Benchmark Study Inves- tigation of layer-wise speech representations in self-supervised learning models: A cross-lingual study in detecting depression,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-12T06:14:42.321767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:14:42.321767Z digest=sha256:116124e4cdd8d47e66e4a032e68125c4f7ea2b83c8613fb8197abb377d92b232

Observation 216bf54f-d6ec-4ee1-8a76-8e3e26cc3b3b · outbound

This paper cites SUPERB: Speech processing universal performance benchmark,.

Speaker-Aware Temporal Aggregation Strategies on Segment Representations for Depression Detection in Dyadic Interaction: A Benchmark Study SUPERB: Speech processing universal performance benchmark,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-12T06:14:42.321767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:14:42.321767Z digest=sha256:473daac4517bbf7a79cc2f1403ac1b88d045a387a6dc56894c7a962448de12d3

Observation 2499a599-837b-437c-b8b3-c8f878df8995 · outbound

This paper cites EMO-SUPERB: An In-depth Look at Speech Emotion Recognition.

Speaker-Aware Temporal Aggregation Strategies on Segment Representations for Depression Detection in Dyadic Interaction: A Benchmark Study EMO-SUPERB: An In-depth Look at Speech Emotion Recognition

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-12T06:14:42.321767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:14:42.321767Z digest=sha256:8c74d1fc31daed37b868ddd5777af3de7e4a05b1d00a0baedfe9115a5b9d8109

Observation be9de19d-e7db-4d1b-b876-96b7518b54f8 · outbound

This paper cites Cross-Lingual Speech Emotion Recognition: Humans vs. Self-Supervised Models.

Speaker-Aware Temporal Aggregation Strategies on Segment Representations for Depression Detection in Dyadic Interaction: A Benchmark Study Cross-Lingual Speech Emotion Recognition: Humans vs. Self-Supervised Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-12T06:14:42.321767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:14:42.321767Z digest=sha256:bf54cd070aee2ce7ddcfb6e24627a117e09f76df87526214e612a672a8b8fd66

Observation fac8b026-cda0-476c-815a-b3e279d781a1 · outbound

This paper cites Association between acoustic speech features and non-severe levels of anxiety and depression symptoms across lifespan,.

Speaker-Aware Temporal Aggregation Strategies on Segment Representations for Depression Detection in Dyadic Interaction: A Benchmark Study Association between acoustic speech features and non-severe levels of anxiety and depression symptoms across lifespan,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-12T06:14:42.321767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:14:42.321767Z digest=sha256:4efd2bbac9a4f70faa7e59425e2c74e865edbf8d066efdba48608a6aa0194293

Observation b803d68e-052b-43a7-8497-a8c47c319494 · outbound

This paper cites Detecting depression with audio/text sequence modeling of interviews,.

Speaker-Aware Temporal Aggregation Strategies on Segment Representations for Depression Detection in Dyadic Interaction: A Benchmark Study Detecting depression with audio/text sequence modeling of interviews,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-12T06:14:42.321767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:14:42.321767Z digest=sha256:7a53626b2ab691290f79fc5b5186bd7381129c03371712ba0e314476df533945

Observation aac68f63-a811-498c-a970-1d96ab5b18c9 · outbound

This paper cites DEPA: Self-supervised audio embedding for depression detection,.

Speaker-Aware Temporal Aggregation Strategies on Segment Representations for Depression Detection in Dyadic Interaction: A Benchmark Study DEPA: Self-supervised audio embedding for depression detection,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-12T06:14:42.321767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:14:42.321767Z digest=sha256:39d6e37edb196f32e41ffff317aaf041ec0b99a88d3b8d4b7497272162de9413

Observation 1c122ce8-1fa7-490d-8cb0-98cddfc424c3 · outbound

This paper cites Emotion Recognition from Speech Using Wav2vec 2.0 Embeddings.

Speaker-Aware Temporal Aggregation Strategies on Segment Representations for Depression Detection in Dyadic Interaction: A Benchmark Study Emotion Recognition from Speech Using Wav2vec 2.0 Embeddings

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-12T06:14:42.321767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:14:42.321767Z digest=sha256:a2daf481945501eecfbda3bb0d50b59e51098c0507a07ea9199270cd46ab7664

Observation 3e669324-d417-452c-8453-f04e5dc5c7f2 · outbound

This paper cites Tracking depression-related mental state using multimodal analysis,.

Speaker-Aware Temporal Aggregation Strategies on Segment Representations for Depression Detection in Dyadic Interaction: A Benchmark Study Tracking depression-related mental state using multimodal analysis,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-12T06:14:42.321767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:14:42.321767Z digest=sha256:068d950a4c3ca193c3652dfe81c6adf7d58bc8f4148a3156abea8d9b6bdb9195

Observation 5839823f-72e5-4472-a530-a8c012379279 · outbound

This paper cites A VEC 2019 Workshop and Challenge: State-of-Mind, Detecting Depression with AI, and Cross-Cultural Affect Recognition,.

Speaker-Aware Temporal Aggregation Strategies on Segment Representations for Depression Detection in Dyadic Interaction: A Benchmark Study A VEC 2019 Workshop and Challenge: State-of-Mind, Detecting Depression with AI, and Cross-Cultural Affect Recognition,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-12T06:14:42.321767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:14:42.321767Z digest=sha256:32cc24a24fb949c06942e3ac0fd59876e194dc6641a497cd8407b0402740a5ea

Observation 5d493f89-8b4c-4140-9310-868da5f53d7e · outbound

This paper cites Speaker normalization for speech-based depression detection,.

Speaker-Aware Temporal Aggregation Strategies on Segment Representations for Depression Detection in Dyadic Interaction: A Benchmark Study Speaker normalization for speech-based depression detection,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-12T06:14:42.321767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:14:42.321767Z digest=sha256:463d04385a41f73c92c12ca000b8ba4efbd3346e287494908dc485505472f081

Observation 48c12994-19b5-4d04-9eb9-a327c8075211 · outbound

This paper cites DepAudioNet: An efficient deep model for audio based depression classification,.

Speaker-Aware Temporal Aggregation Strategies on Segment Representations for Depression Detection in Dyadic Interaction: A Benchmark Study DepAudioNet: An efficient deep model for audio based depression classification,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-12T06:14:42.321767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:14:42.321767Z digest=sha256:417d2de354296aa38a57cc3cc11a8907522082353101fdd0e31a7c8b4648dec9

Observation 9821d2d4-1c28-47d2-ad7b-327fde51af75 · outbound

This paper cites Automatic speech emotion recognition using recurrent neural networks with local attention,.

Speaker-Aware Temporal Aggregation Strategies on Segment Representations for Depression Detection in Dyadic Interaction: A Benchmark Study Automatic speech emotion recognition using recurrent neural networks with local attention,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-12T06:14:42.321767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:14:42.321767Z digest=sha256:3ef951e30d98d33df9e3adb1b3e8f1ef9c2d491443309462b05adbf9b8654282

Observation d65f4b48-aaf3-4486-bf76-0bd27993ef42 · outbound

This paper cites NetVLAD: CNN architecture for weakly supervised place recognition,.

Speaker-Aware Temporal Aggregation Strategies on Segment Representations for Depression Detection in Dyadic Interaction: A Benchmark Study NetVLAD: CNN architecture for weakly supervised place recognition,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-12T06:14:42.321767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:14:42.321767Z digest=sha256:71f08f14fccf704895d59254a1b3a449084b41d6e2217c7f13136e7841deec0d

Observation b984236a-e221-4729-b69d-ddc92bdf11f7 · outbound

This paper cites Attentive statistics pooling for deep speaker embedding,.

Speaker-Aware Temporal Aggregation Strategies on Segment Representations for Depression Detection in Dyadic Interaction: A Benchmark Study Attentive statistics pooling for deep speaker embedding,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-12T06:14:42.321767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:14:42.321767Z digest=sha256:8b3432267b215b4a1952a53bec0b0e53ace37cc3bdc9be581bd767760dbac29e

Observation 451d3020-6fba-43ce-8160-632385d0d0c3 · outbound

This paper cites BERT: Pre-training of deep bidirectional transformers for language understanding,.

Speaker-Aware Temporal Aggregation Strategies on Segment Representations for Depression Detection in Dyadic Interaction: A Benchmark Study BERT: Pre-training of deep bidirectional transformers for language understanding,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-12T06:14:42.321767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:14:42.321767Z digest=sha256:c79ed8cb438cc3544644b81fcb8e9599a5ab59e070f827e76fe706895695f169

Observation 97bfd86e-6960-4b42-9e0e-37743235eece · outbound

This paper cites The INTERSPEECH 2021 computational paralinguistics challenge,.

Speaker-Aware Temporal Aggregation Strategies on Segment Representations for Depression Detection in Dyadic Interaction: A Benchmark Study The INTERSPEECH 2021 computational paralinguistics challenge,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-12T06:14:42.321767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:14:42.321767Z digest=sha256:c07a91918a26cdf2b8474518a2ee33bbb86692c90aba02a6e5cf10913331e2a8

Observation 5da048d9-2d2d-46be-8479-e5e80e35bf29 · outbound

This paper cites Prevalence of neural collapse during the terminal phase of deep learning training,.

Speaker-Aware Temporal Aggregation Strategies on Segment Representations for Depression Detection in Dyadic Interaction: A Benchmark Study Prevalence of neural collapse during the terminal phase of deep learning training,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-12T06:14:42.321767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:14:42.321767Z digest=sha256:01f805c07578cfb40649a885d66d997b6f4f4e289108df26d3a66b34df5fab68

Observation 81f31ee3-a202-4568-81d2-ee1d0e7560b2 · outbound

This paper cites Scikit-learn: Machine learning in Python,.

Speaker-Aware Temporal Aggregation Strategies on Segment Representations for Depression Detection in Dyadic Interaction: A Benchmark Study Scikit-learn: Machine learning in Python,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-12T06:14:42.321767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:14:42.321767Z digest=sha256:4a9a97a2bee034d94584641b98fe213a587570246a73a1b82d04537a9add841d

Pith citing papers

No inbound Pith citation observations are available.