Pith. sign in

Paper Citation Record · LEDGER

CoLMbo: Speaker Language Model for Descriptive Profiling

As of 10 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 2 inbound Pith citation observations for arXiv:2506.09375.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.09375 v2

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:54:16.523708Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T01:12:30.704449Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-19T03:49:09.587857Z

Reference resolution

29 of 29 outbound references displayed

  • verified exact2
  • verified fuzzy19
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1e55df75-f254-4d7d-b201-a4b37b63cfd7 · outbound

This paper cites Singh, Profiling humans from their voice.

CoLMbo: Speaker Language Model for Descriptive Profiling Singh, Profiling humans from their voice

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:54:22.083119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T04:54:12.573959Z digest=sha256:f7005f8f1bb1a243fe71db46f1ea371d437f60c0fc4e70c4e9f57e618b742956

Observation 52e474f0-1649-4bfb-be8c-8b85b5173e33 · outbound

This paper cites SEGAA: A Unified Approach to Predicting Age, Gender, and Emotion in Speech.

CoLMbo: Speaker Language Model for Descriptive Profiling SEGAA: A Unified Approach to Predicting Age, Gender, and Emotion in Speech

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:54:17.065427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T04:54:12.714022Z digest=sha256:e35150d7c8569d4eda49cca064ce6d348b8cb74e4fcdcbd7dc0be3731f43a8a2

Observation 9a7357ea-5862-4f48-b372-b4e84c051c72 · outbound

This paper cites Prediction of age from speech features using a multi-layer perceptron model,.

CoLMbo: Speaker Language Model for Descriptive Profiling Prediction of age from speech features using a multi-layer perceptron model,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:54:21.806777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T04:54:12.829754Z digest=sha256:c447b4345f53c90899dac101969b9bae2a32fbbde8c47fed0f1666b50df2032b

Observation af666c14-a783-4cc7-a266-e9f4a27f04ce · outbound

This paper cites Salmonn: Towards generic hearing abilities for large language models,.

CoLMbo: Speaker Language Model for Descriptive Profiling Salmonn: Towards generic hearing abilities for large language models,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:54:21.586191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T04:54:12.984271Z digest=sha256:27d508225d73f05bfcd0ee80e385ed0d30aecb8a55ca53349d41956f824ae96c

Observation 622b8839-2497-4a35-b351-39a993fdd582 · outbound

This paper cites Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models.

CoLMbo: Speaker Language Model for Descriptive Profiling Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T04:54:13.128574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:54:13.128574Z digest=sha256:9de4097e36422cae60fdf77dc04172924e6468ed527a2c055c94b51b1eefca97

Observation a0b56506-e4b4-47cf-9bc8-9d7a2890f29a · outbound

This paper cites Pengi: An audio language model for audio tasks,.

CoLMbo: Speaker Language Model for Descriptive Profiling Pengi: An audio language model for audio tasks,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T04:54:13.257810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:54:13.257810Z digest=sha256:41e6579d85852a956dfcf9267e0ebf60c470ea6ba06bf1380f4e86a759c795e2

Observation ed050362-11eb-4766-9b85-376437709e6e · outbound

This paper cites Listen, Think, and Understand.

CoLMbo: Speaker Language Model for Descriptive Profiling Listen, Think, and Understand

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T04:54:13.382210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:54:13.382210Z digest=sha256:d306e879083db372f8ec6479144ccb57ca0b05e85462af051b572e38f538d776

Observation 9c967c50-3543-44f4-b64c-13e5197827fb · outbound

This paper cites GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities.

CoLMbo: Speaker Language Model for Descriptive Profiling GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T04:54:13.519625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:54:13.519625Z digest=sha256:3c3630bd2b5f93a1da9089baa13506928e3e7317bdcc0c7cef4be4d44706fe82

Observation 884ca67a-6d6c-43d9-9744-6ac44965e930 · outbound

This paper cites Speaker recognition for multi-speaker conversations using x-vectors,.

CoLMbo: Speaker Language Model for Descriptive Profiling Speaker recognition for multi-speaker conversations using x-vectors,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:54:21.412124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T04:54:13.637259Z digest=sha256:a3c10393e17fa1a603205992e9d0b111825f6a16e9b6ab4daca3837bffc7c406

Observation 714c89ab-903b-47e1-b36d-78d40589ef49 · outbound

This paper cites Front-end factor analysis for speaker verification,.

CoLMbo: Speaker Language Model for Descriptive Profiling Front-end factor analysis for speaker verification,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:54:21.171886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T04:54:13.770908Z digest=sha256:e166f5ff98b1ed151cd2df30b8594af59f56dcda1a99a03d623ee09c97a7027d

Observation dab474bf-85c3-4fa2-9b32-7ab971586295 · outbound

This paper cites Deep neural networks for small foot- print text-dependent speaker verification,.

CoLMbo: Speaker Language Model for Descriptive Profiling Deep neural networks for small foot- print text-dependent speaker verification,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:54:20.906038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T04:54:13.944604Z digest=sha256:737ba92a581a7728f6b1e187e15fc9a6d316dfc3be5e266a66bb5d9139803aff

Observation aa719a45-2e84-40eb-aa1d-c24c9e641d01 · outbound

This paper cites Generalized end-to-end loss for speaker verification,.

CoLMbo: Speaker Language Model for Descriptive Profiling Generalized end-to-end loss for speaker verification,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:54:20.679258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T04:54:14.063835Z digest=sha256:61752946932bf10be47a7375bef480cf9845c257ac3a6c79176bf36c12b9535c

Observation 5fe4dcfe-dcb8-4f34-a86c-d954bbd1ba6b · outbound

This paper cites Front-end factor analysis for speaker verification,.

CoLMbo: Speaker Language Model for Descriptive Profiling Front-end factor analysis for speaker verification,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:54:20.438171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T04:54:14.186302Z digest=sha256:b419c80c078759a291a43017755bbc77acbd8f90913ef69e1838973e8c017782

Observation e84767c0-373c-4828-8add-3905da63afec · outbound

This paper cites Conformer: Convolution-augmented transformer for speech recognition,.

CoLMbo: Speaker Language Model for Descriptive Profiling Conformer: Convolution-augmented transformer for speech recognition,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:54:20.219859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T04:54:14.348254Z digest=sha256:167a81077e7797dfcf3713aa25dc2342f35e633fffbf29d8c8f0399971c5573b

Observation ad2fb05a-1dbb-4e8c-85cb-d3e9d1612847 · outbound

This paper cites Mfa-conformer: Multi-scale feature aggregation conformer for automatic speaker verification,.

CoLMbo: Speaker Language Model for Descriptive Profiling Mfa-conformer: Multi-scale feature aggregation conformer for automatic speaker verification,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:54:19.955951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T04:54:14.495507Z digest=sha256:6a43c49d268a09b039ef27c057b3544462cd184012fb188ed34d201fe1a5ca74

Observation 48d2153f-20df-4719-96d5-5b086db3dc36 · outbound

This paper cites A novel scheme for speaker recognition using a phonetically-aware deep neural network,.

CoLMbo: Speaker Language Model for Descriptive Profiling A novel scheme for speaker recognition using a phonetically-aware deep neural network,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:54:19.718881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T04:54:14.692566Z digest=sha256:d1e532fffd2e0d05288c60bfa24a5ab2354eaa7db6ea444a848eca12cc622eee

Observation 460728b4-b63d-4ecb-997a-3d05888fbab3 · outbound

This paper cites Employing phonetic information in dnn speaker embeddings to improve speaker recognition performance,.

CoLMbo: Speaker Language Model for Descriptive Profiling Employing phonetic information in dnn speaker embeddings to improve speaker recognition performance,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:54:19.409779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T04:54:14.824249Z digest=sha256:f65aa1d730e636bf7b96b6d9c2526098ed8b1be9de38362e4b5a66e22a92c590

Observation b4f9717d-1e1f-4572-bde9-4f699a9d9dab · outbound

This paper cites Speaker Embedding Extraction with Phonetic Information.

CoLMbo: Speaker Language Model for Descriptive Profiling Speaker Embedding Extraction with Phonetic Information

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:54:16.793671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T04:54:15.003832Z digest=sha256:29f38f5d78084c38adcbe2ec0e343e1131cd3f145f8acf78a9f74a65c748dcff

Observation 2b5087c3-98b3-4329-bb21-02284159fb25 · outbound

This paper cites Gender and age estimation methods based on speech using deep neural networks,.

CoLMbo: Speaker Language Model for Descriptive Profiling Gender and age estimation methods based on speech using deep neural networks,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:54:19.123983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T04:54:15.125576Z digest=sha256:81178444fe1d1edb1a16d251117592f4fb720090f725436673ea8c2fa05ca672

Observation ebce4b4a-d3d7-43bd-80a4-ba29b5b2052c · outbound

This paper cites Explainable Attribute-Based Speaker Verification.

CoLMbo: Speaker Language Model for Descriptive Profiling Explainable Attribute-Based Speaker Verification

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T04:54:15.253430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:54:15.253430Z digest=sha256:309ecb10f69fdf6ea6f7d3cba4485e220903ec801ad25f6b211a40abac928172

Observation bce57a1b-a809-4bea-9796-5cb1d0de19fb · outbound

This paper cites Pdaf: A phonetic debiasing attention framework for speaker verification,.

CoLMbo: Speaker Language Model for Descriptive Profiling Pdaf: A phonetic debiasing attention framework for speaker verification,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:54:18.816521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T04:54:15.419608Z digest=sha256:6781611f6700cfdc7b3a68434c17cbc743879217965288c0ef832e286a0a791a

Observation f8c28283-b524-4c6a-bf1d-4adbcaf03595 · outbound

This paper cites ClipCap: CLIP Prefix for Image Captioning.

CoLMbo: Speaker Language Model for Descriptive Profiling ClipCap: CLIP Prefix for Image Captioning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T04:54:15.589604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:54:15.589604Z digest=sha256:3d633b5aff1920b320d6bfb948d9c4143f1509f416483de45ec591dccef2302c

Observation b80c9514-0a64-4ffe-aec1-0c05fd711fa3 · outbound

This paper cites Multimodal few-shot learning with frozen lan- guage models,.

CoLMbo: Speaker Language Model for Descriptive Profiling Multimodal few-shot learning with frozen lan- guage models,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:54:18.489518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T04:54:15.751445Z digest=sha256:617478ba052e5ae0c186f97c43e6e042634e73708e9a15993451899dd5f203e5

Observation 5ea11c10-beb9-4c77-9595-8e602a406a51 · outbound

This paper cites EARS: An anechoic fullband speech dataset benchmarked for speech enhancement and dereverberation,.

CoLMbo: Speaker Language Model for Descriptive Profiling EARS: An anechoic fullband speech dataset benchmarked for speech enhancement and dereverberation,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:54:18.202577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T04:54:15.862316Z digest=sha256:d4d04a9017fc3e5164592b6e4083953e2fc9824e681fcb03a620014977a8a776

Observation 0ccf614f-994e-4089-8054-e69a8d5e55a5 · outbound

This paper cites Darpa timit acoustic-phonetic continuous speech corpus cd-rom. nist speech disc 1-1.1,.

CoLMbo: Speaker Language Model for Descriptive Profiling Darpa timit acoustic-phonetic continuous speech corpus cd-rom. nist speech disc 1-1.1,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:54:17.892657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T04:54:15.969339Z digest=sha256:b9832bf18c1cbcb206ea9fd7363c9dddcc6df8fb02f3cc39347084130e11c96a

Observation 28bb011e-b81e-47b8-a31d-27847cbbbd50 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

CoLMbo: Speaker Language Model for Descriptive Profiling LLaMA: Open and Efficient Foundation Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T04:54:16.137326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:54:16.137326Z digest=sha256:db5596f10705dbb9a4df5b278fc5e6a12b53b13a7d466e07001552745624ae20

Observation 8d60644f-f691-4de9-a987-93d0a1976863 · outbound

This paper cites Audio Entailment: Assessing Deductive Reasoning for Audio Understanding.

CoLMbo: Speaker Language Model for Descriptive Profiling Audio Entailment: Assessing Deductive Reasoning for Audio Understanding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T04:54:16.288107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:54:16.288107Z digest=sha256:4a3d6b17896df7e5afbabb4b88ca4442c2f131c7a8c28f79545844efe67b3153

Observation b94ec1ce-8733-415f-a8b1-afbc1565e7b0 · outbound

This paper cites Adam: A method for stochastic optimization,.

CoLMbo: Speaker Language Model for Descriptive Profiling Adam: A method for stochastic optimization,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:54:17.641108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T04:54:16.429096Z digest=sha256:595983bb3fe3e8a255d0430cf7754323f0d38499275fc7d5faa543b28716c7ef

Observation 3ec8cfe9-c01c-41dd-84db-f67641700744 · outbound

This paper cites Pytorch: An imperative style, high-performance deep learning library,.

CoLMbo: Speaker Language Model for Descriptive Profiling Pytorch: An imperative style, high-performance deep learning library,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:54:17.336576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T04:54:16.523708Z digest=sha256:13047088d2f02d2da4bc1fa91be0ae38c32f819d1f124b713fa9c770a9cd8b73

Pith citing papers

Observation d2701b38-547c-4780-ad59-9e18bda8fe7e · inbound

SpeakerLLM: A Speaker-Specialized Audio-LLM for Speaker Understanding and Verification Reasoning cites this paper.

SpeakerLLM: A Speaker-Specialized Audio-LLM for Speaker Understanding and Verification Reasoning CoLMbo: Speaker Language Model for Descriptive Profiling

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-15T03:24:55.781236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T03:20:37.488507Z digest=sha256:c7900e746f1f59940e2dda345e1611cc45cbd860eaea88ea95d3baffa3414560

Observation 00bc2fc1-edc9-4bd5-87d5-2f6625cd7dfb · inbound

Large Audio Language Models for Spoofing-Aware Speaker Verification cites this paper.

Large Audio Language Models for Spoofing-Aware Speaker Verification CoLMbo: Speaker Language Model for Descriptive Profiling

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-02T01:12:30.704449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:12:30.704449Z digest=sha256:d9f77b55b86afa6ad2c7dab3b043a901a07cef025e79ee745e3911f59f7dce2a