Pith. sign in

Paper Citation Record · LEDGER

Speaker Embedding Informed Audiovisual Active Speaker Detection for Egocentric Recordings

As of 10 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 0 inbound Pith citation observations for arXiv:2502.06012.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.06012 v1

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T17:09:18.680337Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

28 of 28 outbound references displayed

  • verified exact1
  • verified fuzzy12
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c9307a56-36dc-43b4-add9-feb4284afeb0 · outbound

This paper cites Is Someone Speaking? Exploring Long-term Temporal Features for Audio-visual Active Speaker Detection,.

Speaker Embedding Informed Audiovisual Active Speaker Detection for Egocentric Recordings Is Someone Speaking? Exploring Long-term Temporal Features for Audio-visual Active Speaker Detection,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:09:19.091720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T17:09:18.540723Z digest=sha256:de449a18218042241bb1e901bb30cc4a7ba54713deb16536e06b2416be9637a9

Observation 47ddadcb-d087-4eb1-86d6-06434784b766 · outbound

This paper cites Combining Residual Networks with LSTMs for Lipreading,.

Speaker Embedding Informed Audiovisual Active Speaker Detection for Egocentric Recordings Combining Residual Networks with LSTMs for Lipreading,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:09:19.074488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T17:09:18.547355Z digest=sha256:5682bb197973b83f6c167959b95381cf2bd5249cd7c523a0688acf9710e65865

Observation 4b4a39e7-a127-4a86-8281-b656e5e37680 · outbound

This paper cites Active Speakers in Context,.

Speaker Embedding Informed Audiovisual Active Speaker Detection for Egocentric Recordings Active Speakers in Context,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:09:19.056810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T17:09:18.552428Z digest=sha256:7857b6f961d926dda02e2fd7ba374827842ab3c3a26f57bc17a01028900c65c5

Observation c1136810-d3a3-402f-819d-de2d8f289c8e · outbound

This paper cites ASD-Transformer: Efficient Active Speaker Detection Using Self And Multimodal Transformers,.

Speaker Embedding Informed Audiovisual Active Speaker Detection for Egocentric Recordings ASD-Transformer: Efficient Active Speaker Detection Using Self And Multimodal Transformers,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T17:09:18.557553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:09:18.557553Z digest=sha256:4ebcfe6b924800fa0eff6d941bed6c7071f57cd4bbb39b62b700b6bafd635c98

Observation 68a840b7-088d-4d2b-8965-d29d9c87d898 · outbound

This paper cites Hello! My name is... Buffy.

Speaker Embedding Informed Audiovisual Active Speaker Detection for Egocentric Recordings Hello! My name is... Buffy

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T17:09:18.564019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:09:18.564019Z digest=sha256:f01a20fffeafa83a0305d2e82c943a3dd027239dcba3c0fb5cc60ee2c37a75cb

Observation 09db0d23-3d42-46be-82b3-4652a159bc36 · outbound

This paper cites MAAS: Multi-modal Assignation for Active Speaker Detection,.

Speaker Embedding Informed Audiovisual Active Speaker Detection for Egocentric Recordings MAAS: Multi-modal Assignation for Active Speaker Detection,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T17:09:18.569695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:09:18.569695Z digest=sha256:df191d37806b8224d8cc893cab3f1a7814cf444eced27718d965bbfd1e0b06ce

Observation d87e82d8-2b51-441b-bd18-8c5834c0465f · outbound

This paper cites Target Active Speaker Detection with Audio-visual Cues,.

Speaker Embedding Informed Audiovisual Active Speaker Detection for Egocentric Recordings Target Active Speaker Detection with Audio-visual Cues,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T17:09:18.576552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:09:18.576552Z digest=sha256:d4d4b39a5b3862005be043b1b2947537e14556f197e2df7b80692733b0f22530

Observation 6f99e09c-4b40-4109-b829-75bf046371a9 · outbound

This paper cites Ava Active Speaker: An Audio-Visual Dataset for Active Speaker De- tection,.

Speaker Embedding Informed Audiovisual Active Speaker Detection for Egocentric Recordings Ava Active Speaker: An Audio-Visual Dataset for Active Speaker De- tection,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T17:09:18.581381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:09:18.581381Z digest=sha256:b056afa42759b9f818973027db17dcf87909806ea8275156c59add1448ff9a23

Observation ac4f4dfd-872a-45dc-891f-bf0e2010181a · outbound

This paper cites Improving Audiovisual Active Speaker Detection in Egocentric Recordings with the Data-Efficient Image Transformer,.

Speaker Embedding Informed Audiovisual Active Speaker Detection for Egocentric Recordings Improving Audiovisual Active Speaker Detection in Egocentric Recordings with the Data-Efficient Image Transformer,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:09:18.988225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T17:09:18.585590Z digest=sha256:4162a3565bd9a3b511c63d9c479c2213b33a11b32676022caf1d9b3ce794ee3c

Observation 21ef0601-6f5b-49da-a5f2-e9ad0d680dfc · outbound

This paper cites Ego4D: Around the World in 3,000 Hours of Egocentric Video,.

Speaker Embedding Informed Audiovisual Active Speaker Detection for Egocentric Recordings Ego4D: Around the World in 3,000 Hours of Egocentric Video,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T17:09:18.590424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:09:18.590424Z digest=sha256:959c0235cb2a2264b53f1298be6d606fbcf66f79efdcb5bc648cf8f00543c981

Observation 36daf742-d0ff-45dd-9b53-f0257df901ad · outbound

This paper cites End-to-End Active Speaker Detection, author=Juan Leon Alcazar and Moritz Cordes and Chen Zhao and Bernard Ghanem,.

Speaker Embedding Informed Audiovisual Active Speaker Detection for Egocentric Recordings End-to-End Active Speaker Detection, author=Juan Leon Alcazar and Moritz Cordes and Chen Zhao and Bernard Ghanem,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:09:18.962200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T17:09:18.594992Z digest=sha256:a7fedb17436f6e2e2e9773bbf55679fd617984b4987a6694749316c8424ce930

Observation 04974023-5009-4843-8fbd-f24c2d8d0f4f · outbound

This paper cites How to Design a Three-Stage Architecture for Audio-Visual Active Speaker Detection in the Wild,.

Speaker Embedding Informed Audiovisual Active Speaker Detection for Egocentric Recordings How to Design a Three-Stage Architecture for Audio-Visual Active Speaker Detection in the Wild,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T17:09:18.600233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:09:18.600233Z digest=sha256:497968602f2853a55bdf17c009dd69c51d0d265fb6e9935c74f8385d0af4167d

Observation eb15f594-52b5-4eb8-a1a7-cfaff91a8646 · outbound

This paper cites A Light Weight Model for Active Speaker Detection,.

Speaker Embedding Informed Audiovisual Active Speaker Detection for Egocentric Recordings A Light Weight Model for Active Speaker Detection,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T17:09:18.605275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:09:18.605275Z digest=sha256:b03caf9681577c2f7968462b7ced52eaf9b02f7b25de531f531ea0823b851bcb

Observation 672bc6f7-3ec5-4460-969d-ccaff54edcc4 · outbound

This paper cites LoCoNet: Long-Short Context Network for Active Speaker Detection,.

Speaker Embedding Informed Audiovisual Active Speaker Detection for Egocentric Recordings LoCoNet: Long-Short Context Network for Active Speaker Detection,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T17:09:18.609862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:09:18.609862Z digest=sha256:9fd533f3c42fcae306c436d99a88fec955c32d14ba7f7ba72ac6e34ea385dd99

Observation 1ba33c77-cfdf-4790-afb3-92d924c42c6a · outbound

This paper cites Learning Long-Term Spatial-Temporal Graphs for Active Speaker Detection,.

Speaker Embedding Informed Audiovisual Active Speaker Detection for Egocentric Recordings Learning Long-Term Spatial-Temporal Graphs for Active Speaker Detection,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T17:09:18.615111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:09:18.615111Z digest=sha256:25229bd5d0846db85ffccbd2a644859724c5c9c2b309087ce3e987725f6276f7

Observation 7985d78f-7386-4bc4-bfe9-64a21de0d797 · outbound

This paper cites ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification,.

Speaker Embedding Informed Audiovisual Active Speaker Detection for Egocentric Recordings ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T17:09:18.619710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:09:18.619710Z digest=sha256:424fb07ebb13f4858c33abd6148ee2237a0d2af283c91ca57863247a94ade70f

Observation a9c9d624-29e9-44ae-9f97-53500061f0d5 · outbound

This paper cites Personal VAD 2.0: Optimizing Personal Voice Activity Detection for On-Device Speech Recognition.

Speaker Embedding Informed Audiovisual Active Speaker Detection for Egocentric Recordings Personal VAD 2.0: Optimizing Personal Voice Activity Detection for On-Device Speech Recognition

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-08T17:09:18.766401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T17:09:18.624973Z digest=sha256:1511acb4e61e9110af5e5c22b03def876e9478df953a29574e97dacd24caa1b4

Observation 7eda9495-fd1b-4522-bbbf-c9191089d911 · outbound

This paper cites Efficient Personal V oice Activity Detection with Wake Word Reference Speech,.

Speaker Embedding Informed Audiovisual Active Speaker Detection for Egocentric Recordings Efficient Personal V oice Activity Detection with Wake Word Reference Speech,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:09:18.904791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T17:09:18.631171Z digest=sha256:c2ac9b12da92931ab053a075ca85e58207d76686520283b18ba7b74f1c4416e3

Observation 9fe9b316-a7de-4906-a1bf-4b9883193f88 · outbound

This paper cites V oxCeleb: A Large-Scale Speaker Identification Dataset,.

Speaker Embedding Informed Audiovisual Active Speaker Detection for Egocentric Recordings V oxCeleb: A Large-Scale Speaker Identification Dataset,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T17:09:18.636126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:09:18.636126Z digest=sha256:ef9d19adf08bea12b85b07997b8737890d28a8a86e7f1b1025b0e9a80720d422

Observation b8792db4-8364-4ab7-9d03-35e7e6946e63 · outbound

This paper cites Ring Loss: Convex Feature Normalization for Face Recognition,.

Speaker Embedding Informed Audiovisual Active Speaker Detection for Egocentric Recordings Ring Loss: Convex Feature Normalization for Face Recognition,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:09:18.887481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T17:09:18.641060Z digest=sha256:4c5454433e93593d67d6d62110295dfcea78200bb7658522341e89212e5cfe84

Observation 1c34fe6a-207f-4beb-a411-a14f2c7ce986 · outbound

This paper cites ArcFace: Additive Angular Margin Loss for Deep Face Recognition,.

Speaker Embedding Informed Audiovisual Active Speaker Detection for Egocentric Recordings ArcFace: Additive Angular Margin Loss for Deep Face Recognition,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:09:18.870149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T17:09:18.645922Z digest=sha256:133f5f404a89597ae7a85e1aa6370e2b3d24ecaa5ca400720209618d5bd8c350

Observation 786bcc82-f5bf-4bc3-acaa-a9ff2a5cfed2 · outbound

This paper cites SphereFace: Deep Hypersphere Embedding for Face Recognition,.

Speaker Embedding Informed Audiovisual Active Speaker Detection for Egocentric Recordings SphereFace: Deep Hypersphere Embedding for Face Recognition,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:09:18.853270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T17:09:18.650052Z digest=sha256:a04e7139476401c324493b60877bf822faf96d47290a3f1937b2e1ad1abed22d

Observation e72199d4-0acd-4d42-85b9-72cdffc84cec · outbound

This paper cites CosFace: Large Margin Cosine Loss for Deep Face Recognition,.

Speaker Embedding Informed Audiovisual Active Speaker Detection for Egocentric Recordings CosFace: Large Margin Cosine Loss for Deep Face Recognition,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:09:18.835514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T17:09:18.655038Z digest=sha256:27416698ed7cfbe774916f3af4b8e3faf10d5d87cea95950b8cca2399ba58075

Observation 599cbf32-1274-4f73-a393-10805583039d · outbound

This paper cites Partial FC: Training 10 Million Identities on a Single Ma- chine,.

Speaker Embedding Informed Audiovisual Active Speaker Detection for Egocentric Recordings Partial FC: Training 10 Million Identities on a Single Ma- chine,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:09:18.819850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T17:09:18.660952Z digest=sha256:0be4f511d3dec25415e028720350542c11e12f2508215eab53af6e92ccfbaa36

Observation 72c51122-fd47-4229-8770-65d107bced6b · outbound

This paper cites Attention is All you Need,.

Speaker Embedding Informed Audiovisual Active Speaker Detection for Egocentric Recordings Attention is All you Need,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:09:18.803865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T17:09:18.665695Z digest=sha256:53ccc56352f5f890711dda721286a9d22299f44a6266249fb2fc209b0822c3a6

Observation 497d1872-2ede-461d-b85a-85a0398a7510 · outbound

This paper cites Robust Object Recogni- tion Through Symbiotic Deep Learning In Mobile Robots,.

Speaker Embedding Informed Audiovisual Active Speaker Detection for Egocentric Recordings Robust Object Recogni- tion Through Symbiotic Deep Learning In Mobile Robots,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T17:09:18.670905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:09:18.670905Z digest=sha256:7a3297a4728521c0dce66315b1debd02b159883020f96b2e63ced0d9aefeb2e8

Observation a683ce62-4341-4bca-9986-5edeae446089 · outbound

This paper cites The PASCAL Visual Object Classes Challenge 2012 (VOC2012) Results,.

Speaker Embedding Informed Audiovisual Active Speaker Detection for Egocentric Recordings The PASCAL Visual Object Classes Challenge 2012 (VOC2012) Results,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T17:09:18.675378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:09:18.675378Z digest=sha256:8b161ae65276f39f9dc50c55f14c80f3ed39fc40f1d1019e28c909d9b1f0a45a

Observation 23a06e36-1d2c-49dd-80aa-aab1d90798dc · outbound

This paper cites Technical Report for Ego4D Long Term Action Anticipation Challenge 2023.

Speaker Embedding Informed Audiovisual Active Speaker Detection for Egocentric Recordings Technical Report for Ego4D Long Term Action Anticipation Challenge 2023

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T17:09:18.680337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:09:18.680337Z digest=sha256:9231760234a68f51d2bcf0cf995a947109558147c1e75b82afe97b57a35dc117

Pith citing papers

No inbound Pith citation observations are available.