Pith. sign in

Paper Citation Record · LEDGER

Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition

As of 8 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2507.09323.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.09323 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:03:03.579463Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

31 of 31 outbound references displayed

  • verified exact2
  • verified fuzzy22
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 81fd07c6-b149-4500-99ae-b80d39c05c51 · outbound

This paper cites Vatt: Transformers for multimodal self-supervised learning from raw video, audio and text.

Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition Vatt: Transformers for multimodal self-supervised learning from raw video, audio and text

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:03:04.021804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:03:02.905926Z digest=sha256:59874c539d0e02f77a3ff314f8ff33e17294364c8b97be025af9f5c794ed18f6

Observation cefb466b-5949-46af-bcea-ebb65e9331cc · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition Flamingo: a visual language model for few-shot learning

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:03:04.007900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:03:02.983484Z digest=sha256:f418ebffde0bca099566794259c4485a0c38d9fecb9d66db6860c583552d0744

Observation b0797321-987a-43db-877a-5e0851227593 · outbound

This paper cites Multimodal machine learning: A survey and tax- onomy.

Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition Multimodal machine learning: A survey and tax- onomy

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:03:03.994738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:03:03.160526Z digest=sha256:6750df0eb0b36e9bc03816cf69d23a44f4dd76650096765408f5995609f9bb62

Observation 8ee51abf-9a5d-485d-9cd4-b7c4d669e9d0 · outbound

This paper cites Vggsound: A large-scale audio-visual dataset.

Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition Vggsound: A large-scale audio-visual dataset

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:03:03.980373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:03:03.298243Z digest=sha256:633c03f91c66e5cad1908850fbdc123d4d7d6f072b3f7657c53caf270de6b5b3

Observation 7537496e-dbc4-4ce6-b0a8-25fb935740eb · outbound

This paper cites A simple framework for contrastive learning of visual representations.

Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition A simple framework for contrastive learning of visual representations

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:03:03.965880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:03:03.420105Z digest=sha256:dfdc5e638e5712c9bfeb1e3e5e17971e16fc86b04d8a8a0713d0ad374865d045

Observation 87a38546-9c25-4da7-b423-bcd7692e4e77 · outbound

This paper cites Scaling egocentric vision: The epic-kitchens dataset.

Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition Scaling egocentric vision: The epic-kitchens dataset

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:03:03.951067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:03:03.475163Z digest=sha256:087d3e2817e9c4e5438980dcf676828798ff5e8d5bcd4812a1fb5771df2c8ad0

Observation 08276ac1-58f2-4dcf-be21-fe5fceb7038d · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition An image is worth 16x16 words: Transformers for image recognition at scale

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:03:03.936862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:03:03.480607Z digest=sha256:aeb2527ead65cd741b8b7990a731c40270f455f3b924efb85b2ca0911b267e4c

Observation cda91782-4f55-477d-a412-4ffe5229e882 · outbound

This paper cites Large Scale Audiovisual Learning of Sounds with Weakly Labeled Data.

Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition Large Scale Audiovisual Learning of Sounds with Weakly Labeled Data

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:03:03.685469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:03:03.485282Z digest=sha256:bdcaa83702289c37829de7a94480c54ea94f758633ec73a1a908f61f440036d0

Observation db469e9c-eab5-497f-84ec-e9d938026c04 · outbound

This paper cites Audio set: An ontology and human- labeled dataset for audio events.

Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition Audio set: An ontology and human- labeled dataset for audio events

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:03:03.922445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:03:03.490036Z digest=sha256:e31e10d4b6023c5b4c77ca7329d228f9ea514603844322a53be143bc32716f53

Observation 13d93ebc-8506-4e7e-ac6c-8d0f6ba74705 · outbound

This paper cites Audiovisual masked autoencoders.

Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition Audiovisual masked autoencoders

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:03:03.909656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:03:03.494154Z digest=sha256:115278dfed8e1bd3bc6f989e5673b6da66db40b2d2e2122ade22e6c66c6a6ea3

Observation 8eeb24fa-c3e4-4bf2-bce0-6e3ca5a37666 · outbound

This paper cites AST: Audio Spectrogram Transformer.

Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition AST: Audio Spectrogram Transformer

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:03.498534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:03.498534Z digest=sha256:cb799b1ab0ca6c403b967077cf361a81f0d767cab3a5afc9fd6d0b39b90d8049

Observation 75677a28-1a87-4e6a-8b4c-a88b97687b0e · outbound

This paper cites Uavm: Towards unifying audio and visual models.

Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition Uavm: Towards unifying audio and visual models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:03:03.896484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:03:03.503076Z digest=sha256:07c194dccdfaac40ba75b30359cbe20ce0805916afce1de622df2ca89264a061

Observation 8696b57a-3e26-4303-8f96-eccab980b3f1 · outbound

This paper cites Contrastive Audio-Visual Masked Autoencoder.

Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition Contrastive Audio-Visual Masked Autoencoder

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:03.507305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:03.507305Z digest=sha256:417d4ffbf2317e74c9b965fa72a067669e25b70f4eec18ad587688f206a06a2b

Observation 43fe541e-1486-4185-b0fb-1a857cf9e9d8 · outbound

This paper cites Deep residual learning for image recognition.

Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition Deep residual learning for image recognition

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:03.511781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:03.511781Z digest=sha256:da7f48c982bc51cf749481190de239e8d7a00e9a1562b600eca2bd1b98aa673e

Observation 34a4a523-6847-4011-a690-2d8c5743b512 · outbound

This paper cites Hubert: Self-supervised speech representation learning by masked prediction of hidden units.

Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition Hubert: Self-supervised speech representation learning by masked prediction of hidden units

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:03:03.875713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:03:03.515621Z digest=sha256:ee0fc1fed5ed6ac08d024bf36a1b68443508797a610c8a473a238634994ec75b

Observation 66f5017d-ab57-4e8a-9d3d-ba3f78cce6a2 · outbound

This paper cites Perceiver io: A general architecture for structured inputs & outputs.

Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition Perceiver io: A general architecture for structured inputs & outputs

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:03:03.862821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:03:03.519693Z digest=sha256:5cb613997794cd21e386428dda5f99f921c77be47f607f17c3a7ae044041f1cf

Observation 9fd1149d-b653-4743-b5b1-1c9cda6cd8de · outbound

This paper cites Supervised contrastive learning.

Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition Supervised contrastive learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:03.523260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:03.523260Z digest=sha256:08aaa2203eced6d170a001ca594f6e17700f7f97c2dcb9164e3c54f2872848e8

Observation 3dfc1e15-bec4-4adf-9842-a7c36837af65 · outbound

This paper cites Kipf and Max Welling.

Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition Kipf and Max Welling

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:03:03.840950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:03:03.527353Z digest=sha256:a99236f2bac13eb4859b4a1e196599ae11a1626d920e52503638754601f24cfc

Observation 91f8c0f3-0761-4cd8-88b6-c75018c4447e · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition Swin transformer: Hierarchical vision transformer using shifted windows

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:03.531207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:03.531207Z digest=sha256:0206a0dc824a86f2e39f65f731b37eddc288950c02df9d9e1809064411873b6b

Observation 8884ea43-29e0-40e3-97c3-353b84b113d4 · outbound

This paper cites Attention bottlenecks for multimodal fusion.

Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition Attention bottlenecks for multimodal fusion

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:03:03.817741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:03:03.534986Z digest=sha256:50640787e4233cae4615ac8ca43e2cd3e51e2ab3e90c1d3204d69ea7c8aec434

Observation d8e499b7-5b22-433e-a139-d600351e428f · outbound

This paper cites OmniNet: A unified architecture for multi-modal multi-task learning.

Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition OmniNet: A unified architecture for multi-modal multi-task learning

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:03:03.635985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:03:03.538749Z digest=sha256:472e455b332934359db26652dc52a6c2919cc08f0e08eb91f8a119b0c98c283e

Observation cbb55465-42fc-4aad-af02-d8caa9f7e5b2 · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition Robust speech recognition via large-scale weak supervision

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:03:03.804017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:03:03.542837Z digest=sha256:7e4380d5e2a1c9b90c3a7ca7972f89f9eb78cc68cb1e551123cd58e996a13223

Observation 45308ca1-9e25-49cf-8ec7-039260699ba7 · outbound

This paper cites Multimodal fusion for audio-image and video action recognition.

Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition Multimodal fusion for audio-image and video action recognition

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:03:03.789208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:03:03.546642Z digest=sha256:e103a5fda159c984587404e9619ff96a7f7d37679c84333a64b7ce1270cddc34

Observation 93092bbe-2d4a-4d8b-9715-0b47c7270d60 · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:03.550744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:03.550744Z digest=sha256:9fdbc0574efc1dd8e8e4eab7b3c489214f680cccd819737d8ced6f6f30ee5915

Observation 69247b40-9d48-4fe2-82f4-5a69517c00d3 · outbound

This paper cites Omnivec: Learn- ing robust representations with cross modal sharing.

Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition Omnivec: Learn- ing robust representations with cross modal sharing

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:03:03.776085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:03:03.555128Z digest=sha256:a6b1e511baf3262b011b6daf0f97ee3658f314c705c71a00a060df695f5d15f9

Observation 43b5864b-2380-435f-9838-5310b5699aad · outbound

This paper cites One-peace: Ex- ploring one general representation model toward unlimited modalities.

Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition One-peace: Ex- ploring one general representation model toward unlimited modalities

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:03:03.762263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:03:03.559513Z digest=sha256:907bb823464afb42d1e2a2486a243697be4928de527e5db3706f7cc6936450de

Observation 7db7b7c4-35da-46b2-9112-372c202ae259 · outbound

This paper cites What makes train- ing multi-modal classification networks hard? In Proceed- ings of the IEEE/CVF conference on computer vision and pattern recognition, pages 12695–12705, 2020.

Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition What makes train- ing multi-modal classification networks hard? In Proceed- ings of the IEEE/CVF conference on computer vision and pattern recognition, pages 12695–12705, 2020

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:03:03.749426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:03:03.563766Z digest=sha256:824e94a32c77436bd58871a7a9fe1379539c137ae728f43757818cab07e101f6

Observation 1ee6d336-c67d-4573-95c9-601d530a2871 · outbound

This paper cites Multi-stream multi-class fusion of deep net- works for video classification.

Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition Multi-stream multi-class fusion of deep net- works for video classification

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:03:03.736313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:03:03.567838Z digest=sha256:7b0c9a65da69ada206bf5fb16d4957fc6e1916d458d52316b2e7f8dc00994ed4

Observation 430b57eb-fd29-4b5e-a891-bc128515017a · outbound

This paper cites Multimodal learning with transformers: A survey.

Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition Multimodal learning with transformers: A survey

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:03.571745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:03.571745Z digest=sha256:ad1930a41fa8256c4ac5a9c6b95a63ebec962b22ad75bf96bc8d8661bf8b0a0e

Observation 386e133a-4d35-425b-b3ee-13d6a532b2b8 · outbound

This paper cites Peters, and Yejin Choi.

Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition Peters, and Yejin Choi

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:03:03.712763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:03:03.575757Z digest=sha256:8a8044f0e57a85ffb1a9a00a525d8f4d740eb7a4bd8180a6f979bb194cfaae16

Observation 43cce9bc-5253-4dba-a12c-ba2f1337c6fd · outbound

This paper cites Multimodal representation learning: Advances, trends and challenges.

Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition Multimodal representation learning: Advances, trends and challenges

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:03:03.699421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:03:03.579463Z digest=sha256:aa6258e2efea00431c5432838d9b80264ece1773e953001fe522a0e1f106f13f

Pith citing papers

No inbound Pith citation observations are available.