Pith. sign in

Paper Citation Record · LEDGER

On Aligning Hierarchical Standardized Embedding for Audio-visual Generalized Zero-shot Learning

As of 21 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 0 inbound Pith citation observations for arXiv:2606.11602.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.11602 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-27T10:39:37.920953Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

29 of 29 outbound references displayed

  • verified exact3
  • verified fuzzy0
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 407f5517-65c5-49dc-9f2c-0ecbacd8af5f · outbound

This paper cites Look, listen, and attend: Co-attention network for self-supervised audio-visual representa- tion learning,.

On Aligning Hierarchical Standardized Embedding for Audio-visual Generalized Zero-shot Learning Look, listen, and attend: Co-attention network for self-supervised audio-visual representa- tion learning,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-27T10:39:37.920953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:39:37.920953Z digest=sha256:53e3177483c9a8772e4ff5b4212f5aece697eedaa8fe32b2c08b02116339d2fb

Observation 427b5d01-357a-46bc-9155-2cea6392bb8c · outbound

This paper cites Self- supervised object detection from audio-visual correspondence,.

On Aligning Hierarchical Standardized Embedding for Audio-visual Generalized Zero-shot Learning Self- supervised object detection from audio-visual correspondence,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-27T10:39:37.920953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:39:37.920953Z digest=sha256:557067fb00ffb8495118c8667ff4ff220d37e7a68fb6d1bffbabcb5b9ca069f2

Observation 7705b846-6f68-461e-8303-2d18d069dbd6 · outbound

This paper cites Detection of audio-video synchronization errors via event detection,.

On Aligning Hierarchical Standardized Embedding for Audio-visual Generalized Zero-shot Learning Detection of audio-video synchronization errors via event detection,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-27T10:39:37.920953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:39:37.920953Z digest=sha256:0f07142596bac7e799eef994a427d49fc3fe74c3859e10c0522c8339fbf4c2e6

Observation 8a8885ba-b003-423e-a10c-cc72a3821ff8 · outbound

This paper cites Audiovisual SlowFast Networks for Video Recognition.

On Aligning Hierarchical Standardized Embedding for Audio-visual Generalized Zero-shot Learning Audiovisual SlowFast Networks for Video Recognition

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-03T08:57:47.661964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T10:39:37.920953Z digest=sha256:7887e2886bf620ed27d1e5577c9b02ae6b6da6f75fd71ee3fa70bb574ca5ed2a

Observation d06aa5fc-5c46-4c6a-818b-64f8fad912a4 · outbound

This paper cites Text-to-feature diffusion for audio-visual few-shot learning,.

On Aligning Hierarchical Standardized Embedding for Audio-visual Generalized Zero-shot Learning Text-to-feature diffusion for audio-visual few-shot learning,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-27T10:39:37.920953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:39:37.920953Z digest=sha256:cae42a93642760baafbb2233976b04cdd594c03aae03f6f392f9053ce6754694

Observation 54404212-e78c-4650-b4f9-5933a0f61669 · outbound

This paper cites Advancing weakly- supervised audio-visual video parsing via segment-wise pseudo label- ing,.

On Aligning Hierarchical Standardized Embedding for Audio-visual Generalized Zero-shot Learning Advancing weakly- supervised audio-visual video parsing via segment-wise pseudo label- ing,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-27T10:39:37.920953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:39:37.920953Z digest=sha256:defe4b49219136d3d863dd7088f6b68f792e60da7e594dd62087c52a081b851b

Observation 5153b275-6677-40c3-a6cb-48ccbd299d6a · outbound

This paper cites Aloha: Adapting local spatio-temporal context to enhance the audio-visual semantic segmentation,.

On Aligning Hierarchical Standardized Embedding for Audio-visual Generalized Zero-shot Learning Aloha: Adapting local spatio-temporal context to enhance the audio-visual semantic segmentation,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-27T10:39:37.920953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:39:37.920953Z digest=sha256:0174fe8e6ba0ea111cd4568420f56c6175e3c50e88b5caea23392f30ccd62199

Observation 7d3c7e45-ce3d-4c2c-938d-a7f0aed9b531 · outbound

This paper cites Audio-visual event localization with cross co-attention and dynamic audio-object semantic alignment,.

On Aligning Hierarchical Standardized Embedding for Audio-visual Generalized Zero-shot Learning Audio-visual event localization with cross co-attention and dynamic audio-object semantic alignment,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-27T10:39:37.920953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:39:37.920953Z digest=sha256:b50630b6f7a6fa3631fba2a1de8017d36e6231074a2555e3dfabef5ca53523b8

Observation 2959b878-b23b-4c2f-8b06-ca6442f5a6cf · outbound

This paper cites X-sta: Cross-modal spatial-temporal alignment network for unified audio-visual segmenta- tion,.

On Aligning Hierarchical Standardized Embedding for Audio-visual Generalized Zero-shot Learning X-sta: Cross-modal spatial-temporal alignment network for unified audio-visual segmenta- tion,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-27T10:39:37.920953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:39:37.920953Z digest=sha256:d8481ff9646f88ede2575d6c8c516a7572be58870b90e54bd97d00b1d9e6ffdf

Observation fe56af1e-d90c-4690-b982-ded29839d47e · outbound

This paper cites Coordinated joint multimodal embeddings for generalized audio-visual zero-shot classifi- cation and retrieval of videos,.

On Aligning Hierarchical Standardized Embedding for Audio-visual Generalized Zero-shot Learning Coordinated joint multimodal embeddings for generalized audio-visual zero-shot classifi- cation and retrieval of videos,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-27T10:39:37.920953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:39:37.920953Z digest=sha256:4e6a7b5c25b8c40f340ea97d19fd06079dd5b420c6262347edf64b0bdbf7d016

Observation 8b59b25f-814a-4d8c-8653-d6283f5eacd5 · outbound

This paper cites Avgzslnet: Audio-visual generalized zero-shot learning by reconstructing label fea- tures from multi-modal embeddings,.

On Aligning Hierarchical Standardized Embedding for Audio-visual Generalized Zero-shot Learning Avgzslnet: Audio-visual generalized zero-shot learning by reconstructing label fea- tures from multi-modal embeddings,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-27T10:39:37.920953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:39:37.920953Z digest=sha256:7401a8dae4cdbec27b92870c40a249aac1b2c356f9eb11fcbdee14545d5953d5

Observation 16a33141-9b62-4e3b-92f1-47baa334d382 · outbound

This paper cites Audio-visual generalised zero-shot learning with cross-modal attention and language,.

On Aligning Hierarchical Standardized Embedding for Audio-visual Generalized Zero-shot Learning Audio-visual generalised zero-shot learning with cross-modal attention and language,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-27T10:39:37.920953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:39:37.920953Z digest=sha256:a4d35b2ff54859f05c55ab14bdabebc7042bbae233e6571840a8669d9c83e3fc

Observation 5a0dd993-46ba-4cd9-9439-0e3d68b41b26 · outbound

This paper cites Temporal and cross-modal attention for audio-visual zero-shot learning,.

On Aligning Hierarchical Standardized Embedding for Audio-visual Generalized Zero-shot Learning Temporal and cross-modal attention for audio-visual zero-shot learning,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-27T10:39:37.920953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:39:37.920953Z digest=sha256:15350eeda0856d8c29cace237f7063a2f00534018c20708df7cb7ef9188c9ba3

Observation 1d9a12c0-d7d0-48e7-8569-fd45e1995514 · outbound

This paper cites Hyperbolic audio-visual zero-shot learning,.

On Aligning Hierarchical Standardized Embedding for Audio-visual Generalized Zero-shot Learning Hyperbolic audio-visual zero-shot learning,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-27T10:39:37.920953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:39:37.920953Z digest=sha256:875a840affb6cf5191f78cb15e2a90e645d35cc31057746f38084bdf8c4c7e7b

Observation 0fd3846c-9034-4dba-b5f0-500a7bb01383 · outbound

This paper cites Motion- decoupled spiking transformer for audio-visual zero-shot learning,.

On Aligning Hierarchical Standardized Embedding for Audio-visual Generalized Zero-shot Learning Motion- decoupled spiking transformer for audio-visual zero-shot learning,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-27T10:39:37.920953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:39:37.920953Z digest=sha256:a033fbf46cde215cca5ce3c175091bf0683df655418406380edf2a5502b8a878

Observation 3da2b60a-7be1-4598-bb26-53eee3b2f6c1 · outbound

This paper cites A generative approach to audio- visual generalized zero-shot learning: Combining contrastive and dis- criminative techniques,.

On Aligning Hierarchical Standardized Embedding for Audio-visual Generalized Zero-shot Learning A generative approach to audio- visual generalized zero-shot learning: Combining contrastive and dis- criminative techniques,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-27T10:39:37.920953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:39:37.920953Z digest=sha256:5e25b02c5e7a8d79a752382cc7fbc7fa8d58289f17ab3d8bbd73b03ffaddb162

Observation 70b2dc0a-9254-43b3-89da-0204dc39aa1d · outbound

This paper cites Spiking tucker fusion trans- former for audio-visual zero-shot learning,.

On Aligning Hierarchical Standardized Embedding for Audio-visual Generalized Zero-shot Learning Spiking tucker fusion trans- former for audio-visual zero-shot learning,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-27T10:39:37.920953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:39:37.920953Z digest=sha256:a874dfbcfa938da38e071631dab911f846cc7170c788db84856acb21e20100cc

Observation f343b011-37cc-45a4-b789-b76f70a541bc · outbound

This paper cites Audio- visual generalized zero-shot learning using pre-trained large multi-modal models,.

On Aligning Hierarchical Standardized Embedding for Audio-visual Generalized Zero-shot Learning Audio- visual generalized zero-shot learning using pre-trained large multi-modal models,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-27T10:39:37.920953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:39:37.920953Z digest=sha256:358039f57ad0031986df882f139e06602e488cdc9867b9129bdd763948cab17a

Observation 3b6ef146-2b3f-4e39-8d5c-9265421c2bb5 · outbound

This paper cites Audio-visual generalized zero-shot learning the easy way,.

On Aligning Hierarchical Standardized Embedding for Audio-visual Generalized Zero-shot Learning Audio-visual generalized zero-shot learning the easy way,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-27T10:39:37.920953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:39:37.920953Z digest=sha256:6bee3be4d6fc7bba46f13dd2ff0b1de71d3d8197858ac6ea73c6b622dd9d5f92

Observation 58a3d039-e239-4f10-8228-e408eb613dc7 · outbound

This paper cites Discrepancy-aware attention network for enhanced audio-visual generalized zero-shot learn- ing,.

On Aligning Hierarchical Standardized Embedding for Audio-visual Generalized Zero-shot Learning Discrepancy-aware attention network for enhanced audio-visual generalized zero-shot learn- ing,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-27T10:39:37.920953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:39:37.920953Z digest=sha256:2e7d5c59690c4050ccd74dd40e9b086f9579d8dd618a738b8e260011669e7d68

Observation b1f66517-45b1-4e3a-b415-0024d349a36d · outbound

This paper cites Fusion-regularized alignment modality-adaptive audio-visual network for audio-visual zero- shot learning,.

On Aligning Hierarchical Standardized Embedding for Audio-visual Generalized Zero-shot Learning Fusion-regularized alignment modality-adaptive audio-visual network for audio-visual zero- shot learning,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-06-27T10:39:37.920953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:39:37.920953Z digest=sha256:5d0bf3b6de558592b00c26776b4dee9a4dbfe7fa03b5bf636509e94fa8376e0b

Observation f50ba71e-18b2-47c4-8a6b-5c24760609e0 · outbound

This paper cites Z-score normalization, hubness, and few-shot learning,.

On Aligning Hierarchical Standardized Embedding for Audio-visual Generalized Zero-shot Learning Z-score normalization, hubness, and few-shot learning,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-27T10:39:37.920953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:39:37.920953Z digest=sha256:a2af7f200f3a76dcdc359137cff3c5e19bf8e820242c51781b3df95ea748dffd

Observation b1d40317-1968-4647-b5f6-d577c1b754b8 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

On Aligning Hierarchical Standardized Embedding for Audio-visual Generalized Zero-shot Learning Learning transferable visual models from natural language supervision,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-06-27T10:39:37.920953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:39:37.920953Z digest=sha256:8ca2b2a2432385b20d26600a86abf94be8ea8e85362e0f030c94c2428d91fce0

Observation 05d828a6-f434-4c1c-8347-61637631926e · outbound

This paper cites Wavcaps: A chatgpt-assisted weakly-labelled audio captioning dataset for audio-language multimodal research,.

On Aligning Hierarchical Standardized Embedding for Audio-visual Generalized Zero-shot Learning Wavcaps: A chatgpt-assisted weakly-labelled audio captioning dataset for audio-language multimodal research,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-27T10:39:37.920953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:39:37.920953Z digest=sha256:73a29ba752d4d7eb7bef7d600637f05cb586b8aca3527933c3b86fd854fc60fa

Observation 4d8135ae-73c5-4da3-84e0-ad35475fea95 · outbound

This paper cites Tackling uncertain correspondences for multi-modal entity alignment,.

On Aligning Hierarchical Standardized Embedding for Audio-visual Generalized Zero-shot Learning Tackling uncertain correspondences for multi-modal entity alignment,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-27T10:39:37.920953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:39:37.920953Z digest=sha256:25ff0e48877eee457505b3202713044d8e9f894de3a2540ddcdec303c4e10e6a

Observation 22b77b5c-c76d-4c97-85a9-c2d1d97125c5 · outbound

This paper cites Unialign: Scaling multimodal alignment within one unified model,.

On Aligning Hierarchical Standardized Embedding for Audio-visual Generalized Zero-shot Learning Unialign: Scaling multimodal alignment within one unified model,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-06-27T10:39:37.920953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:39:37.920953Z digest=sha256:d27e268c744abafd81d5729a7cc42951b293f9e14fe0c133e4b30614224b934a

Observation a6312b14-7efb-4631-9251-229ff764654a · outbound

This paper cites Relational knowledge distilla- tion,.

On Aligning Hierarchical Standardized Embedding for Audio-visual Generalized Zero-shot Learning Relational knowledge distilla- tion,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-06-27T10:39:37.920953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:39:37.920953Z digest=sha256:d79435d6d31ef52270a161a51030417a3cc6f4c05aa6904563b3b4bd009742d7

Observation f54605b0-cedf-401e-ba73-e45c41f2ca84 · outbound

This paper cites On Distilling the Displacement Knowledge for Few-Shot Class-Incremental Learning.

On Aligning Hierarchical Standardized Embedding for Audio-visual Generalized Zero-shot Learning On Distilling the Displacement Knowledge for Few-Shot Class-Incremental Learning

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-03T08:57:47.664025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T10:39:37.920953Z digest=sha256:bb390d5a552646acd84737017d75369ddd5892b4abf7d18ee65ce2a9bed05b7f

Observation 8ad004f7-4a2a-40d2-beff-f16e962c6e0c · outbound

This paper cites Extremely Simple Out-of-distribution Detection for Audio-visual Generalized Zero-shot Learning.

On Aligning Hierarchical Standardized Embedding for Audio-visual Generalized Zero-shot Learning Extremely Simple Out-of-distribution Detection for Audio-visual Generalized Zero-shot Learning

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-03T08:57:47.661468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T10:39:37.920953Z digest=sha256:2fc6f2b3efb3d85f34f12153361976c86a0d0cb5b48fd7c64cb5dab2495cc9e1

Pith citing papers

No inbound Pith citation observations are available.