Pith. sign in

Paper Citation Record · LEDGER

Exploring Audio Cues for Enhanced Test-Time Video Model Adaptation

As of 10 August 2026, this Paper Citation Record lists 67 of 67 outbound references and 0 inbound Pith citation observations for arXiv:2506.12481.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.12481 v1

Coverage vector

measured 67 of 67 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:54:10.310454Z

measured 67 of 67 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

67 of 67 outbound references displayed

  • verified exact4
  • verified fuzzy51
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 33ae0cbb-4038-4f82-abc3-4610f115b6e6 · outbound

This paper cites Action-net: Multipath excitation for action recognition,.

Exploring Audio Cues for Enhanced Test-Time Video Model Adaptation Action-net: Multipath excitation for action recognition,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:24.615139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:54:04.567844Z digest=sha256:5fc238bedfafb7d1f8b6c3ad8f9333daf744f66a021ff7a7a3502e2391ba5c32

Observation ee869c44-cb35-42aa-9f14-7e0d3d9e2a1c · outbound

This paper cites Spa- tiotemporal self-attention modeling with temporal patch shift for action recognition,.

Exploring Audio Cues for Enhanced Test-Time Video Model Adaptation Spa- tiotemporal self-attention modeling with temporal patch shift for action recognition,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:24.240040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:54:04.618766Z digest=sha256:69a8f53f50737328153c615370c80fa6316a9921fe3b6d687555bd96549e87e2

Observation 10458ee1-fc22-4035-b16f-b98a2be23da6 · outbound

This paper cites Recurring the transformer for video action recognition,.

Exploring Audio Cues for Enhanced Test-Time Video Model Adaptation Recurring the transformer for video action recognition,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:23.993701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:54:04.717366Z digest=sha256:3bf9d08b9d316422e13db87a1a0268bfcba373252453279e123b36527a5eaacb

Observation 8c3fbca8-55b2-4f3a-ae88-4588751d188b · outbound

This paper cites Graph convolutional module for temporal action localization in videos,.

Exploring Audio Cues for Enhanced Test-Time Video Model Adaptation Graph convolutional module for temporal action localization in videos,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:23.705035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:54:04.799608Z digest=sha256:1bcbbc8d58ffc6c0246440721ceee12d966094f7b65b25a2435b517890922442

Observation d8c8c786-7de4-4842-b7c0-c913004154ff · outbound

This paper cites Tent: Fully test-time adaptation by entropy minimization,.

Exploring Audio Cues for Enhanced Test-Time Video Model Adaptation Tent: Fully test-time adaptation by entropy minimization,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T00:54:04.889542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:54:04.889542Z digest=sha256:5b431faeb4b402e0c8ae17d9b4029fdf79f83e887820bcdd67691f0877552527

Observation d6d6f54d-38c6-45c3-9c36-2a2d30a9815a · outbound

This paper cites Memo: Test time robustness via adaptation and augmentation,.

Exploring Audio Cues for Enhanced Test-Time Video Model Adaptation Memo: Test time robustness via adaptation and augmentation,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:23.485877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:54:04.993798Z digest=sha256:8477dc61a1c266414c20e40481ff703524f613695f8bbaf68b3a7957765f1bd5

Observation 2c7e5712-aaae-4611-b40f-f39171f40063 · outbound

This paper cites Test-time classifier adjustment module for model-agnostic domain generalization,.

Exploring Audio Cues for Enhanced Test-Time Video Model Adaptation Test-time classifier adjustment module for model-agnostic domain generalization,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:23.195994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:54:05.066115Z digest=sha256:b21b931a130259a795f770af9310c9b13392d73ad7742b1e74bd92bb70925798

Observation bc74be26-e558-4bef-af5b-03f3aba10f92 · outbound

This paper cites The norm must go on: dynamic unsupervised domain adaptation by normalization,.

Exploring Audio Cues for Enhanced Test-Time Video Model Adaptation The norm must go on: dynamic unsupervised domain adaptation by normalization,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:22.878578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:54:05.135072Z digest=sha256:58fbab6aa61d283ae27de365d3fcfa3e2eb0d353b80949f5dd5a44dc8a62cae8

Observation 2e9a76d2-9df5-4787-a2c8-ed762a4f47ff · outbound

This paper cites Test-time training with masked autoencoders,.

Exploring Audio Cues for Enhanced Test-Time Video Model Adaptation Test-time training with masked autoencoders,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:22.349948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:54:05.254178Z digest=sha256:0815e771357b931ccff11fef20067dd9936274b26b17d554591b689661c961cb

Observation 3495722d-2572-4367-aa3f-5915c6aa5dc8 · outbound

This paper cites Efficient test-time model adaptation without forgetting,.

Exploring Audio Cues for Enhanced Test-Time Video Model Adaptation Efficient test-time model adaptation without forgetting,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:22.061736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:54:05.322903Z digest=sha256:cfb3f861f613ac1778706984d5348afda250ad0440ba6f6655de530472a23550

Observation c1ea4377-336a-4605-b6d8-3c1632d76d71 · outbound

This paper cites Test- time training with self-supervision for generalization under distribution shifts,.

Exploring Audio Cues for Enhanced Test-Time Video Model Adaptation Test- time training with self-supervision for generalization under distribution shifts,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:21.785630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:54:05.375815Z digest=sha256:b26ca748139b8eee502239e78d70d3555e6030189bc6e132230db874e56d43b7

Observation 84dd50fd-d4c4-4f41-899a-8a87e855d20c · outbound

This paper cites Ttt++: When does self-supervised test-time training fail or thrive?.

Exploring Audio Cues for Enhanced Test-Time Video Model Adaptation Ttt++: When does self-supervised test-time training fail or thrive?

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:21.566451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:54:05.448681Z digest=sha256:ecd6edb418a679b99945e3423d5ecfa6e6ce7a0d2c001a593bbd926eb016ce3d

Observation c10440b1-1872-475a-aafb-534383584764 · outbound

This paper cites Video test-time adaptation for action recognition,.

Exploring Audio Cues for Enhanced Test-Time Video Model Adaptation Video test-time adaptation for action recognition,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:21.264182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:54:05.518462Z digest=sha256:04e0c4d2012f4278e6b39f73b217be8df85c5b91cd46e78a3d9331e31849daf9

Observation 39240480-f0e0-46aa-bf14-37bc0fb45207 · outbound

This paper cites Exploring motion cues for video test-time adaptation,.

Exploring Audio Cues for Enhanced Test-Time Video Model Adaptation Exploring motion cues for video test-time adaptation,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:21.000344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:54:05.610626Z digest=sha256:9de3f1f016eb7517238c13d18d6911f1d3aaba20dad71511e2b3b1c95a53196b

Observation 868f3f2a-d989-4b31-9e57-31d4e654e113 · outbound

This paper cites Temporal coherent test time optimization for robust video classification,.

Exploring Audio Cues for Enhanced Test-Time Video Model Adaptation Temporal coherent test time optimization for robust video classification,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:20.721022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:54:05.690630Z digest=sha256:e279f4fb97de09b8a0b3b3b06b2b5fe2235f7685855a9a49d8a3b67f104edae8

Observation 91fc1409-706e-4903-9822-ac82210290cd · outbound

This paper cites Ast: Audio spectrogram trans- former,.

Exploring Audio Cues for Enhanced Test-Time Video Model Adaptation Ast: Audio spectrogram trans- former,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:20.447973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:54:05.784542Z digest=sha256:b619786afa7cbc58ad869270492f7b7725528e22140ff06bfe31b51b02edfcb6

Observation b65251ea-1f18-4f5e-b089-cf9994bdeab2 · outbound

This paper cites Bert: Pre-training of deep bidirectional transformers for language understanding,.

Exploring Audio Cues for Enhanced Test-Time Video Model Adaptation Bert: Pre-training of deep bidirectional transformers for language understanding,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:20.157611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:54:05.838913Z digest=sha256:38cd1066a48b4a9d98bc40139dc8e726b3d3cc2b5c612af1ef2c39e68eb85524

Observation c4706b51-6c68-4ebd-8e0b-3fdcc501bc46 · outbound

This paper cites Language models are few-shot learners,.

Exploring Audio Cues for Enhanced Test-Time Video Model Adaptation Language models are few-shot learners,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:19.851606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:54:05.948670Z digest=sha256:a0ff44b5a704fc881577f1cc3c7e76c56a1e2c8b35fee1a98c76802fbfc47c68

Observation 70221e23-df80-470b-a2b2-58c44d3425b0 · outbound

This paper cites Benchmarking micro-action recognition: Dataset, methods, and applications,.

Exploring Audio Cues for Enhanced Test-Time Video Model Adaptation Benchmarking micro-action recognition: Dataset, methods, and applications,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T00:54:06.000889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:54:06.000889Z digest=sha256:30f05c8c0cd9ce728540719dbe507b08a8ce1bc337aec9dc690386f3d9f4cd4a

Observation c6ff85e7-5055-4b09-a6ab-c043df2f96c8 · outbound

This paper cites Repetitive action counting with hybrid temporal relation modeling,.

Exploring Audio Cues for Enhanced Test-Time Video Model Adaptation Repetitive action counting with hybrid temporal relation modeling,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:19.626432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:54:06.065746Z digest=sha256:ab84d45daac98d549618ef8eb5924b6bd90903c95e616f162bdb96c5b583e51e

Observation 7520c981-33b3-4f84-a967-38a94ee653a4 · outbound

This paper cites Prototypical Calibrating Ambiguous Samples for Micro-Action Recognition.

Exploring Audio Cues for Enhanced Test-Time Video Model Adaptation Prototypical Calibrating Ambiguous Samples for Micro-Action Recognition

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:54:11.422232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:54:06.134824Z digest=sha256:66ca3212f9fda021b528774c906e44333610179be41f0f086f612d85256ca755

Observation 178c1827-fb16-4470-8a43-a91f2e495c3a · outbound

This paper cites Test-time classifier adjustment module for model-agnostic domain generalization,.

Exploring Audio Cues for Enhanced Test-Time Video Model Adaptation Test-time classifier adjustment module for model-agnostic domain generalization,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:19.444206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:54:06.202380Z digest=sha256:afa19ddbdcaff36a45ee5ee05e35122a73384f861d2020336744aa5d27e76993

Observation ef41e3dd-1d18-4f5b-b14e-bb2a1a9827a9 · outbound

This paper cites Revisiting realistic test-time training: Se- quential inference and adaptation by anchored clustering,.

Exploring Audio Cues for Enhanced Test-Time Video Model Adaptation Revisiting realistic test-time training: Se- quential inference and adaptation by anchored clustering,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:19.222262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:54:06.281441Z digest=sha256:addc8795e16aa74adbeca401e76d4d6d10474a042dc2907efb5addcab503784f

Observation e7b9f108-df4e-422d-aee0-b29fe0e605e0 · outbound

This paper cites SITA: Single Image Test-time Adaptation.

Exploring Audio Cues for Enhanced Test-Time Video Model Adaptation SITA: Single Image Test-time Adaptation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T00:54:06.445024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:54:06.445024Z digest=sha256:de8650dd65f7632864478b9ba43cc5a30ee36373f14d555666c1ad0f9c381718

Observation 98665da7-8afb-46f2-85cd-083c4b518967 · outbound

This paper cites Parameter- free online test-time adaptation,.

Exploring Audio Cues for Enhanced Test-Time Video Model Adaptation Parameter- free online test-time adaptation,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:22.606696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:54:06.515578Z digest=sha256:2a6731338121b49d87e48c19cb7783c9365323cb6c83548596ffc6eb097e36b1

Observation d9f47279-cac6-4db3-ae65-aabd91af77d6 · outbound

This paper cites Camera-aware recurrent learn- ing and earth mover’s test-time adaption for generalizable person re- identification,.

Exploring Audio Cues for Enhanced Test-Time Video Model Adaptation Camera-aware recurrent learn- ing and earth mover’s test-time adaption for generalizable person re- identification,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:19.006423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:54:06.595754Z digest=sha256:aca3988a609c2e9fc72fd85fe6f6fc26b244bc01586bbd3185c1c8d5c43483d6

Observation 4af02cc9-a9bb-41b2-8a9f-5acf0a4dd16e · outbound

This paper cites Question type-aware debiasing for test-time visual question answering model adaptation,.

Exploring Audio Cues for Enhanced Test-Time Video Model Adaptation Question type-aware debiasing for test-time visual question answering model adaptation,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:18.793452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:54:06.646156Z digest=sha256:5ce8570f2b28c191c35a78300c23642f08204bcfde0ae3982756b0f037e57b6e

Observation 426ef29f-b9f3-4086-95eb-a0d32b81d8f0 · outbound

This paper cites Ttagaze: Self- supervised test-time adaptation for personalized gaze estimation,.

Exploring Audio Cues for Enhanced Test-Time Video Model Adaptation Ttagaze: Self- supervised test-time adaptation for personalized gaze estimation,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:18.520748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:54:06.746865Z digest=sha256:0e238c4dff98108e8ac9203bf29e42ad6b618f453f777bd497ba65743c69dc98

Observation 174aeb4e-61c4-42e6-9455-24baaf35f24f · outbound

This paper cites What makes training multi-modal classification networks hard?.

Exploring Audio Cues for Enhanced Test-Time Video Model Adaptation What makes training multi-modal classification networks hard?

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:18.232489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:54:06.826776Z digest=sha256:5478d613af5710ac5bf0e623146db3fad216880913c945b0ea8432da49ce0a3a

Observation 255777d1-a2c8-4efb-8993-09c7139213b3 · outbound

This paper cites Self-supervised learning by cross-modal audio-video cluster- ing,.

Exploring Audio Cues for Enhanced Test-Time Video Model Adaptation Self-supervised learning by cross-modal audio-video cluster- ing,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:17.967929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:54:06.903443Z digest=sha256:87666750b54bfdb15ceeb08934199cd6c8b8734d819646679d42d2a685c40d34

Observation 32c2113e-0cb1-47d9-9873-d608507ab27e · outbound

This paper cites Look, listen and learn,.

Exploring Audio Cues for Enhanced Test-Time Video Model Adaptation Look, listen and learn,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T00:54:06.978167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:54:06.978167Z digest=sha256:f158285cb6c8e2777c2263c9199867eb0eda81d51d4aee9b2472268f7fb1db34

Observation 91176f35-89b1-40de-9e1b-290dcc898439 · outbound

This paper cites Epic-fusion: Audio-visual temporal binding for egocentric action recognition,.

Exploring Audio Cues for Enhanced Test-Time Video Model Adaptation Epic-fusion: Audio-visual temporal binding for egocentric action recognition,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:17.726461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:54:07.054009Z digest=sha256:4ade2acddecf6daf4a6ba195c4b787a984a60460c8c0714cb8369fff3f3bd788

Observation 085ab076-13e7-4e6a-8a21-931541ac745c · outbound

This paper cites Exploring multimodal video representation for action recognition,.

Exploring Audio Cues for Enhanced Test-Time Video Model Adaptation Exploring multimodal video representation for action recognition,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:17.454728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:54:07.124278Z digest=sha256:9abc69ca859af513fc147ce9e26cde26a3d0f49f35e5c913d4df0a6b2fb6e98c

Observation 23b8b71f-fc10-4aae-ae61-be9c24406bd2 · outbound

This paper cites Slowfast networks for video recognition,.

Exploring Audio Cues for Enhanced Test-Time Video Model Adaptation Slowfast networks for video recognition,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:17.205407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:54:07.191632Z digest=sha256:dc3bc0dca735c4ea0f0d4e2f75a7e43360dd059d31636c7018930c8b4eeee94e

Observation 6fd2c850-6ad3-497b-9e49-8f6d456f83a1 · outbound

This paper cites Attention is all you need,.

Exploring Audio Cues for Enhanced Test-Time Video Model Adaptation Attention is all you need,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:16.943553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:54:07.258909Z digest=sha256:ac0463d2dea8792ea3454bf9220e174c3bd942da1cf91316b799faeb1c461ab6

Observation fb303fa4-2626-4463-a998-61c25ca33da5 · outbound

This paper cites Polyvit: Co-training vision transformers on images, videos and audio,.

Exploring Audio Cues for Enhanced Test-Time Video Model Adaptation Polyvit: Co-training vision transformers on images, videos and audio,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:16.667500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:54:07.337614Z digest=sha256:8f8d63757142122990cc49e870590c2819fe69f0953b72c659a52de08a868d5f

Observation b16713c7-918d-429d-855c-720015e773d1 · outbound

This paper cites Attention bottlenecks for multimodal fusion,.

Exploring Audio Cues for Enhanced Test-Time Video Model Adaptation Attention bottlenecks for multimodal fusion,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:16.446415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:54:07.455568Z digest=sha256:ccc47f234fa8aaf577e968bc7d4ed6784700db690eb717ea9ddf1674225adf56

Observation dcf141e3-719f-4d7e-a95d-e6b3f163ad40 · outbound

This paper cites Mm-vit: Multi-modal video transformer for compressed video action recognition,.

Exploring Audio Cues for Enhanced Test-Time Video Model Adaptation Mm-vit: Multi-modal video transformer for compressed video action recognition,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:16.200539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:54:07.518122Z digest=sha256:e482d3720dcdad2ef9f68c5167f1b98b5e810e32c24f7ebc32b2ebc90b392a9c

Observation e23503e9-c259-49d7-b4b0-9f13b0e6a016 · outbound

This paper cites Multimodal video summa- rization via time-aware transformers,.

Exploring Audio Cues for Enhanced Test-Time Video Model Adaptation Multimodal video summa- rization via time-aware transformers,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:15.879756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:54:07.586329Z digest=sha256:0620d3cb2f76ab1cff98e06a8d35cbd2b1291dc1fb0e3af3d1aa53560e510a00

Observation 5bf451c0-f468-4d10-af81-edc6f728bcc2 · outbound

This paper cites Bootstrapping audio-visual video segmentation by strengthening audio cues,.

Exploring Audio Cues for Enhanced Test-Time Video Model Adaptation Bootstrapping audio-visual video segmentation by strengthening audio cues,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:15.684124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:54:07.655498Z digest=sha256:bd8248a25dfe45246a1f5d0fbd9dc3af9016efbf1af95f953a1fa6628a291213

Observation 9bdd4f10-6020-4ed5-b822-df7c825746f8 · outbound

This paper cites Multimodal imbalance-aware gradient modulation for weakly-supervised audio-visual video parsing,.

Exploring Audio Cues for Enhanced Test-Time Video Model Adaptation Multimodal imbalance-aware gradient modulation for weakly-supervised audio-visual video parsing,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:15.364959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:54:07.721031Z digest=sha256:96de1f0910a4fe525094b108c79591527651eb27e06824bac3f85800b4362db6

Observation 9d225224-9080-4aae-9bab-c8c718e3a678 · outbound

This paper cites Cross-Domain First Person Audio-Visual Action Recognition through Relative Norm Alignment.

Exploring Audio Cues for Enhanced Test-Time Video Model Adaptation Cross-Domain First Person Audio-Visual Action Recognition through Relative Norm Alignment

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:54:11.168964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:54:07.779728Z digest=sha256:3dc83280cb9f6b2cce94265b7e1a2f0e235e8da47475643aa45c899304361f4b

Observation aad33de7-7250-4584-acf3-6102305d0e99 · outbound

This paper cites Domain generalization through audio-visual relative norm align- ment in first person action recognition,.

Exploring Audio Cues for Enhanced Test-Time Video Model Adaptation Domain generalization through audio-visual relative norm align- ment in first person action recognition,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:15.133343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:54:07.851954Z digest=sha256:e04f0b2a023835c82456b08b9701bc426e516e965abfcef6aee22223aa390f50

Observation 59d7580c-d9ac-49f8-850d-120223135aa0 · outbound

This paper cites EPIC-KITCHENS-100 Unsupervised Domain Adaptation Challenge for Action Recognition 2021: Team M3EM Technical Report.

Exploring Audio Cues for Enhanced Test-Time Video Model Adaptation EPIC-KITCHENS-100 Unsupervised Domain Adaptation Challenge for Action Recognition 2021: Team M3EM Technical Report

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:54:10.875629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:54:07.929631Z digest=sha256:674b714f747f0e63062455a9f58e5ef4e65ca7af5c5c41120b3a81c722caf366

Observation d358d349-6e20-44f5-a9dd-9e2524f5621e · outbound

This paper cites Audio-adaptive activity recognition across video domains,.

Exploring Audio Cues for Enhanced Test-Time Video Model Adaptation Audio-adaptive activity recognition across video domains,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:14.886959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:54:07.986874Z digest=sha256:424eee87f3cca27c72d16255f3321dc744d50cce4165a20cf05f8b2b87ecd712

Observation 490a2777-9fd1-498a-b718-385ebbfdf2b5 · outbound

This paper cites Evaluating Prediction-Time Batch Normalization for Robustness under Covariate Shift.

Exploring Audio Cues for Enhanced Test-Time Video Model Adaptation Evaluating Prediction-Time Batch Normalization for Robustness under Covariate Shift

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T00:54:08.059471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:54:08.059471Z digest=sha256:b02ba724e0570ea65203b874fcc6f81ee66ff68ab156cfd61f9e903b24cc239e

Observation 07ed5fe3-4fa2-4c93-9df1-22ac58be4d8f · outbound

This paper cites Do we really need to access the source data? source hypothesis transfer for unsupervised domain adaptation,.

Exploring Audio Cues for Enhanced Test-Time Video Model Adaptation Do we really need to access the source data? source hypothesis transfer for unsupervised domain adaptation,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:14.624959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:54:08.145311Z digest=sha256:66326880a9b71ca9c72d587fb8bee78699a3b38128cc02c71bf0851e9acfde33

Observation dc90267e-2b60-4a94-b076-e9ad0be78f72 · outbound

This paper cites Improving robustness against common corruptions by co- variate shift adaptation,.

Exploring Audio Cues for Enhanced Test-Time Video Model Adaptation Improving robustness against common corruptions by co- variate shift adaptation,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:14.417289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:54:08.228981Z digest=sha256:6eaa758f8e393eddaea6ec42cf23f10bf0996668edd5ecc502f2c7b0d83e4f5b

Observation d93c796a-0242-4373-beae-a396379c58d4 · outbound

This paper cites Entropy is not enough for test-time adaptation: From the perspective of disentangled factors,.

Exploring Audio Cues for Enhanced Test-Time Video Model Adaptation Entropy is not enough for test-time adaptation: From the perspective of disentangled factors,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:14.240539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:54:08.301199Z digest=sha256:7d8a8250a269cf6373ecf80ae04415ffb2561f591955e18269397a13ab6c3ac2

Observation 0476f983-fb1a-487f-b7f6-fdf12a861e8a · outbound

This paper cites Universal test-time adaptation through weight ensembling, diversity weighting, and prior correction,.

Exploring Audio Cues for Enhanced Test-Time Video Model Adaptation Universal test-time adaptation through weight ensembling, diversity weighting, and prior correction,

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T00:54:08.407003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:54:08.407003Z digest=sha256:7bf8c11efe12b2f6a47ac5aa340eeb279963789efbff1cc3996709c9bd362bd2

Observation e768709d-a74b-4c7f-a212-9d9dd48ce3f1 · outbound

This paper cites Audiovisual Moments in Time: A Large-Scale Annotated Dataset of Audiovisual Actions.

Exploring Audio Cues for Enhanced Test-Time Video Model Adaptation Audiovisual Moments in Time: A Large-Scale Annotated Dataset of Audiovisual Actions

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:54:10.599730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:54:08.481578Z digest=sha256:5aeca1ae88e1e55bc4377b4eb84619f73410fe0e1d892b2d87507a73677721f6

Observation 7c8791cc-cb44-4098-89e3-c7f0b9e6e398 · outbound

This paper cites Audio-visual event localization in unconstrained videos,.

Exploring Audio Cues for Enhanced Test-Time Video Model Adaptation Audio-visual event localization in unconstrained videos,

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T00:54:08.563466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:54:08.563466Z digest=sha256:74f680dac44be09d399934f381d0e162d390e6cabc513af7a97a4e8bee4698ca

Observation b130993e-00e3-4e84-93ae-9c0080fb87fc · outbound

This paper cites Audio set: An ontology and human- labeled dataset for audio events,.

Exploring Audio Cues for Enhanced Test-Time Video Model Adaptation Audio set: An ontology and human- labeled dataset for audio events,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:14.034428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:54:08.636158Z digest=sha256:92fb2e76c0987fe684e4c9e8db4cd5f9c865212c91dda485d12e1d8a671e08ed

Observation 36651001-5f58-4677-9828-bb06c1502279 · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

Exploring Audio Cues for Enhanced Test-Time Video Model Adaptation UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T00:54:08.763811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:54:08.763811Z digest=sha256:495a571de0f7f880531b4fffca7032afc5083a36f0e2bc319748a25256391d89

Observation 3574da86-e5cc-46e9-8e52-1f40c592e625 · outbound

This paper cites The Kinetics Human Action Video Dataset.

Exploring Audio Cues for Enhanced Test-Time Video Model Adaptation The Kinetics Human Action Video Dataset

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T00:54:08.874815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:54:08.874815Z digest=sha256:8d74b2610f00764a6e5b47196e03e2066103d5d9d1e94cf8977e233c2de015a8

Observation b8e4cf70-b82d-4853-8bd2-105440ec9aaf · outbound

This paper cites Large-scale robustness analysis of video action recognition models,.

Exploring Audio Cues for Enhanced Test-Time Video Model Adaptation Large-scale robustness analysis of video action recognition models,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:13.827545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:54:09.001476Z digest=sha256:cca4db58f09df15c0c4809227714e4c8abff3983b313d4eb8d0256b01b2726cb

Observation 4e846e41-3dc4-440c-8be9-eb0aaa189ba4 · outbound

This paper cites Benchmarking the ro- bustness of spatial-temporal models against corruptions,.

Exploring Audio Cues for Enhanced Test-Time Video Model Adaptation Benchmarking the ro- bustness of spatial-temporal models against corruptions,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:13.597851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:54:09.163831Z digest=sha256:d216e8cf10a7ec83723159e46c648b05ce5eb1c8f2c5178f220b53ba3a112519

Observation 2a2a7b6d-6b0b-4dd2-ab25-f9a9fc0582ef · outbound

This paper cites Tam: Temporal adaptive module for video recognition,.

Exploring Audio Cues for Enhanced Test-Time Video Model Adaptation Tam: Temporal adaptive module for video recognition,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:13.364637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:54:09.279285Z digest=sha256:4f02d22fd93600f6524add38fd8ceabbccde5cc1dc583a6ae7dd96b3b6837e29

Observation 46ee0324-101f-42ee-aab1-47ce005d4311 · outbound

This paper cites Tsm: Temporal shift module for efficient video understanding,.

Exploring Audio Cues for Enhanced Test-Time Video Model Adaptation Tsm: Temporal shift module for efficient video understanding,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:13.073191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:54:09.361593Z digest=sha256:1a544eaf258cb8cdff3a4c101474f4f61e1a1127e641ca492123ca050713bcfb

Observation 87006b49-a61e-4ac3-9280-5a8f1af116ed · outbound

This paper cites Deep residual learning for image recognition,.

Exploring Audio Cues for Enhanced Test-Time Video Model Adaptation Deep residual learning for image recognition,

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T00:54:09.442939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:54:09.442939Z digest=sha256:c579e7a99d810bafa7c7bd9ee472a3532fb6d581191d6241122cd00b344ff8e7

Observation 63070090-db99-4110-ac56-e39954512b6b · outbound

This paper cites GPT-4 Technical Report.

Exploring Audio Cues for Enhanced Test-Time Video Model Adaptation GPT-4 Technical Report

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T00:54:09.583647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:54:09.583647Z digest=sha256:2619c239d3d7dbf3c5019a30dbed0f2d3646a780b80bd6e95372ef895824e9b8

Observation 509beed8-46ef-43fc-b9db-5dcc97aeaea0 · outbound

This paper cites The claude 3 model family: Opus, sonnet, haiku.

Exploring Audio Cues for Enhanced Test-Time Video Model Adaptation The claude 3 model family: Opus, sonnet, haiku

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:12.826462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:54:09.675286Z digest=sha256:24d835e785f1c112fd00c72d3f8005df985e5ee98a7f6584f0a3f58279fad025

Observation 8f407caa-0ac6-4087-82e7-a198afb82056 · outbound

This paper cites The Llama 3 Herd of Models.

Exploring Audio Cues for Enhanced Test-Time Video Model Adaptation The Llama 3 Herd of Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T00:54:09.790417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:54:09.790417Z digest=sha256:fe4f111a181a3f2e42868c69e282cb003e6ada87947118892c31b9650e1bc3c8

Observation ae8a3d41-acce-48ab-98aa-75f321b8c5bd · outbound

This paper cites Binggpt: Desktop application of new bing’s ai-powered chat,.

Exploring Audio Cues for Enhanced Test-Time Video Model Adaptation Binggpt: Desktop application of new bing’s ai-powered chat,

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:12.566777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:54:09.936000Z digest=sha256:8abaf5825820eba231a1fa54ed7a803b405ddfaeac2a95e5671d4cb0945d4ac7

Observation 11f5fe04-e0d2-4495-9474-6b730c5133c5 · outbound

This paper cites Audio mamba: Bidirectional state space model for audio representation learning,.

Exploring Audio Cues for Enhanced Test-Time Video Model Adaptation Audio mamba: Bidirectional state space model for audio representation learning,

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:12.285855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:54:10.052087Z digest=sha256:76bb0a238557a0893119cc3886ca049c26df7b62d6eb0d56873589aab93aa009

Observation 74beaab0-40bb-4086-b1e8-2dedecce1868 · outbound

This paper cites Beats: Audio pre-training with acoustic tokenizers,.

Exploring Audio Cues for Enhanced Test-Time Video Model Adaptation Beats: Audio pre-training with acoustic tokenizers,

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:12.035490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:54:10.186685Z digest=sha256:b476b32f9a531ae6936839ef91d1fcfdd756a7c241c92fcb7b639d76e46b98b8

Observation 1976f334-fbb2-432d-bf83-acf962a6dd6f · outbound

This paper cites Tim: A time interval machine for audio-visual action recognition,.

Exploring Audio Cues for Enhanced Test-Time Video Model Adaptation Tim: A time interval machine for audio-visual action recognition,

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:11.746977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:54:10.310454Z digest=sha256:279124acf87ae665e0f420256e7c4f0c7fc74abaa2837441a5af631c32a4993c

Pith citing papers

No inbound Pith citation observations are available.