Pith. sign in

Paper Citation Record · LEDGER

CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization

As of 14 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 1 inbound Pith citation observation for arXiv:2505.23524.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.23524 v2

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:48:48.549041Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:48:45.412605Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T12:48:48.744686Z

Reference resolution

22 of 22 outbound references displayed

  • verified exact1
  • verified fuzzy18
  • unresolved2
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation be7dd9f4-c7b6-44a7-9938-a620c55f2b37 · outbound

This paper cites Tradi- tional methods primarily focus on identifying the most rele- vant videos based on a given query [1].

CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization Tradi- tional methods primarily focus on identifying the most rele- vant videos based on a given query [1]

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:55.060702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:48:45.288073Z digest=sha256:f5c0672e21be92da973ce83bc43c00e2ef6127188c879d3776c141d465b8d008

Observation b8a62b15-6959-49e3-a1fd-6abd14e3d672 · outbound

This paper cites CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization.

CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:48:48.858700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:48:45.412605Z digest=sha256:107656d835e6ac3a5763dc1acb29b054dc2fae878950e14fb6eaf432f5f95f1a

Observation b4972a19-a7db-48a8-b0dc-916c2c9ecc51 · outbound

This paper cites Pre-trained audio, VLP, and CBP feature extractors first extractF audio,F CBP , andF V LP features.

CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization Pre-trained audio, VLP, and CBP feature extractors first extractF audio,F CBP , andF V LP features

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:54.769020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:48:45.570821Z digest=sha256:e2cdc38ef5f3fd5e46f32345b8ab5695aa2af9e0eb20df638d84f14fe19e13a4

Observation 7fc4dc23-0ae0-46ec-8056-d21829e2a281 · outbound

This paper cites Datasets THUMOS14consists of 200 validation and 213 test videos across 20 action classes, averaging 15 action segments per video.

CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization Datasets THUMOS14consists of 200 validation and 213 test videos across 20 action classes, averaging 15 action segments per video

Reference 4

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T12:48:54.499112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:48:45.706456Z digest=sha256:0ef828c03e1bd8047a95fd70ed377b0ffe0563f65789b3455d6c7f703dcef56f

Observation 9d21f908-8417-4162-9ee9-31e60b9c2cb9 · outbound

This paper cites This is the first time CLIP and audio are incorporated into UTAL.

CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization This is the first time CLIP and audio are incorporated into UTAL

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:54.205295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:48:45.858643Z digest=sha256:2843658e5d9b1faa4f74478c798df7fe9118244882a285fc629106cd0c5c15ee

Observation 972e97a1-c7f6-466f-8b0d-dd926d4e3e0a · outbound

This paper cites an unresolved cited work.

CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:48:53.972067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:48:45.938182Z digest=sha256:4829b943bbf906ccf05275334c75ae93303a6ebe4d9a7fa009714c23932aa9b6

Observation 4ab952d0-c16a-479a-86a9-f0da1962b537 · outbound

This paper cites Two-stream consensus network for weakly-supervised temporal action local- ization,.

CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization Two-stream consensus network for weakly-supervised temporal action local- ization,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:53.658122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:48:46.145963Z digest=sha256:0a833cf1b64ee9129f608d4df34d5fd25ec054f7f7ca8bede72edadf68f245ac

Observation 0d566cc5-1f04-4187-a95c-7326fd47617b · outbound

This paper cites Weakly-supervised temporal action lo- calization by inferring salient snippet-feature,.

CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization Weakly-supervised temporal action lo- calization by inferring salient snippet-feature,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:53.362935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:48:46.302868Z digest=sha256:bd5f2350beab97c345852714d488af7c467d7d0b2fed8295fa4d53bf74684ce1

Observation 9f980152-557e-422b-acdd-578e05553ac2 · outbound

This paper cites Two-stream networks for weakly-supervised temporal action local- ization with semantic-aware mechanisms,.

CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization Two-stream networks for weakly-supervised temporal action local- ization with semantic-aware mechanisms,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:53.071181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:48:46.477326Z digest=sha256:9c41b0aa149b58d3188a8823f20c252117ebd0c08e01f21ee1bd3be98953b7d3

Observation bedf9601-f849-45c1-893d-4433e4ec5471 · outbound

This paper cites Weakly Supervised Temporal Action Localization via Dual-Prior Collaborative Learning Guided by Multimodal Large Language Models.

CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization Weakly Supervised Temporal Action Localization via Dual-Prior Collaborative Learning Guided by Multimodal Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:46.650668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:48:46.650668Z digest=sha256:65f468fb6289b6464896e648097ccde8307b1c23a564ed8413f561b2c3bf5432

Observation 526873f1-ab5e-4d40-ad72-37b7d1b01cb6 · outbound

This paper cites Rethinking pseudo-label guided learning for weakly supervised temporal action localization from the perspective of noise correction,.

CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization Rethinking pseudo-label guided learning for weakly supervised temporal action localization from the perspective of noise correction,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:52.740062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:48:46.819500Z digest=sha256:93230ec9489e100b93fcd5e92ecd8b52a7fdb9fab12b4095aaf840976988e65c

Observation 1eb1bdc3-2c6b-4836-badc-0f02d60d79f2 · outbound

This paper cites Apsl: Action-positive separation learning for unsupervised temporal action localization,.

CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization Apsl: Action-positive separation learning for unsupervised temporal action localization,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:52.443787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:48:46.937683Z digest=sha256:98fd718d49381c393b3a49cc51fb83e85f24619838609ce9f4f3a576df11db2e

Observation e54e0186-25fe-479a-acec-d8f5c2c77a5d · outbound

This paper cites Learning temporal co-attention models for unsupervised video action localization,.

CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization Learning temporal co-attention models for unsupervised video action localization,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:52.029311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:48:47.187380Z digest=sha256:6612addede9bc91b0996850b5f988ea75e0ba0fa338fd7e72151e752260d0de4

Observation d4322ed8-362d-4821-a264-28948253a0cc · outbound

This paper cites Uncertainty guided collaborative training for weakly supervised and unsupervised temporal action localization,.

CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization Uncertainty guided collaborative training for weakly supervised and unsupervised temporal action localization,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:51.717507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:48:47.327506Z digest=sha256:2512d99c084306f8f8234b8318d5559046cb9102c60ad9d9f0e6e4391fe4b7de

Observation 1203afcb-20bb-4cbc-ae52-0b04790e2c0b · outbound

This paper cites Revisiting foreground and background separation in weakly-supervised temporal action local- ization: A clustering-based approach,.

CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization Revisiting foreground and background separation in weakly-supervised temporal action local- ization: A clustering-based approach,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:51.345450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:48:47.485299Z digest=sha256:f5e2880d27d9c8987a835326eeb9fcf5f3ef459753b1c70db110fded147191a9

Observation 2b287037-7dee-41e7-b1c1-deed347d2a52 · outbound

This paper cites Weakly supervised action localization by sparse temporal pooling network,.

CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization Weakly supervised action localization by sparse temporal pooling network,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:50.957971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:48:47.600169Z digest=sha256:ca5adb72d9069e343a2160029c9bb3b43338826d1d935b8208281910f9729d11

Observation 571033fb-cc77-4cb0-ae8d-758dfd04c236 · outbound

This paper cites Weakly-supervised action localization with background modeling,.

CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization Weakly-supervised action localization with background modeling,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:50.754831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:48:47.796538Z digest=sha256:968b4a7ffe93812e7e890e7a43a23b1a18215b28f6321dc1712a41573a793a90

Observation 3c64e08f-d674-44af-a81b-f7b95e026678 · outbound

This paper cites Proposal-based multiple instance learning for weakly-supervised temporal action localization,.

CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization Proposal-based multiple instance learning for weakly-supervised temporal action localization,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:50.527963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:48:47.946877Z digest=sha256:880f4a4b6510f5aba2a5a0ebf96fb4a35d384544e6dd966af9f1e5bc0f9de598

Observation 3d868298-5ba4-48a8-aeff-daaad0a00920 · outbound

This paper cites IMDPrompter: Adapting SAM to image manipulation detection by cross-view au- tomated prompt learning,.

CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization IMDPrompter: Adapting SAM to image manipulation detection by cross-view au- tomated prompt learning,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:50.257347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:48:48.045654Z digest=sha256:5a1534c610fdc6aaa7ba69f1702bbe3195ac11c616d59b75d39671536e50da4d

Observation 1ce0dbff-495f-464a-9a6d-e5fbfb64174c · outbound

This paper cites Gim: A million-scale benchmark for generative image manipulation detection and localization,.

CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization Gim: A million-scale benchmark for generative image manipulation detection and localization,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:49.843937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:48:48.218893Z digest=sha256:7480e3491060104b3ef48fd9ecef7d413c3c5b3ee5468b2ee9328a865ccc23a6

Observation af09fa38-d8d3-47cb-8882-7144e755d5c9 · outbound

This paper cites Seeing beyond noise: Joint graph structure evaluation and denoising for mul- timodal recommendation,.

CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization Seeing beyond noise: Joint graph structure evaluation and denoising for mul- timodal recommendation,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:49.414467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:48:48.390480Z digest=sha256:b27dae470c60d23a1f6b5b234a2571824a82c65530ee03697808ee477c17047b

Observation e236e4ff-3518-4f22-8806-362f13497fa1 · outbound

This paper cites Distilling semantic priors from sam to effi- cient image restoration models,.

CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization Distilling semantic priors from sam to effi- cient image restoration models,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:49.123969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:48:48.549041Z digest=sha256:94b8d87b60f4cc8dedf30ddb5bc936f2ad7c7217d32141b6717c5ba1f7a43727

Pith citing papers

Observation b8a62b15-6959-49e3-a1fd-6abd14e3d672 · inbound

CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization cites this paper.

CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:48:48.858700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:48:45.412605Z digest=sha256:107656d835e6ac3a5763dc1acb29b054dc2fae878950e14fb6eaf432f5f95f1a