Pith. sign in

Paper Citation Record · LEDGER

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment

As of 8 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 0 inbound Pith citation observations for arXiv:2507.21945.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.21945 v1

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T12:18:26.355829Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

46 of 46 outbound references displayed

  • verified exact1
  • verified fuzzy40
  • unresolved5
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 293d6fb2-0392-45ce-9586-c84c1a239967 · outbound

This paper cites Finediving: A fine-grained dataset for procedure-aware action quality assessment.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Finediving: A fine-grained dataset for procedure-aware action quality assessment

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:35.144656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:18:21.844647Z digest=sha256:c1195685e4c5dccfd675463e5bab90c3d8f3a661bbd43c4832a88cc87aae16c3

Observation e9c96844-2699-418e-9bf9-aea2a583ef50 · outbound

This paper cites Fine-grained spatio-temporal parsing network for action quality assess- ment.IEEE Transactions on Image Processing, 32:6386–6400, 2023.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Fine-grained spatio-temporal parsing network for action quality assess- ment.IEEE Transactions on Image Processing, 32:6386–6400, 2023

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:34.898184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:18:21.954456Z digest=sha256:8ec53e906b2dc49301e1fdd27fb476797a21c39930810bcb4ecb08fa24650fbf

Observation bce7b54f-aafb-4f97-b38d-1f2e3ad1d1ca · outbound

This paper cites Learning to score figure skating sport videos.IEEE transactions on circuits and systems for video technology, 30(12):4578– 4590, 2019.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Learning to score figure skating sport videos.IEEE transactions on circuits and systems for video technology, 30(12):4578– 4590, 2019

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:34.671458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:18:22.098946Z digest=sha256:4926df4580c001f6ffa59e8742cc94835f538b23ea4f3b5cdb9553db13ce3687

Observation b9c9d0af-3ffd-43c5-bfa0-60fba00ce5e0 · outbound

This paper cites Hybrid dynamic-static context-aware attention network for action assessment in long videos.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Hybrid dynamic-static context-aware attention network for action assessment in long videos

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:34.388543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:18:22.178255Z digest=sha256:36fa585b509416ea7a29d8ef44fa0c5c9b5c8e449d92860432890ef67e11a761

Observation f703451b-5f11-4694-a114-ee6f3d629405 · outbound

This paper cites End-to-end object detection with transformers.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment End-to-end object detection with transformers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:22.271513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:22.271513Z digest=sha256:57219ef4cbc2f28559f5c6e3b4fea52c757de257900c3d18664eb64682df04bd

Observation fae1c94c-ff63-46bd-b38d-4890b04186f0 · outbound

This paper cites Likert scoring with grade decoupling for long-term action assessment.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Likert scoring with grade decoupling for long-term action assessment

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:34.116411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:18:22.358013Z digest=sha256:86d439bbc28af0ea9d6d64324ff2788ff86416058a6d5d1865a098f58a45cf20

Observation 0355ddde-df52-4237-93d8-7814f517d19c · outbound

This paper cites Localization-assisted uncertainty score disentanglement net- work for action quality assessment.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Localization-assisted uncertainty score disentanglement net- work for action quality assessment

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:33.861720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:18:22.478760Z digest=sha256:5d235bea8aea78f54ee803d2386aaf2d30f88141a97504ff7765df9c57fd60ee

Observation 10f02535-4362-44a8-a5d1-6e4330fe6d0a · outbound

This paper cites Skating-mixer: Long-term sport audio- visual modeling with mlps.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Skating-mixer: Long-term sport audio- visual modeling with mlps

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:33.569918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:18:22.611670Z digest=sha256:e2449891c787a5e4b3a2207d1e4dff5a224ce8f89427dc57fbb2fb5ce0217be3

Observation a137ae47-4f6a-4496-a78a-dbd49b3bc654 · outbound

This paper cites Multimodal action quality assess- ment.IEEE Transactions on Image Processing, 2024.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Multimodal action quality assess- ment.IEEE Transactions on Image Processing, 2024

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:33.296808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:18:22.702717Z digest=sha256:8248e8fee2a359af00edfea9a0545102ff6bea50edcfdb953c90ea7b72b032e6

Observation 42c0b6d9-8172-420e-9538-900c4e13dc03 · outbound

This paper cites Audio-visual scene analysis with self-supervised multisensory features.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Audio-visual scene analysis with self-supervised multisensory features

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:33.080761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:18:22.758132Z digest=sha256:523ac8fdcd53c48bd61832baabe77e0fcf89297c50074f6db2597f1101b66adf

Observation 3aa4d9d9-625d-42c8-8ca1-a744569cd76a · outbound

This paper cites Dual attention matching for audio-visual event localization.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Dual attention matching for audio-visual event localization

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:32.859041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:18:22.823789Z digest=sha256:eff4815193ce1b218978a170e079435cb5abea64dd8be78fec70c72a32a8dbde

Observation c5da39c2-3d3a-4cef-98d2-fadf74989a0b · outbound

This paper cites Egocentric deep multi-channel audio-visual active speaker localization.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Egocentric deep multi-channel audio-visual active speaker localization

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:32.408078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:18:23.069552Z digest=sha256:79451114b9bec0ab331c2fef64a37f181b62aa0a4f87f1bd33995124e3aaf12d

Observation d69bf16c-1e26-4758-83d9-628fe87471fd · outbound

This paper cites Assessing the quality of actions.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Assessing the quality of actions

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:32.134002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:18:23.128630Z digest=sha256:43bf728d281819c2d26557fd8de80ffcae6da30246fca30065e732170330478f

Observation ee7b2f01-1aec-417b-b3bb-efc9d3d894c1 · outbound

This paper cites Learning to score olympic events.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Learning to score olympic events

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:31.939914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:18:23.200363Z digest=sha256:febe7e46337005cc71e20c0a1883ca3c60d48c368ac74b8f067c3b224207ae3c

Observation 6133f46d-31a4-4a7e-b90b-a5545b9c8fb3 · outbound

This paper cites Scoringnet: Learning key fragmentforactionqualityassessmentwithrankinglossinskilledsports.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Scoringnet: Learning key fragmentforactionqualityassessmentwithrankinglossinskilledsports

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:31.780758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:18:23.255859Z digest=sha256:9d002f0759f363ed05dd3e886dcd73e23fa2934547e6f137ab24a0a63c853f58

Observation 2c24ff34-4a45-44be-8408-6740c5944ac1 · outbound

This paper cites S3d: Stacking segmental p3d for action quality assessment.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment S3d: Stacking segmental p3d for action quality assessment

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:31.471727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:18:23.326548Z digest=sha256:a640fb260037622f60059345b1a80cf3018be24a227d8c50d8a2d12ba4c34eac

Observation cbc69f9b-b598-4031-884c-9857ee33c4da · outbound

This paper cites What and how well you performed? a multitask learning approach to action quality assessment.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment What and how well you performed? a multitask learning approach to action quality assessment

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:31.211318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:18:23.377327Z digest=sha256:ca44cb442d37e5101e8667f422e7d9a3cfaa0baaf52686c6feffd5c092349cfb

Observation 86725ee6-d879-47df-aea2-21bae0051371 · outbound

This paper cites Action quality assess- ment using siamese network-based deep metric learning.IEEE Trans- actions on Circuits and Systems for Video Technology, 31(6):2260–2273, 2020.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Action quality assess- ment using siamese network-based deep metric learning.IEEE Trans- actions on Circuits and Systems for Video Technology, 31(6):2260–2273, 2020

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:30.927185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:18:23.455277Z digest=sha256:55c103a35af9d57b1d4169fd3d76f0e32c52a5b17ff8a4bf90a0d1978973f297

Observation 05c56ede-cb32-4a2a-9325-622501bcd876 · outbound

This paper cites Group-aware contrastive regression for action quality assessment.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Group-aware contrastive regression for action quality assessment

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:30.770604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:18:23.594595Z digest=sha256:ad3e11cb551546c3bf700a01c7167125f9a83da1be54a72675c28db4721d4337

Observation 1bf7300b-4013-408e-8638-66c6b208df1e · outbound

This paper cites Tsa-net: Tube self-attention network for action quality assess- ment.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Tsa-net: Tube self-attention network for action quality assess- ment

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:30.587767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:18:23.704637Z digest=sha256:4d7ad5926d8e75819628452e941a6ccecf7520154b53d313d7bea1b3b0e5b8dc

Observation 5b1f2cf4-99be-4007-b2fa-335733409e54 · outbound

This paper cites The pros and cons: Rank-aware temporal attention for skill determination in long videos.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment The pros and cons: Rank-aware temporal attention for skill determination in long videos

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:30.370377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:18:23.751508Z digest=sha256:845c38ff7953cbc75b78a095c61bc93c8f05bf1b6a1bff33166413f7ab513952

Observation ce7576cd-de75-4a86-b20d-eff8782e34aa · outbound

This paper cites Logo: A long-form video dataset for group action quality assessment.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Logo: A long-form video dataset for group action quality assessment

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:30.174716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:18:23.882275Z digest=sha256:190cf0fbc01296e31428e17b214ca866165d8b6ba26809b1fc83fc9720b5bf14

Observation 54681c95-aeda-4cf9-aca8-04ec0f6c8b98 · outbound

This paper cites Audiovisual SlowFast Networks for Video Recognition.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Audiovisual SlowFast Networks for Video Recognition

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:23.939472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:23.939472Z digest=sha256:65db1b3072848156cf1e33924e9e906e5f323b19f47a45a7e9e5f28f64307628

Observation 81734ab5-1ad0-4503-83d4-00709ab74d6b · outbound

This paper cites Listentolook: Actionrecognitionbypreviewingaudio.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Listentolook: Actionrecognitionbypreviewingaudio

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:30.005897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:18:24.019897Z digest=sha256:f40cee11ec03a0f7fe419615cb07c15471057f83886bef25bb4e4012681b3c8f

Observation 565f5d5d-8e86-4cb2-9ba0-be13bcebb5d9 · outbound

This paper cites Cross- attentional audio-visual fusion for weakly-supervised action localization.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Cross- attentional audio-visual fusion for weakly-supervised action localization

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:32.617834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:18:24.162198Z digest=sha256:984d0906d1e946e01f20eba6d4e4e9d976756f26774c376774c6eb7782a48159

Observation e52dcf7d-07b4-465c-be09-7cb7d4688d30 · outbound

This paper cites Cross-modal background suppression for audio- visual event localization.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Cross-modal background suppression for audio- visual event localization

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:29.857858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:18:24.275209Z digest=sha256:5439b331f2092c6301af9fdaa35e9c11c1b9c5e3e4feac767fac65135b3d05b0

Observation b365b3dc-c437-428e-b17f-2b8798d840de · outbound

This paper cites Mmw-aqa: Multimodal in-the-wild dataset for action quality assess- ment.IEEE Access, 2024.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Mmw-aqa: Multimodal in-the-wild dataset for action quality assess- ment.IEEE Access, 2024

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:29.724551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:18:24.438832Z digest=sha256:570ddef6562f51cb288b26de3a0b303d4470da8787cca2c623edc05cb9e2bc70

Observation 5316e41a-8695-4090-a504-b1b094029013 · outbound

This paper cites Vision-language action knowledge learning for semantic-aware action quality assessment.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Vision-language action knowledge learning for semantic-aware action quality assessment

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:29.608867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:18:24.554309Z digest=sha256:3fbd87f7ff1bcb2a9e943a1003892926231f87842e2b2955f4c74152c630ffbd

Observation 373a5f44-f617-4a56-826b-dcf0f2b3c98a · outbound

This paper cites Learning semantics- guided representations for scoring figure skating.IEEE Transactions on Multimedia, 26:4987–4997, 2023.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Learning semantics- guided representations for scoring figure skating.IEEE Transactions on Multimedia, 26:4987–4997, 2023

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:29.453029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:18:24.665443Z digest=sha256:8fa388ece26752c424af9dafa42f068bbf4ac6a8fd50dacca72ff9d8f979f112

Observation 6466f5bb-3bea-4602-976b-3fd74bde4804 · outbound

This paper cites Temporal and cross-modal attention for audio-visual zero-shot learning.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Temporal and cross-modal attention for audio-visual zero-shot learning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:29.348648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:18:24.812348Z digest=sha256:191af1be7fcb59b6ebb5bab2e1fdc12de82a884dfeb240376725fa232e661964

Observation 28f464e8-07a7-4ac6-af2c-8f45d0800f4e · outbound

This paper cites Cross-attention is not always needed: Dynamic cross-attention for audio-visual dimensional emotion recognition.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Cross-attention is not always needed: Dynamic cross-attention for audio-visual dimensional emotion recognition

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:29.201321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:18:24.981888Z digest=sha256:db726248f8f1f3d190e2002dca04327458650e2d640163746e56d411e4a7a909

Observation dc635a20-9a1d-41c8-b400-95e65f7d9f20 · outbound

This paper cites AlignVSR: Audio-Visual Cross-Modal Alignment for Visual Speech Recognition.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment AlignVSR: Audio-Visual Cross-Modal Alignment for Visual Speech Recognition

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-08-06T12:18:26.702155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:18:25.094671Z digest=sha256:c175598c9babb82f7b85fb18288a34405b674e9d5f8c2491d4d086522b4a27e2

Observation 81151f82-1911-418a-9873-2a844699de31 · outbound

This paper cites Cross-Modal Global Interaction and Local Alignment for Audio-Visual Speech Recognition.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Cross-Modal Global Interaction and Local Alignment for Audio-Visual Speech Recognition

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:25.205516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:25.205516Z digest=sha256:69d2a66318462b214e444c1d34792ac7c5978c6fd8e95e46963e93950a9aa388

Observation 9a091c2c-f28a-4b37-a51e-d7c82d18db2c · outbound

This paper cites Temporal alignment networks for long-term video.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Temporal alignment networks for long-term video

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:28.997529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:18:25.294446Z digest=sha256:bf6a68f936a0511ffb17e38a2333bc2e2a3673fa4322ba92ce7161030929cb75

Observation 0cce93b2-4bc4-4ec1-b98c-a11a30d0c358 · outbound

This paper cites Temporal and cross-modal attention for audio-visual zero-shot learning.Advances in neural information processing systems, 35:38032–38045, 2022.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Temporal and cross-modal attention for audio-visual zero-shot learning.Advances in neural information processing systems, 35:38032–38045, 2022

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:28.715831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:18:25.400298Z digest=sha256:8876fc9ea2a754472adb59a5fa412629e9de7fdaf3f1f56d689c016583501674

Observation 371a4223-90b5-496a-ad60-888a71e4af2e · outbound

This paper cites Video and accelerometer-based motion analysis for automated surgical skills assessment.International journal of computer assisted radiology and surgery, 13:443–455, 2018.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Video and accelerometer-based motion analysis for automated surgical skills assessment.International journal of computer assisted radiology and surgery, 13:443–455, 2018

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:28.491535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:18:25.469481Z digest=sha256:9916208590b86cf590af469ad952f9057db189797028f01b336f926d4a6322ae

Observation be99137b-cb2b-4394-9f9e-23c44703000e · outbound

This paper cites Audio set: An ontology and human-labeled dataset for audio events.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Audio set: An ontology and human-labeled dataset for audio events

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:28.040623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:18:25.735323Z digest=sha256:13f6f723c371ffaf3a7f838bef14724eccf8222ea6a51ae714d3643c5cc3fdb1

Observation cae90976-fcc6-48fb-9880-545d7827f2bc · outbound

This paper cites Action quality assessment with temporal parsing transformer.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Action quality assessment with temporal parsing transformer

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:27.837087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:18:25.826751Z digest=sha256:1a4978e60823853e21de8b3b5eb482f6621b9b95b5adb8f97b8b8064969e3b51

Observation 47c93c07-361c-486a-bf55-ee1634b7f01a · outbound

This paper cites Learning spatiotemporal features with 3d convolu- tional networks.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Learning spatiotemporal features with 3d convolu- tional networks

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:27.690611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:18:25.887329Z digest=sha256:9c66ee8c0aa1dd7e7d77fd404bcb98f8f120a579c97d89c48844b876d9d54e14

Observation a7d2812a-e2a4-4ec1-bf60-f120a2df2321 · outbound

This paper cites Video and accelerometer-based motion analysis for automated 43 surgical skills assessment.International journal of computer assisted radiology and surgery, 13:443–455, 2018.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Video and accelerometer-based motion analysis for automated 43 surgical skills assessment.International journal of computer assisted radiology and surgery, 13:443–455, 2018

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:27.509449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:18:25.948741Z digest=sha256:6b9faf9767e8643da90309db04f4d883cbe41f5895151c2c282578ab7039d082

Observation be607b59-d4b6-4031-b12e-9f75cf801bc5 · outbound

This paper cites Deep resid- uallearningforimagerecognition.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Deep resid- uallearningforimagerecognition

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:27.296707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:18:26.008721Z digest=sha256:46450ef80609bd7efc0dff09b9353a813af8c5397858ec2767f1576a70cc0c38

Observation b63b5445-ccb8-4bc5-b274-386495198d89 · outbound

This paper cites Quo vadis, action recognition? a new model and the kinetics dataset.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Quo vadis, action recognition? a new model and the kinetics dataset

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:28.289591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:18:26.071958Z digest=sha256:445b428047382631d74a0099c46adb05ede4e726f44a6db751c947a67171037b

Observation 187f7f20-0af3-4e6f-a632-27b512db3adc · outbound

This paper cites AST: Audio Spectrogram Transformer.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment AST: Audio Spectrogram Transformer

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:26.116606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:26.116606Z digest=sha256:56187f19b0c27138c8fffb4d5c342cdd0c9c5aa878af46c55ca40e305b14fca3

Observation 7f0a990d-0458-4ace-99a2-e20b3a350c74 · outbound

This paper cites Joint visual and audio learning for video highlight detection.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Joint visual and audio learning for video highlight detection

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:27.086584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:18:26.176463Z digest=sha256:62161a8dacb0ccfe0a7a5dd2894402bbdb5c2aae5064fe15793515d3ae70dfad

Observation da49cd3c-3a8c-444d-a127-ce5d3ba5b860 · outbound

This paper cites MSAF: Multimodal Split Attention Fusion.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment MSAF: Multimodal Split Attention Fusion

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:26.255413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:26.255413Z digest=sha256:7dd5262834e7702e253dbc69b9e5215c479a30e7b2348e407d2b74e1e1fdc2a0

Observation be1ce44f-b478-4363-9b88-7c5c7f62d816 · outbound

This paper cites Umt: Unified multi-modal transformers for joint video moment retrieval and highlight detection.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Umt: Unified multi-modal transformers for joint video moment retrieval and highlight detection

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:26.885434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:18:26.355829Z digest=sha256:aea21a5fca0a91275585fc7e0fcfb3dcaca4299767b9146d90e45ad924b879eb

Pith citing papers

No inbound Pith citation observations are available.