Pith. sign in

Paper Citation Record · LEDGER

Multiscaled Multi-Head Attention-based Video Transformer Network for Hand Gesture Recognition

As of 11 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 0 inbound Pith citation observations for arXiv:2501.00935.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.00935 v1

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:44:42.468224Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

30 of 30 outbound references displayed

  • verified exact2
  • verified fuzzy23
  • unresolved5
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 01016162-a430-49d0-a8cc-0666398c37ef · outbound

This paper cites Exploiting recurrent neural networks and leap motion controller for the recognition of sign language and semaphoric hand gestures,.

Multiscaled Multi-Head Attention-based Video Transformer Network for Hand Gesture Recognition Exploiting recurrent neural networks and leap motion controller for the recognition of sign language and semaphoric hand gestures,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:44:42.898970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:44:42.338049Z digest=sha256:adeb2cdab095ef9ffd7c25a47c02a3f093167277bda2ff479fc060b657a7d863

Observation 8ea96a26-a8af-4e17-81f0-99f4260994e0 · outbound

This paper cites Attention in convolutional LSTM for gesture recognition,.

Multiscaled Multi-Head Attention-based Video Transformer Network for Hand Gesture Recognition Attention in convolutional LSTM for gesture recognition,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:44:42.886256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:44:42.343199Z digest=sha256:6c2566500fcf08675e40de72d05ad1676c21b40889484e66a53be0c00adae6e8

Observation d7ec6c33-debc-4041-a47a-d7518ab44d3b · outbound

This paper cites Attention-based gated recurrent unit for gesture recognition,.

Multiscaled Multi-Head Attention-based Video Transformer Network for Hand Gesture Recognition Attention-based gated recurrent unit for gesture recognition,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:44:42.873000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:44:42.347751Z digest=sha256:5899e4594ea37682666cc37bf5aae27214dfbc322a39a20651ea7f59e6d3a16c

Observation 17dd7c9e-576d-4da0-9294-a81fa22911ba · outbound

This paper cites Online detection and classification of dynamic hand gestures with recurrent 3D convolutional neural network,.

Multiscaled Multi-Head Attention-based Video Transformer Network for Hand Gesture Recognition Online detection and classification of dynamic hand gestures with recurrent 3D convolutional neural network,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:44:42.860177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:44:42.352633Z digest=sha256:4f56aa6ca33020233955a7a49ecdbcc211e16dd27bb735eace8bf4f9ebfe272b

Observation 22bf9b31-3bf1-4c4d-babb-d9c810005100 · outbound

This paper cites Attention is all you need,.

Multiscaled Multi-Head Attention-based Video Transformer Network for Hand Gesture Recognition Attention is all you need,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:44:42.848104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:44:42.357398Z digest=sha256:0f91b2283af4124ef691ab2f60b58db5020453bdf6eaab261d202ff3dce3121a

Observation bfbb28b8-bd59-4ae7-8b28-33d1cd3b9961 · outbound

This paper cites An image is worth 16 ×16 words: Transformers for image recognition at scale,.

Multiscaled Multi-Head Attention-based Video Transformer Network for Hand Gesture Recognition An image is worth 16 ×16 words: Transformers for image recognition at scale,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:44:42.835575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:44:42.362231Z digest=sha256:bbdaf7655f5921153735c621e432ccf869820fc1b7704fac2ee0d38b92778a8b

Observation 7164164d-1986-427c-98b1-cab1b9334db1 · outbound

This paper cites Training data-efficient image transformers & distillation through attention.

Multiscaled Multi-Head Attention-based Video Transformer Network for Hand Gesture Recognition Training data-efficient image transformers & distillation through attention

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T22:44:42.366245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:44:42.366245Z digest=sha256:815c6d697555e222aa83feaaf797d94d6c4f1499a6cb3de42e37773bf5245c18

Observation 422b1919-b928-445e-a91f-9bc6b623814f · outbound

This paper cites Video Transformer Network.

Multiscaled Multi-Head Attention-based Video Transformer Network for Hand Gesture Recognition Video Transformer Network

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-10T22:44:42.580679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:44:42.370070Z digest=sha256:a4333aa3b6a39c34b7ead8f75af343f55862387cbc1b06f611ec4e0b3a66245d

Observation cd7e5ee1-2f35-433d-ab50-f22063b446a8 · outbound

This paper cites CrossViT: Cross-Attention Multi-Scale Vision Transformer for Image Classification.

Multiscaled Multi-Head Attention-based Video Transformer Network for Hand Gesture Recognition CrossViT: Cross-Attention Multi-Scale Vision Transformer for Image Classification

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T22:44:42.374513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:44:42.374513Z digest=sha256:16d84d452b8bfec3055a2a31c8aacfb097685e3ca4d5e916a3a88465bc674792

Observation 3f4d0af7-febb-42fc-8b62-73b8ddfc7ac5 · outbound

This paper cites Multiscale Vision Transformers.

Multiscaled Multi-Head Attention-based Video Transformer Network for Hand Gesture Recognition Multiscale Vision Transformers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T22:44:42.379339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:44:42.379339Z digest=sha256:a0bd8a98259fc8ce6cbe7be0873ffb3f59cf270523f20e757d179f86b93466cb

Observation 30f1d953-db89-4e5d-841f-44dda78cb4fe · outbound

This paper cites MViTv2: Improved Multiscale Vision Transformers for Classification and Detection.

Multiscaled Multi-Head Attention-based Video Transformer Network for Hand Gesture Recognition MViTv2: Improved Multiscale Vision Transformers for Classification and Detection

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T22:44:42.384108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:44:42.384108Z digest=sha256:bd4ccfb2b0c71cd545c8baa9857b57176a5c0b7b876947dde9154fd60e5d9bc4

Observation 68c2f107-f4f3-42a9-9fb1-678ec968b632 · outbound

This paper cites A Transformer-based network for dynamic hand gesture recognition,.

Multiscaled Multi-Head Attention-based Video Transformer Network for Hand Gesture Recognition A Transformer-based network for dynamic hand gesture recognition,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:44:42.822679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:44:42.388888Z digest=sha256:e73d94502baa2fef6fdf08193c6b3802930e891733dd34bf55303588ca99142b

Observation 2cd495e4-8898-44c3-a2a9-c69fba8f1568 · outbound

This paper cites Incorporating relative position information in transformer-based sign language recognition and translation,.

Multiscaled Multi-Head Attention-based Video Transformer Network for Hand Gesture Recognition Incorporating relative position information in transformer-based sign language recognition and translation,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:44:42.811127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:44:42.393489Z digest=sha256:2c2c4b56ac1a5450baf9a6b66b89d228196d0653549950ac3f86ff6c110b949e

Observation 2790838a-6db1-4ad2-8fb6-9e6bbf06c04b · outbound

This paper cites Searching multi-rate and multi-modal temporal enhanced networks for gesture recognition,.

Multiscaled Multi-Head Attention-based Video Transformer Network for Hand Gesture Recognition Searching multi-rate and multi-modal temporal enhanced networks for gesture recognition,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:44:42.799286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:44:42.397778Z digest=sha256:af1c699b5dde8dfa773691c6dba9db1969663d6abab1571f9a33fa79a2fc0bf3

Observation 68c51c78-318f-4110-a498-055d440cd0c6 · outbound

This paper cites TMMF: Temporal Multi-Modal Fusion for single-stage continuous gesture recog- nition,.

Multiscaled Multi-Head Attention-based Video Transformer Network for Hand Gesture Recognition TMMF: Temporal Multi-Modal Fusion for single-stage continuous gesture recog- nition,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:44:42.787116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:44:42.403181Z digest=sha256:2c9ab20b6a3795fc33550c500a2112a8bb7160a92daa508d5ea401a1a578b00d

Observation 504e1112-2760-4982-ac4f-02749289d282 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

Multiscaled Multi-Head Attention-based Video Transformer Network for Hand Gesture Recognition Imagenet: A large-scale hierarchical image database

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:44:42.774288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:44:42.408169Z digest=sha256:d1b05fdbc2afb0e1b040e8099ef12186826d4fdadc09218d499f2fa459212879

Observation 8366ec62-9116-4e03-88a6-6ea0424ff9a6 · outbound

This paper cites Two-Stream Convolutional Networks for Action Recognition in Videos.

Multiscaled Multi-Head Attention-based Video Transformer Network for Hand Gesture Recognition Two-Stream Convolutional Networks for Action Recognition in Videos

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T22:44:42.412647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:44:42.412647Z digest=sha256:65b18beca9d11d33a5223f35342f30d5b92dfb9fe7e1d3e7aaeaea161b9f4ed5

Observation 733ce869-316a-48c3-9fb7-6d46535b37ee · outbound

This paper cites A robust and efficient video representation for action recognition,.

Multiscaled Multi-Head Attention-based Video Transformer Network for Hand Gesture Recognition A robust and efficient video representation for action recognition,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:44:42.761668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:44:42.417345Z digest=sha256:4e22d840dd05039a7ad4a84246bce09581ba8feb1f573532b5f54bd9475adf68

Observation 24fd4795-6661-4413-b6f8-35b5d5c504b8 · outbound

This paper cites Res3atn-deep 3D residual attention network for hand gesture recognition in videos,.

Multiscaled Multi-Head Attention-based Video Transformer Network for Hand Gesture Recognition Res3atn-deep 3D residual attention network for hand gesture recognition in videos,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:44:42.747742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:44:42.421408Z digest=sha256:87fff2085e4f4830f3273b11660b75be44493323dae74e1d2f1a73849a7158ef

Observation 88a65e5d-42fb-488a-a23c-6af0e4a9156f · outbound

This paper cites Learning spatiotemporal features with 3d convolutional networks,.

Multiscaled Multi-Head Attention-based Video Transformer Network for Hand Gesture Recognition Learning spatiotemporal features with 3d convolutional networks,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:44:42.734802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:44:42.425533Z digest=sha256:1d4216c379a59e13e7e3ffc2b417fa2d53434e7049585ed67482ec833fac38c8

Observation be70412e-52b7-462c-969d-c291b3adf076 · outbound

This paper cites Multi-task and multi-modal learning for RGB dynamic gesture recognition,.

Multiscaled Multi-Head Attention-based Video Transformer Network for Hand Gesture Recognition Multi-task and multi-modal learning for RGB dynamic gesture recognition,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:44:42.721546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:44:42.430604Z digest=sha256:c3c834df0ab0f6839085193dd7866f895bfa6f6af45f02c15bc67ccb7108f121

Observation 0bed9be7-8c61-4583-b2b8-247cc5b08c87 · outbound

This paper cites Making convolutional networks recurrent for visual sequence learning,.

Multiscaled Multi-Head Attention-based Video Transformer Network for Hand Gesture Recognition Making convolutional networks recurrent for visual sequence learning,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:44:42.707408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:44:42.435169Z digest=sha256:43a4d281ce9a69a93cdf42d02b91f83fd672e7347911b362d5432ae5f36ac08f

Observation 8304e128-bf4a-4640-8d07-8373307a7945 · outbound

This paper cites Quo vadis, action recognition? A new model and the kinetics dataset,.

Multiscaled Multi-Head Attention-based Video Transformer Network for Hand Gesture Recognition Quo vadis, action recognition? A new model and the kinetics dataset,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:44:42.693721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:44:42.439168Z digest=sha256:a583a46f8c3960ac127d9c2d01a0deabfd9e0ca643bda7908d7c8ffb958b8750

Observation 0bc6a973-ed18-4152-88e0-7b23947773c7 · outbound

This paper cites Real-time hand ges- ture detection and classification using convolutional neural networks,.

Multiscaled Multi-Head Attention-based Video Transformer Network for Hand Gesture Recognition Real-time hand ges- ture detection and classification using convolutional neural networks,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:44:42.680593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:44:42.444421Z digest=sha256:df11578af5370ed6b44eaccb1ae090ee412f78c748e338437c383ee65fc5253c

Observation f502ad87-70f5-4758-9f93-0e6ab8316c40 · outbound

This paper cites Improving the perfor- mance of unimodal dynamic hand-gesture recognition with multimodal training,.

Multiscaled Multi-Head Attention-based Video Transformer Network for Hand Gesture Recognition Improving the perfor- mance of unimodal dynamic hand-gesture recognition with multimodal training,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:44:42.667176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:44:42.448724Z digest=sha256:8ea5cb0922777634ff3badea42ed8048cfce28c6c3a16b1108f000b7034267d5

Observation 62664a0e-456d-4296-b70d-c4380f177282 · outbound

This paper cites Super normal vector for activity recognition using depth sequences,.

Multiscaled Multi-Head Attention-based Video Transformer Network for Hand Gesture Recognition Super normal vector for activity recognition using depth sequences,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:44:42.654082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:44:42.452677Z digest=sha256:5014f88e861f9e053df155fd404e27fa02e899a9b7349eb4cdd530cb6b4e698c

Observation 5a8e84e2-f1d3-4ed4-b213-bd0204b02d81 · outbound

This paper cites Motion fused frames: Data level fusion strategy for hand gesture recognition,.

Multiscaled Multi-Head Attention-based Video Transformer Network for Hand Gesture Recognition Motion fused frames: Data level fusion strategy for hand gesture recognition,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:44:42.632428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:44:42.456529Z digest=sha256:238e5046da84c9947532b76d95e312a72bc6606787c310fe2f6e8716e2a0a12c

Observation 21cce0f5-8e5b-4067-af29-96b6fb0330bd · outbound

This paper cites Dynamic hand gesture recognition based on short-term sampling neural networks,.

Multiscaled Multi-Head Attention-based Video Transformer Network for Hand Gesture Recognition Dynamic hand gesture recognition based on short-term sampling neural networks,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:44:42.619923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:44:42.460704Z digest=sha256:6b7cc19d6304b300163bdf9d09b4e485e9a392c69e35478fb75664b9ffb76864

Observation 6b435dcb-e792-435c-828e-c5a8e9bb572e · outbound

This paper cites Hand gestures for the human-car interaction: The Briareo dataset,.

Multiscaled Multi-Head Attention-based Video Transformer Network for Hand Gesture Recognition Hand gestures for the human-car interaction: The Briareo dataset,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:44:42.606536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:44:42.464869Z digest=sha256:90df6f9952e49754afce2f876e69f5e876379cb853ae65f10f67472f95e6b0a3

Observation c38f8d55-585f-4dc1-8916-f6425ccb2cf9 · outbound

This paper cites Multimodal hand gesture classification for the human–car interaction,.

Multiscaled Multi-Head Attention-based Video Transformer Network for Hand Gesture Recognition Multimodal hand gesture classification for the human–car interaction,

Reference 30

Resolution
verified exact
doi, observed 2026-08-10T22:44:42.505375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:44:42.468224Z digest=sha256:4ca2ad4bcc40e26ad6f86277c19489e6c69834ce71e420585be537635709b133

Pith citing papers

No inbound Pith citation observations are available.