Pith. sign in

Paper Citation Record · LEDGER

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding

As of 13 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2507.03531.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.03531 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:10:40.870961Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

31 of 31 outbound references displayed

  • verified exact2
  • verified fuzzy21
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3f64a285-5486-4481-a848-41d37450c95b · outbound

This paper cites Learning transferable visual models from natural language supervision.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding Learning transferable visual models from natural language supervision

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:10:41.273653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T20:10:40.500401Z digest=sha256:57059c44581d0940ebd114185a3d2cc95c1dd16bf01a396190a0f09b9774b420

Observation 08f5321e-940a-492b-930e-a8d2926d4753 · outbound

This paper cites Cowen, Stefanos Zafeiriou, Irene Kotsia, Eric Granger, Marco Pedersoli, Simon L.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding Cowen, Stefanos Zafeiriou, Irene Kotsia, Eric Granger, Marco Pedersoli, Simon L

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:10:41.264057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T20:10:40.607543Z digest=sha256:9291dcaa90ccd7a0d239659678e15b6d59d8683d949327ac47218f967bb3ba52

Observation a08ef3e8-8b7d-4929-bdad-2564cdb6b6d9 · outbound

This paper cites Advancements in affective and behavior analysis: The 8th abaw workshop and competition.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding Advancements in affective and behavior analysis: The 8th abaw workshop and competition

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:10:41.254562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T20:10:40.645816Z digest=sha256:c5c90e5118ae2cf56a4a9b8990c4f0b0ae63042315dfded430a4e3a6ecb1efdf

Observation f1e90cbc-c112-4c54-b824-e6368022d9fd · outbound

This paper cites 7th ABAW Competition: Multi-Task Learning and Compound Expression Recognition.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding 7th ABAW Competition: Multi-Task Learning and Compound Expression Recognition

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:10:41.059074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T20:10:40.759840Z digest=sha256:3bef307318f08ef69d54f38a9085b4f2f23583bd6c94752f9dd7e7963955025e

Observation 35fda226-7fbe-475b-bd42-1b768c83a350 · outbound

This paper cites The 6th affective behavior analysis in-the-wild (abaw) competition.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding The 6th affective behavior analysis in-the-wild (abaw) competition

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:10:41.244620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T20:10:40.781120Z digest=sha256:0d0c0c9b1b29a38164f884a3f6eb03aba65a0f0383a7e3605389a1d1059134ad

Observation 7445a316-cce0-437c-9b9a-b346f8832216 · outbound

This paper cites Distribution matching for multi-task learning of classification tasks: A large-scale study on faces & beyond.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding Distribution matching for multi-task learning of classification tasks: A large-scale study on faces & beyond

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:10:41.235719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T20:10:40.784942Z digest=sha256:a7ca1275550c96982db8301af123c052d4cc4c8c975b7cd3c4e0be022be06282

Observation f58196a3-b9e7-48ef-aabb-b4550e736c0a · outbound

This paper cites Abaw: Valence-arousal estimation, expression recognition, action unit detection & emotional reaction intensity estimation challenges.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding Abaw: Valence-arousal estimation, expression recognition, action unit detection & emotional reaction intensity estimation challenges

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:10:41.226399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T20:10:40.788478Z digest=sha256:4364a0cdcc50dc5aa7cc40a9d80fdbe2aca813a2070b9bb59e3a49c7071721ea

Observation 03e8bd93-488b-4c2c-aa4e-c3e8d09a1303 · outbound

This paper cites Multi-label compound expression recognition: C-expr database & network.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding Multi-label compound expression recognition: C-expr database & network

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:10:41.215669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T20:10:40.791965Z digest=sha256:7000dc50cf7d79146329f879fad9b898a73c80e97ac21eebc9cef567c5229791

Observation 072d26f9-c235-417b-97cc-5b06ba8674f6 · outbound

This paper cites Abaw: Valence-arousal estimation, expression recognition, action unit detection & emotional reaction intensity estimation challenges.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding Abaw: Valence-arousal estimation, expression recognition, action unit detection & emotional reaction intensity estimation challenges

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:10:41.206809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T20:10:40.795345Z digest=sha256:bf2bd24bcb719b6ae6d3e2dd5f41c177a23f8f1816fcd423705fd22ca05619d6

Observation f63ccf43-2d0e-4eeb-829b-c0f7dcb30879 · outbound

This paper cites Abaw: Valence-arousal estimation, expression recognition, action unit detection & multi-task learning challenges.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding Abaw: Valence-arousal estimation, expression recognition, action unit detection & multi-task learning challenges

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:10:41.197017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T20:10:40.799224Z digest=sha256:b1b707f548a3f26e89a5b7dd7cbda658d313df55482005072ac862f2e5a7ce3b

Observation a085e703-44d0-4474-9ae1-31069f6f3644 · outbound

This paper cites Analysing affective behavior in the second abaw2 competition.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding Analysing affective behavior in the second abaw2 competition

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:10:41.187225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T20:10:40.803144Z digest=sha256:587ed991738af74f92c6bb4edbe17ced9f1000068dc6a58da28d080795c2287d

Observation 87117d46-f234-483e-a142-2a2469971925 · outbound

This paper cites Analysing affective behavior in the first abaw 2020 competition.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding Analysing affective behavior in the first abaw 2020 competition

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:10:41.177571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T20:10:40.806712Z digest=sha256:01569b2696bb5b538cb766d27000bbba0e374158c3d1bc0627dc033510a329ce

Observation 1689b15b-0d71-49b2-a3cb-f3e06dc87cfa · outbound

This paper cites Distribution Matching for Heterogeneous Multi-Task Learning: a Large-scale Face Study.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding Distribution Matching for Heterogeneous Multi-Task Learning: a Large-scale Face Study

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T20:10:40.810429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:10:40.810429Z digest=sha256:5333aa7693b2133b5f0060cdc62d114e7bfcf32d8ef4c3e770ace73d7c43b6c1

Observation 4b24b0fc-aa14-4630-9a09-26b0a906c5be · outbound

This paper cites Affect Analysis in-the-wild: Valence-Arousal, Expressions, Action Units and a Unified Framework.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding Affect Analysis in-the-wild: Valence-Arousal, Expressions, Action Units and a Unified Framework

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T20:10:40.814343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:10:40.814343Z digest=sha256:2046662196880d3476fbe837c11ed26794d4ca1fbb1b733193e4aaa1a4428325

Observation e6a9dfae-7c77-4cc3-9499-58c56f06e7a9 · outbound

This paper cites Expression, Affect, Action Unit Recognition: Aff-Wild2, Multi-Task Learning and ArcFace.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding Expression, Affect, Action Unit Recognition: Aff-Wild2, Multi-Task Learning and ArcFace

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T20:10:40.817947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:10:40.817947Z digest=sha256:adf720364f7096d14fa64918efd8e9aa52cfe031bf42fa6080b37e682f8ece95

Observation 0045ef01-93c9-4bd8-9c95-7beb30d01b9c · outbound

This paper cites Face Behavior a la carte: Expressions, Affect and Action Units in a Single Network.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding Face Behavior a la carte: Expressions, Affect and Action Units in a Single Network

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T20:10:40.821888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:10:40.821888Z digest=sha256:ce832bb301801346b0e56832391c65d9213c502790c475af7d082200a6bbb83a

Observation 1d5046d1-8763-454e-828a-e274d0327289 · outbound

This paper cites Deep affect prediction in-the-wild: Aff-wild database and challenge, deep architectures, and beyond.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding Deep affect prediction in-the-wild: Aff-wild database and challenge, deep architectures, and beyond

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:10:41.167438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T20:10:40.825643Z digest=sha256:ee981bbe9dcea3d754213f0ccf12ead036efde5d90d927124a0c1b78ba05be55

Observation 2df671bc-8334-47fb-89ab-627009e25e1e · outbound

This paper cites Dvd: A comprehensive dataset for advancing violence detection in real-world scenarios.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding Dvd: A comprehensive dataset for advancing violence detection in real-world scenarios

Reference 18

Resolution
verified exact
raw_fallback, observed 2026-08-06T20:10:41.005456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T20:10:40.828773Z digest=sha256:cccd0420fde5cf526b3a1559fd707fc9f65e21c9845af9158881c8a269ce9f4e

Observation 736a5e90-dd88-4472-88b7-674ad358d992 · outbound

This paper cites Deep residual learning for image recognition.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding Deep residual learning for image recognition

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:10:41.157434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T20:10:40.831875Z digest=sha256:2b3592737c36c899abca2f16e357dacfe2daadbf5098594d1444937965266452

Observation 7691a0c4-a1e8-4862-9972-962336ec3e0e · outbound

This paper cites Efficientnetv2: Smaller models and faster training.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding Efficientnetv2: Smaller models and faster training

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:10:41.148032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T20:10:40.835222Z digest=sha256:21c9e011370c294d5833ba91c52ef5ae4da8ef5ff8427f950c6bb963b7fc8174

Observation 9fd62586-fcd0-458d-a8b2-606daee5c610 · outbound

This paper cites Masked autoencoders are scalable vision learners.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding Masked autoencoders are scalable vision learners

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:10:41.137646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T20:10:40.838238Z digest=sha256:07a1a55f4765db10b880e5a318a81741210276bf46f1c3e3836318c6b49f92d0

Observation c0f5d483-74e6-4ac4-8d94-069b27293e7c · outbound

This paper cites Cnn architectures for large-scale audio classification.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding Cnn architectures for large-scale audio classification

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:10:41.127063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T20:10:40.841654Z digest=sha256:303935fc6b7f12e46abea253cf001f802c86759283fb1587e94a57d5a7407fa7

Observation 635b295a-8207-4728-a989-d738270442ec · outbound

This paper cites wav2vec 2.0: A framework for self- supervised learning of speech representations.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding wav2vec 2.0: A framework for self- supervised learning of speech representations

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:10:41.116831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T20:10:40.844737Z digest=sha256:829f4bbf6745e23d8e6c83f3e55e4179b5a44c0dc5e335fd837f115a796fdfb0

Observation bbd5d9cb-16bf-45c9-aef4-bd7838eb0c62 · outbound

This paper cites Bert: Pre-training of deep bidirectional transformers for language understanding.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding Bert: Pre-training of deep bidirectional transformers for language understanding

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:10:41.106839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T20:10:40.847598Z digest=sha256:4c09f070da4e42ff2f1dc9243009b2127b631ab49331e93c7a75e6db2d0730c4

Observation f6bb8e1d-d169-4914-941a-52d1d0f315db · outbound

This paper cites Attention is all you need.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding Attention is all you need

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:10:41.096282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T20:10:40.850744Z digest=sha256:aed6a48b81c7d4120969e00bf97575c0b541d7fd6cb9aa6e40558f29fe3bf0b9

Observation 38f6ce61-6586-45ce-9f5f-36e4e8be9a1c · outbound

This paper cites Learning phrase representations using rnn encoder-decoder for statistical machine translation.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding Learning phrase representations using rnn encoder-decoder for statistical machine translation

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:10:41.085288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T20:10:40.853807Z digest=sha256:2f3b8b2becaee8999299ebde17eb068dadabab2b8b0e435f8f5ad711c8aa5ff2

Observation fae51b66-8d97-4a4e-a4c1-ff63160464e0 · outbound

This paper cites An Empirical Evaluation of Generic Convolutional and Recurrent Networks for Sequence Modeling.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding An Empirical Evaluation of Generic Convolutional and Recurrent Networks for Sequence Modeling

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T20:10:40.856769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:10:40.856769Z digest=sha256:6fd48fdfa1a012837db02e0fa8ed32ac1471b54abf09ef0a678e287b192bcc66

Observation 8ca3f941-0204-4fcd-8856-70ab839f6d5b · outbound

This paper cites Contrastive Training of Complex-Valued Autoencoders for Object Discovery.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding Contrastive Training of Complex-Valued Autoencoders for Object Discovery

Reference 28

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T20:10:40.917485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T20:10:40.861047Z digest=sha256:5955d6c4f221b09526223aa0bf7454ebba2ede4bfc2d42e7f45c279914b518ed

Observation e2935fdb-3a17-4d9c-8fe2-1c96ee5ed57b · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T20:10:40.864539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:10:40.864539Z digest=sha256:f321fa3e00605a0a933b35d296cfe3e5622bb21e64adca2ad31fed9af6562499

Observation a8410bae-100f-4ba2-b17a-20ffe6b0573e · outbound

This paper cites Focal loss for dense object detection.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding Focal loss for dense object detection

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T20:10:40.867742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:10:40.867742Z digest=sha256:310c7ca5fbc9417ab1ed437542655a7c463d7ce71aff47fed221923ebdcedcf3

Observation cbdc3a54-adea-4007-a5d8-8526b33355a3 · outbound

This paper cites Decoupled weight decay regularization.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding Decoupled weight decay regularization

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:10:41.069656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T20:10:40.870961Z digest=sha256:ae8cf4137a6846257e23b5ab4492bfd847ec01a51fd14cca1681faba0f1e4f01

Pith citing papers

No inbound Pith citation observations are available.