Pith. sign in

Paper Citation Record · LEDGER

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding

As of 9 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2507.03531.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.03531 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:10:40.870961Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

31 of 31 outbound references displayed

  • verified exact2
  • verified fuzzy21
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3f64a285-5486-4481-a848-41d37450c95b · outbound

This paper cites Learning transferable visual models from natural language supervision.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding Learning transferable visual models from natural language supervision

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:10:41.273653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:10:40.500401Z digest=sha256:47b203c9fadf55e9bdaf45b02d8b2ef3d3f425f1a2b7fc96ac5e2dbc1ebb4c1a

Observation 08f5321e-940a-492b-930e-a8d2926d4753 · outbound

This paper cites Cowen, Stefanos Zafeiriou, Irene Kotsia, Eric Granger, Marco Pedersoli, Simon L.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding Cowen, Stefanos Zafeiriou, Irene Kotsia, Eric Granger, Marco Pedersoli, Simon L

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:10:41.264057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:10:40.607543Z digest=sha256:ac8b94d913b5243dbfd84dd5314baa2d1bd5f512335013831f5b870d9fcebe6e

Observation a08ef3e8-8b7d-4929-bdad-2564cdb6b6d9 · outbound

This paper cites Advancements in affective and behavior analysis: The 8th abaw workshop and competition.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding Advancements in affective and behavior analysis: The 8th abaw workshop and competition

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:10:41.254562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:10:40.645816Z digest=sha256:2e63ebb95375e73adebee412e77683c1abed53603ee85561bff1aba571256285

Observation f1e90cbc-c112-4c54-b824-e6368022d9fd · outbound

This paper cites 7th ABAW Competition: Multi-Task Learning and Compound Expression Recognition.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding 7th ABAW Competition: Multi-Task Learning and Compound Expression Recognition

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:10:41.059074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:10:40.759840Z digest=sha256:48a99636f5c5cf4d8ee44e63147494302b478a34706387d1e72bae2acb4db2f3

Observation 35fda226-7fbe-475b-bd42-1b768c83a350 · outbound

This paper cites The 6th affective behavior analysis in-the-wild (abaw) competition.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding The 6th affective behavior analysis in-the-wild (abaw) competition

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:10:41.244620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:10:40.781120Z digest=sha256:44e337e4f55760e66448f95dd0938314026aaa03f41d2384bd6caf07312e20a4

Observation 7445a316-cce0-437c-9b9a-b346f8832216 · outbound

This paper cites Distribution matching for multi-task learning of classification tasks: A large-scale study on faces & beyond.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding Distribution matching for multi-task learning of classification tasks: A large-scale study on faces & beyond

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:10:41.235719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:10:40.784942Z digest=sha256:096693cf5ac2f988886c6adb00ce610424a7e81ce306c50e29b4d029096fb45d

Observation f58196a3-b9e7-48ef-aabb-b4550e736c0a · outbound

This paper cites Abaw: Valence-arousal estimation, expression recognition, action unit detection & emotional reaction intensity estimation challenges.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding Abaw: Valence-arousal estimation, expression recognition, action unit detection & emotional reaction intensity estimation challenges

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:10:41.226399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:10:40.788478Z digest=sha256:9f4597e6a013fc813c25103b9930d1dbc086956b1eafa00b7fa37836b8219f30

Observation 03e8bd93-488b-4c2c-aa4e-c3e8d09a1303 · outbound

This paper cites Multi-label compound expression recognition: C-expr database & network.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding Multi-label compound expression recognition: C-expr database & network

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:10:41.215669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:10:40.791965Z digest=sha256:9919f44ce6e7f1160701400f6fe9b7d21b86ae1d921ef86f10c0ea19ec3c18cd

Observation 072d26f9-c235-417b-97cc-5b06ba8674f6 · outbound

This paper cites Abaw: Valence-arousal estimation, expression recognition, action unit detection & emotional reaction intensity estimation challenges.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding Abaw: Valence-arousal estimation, expression recognition, action unit detection & emotional reaction intensity estimation challenges

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:10:41.206809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:10:40.795345Z digest=sha256:882ee00f5f50c7201aec798c325869375b3f32bd88051b894673ea305a3748a5

Observation f63ccf43-2d0e-4eeb-829b-c0f7dcb30879 · outbound

This paper cites Abaw: Valence-arousal estimation, expression recognition, action unit detection & multi-task learning challenges.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding Abaw: Valence-arousal estimation, expression recognition, action unit detection & multi-task learning challenges

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:10:41.197017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:10:40.799224Z digest=sha256:3fee4606864f4dcf13b3c88bc1ee142bc0543e9cf06d4755bbb2bb878e60523d

Observation a085e703-44d0-4474-9ae1-31069f6f3644 · outbound

This paper cites Analysing affective behavior in the second abaw2 competition.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding Analysing affective behavior in the second abaw2 competition

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:10:41.187225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:10:40.803144Z digest=sha256:674a493cb5be219c14415ef107c357b7a4343ea3d31af8a16e1fd602bac0e310

Observation 87117d46-f234-483e-a142-2a2469971925 · outbound

This paper cites Analysing affective behavior in the first abaw 2020 competition.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding Analysing affective behavior in the first abaw 2020 competition

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:10:41.177571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:10:40.806712Z digest=sha256:602ee2f39b0d672d38ca5c75d56f48a840b7658f3d9014dbf38d69182d424544

Observation 1689b15b-0d71-49b2-a3cb-f3e06dc87cfa · outbound

This paper cites Distribution Matching for Heterogeneous Multi-Task Learning: a Large-scale Face Study.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding Distribution Matching for Heterogeneous Multi-Task Learning: a Large-scale Face Study

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T20:10:40.810429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:10:40.810429Z digest=sha256:895dd297199eba063aad7b5d7027648fe9ed3f5851bf84c4fdd30ae0e0bd2a40

Observation 4b24b0fc-aa14-4630-9a09-26b0a906c5be · outbound

This paper cites Affect Analysis in-the-wild: Valence-Arousal, Expressions, Action Units and a Unified Framework.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding Affect Analysis in-the-wild: Valence-Arousal, Expressions, Action Units and a Unified Framework

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T20:10:40.814343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:10:40.814343Z digest=sha256:06a976f478b6be5c51ce39042200c16777055da9e64111fd26713c69c8ccec43

Observation e6a9dfae-7c77-4cc3-9499-58c56f06e7a9 · outbound

This paper cites Expression, Affect, Action Unit Recognition: Aff-Wild2, Multi-Task Learning and ArcFace.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding Expression, Affect, Action Unit Recognition: Aff-Wild2, Multi-Task Learning and ArcFace

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T20:10:40.817947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:10:40.817947Z digest=sha256:606bcd01ae8e4fce3af06cb2adfd9da994b0f99cacef393b45a65019f64d6a5b

Observation 0045ef01-93c9-4bd8-9c95-7beb30d01b9c · outbound

This paper cites Face Behavior a la carte: Expressions, Affect and Action Units in a Single Network.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding Face Behavior a la carte: Expressions, Affect and Action Units in a Single Network

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T20:10:40.821888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:10:40.821888Z digest=sha256:ce832bb301801346b0e56832391c65d9213c502790c475af7d082200a6bbb83a

Observation 1d5046d1-8763-454e-828a-e274d0327289 · outbound

This paper cites Deep affect prediction in-the-wild: Aff-wild database and challenge, deep architectures, and beyond.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding Deep affect prediction in-the-wild: Aff-wild database and challenge, deep architectures, and beyond

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:10:41.167438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:10:40.825643Z digest=sha256:de72de438d5af407623ccc1b73774c6a5350d2c0725e05b00a3bc694f644a7a2

Observation 2df671bc-8334-47fb-89ab-627009e25e1e · outbound

This paper cites Dvd: A comprehensive dataset for advancing violence detection in real-world scenarios.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding Dvd: A comprehensive dataset for advancing violence detection in real-world scenarios

Reference 18

Resolution
verified exact
raw_fallback, observed 2026-08-06T20:10:41.005456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:10:40.828773Z digest=sha256:1a85684564c1ba8bf6cff4796e29b6559c5230eac41e20115837504aeb5c3f08

Observation 736a5e90-dd88-4472-88b7-674ad358d992 · outbound

This paper cites Deep residual learning for image recognition.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding Deep residual learning for image recognition

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:10:41.157434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:10:40.831875Z digest=sha256:60a03599e1a69caa8b517ca0e08b9c9bd4db781cb59b41d052aaa43ac8995a9b

Observation 7691a0c4-a1e8-4862-9972-962336ec3e0e · outbound

This paper cites Efficientnetv2: Smaller models and faster training.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding Efficientnetv2: Smaller models and faster training

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:10:41.148032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:10:40.835222Z digest=sha256:cc99c1e2a4d9a0cb005a2f13999f4b9f1b02e5dbcb4a14f3879b89a84a5d2cb1

Observation 9fd62586-fcd0-458d-a8b2-606daee5c610 · outbound

This paper cites Masked autoencoders are scalable vision learners.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding Masked autoencoders are scalable vision learners

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:10:41.137646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:10:40.838238Z digest=sha256:9c623f880d4637a440dbb90adc4bae3e0c1b146557142a7c6d8c10c4ed4498fa

Observation c0f5d483-74e6-4ac4-8d94-069b27293e7c · outbound

This paper cites Cnn architectures for large-scale audio classification.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding Cnn architectures for large-scale audio classification

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:10:41.127063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:10:40.841654Z digest=sha256:689f9b12a7b8fbe4897a23c74eaa54ab81cd94cbf447af3515a2439ec2c116b2

Observation 635b295a-8207-4728-a989-d738270442ec · outbound

This paper cites wav2vec 2.0: A framework for self- supervised learning of speech representations.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding wav2vec 2.0: A framework for self- supervised learning of speech representations

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:10:41.116831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:10:40.844737Z digest=sha256:ba8563a897f178d4146cb357864541c759a9869dcdb0b59e26368b0105d3ff5d

Observation bbd5d9cb-16bf-45c9-aef4-bd7838eb0c62 · outbound

This paper cites Bert: Pre-training of deep bidirectional transformers for language understanding.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding Bert: Pre-training of deep bidirectional transformers for language understanding

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:10:41.106839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:10:40.847598Z digest=sha256:6259d56b75a64b4a8df7e35e3bf40e47623d1ca14af3f262467707a006298fd4

Observation f6bb8e1d-d169-4914-941a-52d1d0f315db · outbound

This paper cites Attention is all you need.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding Attention is all you need

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:10:41.096282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:10:40.850744Z digest=sha256:23cf84c31c75dc9557dbc24a8dd703a5b1427702e72d17cd78e0ba4f5e726b8e

Observation 38f6ce61-6586-45ce-9f5f-36e4e8be9a1c · outbound

This paper cites Learning phrase representations using rnn encoder-decoder for statistical machine translation.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding Learning phrase representations using rnn encoder-decoder for statistical machine translation

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:10:41.085288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:10:40.853807Z digest=sha256:f9527a091815545b8e9b70c9bea7dd9411e8e9b068fac862b858e7fec6ba27be

Observation fae51b66-8d97-4a4e-a4c1-ff63160464e0 · outbound

This paper cites An Empirical Evaluation of Generic Convolutional and Recurrent Networks for Sequence Modeling.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding An Empirical Evaluation of Generic Convolutional and Recurrent Networks for Sequence Modeling

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T20:10:40.856769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:10:40.856769Z digest=sha256:6fd48fdfa1a012837db02e0fa8ed32ac1471b54abf09ef0a678e287b192bcc66

Observation 8ca3f941-0204-4fcd-8856-70ab839f6d5b · outbound

This paper cites Contrastive Training of Complex-Valued Autoencoders for Object Discovery.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding Contrastive Training of Complex-Valued Autoencoders for Object Discovery

Reference 28

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T20:10:40.917485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:10:40.861047Z digest=sha256:d5b0d82230bea149b86d2c5c77d64ea09537b4baf63d52d5c66c6db9660d014c

Observation e2935fdb-3a17-4d9c-8fe2-1c96ee5ed57b · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T20:10:40.864539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:10:40.864539Z digest=sha256:f321fa3e00605a0a933b35d296cfe3e5622bb21e64adca2ad31fed9af6562499

Observation a8410bae-100f-4ba2-b17a-20ffe6b0573e · outbound

This paper cites Focal loss for dense object detection.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding Focal loss for dense object detection

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T20:10:40.867742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:10:40.867742Z digest=sha256:310c7ca5fbc9417ab1ed437542655a7c463d7ce71aff47fed221923ebdcedcf3

Observation cbdc3a54-adea-4007-a5d8-8526b33355a3 · outbound

This paper cites Decoupled weight decay regularization.

Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding Decoupled weight decay regularization

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:10:41.069656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:10:40.870961Z digest=sha256:e2c2fd62115193985974810671133d5489e278ba7a09595f09db0a50694ed313

Pith citing papers

No inbound Pith citation observations are available.