Pith. sign in

Paper Citation Record · LEDGER

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition

As of 10 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 0 inbound Pith citation observations for arXiv:2507.11892.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.11892 v1

Coverage vector

measured 65 of 65 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:08:58.980743Z

measured 65 of 65 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

65 of 65 outbound references displayed

  • verified exact0
  • verified fuzzy40
  • unresolved23
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2f80eff5-5e07-4e4f-bc11-824bde58144a · outbound

This paper cites The extended cohn-kanade dataset (ck+): A complete dataset for action unit and emotion-specified expression,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition The extended cohn-kanade dataset (ck+): A complete dataset for action unit and emotion-specified expression,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:10.037724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:08:49.233116Z digest=sha256:6aeb0d824773becc14ce5f240d39c2485c86cb03abb32bb7a79d4766588e52ce

Observation 8f33e461-9f0b-45e0-87f3-fc36ada3e9af · outbound

This paper cites Induced disgust, happiness and surprise: an addition to the mmi facial expression database,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Induced disgust, happiness and surprise: an addition to the mmi facial expression database,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:09.788431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:08:49.371685Z digest=sha256:d7d2b87875cf9b4bc51318737f37db7da1313fbfa162bfe74a72d10698d58dda

Observation c7d2cc0f-ba65-45e7-82ba-d384efa2e394 · outbound

This paper cites Facial expression recognition from near-infrared videos,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Facial expression recognition from near-infrared videos,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:09.507381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:08:49.524779Z digest=sha256:cced59f4895882b44f043390c778500933d989ca4eb96b28bd07b33c63d08a95

Observation 697b9dc7-4c81-4042-ba30-c4b59b5cdca8 · outbound

This paper cites Dfew: A large-scale database for recognizing dynamic facial expressions in the wild,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Dfew: A large-scale database for recognizing dynamic facial expressions in the wild,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:09.232452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:08:49.678700Z digest=sha256:bce1ab6656a1600398317b8f2f7ed83dd8959fb713edd57777ce15b8234ae479

Observation 7b370099-2a7c-449f-acc8-6c8f058cd5f8 · outbound

This paper cites Mafw: A large-scale, multi-modal, compound affective database for dynamic facial expression recognition in the wild,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Mafw: A large-scale, multi-modal, compound affective database for dynamic facial expression recognition in the wild,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:08.923395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:08:49.819097Z digest=sha256:9b2b0704535fef28a4f3cad166a7a69eac9896e8b4cb45a668a6686c9de9b22f

Observation bbf98089-a486-4f67-9f52-4f89fe49eff0 · outbound

This paper cites Ferv39k: A large-scale multi-scene dataset for facial expres- sion recognition in videos,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Ferv39k: A large-scale multi-scene dataset for facial expres- sion recognition in videos,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:08.691148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:08:49.985284Z digest=sha256:aa38a9fd02fc2af14edc32a38f0583497d0c27122c4d705e4c952d622b1cc411

Observation 154984c9-e876-47c2-84ae-4d67398609e4 · outbound

This paper cites Dep-fer: Facial expression recognition in depressed patients based on voluntary facial expression mimicry,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Dep-fer: Facial expression recognition in depressed patients based on voluntary facial expression mimicry,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:08.492490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:08:50.172639Z digest=sha256:800302917b6abf4268c0d2ac4c0009a7b72d6cd0d6d1c4398ec864dbebb3f195

Observation 770416af-30bd-44e0-ab48-7554f99f46f0 · outbound

This paper cites Efficient facial expression recognition with representation reinforcement network and transfer self-training for human–machine interaction,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Efficient facial expression recognition with representation reinforcement network and transfer self-training for human–machine interaction,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:08.367493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:08:50.384854Z digest=sha256:f91bc2809f6b4eb1e1245254729e124db7b34edb5f7b3a3d4b993031c5fbe75e

Observation afccfec2-b81d-41ca-baa9-d4b7a9b331e4 · outbound

This paper cites Predicting personal- ized image emotion perceptions in social networks,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Predicting personal- ized image emotion perceptions in social networks,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:08.225797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:08:50.553598Z digest=sha256:a1941f58bb245ecea45a2222dd11a3f2aaacc47796212d53a6fba4e7172d6c23

Observation 8e3d8a74-1cb2-45fd-9683-5afaff8dc945 · outbound

This paper cites Spatio-temporal convolutional features with nested lstm for facial expression recognition,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Spatio-temporal convolutional features with nested lstm for facial expression recognition,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:08.056382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:08:50.726269Z digest=sha256:6ff8ac5b7165f19e2d24597fe29a8a0a9c6122b51bc80a53f01204dfe0df291f

Observation 5639bf19-85ab-4294-8870-a9829cdffe31 · outbound

This paper cites Saanet: Siamese action-units attention network for improving dynamic facial expression recognition,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Saanet: Siamese action-units attention network for improving dynamic facial expression recognition,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T17:08:50.885960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:08:50.885960Z digest=sha256:2c67d63967d7fd903d4f11e335d42965cfdb20183c6552736f9d7ca111fc55a8

Observation 9fe8fd94-4f81-4457-8460-8a451088566b · outbound

This paper cites Deep residual learning for image recognition,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Deep residual learning for image recognition,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T17:08:51.446133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:08:51.446133Z digest=sha256:90824e22908a45e267e480317676c06e6314a009ee0741d859b1856faff9a134

Observation 76ef193a-40c2-48a4-8ada-3bf341c47643 · outbound

This paper cites Former-dfer: Dynamic facial expression recog- nition transformer,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Former-dfer: Dynamic facial expression recog- nition transformer,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:07.867624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:08:51.579156Z digest=sha256:d8f8bb1233719c661f9830e46bc444ee3f74d00725d716ac3c072cc97ef32068

Observation 80ec1c43-07f3-4a26-b909-4f2dc26d0eb7 · outbound

This paper cites Ex- pression snippet transformer for robust video-based facial expression recognition,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Ex- pression snippet transformer for robust video-based facial expression recognition,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T17:08:51.756681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:08:51.756681Z digest=sha256:392c86ce1db3e35dcc6927b072ecd6bcf5860d400bb43f0768087257a008e6cf

Observation 0b439d16-a4e4-443e-8a0d-f2d8140ab652 · outbound

This paper cites Freq-hd: An interpretable frequency-based high-dynamics affective clip selection method for in-the-wild facial expression recognition in videos,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Freq-hd: An interpretable frequency-based high-dynamics affective clip selection method for in-the-wild facial expression recognition in videos,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T17:08:52.037086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:08:52.037086Z digest=sha256:6a6ec9f1d857846650029846adc5605a87ee683a0b4734970889a9f38e064bed

Observation 18c7a973-d1b5-4155-985c-bc9ab4855e05 · outbound

This paper cites Facial expression recognition with adaptive frame rate based on multiple testing correction,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Facial expression recognition with adaptive frame rate based on multiple testing correction,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:07.651357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:08:52.159837Z digest=sha256:a33e37be67f5dfaa0868256a00b209752bb7bfecbbfaeea5b1bd6681c59fa5fe

Observation a028bf9e-5a31-4acf-b5d9-cbfd4792e5d4 · outbound

This paper cites Video-based facial micro-expression analysis: A survey of datasets, features and algorithms,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Video-based facial micro-expression analysis: A survey of datasets, features and algorithms,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:07.455501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:08:52.268113Z digest=sha256:ba17d198f06ff437f90de06de0314bde5a71edd32e12b65e652a8da507fe660d

Observation 7fd07951-24e8-4737-90d0-a71bd023e60d · outbound

This paper cites Dynamic facial expression recognition under partial occlusion with optical flow reconstruction,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Dynamic facial expression recognition under partial occlusion with optical flow reconstruction,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:07.145228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:08:52.428689Z digest=sha256:012bc75dd1337a5ad2b71dd24daffa40afa5d97576e02c0e49c937ec35796a4d

Observation c2d87dcf-a356-49e0-8c39-70c1a6bfed9d · outbound

This paper cites From static to dynamic: Adapting landmark-aware image models for facial expression recognition in videos,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition From static to dynamic: Adapting landmark-aware image models for facial expression recognition in videos,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:06.532422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:08:52.667666Z digest=sha256:a0458ff1b36e5eb41a118891450aad29483577ad09070fa612bbfa87e8a80cee

Observation e9feacc7-04c8-4699-b675-a01cc7560e19 · outbound

This paper cites A Survey on Facial Expression Recognition of Static and Dynamic Emotions.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition A Survey on Facial Expression Recognition of Static and Dynamic Emotions

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T17:08:52.781534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:08:52.781534Z digest=sha256:637c2c6e44c1da49a414691bf5cc2be4d517706147b739e4c7d7b8a48452cc58

Observation 0e0c819f-6ba6-426b-852d-26ab597a8c06 · outbound

This paper cites Prompting Visual-Language Models for Dynamic Facial Expression Recognition.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Prompting Visual-Language Models for Dynamic Facial Expression Recognition

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T17:08:52.937811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:08:52.937811Z digest=sha256:c521b2db6c8204780b7042a605d8518ef28e7424ec0bacdc1f0ef83206fb227b

Observation 4e4d0b94-72eb-4c57-bbbd-9f175b5202df · outbound

This paper cites Emoclip: A vision-language method for zero-shot video facial expression recognition,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Emoclip: A vision-language method for zero-shot video facial expression recognition,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:06.252373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:08:53.030094Z digest=sha256:00aa1cdeadf0bf3ef5ee55c0a3fc3eb92f0758e6aed062d41714ca14a9d129b3

Observation 964eea29-3d44-4c2b-b79f-40a7c4518bee · outbound

This paper cites Domain knowledge enhanced vision-language pretrained model for dynamic facial expression recognition,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Domain knowledge enhanced vision-language pretrained model for dynamic facial expression recognition,

Reference 25

Resolution
metadata mismatch
raw_fallback, observed 2026-08-06T17:08:59.856215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:08:53.192045Z digest=sha256:d791183e51ac4804b494ecf45f21906f2515568b2e61dfd930a2f4977f007d40

Observation cfdaf705-c9c9-4083-b02b-5a5371ec2dfe · outbound

This paper cites Enhancing zero-shot facial expression recognition by llm knowledge transfer,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Enhancing zero-shot facial expression recognition by llm knowledge transfer,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:05.963193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:08:53.286611Z digest=sha256:8dc5ff8cfd426e1e46e20ef901db607a45e0621174be351ca4522d40c46f83cb

Observation 65ef274e-e86c-4c71-a2ac-6d5b54a12094 · outbound

This paper cites Finecliper: Multi-modal fine-grained clip for dynamic facial expression recognition with adapters,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Finecliper: Multi-modal fine-grained clip for dynamic facial expression recognition with adapters,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T17:08:53.456059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:08:53.456059Z digest=sha256:74b60a3209d741ace558445f91a5e4df84a8b81ab41edc8f280f8e92dbe79451

Observation 1013a341-58d0-45a2-ba81-9eef0ddfdd18 · outbound

This paper cites Describe your facial expressions by linking image encoders and large language models.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Describe your facial expressions by linking image encoders and large language models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:05.528965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:08:53.642988Z digest=sha256:044009d02a76cb33ef97f4a5410a3f40d623666100942d3e5cfdff0ae714f2fc

Observation 4efb4a3a-9c7f-49e8-8ad4-233405174190 · outbound

This paper cites Bert: Pre-training of deep bidirectional transformers for language understanding,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Bert: Pre-training of deep bidirectional transformers for language understanding,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T17:08:53.771017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:08:53.771017Z digest=sha256:a772f5b78b6a0f438bb1c8e80875cdde1f5b6d156b8a9719a57219b5a9adb598

Observation acbe33dc-ae8f-43e1-a5a4-b2d18cd43e08 · outbound

This paper cites Hierarchical Transformers for Multi-Document Summarization.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Hierarchical Transformers for Multi-Document Summarization

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T17:08:53.883130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:08:53.883130Z digest=sha256:320011dc039ab6e4f6b3d1a9887b831ef592055787e51c72b82ca3e2d1bf2af0

Observation 725a3fc7-35fd-49bc-ac14-a37a4d0e6b46 · outbound

This paper cites HDT: Hierarchical Document Transformer.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition HDT: Hierarchical Document Transformer

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T17:08:53.989568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:08:53.989568Z digest=sha256:ef44a5b36d840d462c6f8ec115dc775816d20ef7030aceb216a059e80c51414b

Observation ce50ff7e-9644-4481-b55c-2d587c8e3144 · outbound

This paper cites Multi-task learning of hierarchical vision-language representation,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Multi-task learning of hierarchical vision-language representation,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:05.349116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:08:54.100810Z digest=sha256:d5e85ed3de80233e4b7d7975498079eac6f091cc837b541f4c8dbea936da2926

Observation 0ca66b65-6b30-4633-b844-1cc8f35be18f · outbound

This paper cites HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T17:08:54.214011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:08:54.214011Z digest=sha256:fab968090e437e2f62e42d579fa4eab976d978cbc3feb0858e50f560c0e55bd1

Observation 242d3e97-996a-4f4f-86b9-67bcf9fb00d2 · outbound

This paper cites Hierarchical modular network for video captioning,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Hierarchical modular network for video captioning,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:05.135439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:08:54.323924Z digest=sha256:0e5dbbd9a61ce140b8e91b7323ea403654b66d2f50cebb6fd6ad2d0d422b03ea

Observation cb83ea55-fd32-41b7-87bb-c41d4f594515 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Learning transferable visual models from natural language supervision,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T17:08:54.486076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:08:54.486076Z digest=sha256:c395755edda0d356b8feef85360f230fa1a9d159bfa7eb16d2d9fb935737f9b1

Observation 190d69ec-35eb-46a9-8f1f-3895e2b534c9 · outbound

This paper cites Ceprompt: Cross-modal emotion-aware prompting for facial expression recognition,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Ceprompt: Cross-modal emotion-aware prompting for facial expression recognition,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:05.672039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:08:54.620430Z digest=sha256:0d0e226e871ee5f66afccfd8b91c2d27a0aaa4509f31d5197ead5a583d77f659

Observation 06cef10d-69f3-47cb-a657-62d3cab51910 · outbound

This paper cites Sinkhorn distances: Lightspeed computation of optimal transport,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Sinkhorn distances: Lightspeed computation of optimal transport,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:04.917377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:08:54.819539Z digest=sha256:75b07aba16258df71d5cbc77cb7304b7d165adaa381645e38648610c4156c844

Observation 5110dfd2-d14c-4c5c-9502-d00fe488b898 · outbound

This paper cites Recent advances in optimal transport for machine learning,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Recent advances in optimal transport for machine learning,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:04.570793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:08:54.948095Z digest=sha256:427dd1da349f88d6bab635f05a51b0862986bf08a101ca4e6354ea8c151bc5b8

Observation 7cb898f7-e789-4677-a77b-f6aee9e2dcfe · outbound

This paper cites Reliable weighted optimal transport for unsupervised domain adaptation,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Reliable weighted optimal transport for unsupervised domain adaptation,

Reference 41

Resolution
malformed identifier
no resolver link, observed 2026-08-06T17:08:55.068589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:08:55.068589Z digest=sha256:b33fc86f737be655bb8ebd4f205b13b47c2a0bfc2095f46cf38e73a7d91cc707

Observation ff886e30-8ed9-46c3-8ff7-4223c9bcf03d · outbound

This paper cites Unsupervised learning of visual features by contrasting cluster as- signments,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Unsupervised learning of visual features by contrasting cluster as- signments,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:04.349557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:08:55.184879Z digest=sha256:94e30a8bde27c608da38fa3367c794254eb746cd0c8818e9f4d0ac4c6869adad

Observation 8550292e-d670-4762-98d3-edd840d6d036 · outbound

This paper cites Optimal partial transport based sentence selection for long-form document matching,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Optimal partial transport based sentence selection for long-form document matching,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:04.049201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:08:55.298606Z digest=sha256:2f13369040dfbd06944474210944400be74367f4fcad6d0628c2a61fa0c98846

Observation ee3bd80c-b722-45b9-8bfb-73909537e983 · outbound

This paper cites Learning to align sequential actions in the wild,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Learning to align sequential actions in the wild,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:03.796463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:08:55.446654Z digest=sha256:30d82f4068c4ac4f4915cfb0e083102a63694f9169efcc9efe776af22d493e6b

Observation 55dcb424-4071-4d6f-abbd-814214ae3886 · outbound

This paper cites What when and where? self-supervised spatio-temporal grounding in untrimmed multi-action videos from narrated instructions,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition What when and where? self-supervised spatio-temporal grounding in untrimmed multi-action videos from narrated instructions,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:03.611461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:08:55.583412Z digest=sha256:3960823f39bb7d57cf2c8b951b604e5c76a2e5513d2767ed3df3264173173b4b

Observation a0a922a5-baee-4e19-869d-6cc760f1c08f · outbound

This paper cites Multi-granularity Correspondence Learning from Long-term Noisy Videos.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Multi-granularity Correspondence Learning from Long-term Noisy Videos

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T17:08:55.704482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:08:55.704482Z digest=sha256:428792977844ee45db74743d274d91f3db6c2694afa8deaa6e0937cf79677eb5

Observation 3acb2635-48d1-4bd5-8d6f-82a43a92f8d2 · outbound

This paper cites Spatial- temporal graphs plus transformers for geometry-guided facial expression recognition,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Spatial- temporal graphs plus transformers for geometry-guided facial expression recognition,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:06.842125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:08:55.829988Z digest=sha256:c3a20fb9d934de2adedbe7116a9d6a0f7f12838f6e204e8f2af3c958d19fda8d

Observation 22258bd3-e252-4538-b727-1377306c1864 · outbound

This paper cites Multimodal transformer for unaligned multimodal language sequences,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Multimodal transformer for unaligned multimodal language sequences,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:02.956146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:08:56.109359Z digest=sha256:96e8841d535ed83e89d92b9636c2a982b2bba40a6040b6c0a4253e83bdc2a46d

Observation e3b68206-aaa7-440c-8809-87610c2582a6 · outbound

This paper cites Vilt: Vision-and-language transformer without convolution or region supervision,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Vilt: Vision-and-language transformer without convolution or region supervision,

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T17:08:56.248682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:08:56.248682Z digest=sha256:497ac08fa233b4028ec1184b506e828ac53a8a1ce47c66d74b5a131b32a13f1f

Observation adc6628f-859d-410f-b0bc-9ed5bdce20f6 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T17:08:56.385574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:08:56.385574Z digest=sha256:667eab101e9ca61c494f08f99f526adb26b24e1c982119acb86e3a3e6eed08af

Observation 91adfa9d-5d95-413a-b0dc-543745217bdc · outbound

This paper cites Cliper: A unified vision-language framework for in-the-wild facial expression recognition,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Cliper: A unified vision-language framework for in-the-wild facial expression recognition,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:02.695424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:08:56.546603Z digest=sha256:b22baa29b6406a5cbc33338dee5d3f3de0ec973c8e9177c5cbf4448d5c20abf9

Observation a771a906-1941-46cd-9384-d776e716758d · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T17:08:56.717302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:08:56.717302Z digest=sha256:1e42d431c9fa59af5a9c4cc048462fbc8417000d920d13faef19ee02794d2542

Observation 2ef3652f-b4ba-4a6a-9be8-6020b5bdcb3b · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Exploring the limits of transfer learning with a unified text-to-text transformer,

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T17:08:56.854748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:08:56.854748Z digest=sha256:ed87a4979b5e71b5a506c53674e5420731bde6594dabd9290a77993e48d38e2a

Observation 0ff0bbb6-eb8a-41d8-ba58-e1eb879b040d · outbound

This paper cites VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T17:08:56.989600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:08:56.989600Z digest=sha256:b50b8764594283302cfa2fe355b9962fbb6757b246f2c14343ff9569cdb9868c

Observation d7fdddc8-f420-4089-a31d-40ca86630da1 · outbound

This paper cites Videomae: Masked autoen- coders are data-efficient learners for self-supervised video pre-training,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Videomae: Masked autoen- coders are data-efficient learners for self-supervised video pre-training,

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T17:08:57.095778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:08:57.095778Z digest=sha256:6ad87f9c785c9790f066ca0b23cb398084166ca25cbec17390a7cf90da28212d

Observation c0f13447-3841-4612-bf12-eab2e9546562 · outbound

This paper cites Mae-dfer: Efficient masked autoencoder for self-supervised dynamic facial expression recognition,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Mae-dfer: Efficient masked autoencoder for self-supervised dynamic facial expression recognition,

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T17:08:57.212905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:08:57.212905Z digest=sha256:b581920e8ea09105a008252941022c10c1d9e1025b5e54165880c01e2ebbeb93

Observation 40c53bc4-6176-455b-b879-1db2e02025b8 · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image cap- 14 tioning,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image cap- 14 tioning,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:02.372262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:08:57.346343Z digest=sha256:7d57adc93d8a85f4474a53cd3111117f046cff43b22a1b61469e3e84cd6b9700

Observation 0d0349cd-faf4-4122-9945-7563471fbdbb · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to-end retrieval,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Frozen in time: A joint video and image encoder for end-to-end retrieval,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:02.133934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:08:57.491789Z digest=sha256:59b9f18aab4019b8b050e99068171ba75634e0e40db62313ed34687424d0c9d5

Observation c3208862-0942-423d-9cbc-8b85e139500c · outbound

This paper cites Emotion recognition using imperfect speech recognition,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Emotion recognition using imperfect speech recognition,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:01.900774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:08:57.635608Z digest=sha256:2241a96841e6c103acf3a07c66547d7a918c618185cc3dd7242b9d11410b2f12

Observation 0edd08c9-2483-4ae2-b727-2e19eecae091 · outbound

This paper cites Posterior calibration for multi- class paralinguistic classification,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Posterior calibration for multi- class paralinguistic classification,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:01.577572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:08:57.805680Z digest=sha256:13784c2c704e437f3e05c2e503437dd6cbf4b33e064b02123e6b012f3734bea7

Observation b06a9716-7b97-40fa-a894-d07f957a2a4c · outbound

This paper cites Rethinking the learning paradigm for dynamic facial expression recog- nition,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Rethinking the learning paradigm for dynamic facial expression recog- nition,

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:01.317518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:08:57.917571Z digest=sha256:2c5a2ca055ce8ac3839cbaab6193d00757c6a21c7cedc990bc792c1bf4e73870

Observation e9fc0632-63cf-4362-9066-95f90e68fbd1 · outbound

This paper cites A$^{3}$lign-DFER: Pioneering Comprehensive Dynamic Affective Alignment for Dynamic Facial Expression Recognition with CLIP.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition A$^{3}$lign-DFER: Pioneering Comprehensive Dynamic Affective Alignment for Dynamic Facial Expression Recognition with CLIP

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T17:08:58.053197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:08:58.053197Z digest=sha256:fc4e2bd5530cf0a111c95dfb29a8da3bb3070aee50487bc76c2cd8e28ba9e149

Observation bb674402-af34-407a-86c6-4a7b01908490 · outbound

This paper cites Clip-aware expressive feature learning for video-based facial expression recognition,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Clip-aware expressive feature learning for video-based facial expression recognition,

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:01.028233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:08:58.167769Z digest=sha256:eafda55f11de2a5f8f18c34b2ffe21e10d48715b6c63a4c33b671a6e84166f57

Observation 5c53bea4-72b1-4926-8ec7-ed45a3ceb9b8 · outbound

This paper cites NR-DFERNet: Noise-Robust Network for Dynamic Facial Expression Recognition.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition NR-DFERNet: Noise-Robust Network for Dynamic Facial Expression Recognition

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T17:08:58.322047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:08:58.322047Z digest=sha256:315633e539d369d3a4a5e9c9ab2d3c147d6086a1f020e322a11a0b53e35a358c

Observation c01291c4-fcdc-4459-8ccd-dd3fbabce61f · outbound

This paper cites Logo-former: Local-global spatio-temporal transformer for dynamic facial expression recognition,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Logo-former: Local-global spatio-temporal transformer for dynamic facial expression recognition,

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T17:08:58.457153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:08:58.457153Z digest=sha256:56e594d6c59e673d4afaccb5423b2d85d85b73883012abff1a5aa05a45090b1d

Observation 025acfcc-723b-4024-9ea7-4130c4153e70 · outbound

This paper cites Intensity-aware loss for dynamic facial expression recognition in the wild,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Intensity-aware loss for dynamic facial expression recognition in the wild,

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:03.269880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:08:58.592846Z digest=sha256:fbe5ad4ae6b97a909c2b7cca56f4711b6a3f40ced47dbcfb31fa44f85db5fb85

Observation 44e6292c-dc8c-490b-8ef8-c2126613213e · outbound

This paper cites Transformer-based multimodal emotional perception for dynamic facial expression recogni- tion in the wild,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Transformer-based multimodal emotional perception for dynamic facial expression recogni- tion in the wild,

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:00.714047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:08:58.705759Z digest=sha256:af4d235e954ebfe4a88d115776acf6cd64ac58f51e4b651b7479a445cda4a684

Observation 6d6c81c3-f280-4671-9ef9-d9f219746c94 · outbound

This paper cites Svfap: Self-supervised video facial affect perceiver,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Svfap: Self-supervised video facial affect perceiver,

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:00.483481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:08:58.844014Z digest=sha256:e4ac8813443def3993718702b944c15712a5fbda874076b6d326d6f16bcea1ce

Observation 56f4fb15-b13d-4326-8e4b-6bd0070b72bd · outbound

This paper cites Hicmae: Hierarchical contrastive masked autoencoder for self-supervised audio-visual emotion recogni- tion,.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition Hicmae: Hierarchical contrastive masked autoencoder for self-supervised audio-visual emotion recogni- tion,

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:00.250849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:08:58.980743Z digest=sha256:e2d9941048b827076023828ef684803e54b1feb09e5854b28db52e93f61c999a

Pith citing papers

No inbound Pith citation observations are available.