Pith. sign in

Paper Citation Record · LEDGER

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness

As of 14 August 2026, this Paper Citation Record lists 95 of 95 outbound references and 4 inbound Pith citation observations for arXiv:2501.07978.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.07978 v1

Coverage vector

measured 95 of 95 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T20:33:54.320561Z

measured 99 of 99 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:24:36.541729Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-23T06:02:37.644844Z

Reference resolution

95 of 95 outbound references displayed

  • verified exact0
  • verified fuzzy43
  • unresolved52
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cc93ae32-cc76-4acc-ba60-cbec4923ba3e · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:53.884278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:53.884278Z digest=sha256:b3868b291f8e2503a0a7ee695188109f072b6b1953a86dcc2cf6f1ea27b2eb57

Observation 4abc8b02-edbf-4fe1-a5c3-b75fdb296c4a · outbound

This paper cites Emotion recognition in speech using cross- modal transfer in the wild, 2018.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Emotion recognition in speech using cross- modal transfer in the wild, 2018

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:53.889385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:53.889385Z digest=sha256:91b27bba7093ea3c541b517e670eba59f67d9b8c621691787166901bcaaded06

Observation cdd82447-b527-47b6-8bcc-f1a8e953315f · outbound

This paper cites Claude-3.5, 2024.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Claude-3.5, 2024

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:53.893802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:53.893802Z digest=sha256:bb6f3d5c6af366d1089aef1b3d723f77154b9643723adc284dd4f6cf1f615e42

Observation d6f6185c-f02d-4cf3-a693-a66927114450 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:53.903232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:53.903232Z digest=sha256:1205c9673a275436d4fdb6b524a8bf121a1d38dd6e6d070d09c664cbb9cfccfe

Observation 8cbe17d3-6f5a-421c-ba14-585c11660183 · outbound

This paper cites Collecting highly paral- lel data for paraphrase evaluation.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Collecting highly paral- lel data for paraphrase evaluation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:53.907800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:53.907800Z digest=sha256:3fa49f1ddff605e212b57e2457448a04e064cabe7577ed9ec30a05344e71b2b4

Observation 14417cbe-13db-43dc-8f44-377172cdc49b · outbound

This paper cites FineCLIPER: Multi-modal Fine-grained CLIP for Dynamic Facial Expression Recognition with AdaptERs.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness FineCLIPER: Multi-modal Fine-grained CLIP for Dynamic Facial Expression Recognition with AdaptERs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:53.913411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:53.913411Z digest=sha256:93ba61bd7b0fd69df0e8c660a7cd65a2056a3ba50d54bf8c1ebf645b5ad03f0d

Observation 7ade2e87-b39e-4fb9-a3b9-2a3bacd370d0 · outbound

This paper cites ShareGPT4Video: Improving Video Understanding and Generation with Better Captions.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:53.918678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:53.918678Z digest=sha256:037368356808d341e2724488add849ba5250a1272e962770ed116dee87e379e3

Observation a70c08be-1453-477b-a074-daa4cfb9a060 · outbound

This paper cites Stcam: Spatial-temporal and channel attention module for dynamic facial expression recognition.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Stcam: Spatial-temporal and channel attention module for dynamic facial expression recognition

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:53.923207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:53.923207Z digest=sha256:a012a2d93edc1bb7a1a3adf0b265dcca65a42641b08145803c8228f119a5963f

Observation 058f14fb-70b5-4ece-a08c-26afc70ea1c0 · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:53.927345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:53.927345Z digest=sha256:f81231c015497ade981c891660ba7ca906e8c64135ada0c904ce2f8f0074661a

Observation 8d653a64-b7f1-4921-926f-8b741118b8ba · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:53.930933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:53.930933Z digest=sha256:50bb3bab5554d3a3fd7d95ec6f710408835e2f5b29636da6a4d23f55d059b72e

Observation 184a51db-02cc-49ff-a15b-11671cca3d01 · outbound

This paper cites Transface: Calibrating trans- former training for face recognition from a data-centric per- spective, 2023.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Transface: Calibrating trans- former training for face recognition from a data-centric per- spective, 2023

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:53.934690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:53.934690Z digest=sha256:da91db4bfa538c4717ea48959a9f4c5c48beb88749f85465ea6953dc7a9724d3

Observation 70be547e-fdad-44e1-8463-20e460e9a722 · outbound

This paper cites Diffusionrig: Learning personal- ized priors for facial appearance editing.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Diffusionrig: Learning personal- ized priors for facial appearance editing

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:53.937932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:53.937932Z digest=sha256:78a93bb45f1487dad153906fcff8335dabf427d1587467ed5aef54d1310d9656

Observation 868dbd0f-f384-43a1-8d93-af5d185986ac · outbound

This paper cites InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:53.942983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:53.942983Z digest=sha256:feda528d59930f4698ea3fda2544813f9caca97a13d69aaeedfa93280519d5c0

Observation 6236c81b-5727-4fb4-8b60-8bc7774588a9 · outbound

This paper cites Strongsort: Make deep- sort great again.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Strongsort: Make deep- sort great again

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:53.947008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:53.947008Z digest=sha256:a86cbb3b3f2dfc868d18cfded18c50f7bf9716e1601a76d734877415a20c86e1

Observation 17472c72-cec0-40a2-99b2-84c691b507e7 · outbound

This paper cites EmoCLIP: A Vision-Language Method for Zero-Shot Video Facial Ex- pression Recognition.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness EmoCLIP: A Vision-Language Method for Zero-Shot Video Facial Ex- pression Recognition

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:53.950534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:53.950534Z digest=sha256:a82ed15f8b97912a5947f7c7df6787d15c5587603c1b01af62a31019d82be9be

Observation 8e20eb0f-a12a-465a-9847-5d95dbc2c3b4 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:53.954454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:53.954454Z digest=sha256:5611110561dd6cc38b692fa36a18d79793557a2354015da63ea632e25e649a52

Observation 2ee3927a-96fb-42f8-b2fd-123ecbc2904c · outbound

This paper cites LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:53.958465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:53.958465Z digest=sha256:0960bc9d521c6805f97e25d3b176ede357c00c069cbc15cf7fc46196b1682a94

Observation 289d0ca9-4987-45a4-9795-01f08b933ada · outbound

This paper cites Music Emotion Recognition: Toward new, robust standards in personalized and context-sensitive ap- plications.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Music Emotion Recognition: Toward new, robust standards in personalized and context-sensitive ap- plications

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:53.963217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:53.963217Z digest=sha256:52ff76f1606274883c3133fd9faf67f77bc33f84fcb7f02418f3e8cfa9c85eec

Observation b69d2990-0a6e-4675-a16e-24c27a8da406 · outbound

This paper cites LoRA: Low-rank adaptation of large language models.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness LoRA: Low-rank adaptation of large language models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:53.967225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:53.967225Z digest=sha256:2ee3ee1dced305677da8183db5c09ec04040026b10c3cb4825d7c0e52a0a1ad2

Observation b7602365-cb26-4dfd-ba34-4fa8ff056ba9 · outbound

This paper cites Multimodal Pretraining for Dense Video Captioning.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Multimodal Pretraining for Dense Video Captioning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:53.972371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:53.972371Z digest=sha256:6a73a9cc68d87915b6eb3e4950a216b6a430a28739cc0368709e1f2fce274ce9

Observation 4d501ff5-cfb1-45a9-8ef6-2707817343ae · outbound

This paper cites Video ReCap: Recursive Captioning of Hour-Long Videos.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Video ReCap: Recursive Captioning of Hour-Long Videos

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:53.977080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:53.977080Z digest=sha256:f20bd6ad9a71778df0ca30ed8c6f515f63fc931f8b44b27a317539a047d53b73

Observation 33d0f7e5-2d02-4c8f-9663-7bed3fa6c57e · outbound

This paper cites an unresolved cited work.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:53.982395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:53.982395Z digest=sha256:9dbbc773f5830b93010a69b95a47d39ba258492ff3ca3ef73810a8d12cbdc0b2

Observation 0d82b7f6-3584-472c-9619-29f1ddd9bf3a · outbound

This paper cites Dfew: A large-scale database for recognizing dynamic facial expres- sions in the wild, 2020.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Dfew: A large-scale database for recognizing dynamic facial expres- sions in the wild, 2020

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:53.986831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:53.986831Z digest=sha256:86460ea2653afe84d46d90ade8f80412da5354c8602738183710150fbfda654c

Observation ac506810-02a2-41f5-af59-04498bbaecb6 · outbound

This paper cites Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:53.991893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:53.991893Z digest=sha256:9536b083ed0ebf4b304fc1b2fe3500b9391117c28dbcd1899f0bf3980b3f4fc2

Observation b19a64f9-dbfe-445d-8e6b-174955e70bb1 · outbound

This paper cites Expression, affect, action unit recognition: Aff-wild2, multi-task learning and arcface, 2019.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Expression, affect, action unit recognition: Aff-wild2, multi-task learning and arcface, 2019

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:53.996900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:53.996900Z digest=sha256:8c3a2d5dc303f65f01dfd1d6b4c6a47db52d4bed4b71cbb00fdcce34bef6fa7a

Observation 49941daa-11fa-4191-86d9-ffe12cf973be · outbound

This paper cites Afew-va database for valence and arousal estimation in-the-wild.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Afew-va database for valence and arousal estimation in-the-wild

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:54.000691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:54.000691Z digest=sha256:cce49fc3ae851ac29737b2e0cc8c7f5e68cf9bd8e4c26aac1decd0c09961546b

Observation 61377989-453c-4780-bcc3-f74d63bdc9c6 · outbound

This paper cites Dense-captioning events in videos.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Dense-captioning events in videos

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:54.005726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:54.005726Z digest=sha256:860ae271d35c0f3e2f1bab27bd8d2348cc33b66d68e8121b68816c085bd334c7

Observation b64a0f55-59a7-4ea3-b690-8dfe4fe2cbb8 · outbound

This paper cites Context-aware emotion recognition net- works.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Context-aware emotion recognition net- works

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.430420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:33:54.016846Z digest=sha256:b463f3d9f4081d446de5063580df972335993ff2ffa69ec8433930273447a07d

Observation c2632afb-5b1a-4883-b728-12af4c6b4f4e · outbound

This paper cites Llava-next: What else influences visual instruction tun- ing beyond data?, 2024.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Llava-next: What else influences visual instruction tun- ing beyond data?, 2024

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.416512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:33:54.021962Z digest=sha256:cc8d6be56236a3cba72c82166a2fd8780adb1e6a430231fbff5e1029f353dd17

Observation 7cff97f2-83c7-460b-8ba4-38f3988fdfee · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness LLaVA-OneVision: Easy Visual Task Transfer

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:54.026558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:54.026558Z digest=sha256:e1b3699092de069f3eb3aa586a825a08e34b12353beaf9ab091873d80e287494

Observation e26760f4-3fcf-4b4e-ac56-e16b76841d14 · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:54.031175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:54.031175Z digest=sha256:e98b75d27ab2e1e5775f8a013798b5da6859e3f9fd67dee6ff363c18b3ec1407

Observation 7f8c1543-a945-4910-af7e-409144b7f88f · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:54.035535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:54.035535Z digest=sha256:a401e3a375df0565ddad4fae188f46ee3c18b96e25c5fc03e2274e1e2fb0eab6

Observation 2e33504e-cdb5-4ea3-9099-cc2a3bfd6e84 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness VideoChat: Chat-Centric Video Understanding

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:54.040489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:54.040489Z digest=sha256:8b2bbda125e565cde03c1b539e1afe669ba6537dc0952c770dfc7940c807f59b

Observation 9094410c-9abc-4c79-bf0a-6f01cc769b6e · outbound

This paper cites Dual-sti: Dual-path spatial-temporal interac- tion learning for dynamic facial expression recognition.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Dual-sti: Dual-path spatial-temporal interac- tion learning for dynamic facial expression recognition

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.395528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:33:54.045144Z digest=sha256:080701de29c81385f9c899a485fb01ab286bc6fe006d1dc4bd9c5705b010c0e9

Observation b5f712da-2f25-4616-844c-628f256cccd3 · outbound

This paper cites Facial affective behavior analysis with instruction tuning, 2024.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Facial affective behavior analysis with instruction tuning, 2024

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.384319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:33:54.049977Z digest=sha256:337f4b5052350510c5db875bcaa0ffd7f5bc7744b37ca75306b022868e32a6fc

Observation 88837485-f367-4190-a78b-1cc280f44245 · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Llama-vid: An image is worth 2 tokens in large language models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.373429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:33:54.054027Z digest=sha256:b340041f00720725a8841ca08495f895efb65db0f2bf1e111238a53a4b4fda28

Observation 231af9a6-5224-4f1f-9aa0-310260037293 · outbound

This paper cites Photomaker: Customizing realistic human photos via stacked id embedding.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Photomaker: Customizing realistic human photos via stacked id embedding

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.358412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:33:54.058317Z digest=sha256:e9c8e3d2907e8a6e74cbfb43bd41489f979293374efa17fce9be704c2d3e3ca6

Observation fa8073b5-3bd1-43cd-963c-2284534b242b · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:54.063000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:54.063000Z digest=sha256:b95c6d76f171e4c27ca97bc28dd531c9e3b3f47e974b303beda08018104bf3b5

Observation c2d2eb12-c3d7-4f98-85e0-64962d28268b · outbound

This paper cites Saanet: Siamese action-units attention network for improving dynamic facial expression recogni- tion.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Saanet: Siamese action-units attention network for improving dynamic facial expression recogni- tion

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.342396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:33:54.068366Z digest=sha256:081869b0836e0d5eebd21231a0d29160eaeb7ccdddc7d50daba8ca377a24dddf

Observation b1ae5ed9-7d2a-466e-baf4-923d799a0625 · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Improved Baselines with Visual Instruction Tuning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:54.072952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:54.072952Z digest=sha256:711704f760ef2ed8f0d3b63f9f448073d489d7da2c43c6b302442181abddcba2

Observation deaaf7f4-3e1b-4922-a469-7e7c51f095d5 · outbound

This paper cites Visual instruction tuning.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Visual instruction tuning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:54.079056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:54.079056Z digest=sha256:da52cb11eccb6b103ef2e56580684967b69671cc30fee1d5f1bcfdf65f41c12a

Observation 2a8b8f70-e38a-45c8-8359-777e250eb065 · outbound

This paper cites World Model on Million-Length Video And Language With Blockwise RingAttention.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness World Model on Million-Length Video And Language With Blockwise RingAttention

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:54.083552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:54.083552Z digest=sha256:c3c7e26ac2fbaee81cd06941a4b7d6df65101addce5b3ce64e9d8fd39833a431

Observation c6196b34-0323-4b0c-a59a-c4796f0303b6 · outbound

This paper cites BT-Adapter: Video Conversation is Feasible Without Video Instruction Tuning.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness BT-Adapter: Video Conversation is Feasible Without Video Instruction Tuning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:54.087918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:54.087918Z digest=sha256:3474cff3e3d9a681817aef4564313ee00cb25281a18d87cd9444069624e60ff2

Observation 4d075e73-3a34-4aa7-874c-120ee1be9681 · outbound

This paper cites ST-LLM: Large Language Models Are Effective Temporal Learners.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness ST-LLM: Large Language Models Are Effective Temporal Learners

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:54.092524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:54.092524Z digest=sha256:dbf93682462b91edb5d81c4d758578a3d6f52206680db58520d6cb0e15de6b78

Observation 803469f6-deef-4297-b87b-92833edf75ad · outbound

This paper cites Mafw: A large-scale, multi-modal, compound affective database for dynamic facial expression recognition in the wild, 2023.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Mafw: A large-scale, multi-modal, compound affective database for dynamic facial expression recognition in the wild, 2023

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.321731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:33:54.098492Z digest=sha256:c45cfe5e0fae2606cfc99f1a7eb36d5ce9421f05eb2373bc23757081e81b7892

Observation e8f43fe3-bbf1-4cc1-bce9-5e410716a276 · outbound

This paper cites DamoFD: Digging into backbone de- sign on face detection.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness DamoFD: Digging into backbone de- sign on face detection

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.309781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:33:54.102402Z digest=sha256:4c51ef2b7aadab8fb520875e7fa3454b9ca9889c9f902bd472be061d277cb263

Observation 1a21361f-515c-4376-a7f1-c07d45c73d82 · outbound

This paper cites Livingstone and Frank A.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Livingstone and Frank A

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.297610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:33:54.109723Z digest=sha256:e317e80d513185947850c201859e130a634a63cd68c874e0c7a80a0b4de0f076

Observation c9ccb828-d05c-4b47-ae5e-5b61b57bade5 · outbound

This paper cites Cohn, Takeo Kanade, Jason Saragih, Zara Ambadar, and Iain Matthews.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Cohn, Takeo Kanade, Jason Saragih, Zara Ambadar, and Iain Matthews

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.286051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:33:54.114235Z digest=sha256:7825701c0caee05d4d194d4ef1ccddc6a0d0b121456e4f8d8facd44305bdb432

Observation 2d697a4f-8b27-4320-b587-e8742ea614d5 · outbound

This paper cites The extended cohn- kanade dataset (ck+): A complete dataset for action unit and emotion-specified expression.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness The extended cohn- kanade dataset (ck+): A complete dataset for action unit and emotion-specified expression

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.273230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:33:54.118502Z digest=sha256:15cbb02d41daacb5738b6a16da0e1fb94583b41c07ad3e16219e066756976227

Observation b5c470e0-abcf-40aa-ba94-8c820facbbd1 · outbound

This paper cites Learning multi-dimensional edge feature- based au relation graph for facial action unit recognition.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Learning multi-dimensional edge feature- based au relation graph for facial action unit recognition

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.257974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:33:54.123016Z digest=sha256:6955107bd67dbe4d020ddd97fe9a5333ddd9ff04ac42fb5b30d1b772fe5d9875

Observation d0e643a4-7903-4f7c-91ee-b39f6e6d84e9 · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:54.127444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:54.127444Z digest=sha256:336a3c995cf210c2caf107731c2af62b9385e06298b4e2163928fa8047187edc

Observation 8ffd7c41-e2d5-4641-a30c-47452385392f · outbound

This paper cites Video-chatgpt: Towards detailed video 10 understanding via large vision and language models, 2024.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Video-chatgpt: Towards detailed video 10 understanding via large vision and language models, 2024

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.246121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:33:54.131730Z digest=sha256:c461967cbae13bc0319b7c65a1430b4233ff4d41809f2a0ba915a40547d475ab

Observation e82a982a-ec88-45ab-9a48-d0a347527cc1 · outbound

This paper cites Egoschema: A diagnostic benchmark for very long- form video language understanding.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Egoschema: A diagnostic benchmark for very long- form video language understanding

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:54.135546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:54.135546Z digest=sha256:3646288547f9d5bb8efafae35fff901849a5a6138e507af4733a63838274ff6a

Observation 710122ae-8fba-4ba8-aa6f-3f5913cf3f7b · outbound

This paper cites The importance of emotional regulation in mental health.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness The importance of emotional regulation in mental health

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.227238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:33:54.139183Z digest=sha256:95690056803dcef431f14779c8aeadece7c9ee45e34e3b00a83c45e4bbfbaeaf

Observation 719aa5fe-2e0a-4e13-9845-7ea75642b62f · outbound

This paper cites FaceXFormer: A Unified Transformer for Facial Analysis.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness FaceXFormer: A Unified Transformer for Facial Analysis

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:54.143445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:54.143445Z digest=sha256:dd8b6418c6bcdeb77cd949fde5a6b9090db92277558774a845d25ba00e90c26e

Observation 9f37c381-ba5c-4eb6-8586-9ccfaf984da2 · outbound

This paper cites Repre- sentation learning and identity adversarial training for facial behavior understanding, 2024.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Repre- sentation learning and identity adversarial training for facial behavior understanding, 2024

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.215683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:33:54.148000Z digest=sha256:2b482df678fd2afb52c15cb2b4b847488fee616c6b6a647e383d09e493db7ef8

Observation d662422b-4de8-4bca-bc29-6399216d3fe9 · outbound

This paper cites an unresolved cited work.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:33:55.202907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:33:54.152472Z digest=sha256:fe9c55c346c65e96068038ee9c81321243260461e2adb80de88104ad50e488cd

Observation f6f9678f-c6e2-4ba8-b2e7-9df8a433944e · outbound

This paper cites Gpt-4v(ision) system card.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Gpt-4v(ision) system card

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.189876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:33:54.156778Z digest=sha256:bb98f45b373ade5e0b9b5c79a65f4aee023a0fb38fadb295b366eee202918857

Observation feebdaca-2d58-4d68-b35f-94a596c1ce9d · outbound

This paper cites Gpt-4 technical report, 2023.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Gpt-4 technical report, 2023

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.178240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:33:54.160875Z digest=sha256:4b2d8c12bcef07fd3be197dd9554d722e5f7c39168d8b3237f0b9f907259b19b

Observation 68ab66d2-4688-4f8d-8a79-199d155881f5 · outbound

This paper cites Gpt-4o system card, 2024.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Gpt-4o system card, 2024

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.166879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:33:54.167146Z digest=sha256:a963a2d11a8e39f3cc78ddfde4b9a807f2f8b11387a21c1815cbb2b8fd1f700b

Observation fed2f313-d31c-47f1-a744-9a6b3e0308f9 · outbound

This paper cites Digihuman: A con- versational digital human with facial expressions.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Digihuman: A con- versational digital human with facial expressions

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.152984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:33:54.171845Z digest=sha256:e50849fabceb0642717f8d121435e43c21fe9d835a05ffb04a29bc68fa687703

Observation 1d57b4a1-9157-4f8f-a53c-d32938b2a18e · outbound

This paper cites Pantic, M.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Pantic, M

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.140334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:33:54.176751Z digest=sha256:c0754b45bf3a23660de9835fcebb57a24c0f9e407b14af43caa7e2729a2473f9

Observation 5af1b47d-f09a-4c75-9ece-bce449cfd8f5 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Learning transferable visual models from natural language supervi- sion

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:54.181114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:54.181114Z digest=sha256:c773bbf13d03deb0c246c233bfa598f4db5087f0f06e0808ac1daf04b447e7c1

Observation 113794d7-32e1-42f4-adb7-8d05ccabfe0a · outbound

This paper cites Movie description.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Movie description

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.121966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:33:54.186009Z digest=sha256:c501372b265c1222da72f3ad1b7c69405e629b3f589b0e4ec75fb4f82daa7014

Observation d423450c-b4bf-45a3-a3a2-63e59e6853fb · outbound

This paper cites Multi-view dynamic facial action unit detection.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Multi-view dynamic facial action unit detection

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.111532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:33:54.190865Z digest=sha256:98a6a98c7d308381f3bb20fd0cb33d79c07a1b2cb8fc4349c0032239bca07a80

Observation adc5adc1-c310-4d96-94c4-3e30f01709c8 · outbound

This paper cites Deep adaptive attention for joint facial action unit detection and face alignment.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Deep adaptive attention for joint facial action unit detection and face alignment

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.098590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:33:54.195507Z digest=sha256:60b056dc2e6cb5cddf79cce52658b3e8033db9d855b712f790e601ba787d4436

Observation d57afeef-1816-4ffa-ab42-5dd4bc8d113b · outbound

This paper cites Driver’s emotion and behavior classification system based on internet of things and deep learning for advanced driver assistance system (adas).

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Driver’s emotion and behavior classification system based on internet of things and deep learning for advanced driver assistance system (adas)

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.084813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:33:54.200363Z digest=sha256:f4798687ccf565ba3a3f637109583d333f72a5c41bf54d90b77de60aedb5c2f9

Observation e9a31929-d21b-4728-9114-4d844a9aa896 · outbound

This paper cites Gemini: A family of highly capable multi- modal models, 2024.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Gemini: A family of highly capable multi- modal models, 2024

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.073989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:33:54.204861Z digest=sha256:25ae5ae829e1041d63c3a85940a99bc19fdc54e6301ebe3e1c9fb3d11bb4bdca

Observation 2954039a-fbb4-45d5-a87f-b04dd8aaa782 · outbound

This paper cites Qwen2.5: A party of foundation models, 2024.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Qwen2.5: A party of foundation models, 2024

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.063870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:33:54.209214Z digest=sha256:827ac2f10370add21a7941ad358f311114f920c841acc569d5afd25d7bb1953a

Observation e2ee8a76-9ba8-42df-b627-373ccf0dfbdd · outbound

This paper cites Induced disgust, hap- piness and surprise: an addition to the mmi facial expres- sion database.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Induced disgust, hap- piness and surprise: an addition to the mmi facial expres- sion database

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.053978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:33:54.213735Z digest=sha256:9e017b29f94b0eacde390947ed69b6f57aba0c4a6c71786c30385fd9c92e4634

Observation 96b16d77-d1c1-42b2-9937-6c6580196ac4 · outbound

This paper cites Cider: Consensus-based image description evalua- tion.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Cider: Consensus-based image description evalua- tion

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.042966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:33:54.218449Z digest=sha256:8ecd7a936c7a11055660bdd4b31243fbdc074cd8b28d13dd42b388cc602d241b

Observation fe9dd5c5-1566-4bc3-808d-f1c7e92a8f29 · outbound

This paper cites A survey on the pipeline evolution of facial capture and tracking for digital humans.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness A survey on the pipeline evolution of facial capture and tracking for digital humans

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.028702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:33:54.222398Z digest=sha256:fc6f96b55d67971965c9c2caeca7b2717c68ee43962a85721e3dfc1a1b59903a

Observation f278228a-cd22-4c69-906e-39655cb90a70 · outbound

This paper cites Gross, Kristina H ¨o¨ok, Regan Mandryk, and Petr Slovak.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Gross, Kristina H ¨o¨ok, Regan Mandryk, and Petr Slovak

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.017717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:33:54.226871Z digest=sha256:c697fe1cc4c56fc54910dd2a85933aa1e2cbfb07809c3bc0c9878d6839c6e955

Observation 0b0fa2dc-26a1-483e-b33e-485eb0196ed9 · outbound

This paper cites Tarsier: Recipes for training and evaluating large video description models, 2024.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Tarsier: Recipes for training and evaluating large video description models, 2024

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:55.007121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:33:54.231149Z digest=sha256:caf3d111625f76a3776554bff1b16828ded979908065d24540035a5217b4623c

Observation 329bd536-d2db-4793-a525-e6978bf8566a · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:54.235456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:54.235456Z digest=sha256:d47babb1d91f8f2934ae7cf47aeb1c5012adf4fbba691028d80c953619b0f463

Observation f7cb10f3-2000-433b-95c5-bdfcf0a0d9a0 · outbound

This paper cites Vatex: A large-scale, high- quality multilingual dataset for video-and-language research.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Vatex: A large-scale, high- quality multilingual dataset for video-and-language research

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:54.996676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:33:54.239911Z digest=sha256:e87b6aad45271fea74a3827973c0fdc09dd4a57e287aca54a8e5dee820223e2b

Observation 8365d89e-3c8f-480b-913e-31a5ee19eee5 · outbound

This paper cites Ferv39k: A large-scale multi-scene dataset for fa- cial expression recognition in videos, 2022.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Ferv39k: A large-scale multi-scene dataset for fa- cial expression recognition in videos, 2022

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:54.985798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:33:54.244178Z digest=sha256:a04ec964d3384441e0a9fc6382d5e1d368143f8575cbcf3eaea0498ebd2b2871

Observation 17a4cd68-123f-4ab6-a9b6-59c5b94b9864 · outbound

This paper cites InternVideo2: Scaling Foundation Models for Multimodal Video Understanding.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:54.248879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:54.248879Z digest=sha256:934e5cebccc66fe1780f84e0b9c4fa0b970dbb5b8c90aa29bc7fe5e985580be4

Observation 78a57c11-57c5-49b9-a415-014a2697bd64 · outbound

This paper cites Msr-vtt: A large video description dataset for bridging video and language.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Msr-vtt: A large video description dataset for bridging video and language

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:54.253320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:54.253320Z digest=sha256:618bcddda56af3394f20380c907fee2b4399b6274c7fba15006cf20eb2424f5b

Observation b9baf577-4213-4296-898f-e4561352466e · outbound

This paper cites Pllava : Parameter-free llava extension from images to videos for video dense captioning, 2024.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Pllava : Parameter-free llava extension from images to videos for video dense captioning, 2024

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:54.842118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:33:54.258279Z digest=sha256:27a158142d52d98d2b0be3cef30190d119c041f69000e7b29989884a4d75ba20

Observation b86eb911-55bb-4a55-ad71-570e44686c28 · outbound

This paper cites xgen-mm (blip-3): A family of open large multimodal models, 2024.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness xgen-mm (blip-3): A family of open large multimodal models, 2024

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:54.262867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:54.262867Z digest=sha256:4d9bb248d2a0618598307868b2b6f6edbc88507843abf9c8bb4581a0fc7e8ab6

Observation a80bd366-e3b8-4c07-bc84-872b9aa404ef · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:54.270705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:54.270705Z digest=sha256:05066c1aa25c6c8887dee1f75116e7d760327a65e939710863566ab541e1ed47

Observation 7120bbcc-1364-4dbf-9b61-41dc9cd72a46 · outbound

This paper cites mplug- owl2: Revolutionizing multi-modal large language model with modality collaboration, 2023.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness mplug- owl2: Revolutionizing multi-modal large language model with modality collaboration, 2023

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:54.814951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:33:54.274065Z digest=sha256:bebcbf78e89530913d081100c54a40ef0981b597c6fdf902b5ef25f2c1fbbf21

Observation 99cbcbe2-93a1-4092-83f4-6ed80486ba15 · outbound

This paper cites Spatio-temporal convolutional features with nested lstm for facial expression recognition.Neurocomputing, 317: 50–57, 2018.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Spatio-temporal convolutional features with nested lstm for facial expression recognition.Neurocomputing, 317: 50–57, 2018

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:54.805828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:33:54.277803Z digest=sha256:4f8629afce2ef15b5108cbfd85f3b6d264762c5c4b028b519784a2a647b781fd

Observation 1d5ac5bb-0e4d-495e-9633-9bcc71042d74 · outbound

This paper cites Auformer: Vision transformers are parameter-efficient facial action unit detectors.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Auformer: Vision transformers are parameter-efficient facial action unit detectors

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:54.796041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:33:54.281283Z digest=sha256:e40509e471bb0a506d1f6a583deaef76c8fa0a45e14cfeb98c494490d99c07e2

Observation 262d4719-76df-40e9-88b2-d0e486e2a57e · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:54.285100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:54.285100Z digest=sha256:e4bd228242802a8ec3bb66895d65a329dad57517588576948c2c14a702610718

Observation 1ffd1b68-9ee3-4d03-8634-e20268023f95 · outbound

This paper cites Vision Transformer with Quadrangle Attention.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Vision Transformer with Quadrangle Attention

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:54.289229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:54.289229Z digest=sha256:ea6029ac69c097c95d6315abe634b0490a113786b3a7f61757883d5e40fe9358

Observation 9812e6b7-9531-4f92-aa79-f3b95bdc06a8 · outbound

This paper cites Cohn, Shaun Canavan, Michael Reale, Andy Horowitz, Peng Liu, and Jeffrey M.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Cohn, Shaun Canavan, Michael Reale, Andy Horowitz, Peng Liu, and Jeffrey M

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:54.785326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:33:54.294736Z digest=sha256:9d3434b755fde7fc329808a31231cecb3adce5aa6c80046f0fa1358d5c4a6707

Observation ca82cb7d-5634-4986-921c-5a23b8072967 · outbound

This paper cites Llava- next: A strong zero-shot video understanding model, 2024.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Llava- next: A strong zero-shot video understanding model, 2024

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:54.774625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:33:54.298057Z digest=sha256:e03780143c14531eff41e4819cb229b3fe57fdf36c2c5ceb400cb919dbc82222

Observation e265d731-5a5a-48a6-beac-ca464787be71 · outbound

This paper cites Facial expression recognition from near- infrared videos.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Facial expression recognition from near- infrared videos

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:54.764001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:33:54.301613Z digest=sha256:0a8671dd433e17b16ea5c82d340bd2176087bd76d85ee937544a767de42be8e1

Observation eace1683-3784-4d4d-a46a-932b0a4d921c · outbound

This paper cites LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:54.304945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:54.304945Z digest=sha256:18a744425b326ae3cbac46ec3cb6849e564b05feaaec00f226968395a92503e2

Observation 14e9c2b1-9476-4ab0-b54c-3afe638013e0 · outbound

This paper cites Deep region and multi-label learning for facial action unit detec- tion.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Deep region and multi-label learning for facial action unit detec- tion

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:54.751577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:33:54.308759Z digest=sha256:2d365ea6ef66137aa9d862251187f5c5df719e8cd914205c287a7808cb7481c9

Observation 6d9ec7d2-adf8-4084-9f4a-740fce5b0c11 · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness MLVU: Benchmarking Multi-task Long Video Understanding

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:54.312278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:54.312278Z digest=sha256:2723930ced692076c2fb604a0a9be598c18ccf19a4b0376ac68c705eb22a6d73

Observation 2276db77-201b-497d-a22f-600d3dd25b63 · outbound

This paper cites Towards automatic learning of procedures from web instructional videos.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness Towards automatic learning of procedures from web instructional videos

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:54.739302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:33:54.316060Z digest=sha256:ce47333ba4b66f8b879192fef6fb6fe17dab8f31dbe4fc381f4442746644e17c

Observation 30e8f1e5-610c-4d29-a881-238bc6134e5d · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:54.320561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:54.320561Z digest=sha256:7dde0ccb29d9610015235debd570043bf4c5e63277a917661e5f915a2c52f421

Pith citing papers

Observation dfbf3f28-6334-4a74-834c-57a2e45d06e1 · inbound

LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding cites this paper.

LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness

Reference 93

Resolution
verified exact
arxiv_id, observed 2026-05-23T06:02:37.649306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T06:01:00.775721Z digest=sha256:c4197c3304b8f839899703cc99b8646d2702d9412592a9847a55d99e9a5f4647

Observation 9ba63ce5-3dbf-4d8c-90ca-ef657731fb73 · inbound

FaVChat: Hierarchical Prompt-Query Guided Facial Video Understanding with Data-Efficient GRPO cites this paper.

FaVChat: Hierarchical Prompt-Query Guided Facial Video Understanding with Data-Efficient GRPO Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-23T00:22:18.796437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T00:21:51.621582Z digest=sha256:3494d8b0b98bf75da7ae7b3addf5e6b66a270115c82839a953c7d9e094a45f18

Observation 2f1f05b9-e8db-4b09-a634-0d7783dd5351 · inbound

LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs cites this paper.

LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness

Reference 108

Resolution
unresolved
no resolver link, observed 2026-08-06T22:24:36.541729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:24:36.541729Z digest=sha256:d30b11705c966ff741db65817c7550cfb1b0a376f14b7e03c13f5702df180ac7

Observation 260b9555-6759-4738-8c89-c4f26d70f407 · inbound

See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding cites this paper.

See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness

Reference 107

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:13:16.336074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T12:10:54.874012Z digest=sha256:8bb6080e724451199fba2d2166c0cab4dc8675a0cef0c4c991110c1b6ea95108