Pith. sign in

Paper Citation Record · LEDGER

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets

As of 11 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 0 inbound Pith citation observations for arXiv:2501.03332.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.03332 v1

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T21:57:41.937067Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

41 of 41 outbound references displayed

  • verified exact2
  • verified fuzzy23
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cefb8f7e-1b1c-4885-895d-066ea60b3fd7 · outbound

This paper cites Multimodal personality recognition using cross-attention transformer and behaviour encoding.

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets Multimodal personality recognition using cross-attention transformer and behaviour encoding

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:42.746006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:57:41.747478Z digest=sha256:9fe09611fc392fd0533bf548923e663fca14edb7d2439fdd4a7d75451b61d391

Observation 9a3d884a-bb0a-4e77-989b-edcb82da6259 · outbound

This paper cites Multimodal vision transformers with forced attention for behavior analysis.

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets Multimodal vision transformers with forced attention for behavior analysis

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:42.730891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:57:41.752893Z digest=sha256:5e52941ea8bf5876e89179e2f9f2326c10d6d25187d1e8ad027e27f211eb5f99

Observation 1eaf3175-b942-4fb1-8eb4-6ab022416907 · outbound

This paper cites VATT: Transformers for Multimodal Self-Supervised Learning from Raw Video, Audio and Text.

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets VATT: Transformers for Multimodal Self-Supervised Learning from Raw Video, Audio and Text

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T21:57:41.757556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:57:41.757556Z digest=sha256:22dddd803150ad686e51a7108f569517ff54bcb46f75eda9a7ded40635521f22

Observation 1c420e10-9362-479c-a3a8-236326b87c4a · outbound

This paper cites Vivit: A video vi- sion transformer.

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets Vivit: A video vi- sion transformer

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:42.716215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:57:41.762729Z digest=sha256:ec80e98c6cb36993a61e783183bb5931b93126af548981bc6466cd3f7ff78c7e

Observation 3e79fad3-0e5a-4983-a89b-43454524ad2d · outbound

This paper cites Bodily be- haviors in social interaction: Novel annotations and state-of- the-art evaluation.

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets Bodily be- haviors in social interaction: Novel annotations and state-of- the-art evaluation

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:42.700369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:57:41.767362Z digest=sha256:30d5b3992fb3771971e7403aa87e995e6d130e9b998673cee7d9d87a0ceb7b20

Observation be8de12e-1c86-499f-b53b-51028f9afdec · outbound

This paper cites AdaptFormer: Adapting Vision Transformers for Scalable Visual Recognition.

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets AdaptFormer: Adapting Vision Transformers for Scalable Visual Recognition

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T21:57:41.772310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:57:41.772310Z digest=sha256:e29d353e2b8eafc237878191ff1a623d75ec02f36a0e4892cb65b4c39c993aed

Observation 29b47da0-5ead-46ef-98c8-e46eb20e0c02 · outbound

This paper cites Rescaling Egocentric Vision.

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets Rescaling Egocentric Vision

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T21:57:41.777674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:57:41.777674Z digest=sha256:0aab7a37e2211a2a603bd426032bddacaef61a883edab54962b242672cff32b3

Observation 2d7d4960-2027-4a56-9d79-a5fc39f7bf2c · outbound

This paper cites A transformer-based joint-encoding for emotion recognition and sentiment analysis.

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets A transformer-based joint-encoding for emotion recognition and sentiment analysis

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:42.685448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:57:41.782611Z digest=sha256:a0aaf2e2cc7135963542d68a6129ae20d97a9ea6aa2ce595a8f7311c09df6b33

Observation 517d6aa6-979a-4842-a0e5-e3e409206b80 · outbound

This paper cites PPT: Pre-trained Prompt Tuning for Few-shot Learning.

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets PPT: Pre-trained Prompt Tuning for Few-shot Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T21:57:41.786939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:57:41.786939Z digest=sha256:61622553685f05c1303fa1901376ffd30604d3ddbc7bc2fd3b6a66d71f81c051

Observation 57d12086-2d8c-486e-96e5-af7792988656 · outbound

This paper cites Towards a Unified View of Parameter-Efficient Transfer Learning.

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets Towards a Unified View of Parameter-Efficient Transfer Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T21:57:41.791474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:57:41.791474Z digest=sha256:46e7f3189b82032b58c018abba911859d038e00ae0a8ed085811e8a28efc43e2

Observation 17187cc0-9144-4617-b319-154d44b933fc · outbound

This paper cites Parameter-efficient trans- fer learning for NLP.

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets Parameter-efficient trans- fer learning for NLP

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:42.669455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:57:41.796366Z digest=sha256:276dbb6b9ca3c8632618389a61fc935ee3f9d41a21e43d92c8008e20e53ad6d0

Observation d50bb71c-9b26-4cca-9ca9-3d8dff2a5375 · outbound

This paper cites Lora: Low-rank adaptation of large language models, 2021.

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets Lora: Low-rank adaptation of large language models, 2021

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:42.652332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:57:41.801052Z digest=sha256:b11527157c6971b3d01849501f999676ebc5e88ff1ddd57cf2a68135220409c8

Observation 5f3ad6b7-82a6-4838-98de-3a7c28a17858 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets LoRA: Low-Rank Adaptation of Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T21:57:41.805566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:57:41.805566Z digest=sha256:e3bb00f224028bc00447497516cf206fd83ddf910f5f495a2ce8bdc9cdeb86f8

Observation 47638112-c3bb-4db1-b42e-0a4f6c7fee1f · outbound

This paper cites Mumu: Cooperative mul- titask learning-based guided multimodal fusion.

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets Mumu: Cooperative mul- titask learning-based guided multimodal fusion

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:42.637267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:57:41.810254Z digest=sha256:cb311e78ed2116c04929af179448014c40a730c9cb471fe53996689d6508a9b7

Observation 1e937211-8683-4eed-a255-fd36306eb630 · outbound

This paper cites Vi- sual prompt tuning.

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets Vi- sual prompt tuning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:42.622072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:57:41.814539Z digest=sha256:241c93ea741fd53672c33ad457191105714a37f036015d6cd92a2a0f489b2fcf

Observation 062b865c-3421-4328-b567-87405a285e13 · outbound

This paper cites Compacter: Efficient low-rank hypercomplex adapter layers.

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets Compacter: Efficient low-rank hypercomplex adapter layers

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:42.607216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:57:41.819154Z digest=sha256:204be315bcf9c2af03a9dd43266a19ff5f4b4d4a8d472baecd119242fdee6ef9

Observation 37329977-4498-4e23-aac8-b181ca94c9e5 · outbound

This paper cites Parameter-efficient multi-task fine-tuning for transformers via shared hypernetworks.

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets Parameter-efficient multi-task fine-tuning for transformers via shared hypernetworks

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:42.591462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:57:41.823645Z digest=sha256:e1d6ee65436d7b8f048a42a44255e59df175b74a75025670ea7c5c61b0783e16

Observation 039df51b-eee5-410e-b860-2f411990162c · outbound

This paper cites Transformers in vision: A survey.

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets Transformers in vision: A survey

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T21:57:41.828211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:57:41.828211Z digest=sha256:0ed651f2545867c3cde434ad943f907482af4948347dd695a2907ff56cffff7c

Observation ecc15948-d0c0-46eb-b0a4-d2f7e61e1c0c · outbound

This paper cites Gated Mechanism for Attention Based Multimodal Sentiment Analysis.

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets Gated Mechanism for Attention Based Multimodal Sentiment Analysis

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-10T21:57:42.254175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:57:41.832911Z digest=sha256:df1d8fc4ab040d10823bfd2ddadd2f2049dbd060a41385684f88bc7d16819dee

Observation f7680172-9a00-48f8-b83c-08944022ab9a · outbound

This paper cites The power of scale for parameter-efficient prompt tuning.

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets The power of scale for parameter-efficient prompt tuning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:42.566531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:57:41.837764Z digest=sha256:5dc595de53978db03a8edd4c88fcbf68d238782493ad791e0dd829320d59b985

Observation a2452fab-058f-49ef-a192-af15732c1147 · outbound

This paper cites Prefix-Tuning: Optimizing Continuous Prompts for Generation.

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets Prefix-Tuning: Optimizing Continuous Prompts for Generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T21:57:41.842107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:57:41.842107Z digest=sha256:ce74f8c2b21f0d4c354b586ac6dcde4c4a4767517e4f073e679d148aa34b99e3

Observation 0ec927f3-ed76-41bd-aca2-7d3101e3d2f0 · outbound

This paper cites Video swin transformer.

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets Video swin transformer

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:42.551485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:57:41.847102Z digest=sha256:e6dd70c289203aa48ab0fad7860dd81f3a869ef7f61d848db88c6fa88556c798

Observation c4eee4e8-5c4c-4682-9226-73692234e2f5 · outbound

This paper cites Parameter-efficient Multi-task Fine-tuning for Transformers via Shared Hypernetworks.

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets Parameter-efficient Multi-task Fine-tuning for Transformers via Shared Hypernetworks

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T21:57:41.851696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:57:41.851696Z digest=sha256:152d5b94ce061a0d6470db62a9787503a09623bbd3b67d1b8a74ba68b8122de2

Observation f125b378-f6e3-48ad-9f6f-80238fe6635f · outbound

This paper cites UniPELT: A Unified Framework for Parameter-Efficient Language Model Tuning.

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets UniPELT: A Unified Framework for Parameter-Efficient Language Model Tuning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T21:57:41.856385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:57:41.856385Z digest=sha256:f937662c08e35e1712709ee3698d7024c49db1a4083984355fc49742940b42e9

Observation 791ab3e0-3809-4a6c-a617-1c3875df456e · outbound

This paper cites Tiny adapters for vision transformers, 2023.

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets Tiny adapters for vision transformers, 2023

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:42.536574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:57:41.861544Z digest=sha256:0fc42314af1b121c889a8f63b11e4f9206ddc56484fc37523ddf8ccb6078cda2

Observation 3ab3513c-5b1e-444a-84ab-a79ca64d5114 · outbound

This paper cites Context- aware personality inference in dyadic scenarios: Introducing the udiva dataset.

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets Context- aware personality inference in dyadic scenarios: Introducing the udiva dataset

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:42.521422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:57:41.865979Z digest=sha256:a33a4cd3849fb6daff95d0c83e224fe1efe3e3d8242c38ca57fbe68b6d606a21

Observation f5f113a2-2ebc-4f1c-9de3-62d70bf10728 · outbound

This paper cites ST-Adapter: Parameter-Efficient Image-to-Video Transfer Learning.

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets ST-Adapter: Parameter-Efficient Image-to-Video Transfer Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T21:57:41.870525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:57:41.870525Z digest=sha256:2e8ddf1caa9fcd633875841b2721da3328b5ef3d4b2ed52445a91e7c5f739743

Observation 70d7c14b-1d06-43ed-a5ee-6204b3fdae4c · outbound

This paper cites Dual-path adaptation from image to video transformers.

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets Dual-path adaptation from image to video transformers

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:42.505307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:57:41.875584Z digest=sha256:0f46f6f92a8fd3093aeaad256ee0b77cb500f2127ccca1f8996121c81170fd10

Observation c53a2f50-64df-4e3c-9c6f-09538f2edcfb · outbound

This paper cites Learning transferable visual models from natural language supervision.

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets Learning transferable visual models from natural language supervision

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:42.490340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:57:41.880284Z digest=sha256:b8845b3de9b151cc8f8d904da12d80520aa0784ee7dbbeb3d009675c0a48d3fd

Observation 7dcef711-74ce-4ef1-a1d0-c7859b704e57 · outbound

This paper cites Learning multiple visual domains with residual adapters.Ad- vances in neural information processing systems , 30, 2017.

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets Learning multiple visual domains with residual adapters.Ad- vances in neural information processing systems , 30, 2017

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:42.475392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:57:41.884972Z digest=sha256:c5daf7988f19114a0f33b395a15e91060c4987342fbbfd328118e6279615f000

Observation b3cddac2-21f3-4624-99d5-391385aecde6 · outbound

This paper cites Efficient parametrization of multi-domain deep neural net- works.

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets Efficient parametrization of multi-domain deep neural net- works

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:42.459033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:57:41.889613Z digest=sha256:17223e598e33302f6988430b64fbe6e7cf30a7028f4f0535651df80833fc32e5

Observation 05eb2b85-d972-4d7e-bba8-745c2bd5e3f7 · outbound

This paper cites Multilingual Detection of Check-Worthy Claims using World Languages and Adapter Fusion.

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets Multilingual Detection of Check-Worthy Claims using World Languages and Adapter Fusion

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-10T21:57:42.167868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:57:41.894087Z digest=sha256:b6475814a79591dd46a8380521fd95d13966e76dfa6856bd0f884e25d27f3083

Observation 332d9f91-1186-490b-a9da-593c90a511ec · outbound

This paper cites an unresolved cited work.

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:57:42.442110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:57:41.899007Z digest=sha256:e55124f81a2dcdf9d7088f6c20771b710131aafb53968d5a56a262becb8757e7

Observation 811a7111-1a81-440f-b2b4-57d8cfb7adc8 · outbound

This paper cites Vl-adapter: Parameter-efficient transfer learning for vision-and-language tasks.

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets Vl-adapter: Parameter-efficient transfer learning for vision-and-language tasks

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:42.425548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:57:41.903694Z digest=sha256:df7fa318ec2adb28c03a71752373d0da0572d445bc4e886234f585347668f72a

Observation df61c7e2-bb0c-4dfa-929a-4697adc19e80 · outbound

This paper cites Training neu- ral networks with fixed sparse masks.

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets Training neu- ral networks with fixed sparse masks

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:42.408691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:57:41.909217Z digest=sha256:6921efb73cd45b39a46f78957dc7505e6ba5200e4559081ee3d6d5161deebad3

Observation 1051fd7a-9059-410d-b500-fc5f29f641da · outbound

This paper cites VideoMAE: Masked autoencoders are data-efficient learners for self-supervised video pre-training.

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets VideoMAE: Masked autoencoders are data-efficient learners for self-supervised video pre-training

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:42.393480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:57:41.913810Z digest=sha256:bd2d0fc268d610dd699588bbc2a59efadbc56900c0732de6e7bec36527cbf4fb

Observation 3e61394d-7238-44fd-9e1c-9f45e904f98e · outbound

This paper cites VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training.

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T21:57:41.918538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:57:41.918538Z digest=sha256:874dcd97cffee0627b2ad4daf52c1e02185fa81509dfc75d01feebdff6589d18

Observation a9217669-4064-436c-9976-7f7f8fe6acda · outbound

This paper cites M&M Mix: A Multimodal Multiview Transformer Ensemble.

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets M&M Mix: A Multimodal Multiview Transformer Ensemble

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T21:57:41.923255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:57:41.923255Z digest=sha256:aa03bac7b4a34ecfdcba9b69b42f88cc55495cbd774a75091c1d757c4eca9532

Observation 31ca9c65-cca4-4d16-a0fd-a23491cd5239 · outbound

This paper cites Multiview transformers for video recognition.

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets Multiview transformers for video recognition

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:57:42.376564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:57:41.928345Z digest=sha256:8e112726e7cbc8aeb1c54ae8749d9f88f06e1ae744e647e00c1b2242d7f878db

Observation 80143f1f-e6e2-45f7-9206-914917274a3d · outbound

This paper cites Bitfit: Simple parameter-efficient fine-tuning for transformer-based masked language-models.

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets Bitfit: Simple parameter-efficient fine-tuning for transformer-based masked language-models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T21:57:41.932616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:57:41.932616Z digest=sha256:bc86d81ec21176e68d713e180ca82257d543d401d7d082cc74e7184b212fd274

Observation fab0796f-1759-49e8-a8e4-9bd83f972ab1 · outbound

This paper cites an unresolved cited work.

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:57:42.361307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:57:41.937067Z digest=sha256:8b274425c320904aff5862b8dd7880295db63d4dfa417a56071e5a6f8e437e8c

Pith citing papers

No inbound Pith citation observations are available.