Pith. sign in

Paper Citation Record · LEDGER

Fine-tuning Multimodal Transformers on Edge: A Parallel Split Learning Approach

As of 9 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 1 inbound Pith citation observation for arXiv:2502.06355.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.06355 v3

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T15:45:46.182206Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T16:41:37.241236Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

29 of 29 outbound references displayed

  • verified exact1
  • verified fuzzy20
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 047863be-79c3-4071-8719-f096f8fb55be · outbound

This paper cites Vatt: Transformers for multimodal self-supervised learning from raw video, audio and text,.

Fine-tuning Multimodal Transformers on Edge: A Parallel Split Learning Approach Vatt: Transformers for multimodal self-supervised learning from raw video, audio and text,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:45:46.511457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T15:45:46.074894Z digest=sha256:ad45142d3c32b88deb9c3cedaf0a72eca08c2fbb740b727578f0b64bf67ad7a5

Observation b481ac32-0527-466c-bfb5-1ec656a21d18 · outbound

This paper cites Imagebind: One embedding space to bind them all,.

Fine-tuning Multimodal Transformers on Edge: A Parallel Split Learning Approach Imagebind: One embedding space to bind them all,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:45:46.458426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T15:45:46.099243Z digest=sha256:399b5d097d942e892ab2d7d1803df0f0ccd4bb20274eaac77ac2928cfe55fa88

Observation e69c759e-9c45-43be-b0f2-2f2826b2543d · outbound

This paper cites Dis- tributed learning of deep neural network over multiple agents,.

Fine-tuning Multimodal Transformers on Edge: A Parallel Split Learning Approach Dis- tributed learning of deep neural network over multiple agents,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:45:46.447966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T15:45:46.106043Z digest=sha256:40fd9973ca41cdf4f58146c9fabf17d9875e8909803b6c853f94e7631a125bbd

Observation afc58a40-9189-4937-859e-37974fad4339 · outbound

This paper cites Parameter- efficient transfer learning for nlp.

Fine-tuning Multimodal Transformers on Edge: A Parallel Split Learning Approach Parameter- efficient transfer learning for nlp

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:45:46.425705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T15:45:46.112255Z digest=sha256:036da57edce28a1bd894c1ac31cc381816fa22eceb48c4b342c1aa369f1b7d44

Observation d384fa3a-feb2-4ae9-b880-3e2ffc2774d7 · outbound

This paper cites ViT-Lens: Initiating Omni-Modal Exploration through 3D Insights.

Fine-tuning Multimodal Transformers on Edge: A Parallel Split Learning Approach ViT-Lens: Initiating Omni-Modal Exploration through 3D Insights

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-08T15:45:46.270073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T15:45:46.119825Z digest=sha256:b7cd417c80917fc57fbea222893e397204768617deb6e5dbb0df10707e5b4e0d

Observation c5538d8b-5472-42ba-9f67-dc161476ce5d · outbound

This paper cites Federated learning on non-iid data silos: An experimental study,.

Fine-tuning Multimodal Transformers on Edge: A Parallel Split Learning Approach Federated learning on non-iid data silos: An experimental study,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T15:45:46.123879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:45:46.123879Z digest=sha256:c29308da1624f22f4d4b6c262ab692392d2539a7a11b175341a92923a3fbe343

Observation 88d47e6a-9630-4a15-87ed-9c1a425d8649 · outbound

This paper cites Microsoft coco: Common objects in con- text.

Fine-tuning Multimodal Transformers on Edge: A Parallel Split Learning Approach Microsoft coco: Common objects in con- text

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:45:46.396507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T15:45:46.127552Z digest=sha256:0eeba391e967ce4c92a4caf4635be2cd1a34ac6b3c2f25846e5d55422caab3f6

Observation 6e99a235-226d-40a1-9f01-331a6e6b765e · outbound

This paper cites FedCLIP: Fast Generalization and Personalization for CLIP in Federated Learning.

Fine-tuning Multimodal Transformers on Edge: A Parallel Split Learning Approach FedCLIP: Fast Generalization and Personalization for CLIP in Federated Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T15:45:46.135137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:45:46.135137Z digest=sha256:50262f8b75956239b815d231f365ad80de848580077e8de13f0bac48bc31381b

Observation ff2e8f44-388b-4317-aab3-d1c0534f33e9 · outbound

This paper cites Scalable aggregated split learning for data-driven edge intelligence on internet-of-things.

Fine-tuning Multimodal Transformers on Edge: A Parallel Split Learning Approach Scalable aggregated split learning for data-driven edge intelligence on internet-of-things

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:45:46.374599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T15:45:46.139246Z digest=sha256:41b559c3381250d8733b96061323fb7c33877d762bdea2a6a4b5845a81b5ae03

Observation 76c39fa3-52d1-4317-b6f7-97b9a3ded190 · outbound

This paper cites Communication-Efficient Learning of Deep Networks from De- centralized Data.

Fine-tuning Multimodal Transformers on Edge: A Parallel Split Learning Approach Communication-Efficient Learning of Deep Networks from De- centralized Data

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:45:46.364187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T15:45:46.143029Z digest=sha256:55837f4e5f50d30d2cd597a97537559e1a19ad35f134b836eebfc7f8a2393450

Observation 88096161-94a3-4b34-b601-c554372c80c6 · outbound

This paper cites Mix2sfl: Two-way mixup for scalable, accurate, and communication-efficient split federated learning.

Fine-tuning Multimodal Transformers on Edge: A Parallel Split Learning Approach Mix2sfl: Two-way mixup for scalable, accurate, and communication-efficient split federated learning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:45:46.353629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T15:45:46.146614Z digest=sha256:f3792447b1bc48d0a9d557a801bc9494d6dbca2197b6aefbe85d5a9a1d61bfb2

Observation bc431f59-4408-4cbf-afad-156d1ee9d040 · outbound

This paper cites Server-side local gradient averaging and learning rate acceleration for scalable split learning,.

Fine-tuning Multimodal Transformers on Edge: A Parallel Split Learning Approach Server-side local gradient averaging and learning rate acceleration for scalable split learning,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:45:46.341856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T15:45:46.150833Z digest=sha256:f151e8c1e2b60442421df3abca47c3e05a713c7c6752ae97bca0e5a395f4ebe7

Observation 867ab09c-342e-4aff-97cd-37ca8eb9c915 · outbound

This paper cites Flickr30k entities: Collecting region-to-phrase corre- spondences for richer image-to-sentence models.

Fine-tuning Multimodal Transformers on Edge: A Parallel Split Learning Approach Flickr30k entities: Collecting region-to-phrase corre- spondences for richer image-to-sentence models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:45:46.331970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T15:45:46.154589Z digest=sha256:6af83d8564f39301d4f0dff7bef84b796bfc914d526773275f8cda5e0d2f7d44

Observation e9e6aa05-c7d5-4b4c-b220-c5a596f52fd3 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Fine-tuning Multimodal Transformers on Edge: A Parallel Split Learning Approach Learning transferable visual models from natural language supervision,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:45:46.322614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T15:45:46.162140Z digest=sha256:d12dd979f99a21eb87db6ca21f5125077fa761de1c61228130ced5314fea7446

Observation c6e27adf-2f4b-4aed-96c9-94f094e9df24 · outbound

This paper cites Exploring models and data for image question answering.

Fine-tuning Multimodal Transformers on Edge: A Parallel Split Learning Approach Exploring models and data for image question answering

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:45:46.312961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T15:45:46.166079Z digest=sha256:ff432bf3054fa698653f819e171e88414f9a7883e2cf8bc072263bca51f0cb5a

Observation 068d5ab2-9a8a-42b7-95a8-03906bba3f95 · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

Fine-tuning Multimodal Transformers on Edge: A Parallel Split Learning Approach UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T15:45:46.169912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:45:46.169912Z digest=sha256:b61f5a4c98980b49f52bdafc27c5e7a9826d9b4f7eed580400f6d44a42837eb7

Observation 7ed6be78-5418-4f27-896b-ac766681927e · outbound

This paper cites Permutation Equivariance of Transformers and Its Applications.

Fine-tuning Multimodal Transformers on Edge: A Parallel Split Learning Approach Permutation Equivariance of Transformers and Its Applications

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T15:45:46.177880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:45:46.177880Z digest=sha256:d98d35d25f42a631cc496ab410eb2b031ffaf15c3cb8246b67bfba84baa12cb8

Observation b28a1464-fd51-485c-b3ce-5993f5cd603c · outbound

This paper cites Meta-Transformer: A Unified Framework for Multimodal Learning.

Fine-tuning Multimodal Transformers on Edge: A Parallel Split Learning Approach Meta-Transformer: A Unified Framework for Multimodal Learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T15:45:46.182206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:45:46.182206Z digest=sha256:dcc98049b7e8eb70abf148a674f09a1e3a8fdab55647f366cda93d24d80c8869

Observation b2faa4c5-5bf5-4e57-9a74-ab440ff1a163 · outbound

This paper cites Cross-media learning for image sentiment analysis in the wild.

Fine-tuning Multimodal Transformers on Edge: A Parallel Split Learning Approach Cross-media learning for image sentiment analysis in the wild

Reference 2012

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:45:46.302732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T15:45:46.174092Z digest=sha256:dc8b729d30cd775a6f4bbd2be19cbe0f9ea9659f8b8b174dff367114acfadbe9

Observation 203b684d-ddf3-4051-9698-06c6b2d55a2d · outbound

This paper cites Efficient par- allel split learning over resource-constrained wireless edge net- works.

Fine-tuning Multimodal Transformers on Edge: A Parallel Split Learning Approach Efficient par- allel split learning over resource-constrained wireless edge net- works

Reference 2014

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:45:46.385583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T15:45:46.131345Z digest=sha256:25151fbbec3cf1571fc074a82273615f81b0e235c0f475e13cdecf0aca14ac06

Observation 476a4860-d038-4dbb-988b-e16c8978d0a9 · outbound

This paper cites MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in Conversations.

Fine-tuning Multimodal Transformers on Edge: A Parallel Split Learning Approach MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in Conversations

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-08T15:45:46.158254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:45:46.158254Z digest=sha256:1e8f39b697ed297303209454327380220cf9661b500d3d10a002d718332bbe49

Observation 1919314f-7d34-4b0e-90b4-2d8e4a4a7660 · outbound

This paper cites The iot breaches your household again.

Fine-tuning Multimodal Transformers on Edge: A Parallel Split Learning Approach The iot breaches your household again

Reference 2017

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:45:46.488910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T15:45:46.083782Z digest=sha256:4f533f751f81b2ad3a5e5cba5af5d63cc2c1904ede06e07f7c5cadb821a87ad4

Observation fb285ac0-c69a-48ed-af88-c8e18d841bb3 · outbound

This paper cites Accelerating federated learning with split learning on locally generated losses.

Fine-tuning Multimodal Transformers on Edge: A Parallel Split Learning Approach Accelerating federated learning with split learning on locally generated losses

Reference 2018

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:45:46.436925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T15:45:46.109116Z digest=sha256:0f52d327c157638bf73c98772da9d9860af32bdbebc460e5d2ce622997735459

Observation b61cd8e7-21e4-460c-8769-d1045fc486f5 · outbound

This paper cites Privacy-sensitive parallel split learning.

Fine-tuning Multimodal Transformers on Edge: A Parallel Split Learning Approach Privacy-sensitive parallel split learning

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:45:46.414118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T15:45:46.115887Z digest=sha256:211937ac6abdf130c9002742045d692915e1ab89288d673b234f209ceb996524

Observation a2c380c3-b741-4d10-88a1-fe7d9271c522 · outbound

This paper cites Multi-modal align- ment using representation codebook,.

Fine-tuning Multimodal Transformers on Edge: A Parallel Split Learning Approach Multi-modal align- ment using representation codebook,

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:45:46.478560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T15:45:46.091790Z digest=sha256:3c6e9f6f291a015fef01a602e62344875aa95331fdd155d255993bcf80f700cd

Observation b483d407-0919-4b42-840a-94bbdca68a30 · outbound

This paper cites Look, listen and learn.

Fine-tuning Multimodal Transformers on Edge: A Parallel Split Learning Approach Look, listen and learn

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:45:46.499614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T15:45:46.079585Z digest=sha256:cf5a668655ebf80681b4f393c8656a0bdef59ae5237e039dc5d4fb44da8e5e64

Observation 46967e75-311e-4564-ac23-cebf4f614e40 · outbound

This paper cites Split learn- ing of multi-modal medical image classification.

Fine-tuning Multimodal Transformers on Edge: A Parallel Split Learning Approach Split learn- ing of multi-modal medical image classification

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:45:46.468989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T15:45:46.095406Z digest=sha256:cdf7ac83fb0c92945ce50eecc2057adc15a77c4e223ec0a265abd7b61b8aae7d

Observation 01a0685b-d2c9-4e6d-8743-21b5b06fdc70 · outbound

This paper cites AST: Audio Spectrogram Transformer.

Fine-tuning Multimodal Transformers on Edge: A Parallel Split Learning Approach AST: Audio Spectrogram Transformer

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-08T15:45:46.102491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:45:46.102491Z digest=sha256:75e7f50d4f8f3b46da72f1fb90afdb2980c99138feb1e0d7d1abbc07a699ee11

Observation 633f2018-36aa-45ce-90cd-b1a626ebee3a · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Fine-tuning Multimodal Transformers on Edge: A Parallel Split Learning Approach An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-08T15:45:46.087545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:45:46.087545Z digest=sha256:c9035a75426f42fec69f131803c7ff97e9a246ade0fc68f76856f9eb60abd4a1

Pith citing papers

Observation aebefd0f-aadf-4d82-8629-9886ad8242a6 · inbound

AutoEncoder-Compressed Parallel Split Learning for Pre-trained Model Fine-Tuning cites this paper.

AutoEncoder-Compressed Parallel Split Learning for Pre-trained Model Fine-Tuning Fine-tuning Multimodal Transformers on Edge: A Parallel Split Learning Approach

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T16:41:37.241236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:41:37.241236Z digest=sha256:d8ffb8e5de34718c125905bc5262719a9e23f841beaf6a936bfe0865ebdc9e28