Pith. sign in

Paper Citation Record · LEDGER

Efficient Quantification of Multimodal Interaction at Sample Level

As of 11 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2506.17248.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.17248 v1

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:52:32.281803Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

13 of 13 outbound references displayed

  • verified exact1
  • verified fuzzy4
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6e8a7eaa-f421-42c4-8d4e-9feb6071be88 · outbound

This paper cites Baseline We adopt three primary types of multimodal learning paradigms: Feature-level fusion: Integration of multiple modalities at the feature level.

Efficient Quantification of Multimodal Interaction at Sample Level Baseline We adopt three primary types of multimodal learning paradigms: Feature-level fusion: Integration of multiple modalities at the feature level

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:52:32.409939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:52:32.278424Z digest=sha256:35f3cac93ec441bd7a0405c8cce83fcc7dc40cc907671c2c7c3cec646acd0f79

Observation 06b52f1a-5fa7-44e0-a9eb-3f9877db56c4 · outbound

This paper cites MultiBench: Multiscale Benchmarks for Multimodal Representation Learning.

Efficient Quantification of Multimodal Interaction at Sample Level MultiBench: Multiscale Benchmarks for Multimodal Representation Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:52:32.252230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:52:32.252230Z digest=sha256:701d67e54e24892e3fb621069a5b184381d9751d9625293250d6913f6916b4ca

Observation d08a618e-4352-4159-a23b-d7c550e86355 · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

Efficient Quantification of Multimodal Interaction at Sample Level UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:52:32.260373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:52:32.260373Z digest=sha256:8a773ec53d5d3e8a16216c69f0f9b9404f431a8b40b02805d853bf94b4e3da03

Observation 97c00d96-9983-42c8-8246-12d007582e89 · outbound

This paper cites Understanding Unimodal Bias in Multimodal Deep Linear Networks.

Efficient Quantification of Multimodal Interaction at Sample Level Understanding Unimodal Bias in Multimodal Deep Linear Networks

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T05:52:32.274858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:52:32.274858Z digest=sha256:5d1cf9d0deda26238876a357bf39bee3e88229262bbf3dca6a4cf22f4f6a0fa8

Observation c61b8c16-0215-4891-bcc4-63fa9c6a188b · outbound

This paper cites The specific model variant employed in our experiments features unimodal branches, each consisting of four Transformer layers.

Efficient Quantification of Multimodal Interaction at Sample Level The specific model variant employed in our experiments features unimodal branches, each consisting of four Transformer layers

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:52:32.397280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:52:32.281803Z digest=sha256:c4ae911b59406c2ed8dbe49c33af7ae0ac6ce79b318b2a9cdbad4486d7bca394

Observation c7b47a57-c7c6-4ebb-8f22-277125aad31b · outbound

This paper cites Learning Factorized Multimodal Representations.

Efficient Quantification of Multimodal Interaction at Sample Level Learning Factorized Multimodal Representations

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-07T05:52:32.264142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:52:32.264142Z digest=sha256:b7529fa4e54192ced969954de6bac176886b381d5a0524bf9a71e855b445d4d5

Observation a73c0a1f-38a5-47c7-a49b-2c28abca96b8 · outbound

This paper cites Food-101– mining discriminative components with random forests.

Efficient Quantification of Multimodal Interaction at Sample Level Food-101– mining discriminative components with random forests

Reference 2014

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:52:32.433225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:52:32.235428Z digest=sha256:d7f8750c7ef74da6f74292aeea38b5daab28c237a9cf7b7318f4ac060fc82bb8

Observation b134af1a-5520-438f-af05-d5e204067563 · outbound

This paper cites Nonnegative Decomposition of Multivariate Information.

Efficient Quantification of Multimodal Interaction at Sample Level Nonnegative Decomposition of Multivariate Information

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-07T05:52:32.270929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:52:32.270929Z digest=sha256:954a6e7351ef93b1df7946df2e489360900ee0e2405e54b4dca3034cf521e171

Observation f76dcdcc-cd3a-4c6d-b06b-7760570f1c0f · outbound

This paper cites Feature Interaction Interpretability: A Case for Explaining Ad-Recommendation Systems via Neural Interaction Detection.

Efficient Quantification of Multimodal Interaction at Sample Level Feature Interaction Interpretability: A Case for Explaining Ad-Recommendation Systems via Neural Interaction Detection

Reference 2018

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:52:32.330835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:52:32.267621Z digest=sha256:7cf11a71013c4eab57189391f72b1d6503a895d3033236c88c8f0cba34c3aedf

Observation 5f5153b9-6e44-4bdb-ae19-890ad324f59d · outbound

This paper cites Modality dropout for improved performance-driven talking faces.

Efficient Quantification of Multimodal Interaction at Sample Level Modality dropout for improved performance-driven talking faces

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:52:32.421813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:52:32.248738Z digest=sha256:29061db49197bf0290d498fe167ebdc4f123be890c1658badf4dc5e7cf2684af

Observation 76fc0a54-36e4-440e-bddf-b6fdb68f5579 · outbound

This paper cites Multimodal Learning Without Labeled Multimodal Data: Guarantees and Applications.

Efficient Quantification of Multimodal Interaction at Sample Level Multimodal Learning Without Labeled Multimodal Data: Guarantees and Applications

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T05:52:32.256147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:52:32.256147Z digest=sha256:f7674c06507087a7d2ab6b9472db9b6cfb65f59152d65d40afed8926f4aee181

Observation c7ac0c3c-ab6a-4d69-8447-200c95f06dec · outbound

This paper cites UR-FUNNY: A Multimodal Language Dataset for Understanding Humor.

Efficient Quantification of Multimodal Interaction at Sample Level UR-FUNNY: A Multimodal Language Dataset for Understanding Humor

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T05:52:32.244311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:52:32.244311Z digest=sha256:ffc15a3d2ff85a88e5f5ce97cb7c1d20bd92c5f5bd1b967ad5c0a357e2bcc305

Observation 20d53c37-4dac-4566-8139-7ece982f0470 · outbound

This paper cites Multimodal Compact Bilinear Pooling for Visual Question Answering and Visual Grounding.

Efficient Quantification of Multimodal Interaction at Sample Level Multimodal Compact Bilinear Pooling for Visual Question Answering and Visual Grounding

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T05:52:32.240171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:52:32.240171Z digest=sha256:bb25efcf433b4ffb6a7c2f3ba0df9c4a81c07342df7f977f95b5311388100033

Pith citing papers

No inbound Pith citation observations are available.