Pith. sign in

Paper Citation Record · LEDGER

Efficient Quantification of Multimodal Interaction at Sample Level

As of 18 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2506.17248.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.17248 v1

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:52:32.281803Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

13 of 13 outbound references displayed

  • verified exact1
  • verified fuzzy4
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6e8a7eaa-f421-42c4-8d4e-9feb6071be88 · outbound

This paper cites Baseline We adopt three primary types of multimodal learning paradigms: Feature-level fusion: Integration of multiple modalities at the feature level.

Efficient Quantification of Multimodal Interaction at Sample Level Baseline We adopt three primary types of multimodal learning paradigms: Feature-level fusion: Integration of multiple modalities at the feature level

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:52:32.409939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:52:32.278424Z digest=sha256:293d7d846693653b42b28b60f96b3b1cd06a6d56e702d1c782e08b0c1b6f6bdc

Observation 06b52f1a-5fa7-44e0-a9eb-3f9877db56c4 · outbound

This paper cites MultiBench: Multiscale Benchmarks for Multimodal Representation Learning.

Efficient Quantification of Multimodal Interaction at Sample Level MultiBench: Multiscale Benchmarks for Multimodal Representation Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:52:32.252230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:52:32.252230Z digest=sha256:9b44dea0f23ba5d2cbdd50d8df961b11bce2ae250772471bebb2a0598c258983

Observation d08a618e-4352-4159-a23b-d7c550e86355 · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

Efficient Quantification of Multimodal Interaction at Sample Level UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:52:32.260373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:52:32.260373Z digest=sha256:6ec4144fffa98b342d7eda63da8f2568441d2ec059b21dfa8c14e83397fa1952

Observation 97c00d96-9983-42c8-8246-12d007582e89 · outbound

This paper cites Understanding Unimodal Bias in Multimodal Deep Linear Networks.

Efficient Quantification of Multimodal Interaction at Sample Level Understanding Unimodal Bias in Multimodal Deep Linear Networks

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T05:52:32.274858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:52:32.274858Z digest=sha256:092bcc74cc2dbfb635970b71cbc4302df3f16dfffb9307a726252c22bed5309a

Observation c61b8c16-0215-4891-bcc4-63fa9c6a188b · outbound

This paper cites The specific model variant employed in our experiments features unimodal branches, each consisting of four Transformer layers.

Efficient Quantification of Multimodal Interaction at Sample Level The specific model variant employed in our experiments features unimodal branches, each consisting of four Transformer layers

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:52:32.397280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:52:32.281803Z digest=sha256:3b02f09cbbe1efb5c1b9029fd00f924896a2ba793d4bde124280ca15a6162d8e

Observation c7b47a57-c7c6-4ebb-8f22-277125aad31b · outbound

This paper cites Learning Factorized Multimodal Representations.

Efficient Quantification of Multimodal Interaction at Sample Level Learning Factorized Multimodal Representations

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-07T05:52:32.264142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:52:32.264142Z digest=sha256:636789d9fc8d7edf860a4c889457e24a2de2c7d718069150febd2e3d56e08f70

Observation a73c0a1f-38a5-47c7-a49b-2c28abca96b8 · outbound

This paper cites Food-101– mining discriminative components with random forests.

Efficient Quantification of Multimodal Interaction at Sample Level Food-101– mining discriminative components with random forests

Reference 2014

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:52:32.433225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:52:32.235428Z digest=sha256:fb7f5e5b5d2f639d90c8cc309ac0d5b859311b00f860bcdd9c7513aebaf2071f

Observation b134af1a-5520-438f-af05-d5e204067563 · outbound

This paper cites Nonnegative Decomposition of Multivariate Information.

Efficient Quantification of Multimodal Interaction at Sample Level Nonnegative Decomposition of Multivariate Information

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-07T05:52:32.270929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:52:32.270929Z digest=sha256:2746e05b1cda321db4ce8250bfe58878ee2f9846273dc0d59bc4c85a8cbcf562

Observation f76dcdcc-cd3a-4c6d-b06b-7760570f1c0f · outbound

This paper cites Feature Interaction Interpretability: A Case for Explaining Ad-Recommendation Systems via Neural Interaction Detection.

Efficient Quantification of Multimodal Interaction at Sample Level Feature Interaction Interpretability: A Case for Explaining Ad-Recommendation Systems via Neural Interaction Detection

Reference 2018

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:52:32.330835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:52:32.267621Z digest=sha256:21e64c7458ff7f0bb977a39df59d8c09e976328442c96e59c1d92e6e0dc8055d

Observation 5f5153b9-6e44-4bdb-ae19-890ad324f59d · outbound

This paper cites Modality dropout for improved performance-driven talking faces.

Efficient Quantification of Multimodal Interaction at Sample Level Modality dropout for improved performance-driven talking faces

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:52:32.421813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:52:32.248738Z digest=sha256:0bc4f4b1731f4334ae186043f5bc9aa191d38c44f0b4db0d12209d630b1a34ff

Observation 76fc0a54-36e4-440e-bddf-b6fdb68f5579 · outbound

This paper cites Multimodal Learning Without Labeled Multimodal Data: Guarantees and Applications.

Efficient Quantification of Multimodal Interaction at Sample Level Multimodal Learning Without Labeled Multimodal Data: Guarantees and Applications

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T05:52:32.256147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:52:32.256147Z digest=sha256:51f56bb4b1288f2fd6ffb82da1a0432f2fc0854fb8254338eac199ba69e40450

Observation c7ac0c3c-ab6a-4d69-8447-200c95f06dec · outbound

This paper cites UR-FUNNY: A Multimodal Language Dataset for Understanding Humor.

Efficient Quantification of Multimodal Interaction at Sample Level UR-FUNNY: A Multimodal Language Dataset for Understanding Humor

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T05:52:32.244311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:52:32.244311Z digest=sha256:e5471cc646168144c3f9c6801cdc45b33f009c13f7996ec332681d03e585c7e1

Observation 20d53c37-4dac-4566-8139-7ece982f0470 · outbound

This paper cites Multimodal Compact Bilinear Pooling for Visual Question Answering and Visual Grounding.

Efficient Quantification of Multimodal Interaction at Sample Level Multimodal Compact Bilinear Pooling for Visual Question Answering and Visual Grounding

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T05:52:32.240171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:52:32.240171Z digest=sha256:84ce3e044cd527d7a1b4945888ca2fb09f9cdeb59dfb6f6cfda248b19c518709

Pith citing papers

No inbound Pith citation observations are available.