Pith. sign in

Paper Citation Record · LEDGER

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer

As of 18 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 0 inbound Pith citation observations for arXiv:2506.11465.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.11465 v1

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:09:43.057849Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

61 of 61 outbound references displayed

  • verified exact2
  • verified fuzzy34
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 50dfc4cd-ded7-4194-bd00-11fd38177eae · outbound

This paper cites write newline.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T04:09:38.068767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:09:38.068767Z digest=sha256:79d5d629b2dff7f435e30c52e484008a0584bdf7e5a6cf8a5bed5ea6349f0627

Observation d76e886a-75f2-4af7-a52a-86e4cffbefcb · outbound

This paper cites and Zisserman, A.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer and Zisserman, A

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:49.559290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:09:38.226847Z digest=sha256:47621742c3f505cb9ffe66900e968499684f3574464cf91f0be1df868145e358

Observation 09d7dbc0-1d54-49b5-a6f2-ecb3c3651dfa · outbound

This paper cites Singular value decomposition tutorial.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer Singular value decomposition tutorial

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:49.544180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:09:38.350243Z digest=sha256:0ca8dba3f9c5d95a54ed77f64e74bb3109fbac137fa6c1b76c7bad5cad55d31e

Observation 82bf73e3-e43b-46b8-8094-7ab82b49670d · outbound

This paper cites Multimodal machine learning: A survey and taxonomy.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer Multimodal machine learning: A survey and taxonomy

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:49.530010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:09:38.427179Z digest=sha256:233f4d0716536b8ef6d5733a0db150e953aadccf76e6c315b7a711af193b6476

Observation 2cafc4cc-ed47-42e3-880b-9fcd5e97f458 · outbound

This paper cites Training stochastic model recognition algorithms as networks can lead to maximum mutual information estimation of parameters.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer Training stochastic model recognition algorithms as networks can lead to maximum mutual information estimation of parameters

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:49.515557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:09:38.498054Z digest=sha256:3c6086f0420f1fc4f0828d82d01c52ea93446c2fe7584898820628a617de9b2b

Observation ec1d95fd-8ee2-48a8-accb-0c867feded1f · outbound

This paper cites G., Keutmann, M.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer G., Keutmann, M

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:49.500337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:09:38.594262Z digest=sha256:fdec849380b48e2a152ff248b0f876fa3b871ec2fac502a7e01d6cd74713ba9f

Observation e20a43c3-110e-47fb-afea-7212f0bc8b75 · outbound

This paper cites End-to-end object detection with transformers.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer End-to-end object detection with transformers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T04:09:38.679875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:09:38.679875Z digest=sha256:fe8d7c50c5b1a2b0f575eea1db237ba045f250901b789f21f18e111be81ea563

Observation 7885befd-1227-4d1b-8bef-696afd9da023 · outbound

This paper cites Vggsound: A large-scale audio-visual dataset.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer Vggsound: A large-scale audio-visual dataset

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:49.477032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:09:38.753191Z digest=sha256:2d38a87ffbde722e104457a96353d080375ce62f75b1ed9ee7402e90b178764b

Observation e4c3c1a8-38f2-4f74-8bdd-d088994c2cca · outbound

This paper cites T., Rubanova, Y., Bettencourt, J., and Duvenaud, D.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer T., Rubanova, Y., Bettencourt, J., and Duvenaud, D

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T04:09:38.823412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:09:38.823412Z digest=sha256:4e5c867d765ec8da60ef93ef36889fe7bf9bb15186521b2e4a9434973dcba984

Observation 4786108c-a60a-452b-8d9a-2da6020d58dc · outbound

This paper cites Self-attention fusion for audiovisual emotion recognition with incomplete data.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer Self-attention fusion for audiovisual emotion recognition with incomplete data

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:49.453909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:09:38.893997Z digest=sha256:d5872e6f6b3a476148423fad094db54a8ad6e81e4da09d045ee0ec1235c06a36

Observation cf10ac09-1ebc-4fd9-a007-290cdf6daac3 · outbound

This paper cites What Does BERT Look At? An Analysis of BERT's Attention.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer What Does BERT Look At? An Analysis of BERT's Attention

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T04:09:38.980072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:09:38.980072Z digest=sha256:3059a28c074ee07ef3b8c837af664a6e57f2cc84f825f57fbef53064a4f49600

Observation 0551a8ae-6918-4c74-8dfb-8f47cf9b286c · outbound

This paper cites Addressing failure prediction by learning model confidence.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer Addressing failure prediction by learning model confidence

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:49.438443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:09:39.064181Z digest=sha256:e561b985cbd562b8c998e795f3a5834d2c315abcc10584e5053ad513dd6dd0e0

Observation 55f3978c-c448-4684-ae99-900215ee25cf · outbound

This paper cites MultiOOD: Scaling Out-of-Distribution Detection for Multiple Modalities.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer MultiOOD: Scaling Out-of-Distribution Detection for Multiple Modalities

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:09:43.650849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:09:39.130768Z digest=sha256:2046783a5718666fe9d4bc78104a27932c1f5a5089a0d2149e24b75dea424d06

Observation b7bb59bc-732c-44d1-b58a-d335e21ac48a · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T04:09:39.218605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:09:39.218605Z digest=sha256:eda245bba0331393084d021eff31abfb8f105064ad3686f547a7b237f05f206b

Observation f19b3643-e38e-41ca-813b-463a025a7aaf · outbound

This paper cites Pmr: Prototypical modal rebalance for multimodal learning.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer Pmr: Prototypical modal rebalance for multimodal learning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:49.422317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:09:39.294459Z digest=sha256:23da509d2f2728d93a5916b6a3af9b04f214b11bd57b7d4d8f2ada2370af6c4f

Observation 545a75d6-ae72-469f-a773-b66ee6fd6e5f · outbound

This paper cites A survey on deep learning for multimodal data fusion.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer A survey on deep learning for multimodal data fusion

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:49.341220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:09:39.356065Z digest=sha256:6f6eb2ced0636b137c67f03608ae902798f223a9f74bf53097585d0176c23ddc

Observation a4fd3df9-1b5d-4eb6-8613-e142ab914055 · outbound

This paper cites C., Wang, X., and Li, H.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer C., Wang, X., and Li, H

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:49.199632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:09:39.433527Z digest=sha256:bea791625a939537320603dca7012a0813d6d6d895f65839602f61313b82387a

Observation 8820c915-2613-45bc-ac88-5c32f18a1484 · outbound

This paper cites A mathematical perspective on Transformers.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer A mathematical perspective on Transformers

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T04:09:39.501652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:09:39.501652Z digest=sha256:93e9aa029904ecce2d5a3311d9fa89552cd5c1e5274e8969b1771b2109f67318

Observation 03b11aec-86b4-4fd2-9e88-275a6673e170 · outbound

This paper cites Multimodal dynamics: Dynamical fusion for trustworthy multimodal classification.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer Multimodal dynamics: Dynamical fusion for trustworthy multimodal classification

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:49.139355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:09:39.571302Z digest=sha256:43f107b3e3df7f93afd1b15ace172a27980d77fe0901db6e33b172211d4ebb10

Observation 0074a38c-b600-4ccb-8213-a9b30af6ad1f · outbound

This paper cites Deep residual learning for image recognition.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer Deep residual learning for image recognition

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T04:09:39.650765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:09:39.650765Z digest=sha256:39c211245b115ba42a0bcee2a8b2858db581166954e468061b9cadc4d36c2d70

Observation ea0983aa-2509-4c86-9810-d1956fd73305 · outbound

This paper cites Masked autoencoders are scalable vision learners.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer Masked autoencoders are scalable vision learners

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T04:09:39.718540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:09:39.718540Z digest=sha256:c0020c44af6d34b4658781d8a408ddcf8838b3c99cec88beb383f50dc9b232e6

Observation 6b0631c4-cabe-4c4a-baa7-ce27b0913306 · outbound

This paper cites ReconBoost: Boosting Can Achieve Modality Reconcilement.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer ReconBoost: Boosting Can Achieve Modality Reconcilement

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T04:09:39.795317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:09:39.795317Z digest=sha256:bbb4c6c7b7c77b48ba3b68d715cd0799a63c19fa7c21108759b18844f1396312

Observation 7aaf9aba-4f25-42eb-8a9f-d8ec915f37d2 · outbound

This paper cites Adaptive unimodal regulation for balanced multimodal information acquisition.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer Adaptive unimodal regulation for balanced multimodal information acquisition

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:49.090763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:09:39.862671Z digest=sha256:2e41ad82c5a9f95f15b84e745d85d586763cfab2f7b832bdf3653de99d7bffbb

Observation e306a950-b2aa-40f3-8948-a3b310b95012 · outbound

This paper cites Modality competition: What makes joint training of multi-modal network fail in deep learning?(provably).

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer Modality competition: What makes joint training of multi-modal network fail in deep learning?(provably)

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:49.016974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:09:39.922194Z digest=sha256:2de202c3392137839877c057631acce8b53c2265e364c951dc670215a33976bd

Observation 23ea6d0e-235a-495f-bd8c-67dcabef7be7 · outbound

This paper cites Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T04:09:39.997161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:09:39.997161Z digest=sha256:ac0b9230d5f0785277b2f28f4ba2f3332edf57559a7548069f05183b2757c8ae

Observation 25da82be-d28c-4481-bd7e-c6c802b68b09 · outbound

This paper cites an unresolved cited work.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:09:48.867236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:09:40.072016Z digest=sha256:985dd4b6362859975459c4ff9b9401c287dffcce92fff68ac8495eb309c67c61

Observation 4320cb08-331b-4022-b18d-cbd218295782 · outbound

This paper cites Vilt: Vision-and-language transformer without convolution or region supervision.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer Vilt: Vision-and-language transformer without convolution or region supervision

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:48.736558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:09:40.154442Z digest=sha256:ddf798e33ab4b207c2c3b1b3980d2c8fba025296d7216aa5a84ea207ecf712f5

Observation dd8abe0d-cd06-417d-a8f4-04b7f5f7f8d8 · outbound

This paper cites Revealing the Dark Secrets of BERT.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer Revealing the Dark Secrets of BERT

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T04:09:40.244763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:09:40.244763Z digest=sha256:e73e7313fc35946b0343fbc084954aea387ce510f25c248984d97a27eea32b25

Observation e4c048b0-6165-4175-9943-dff252466a8f · outbound

This paper cites Hmdb: a large video database for human motion recognition.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer Hmdb: a large video database for human motion recognition

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:48.699781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:09:40.308098Z digest=sha256:be1e45390be4a9177a94e9ef3710aef71dc67223f5a4629eceb6a29c1fc890f5

Observation d5a6dd0e-544b-4395-a0b8-db9e4b7cb64e · outbound

This paper cites VisualBERT: A Simple and Performant Baseline for Vision and Language.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer VisualBERT: A Simple and Performant Baseline for Vision and Language

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T04:09:40.379262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:09:40.379262Z digest=sha256:f4c2988938b47c75db70034c14b9a6e9c4efabf40a0d4cbfda8e8b268cc7f479

Observation 9ff55cff-2ff8-4236-b7ff-edc47bc4d7aa · outbound

This paper cites P., Lyu, Y., Fan, X., Wu, Z., Cheng, Y., Wu, J., Chen, L., Wu, P., Lee, M.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer P., Lyu, Y., Fan, X., Wu, Z., Cheng, Y., Wu, J., Chen, L., Wu, P., Lee, M

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:48.581228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:09:40.437438Z digest=sha256:66a4b950a29fffb03e0b3222d012338bef269f261dee2c0da664ae726b287dba

Observation a18a3558-b526-4404-a362-657ea7b69717 · outbound

This paper cites Foundations and Trends in Multimodal Machine Learning: Principles, Challenges, and Open Questions.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer Foundations and Trends in Multimodal Machine Learning: Principles, Challenges, and Open Questions

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T04:09:40.521348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:09:40.521348Z digest=sha256:a44a1746589bfcad48458e010ebde9421ff19c07844c8d051d8ef880e5eafd6d

Observation d7a3ff90-7b6b-4e7e-b8a0-5b3b1f6a6866 · outbound

This paper cites Efficient Low-rank Multimodal Fusion with Modality-Specific Factors.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer Efficient Low-rank Multimodal Fusion with Modality-Specific Factors

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T04:09:40.621399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:09:40.621399Z digest=sha256:ddb0b88cc1bb9fb4e62619c37fd0e7bfdecced25a65c608cda26a52be0a97163

Observation e0e52db5-9606-435e-a9c9-4036ca3f6cdb · outbound

This paper cites Attention bottlenecks for multimodal fusion.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer Attention bottlenecks for multimodal fusion

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:48.420532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:09:40.697091Z digest=sha256:692ed4e03f4280ec04b08ca0562168b2c9b06c1ee684618aa797ef23d1b4296c

Observation 239cea9e-0360-41c7-bea2-1287ac56cd59 · outbound

This paper cites Balanced multimodal learning via on-the-fly gradient modulation.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer Balanced multimodal learning via on-the-fly gradient modulation

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:48.259625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:09:40.821931Z digest=sha256:28172d4492047c1b5fc7f28c22e935bc4aa3e1862509b43fb31b7213f10c1b19

Observation 547cea5c-7798-47df-9452-779a9b14c033 · outbound

This paper cites Dissecting Multimodality in VideoQA Transformer Models by Impairing Modality Fusion.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer Dissecting Multimodality in VideoQA Transformer Models by Impairing Modality Fusion

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T04:09:40.892599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:09:40.892599Z digest=sha256:bb00db75d81837bb6e31f62983bfe270c604ad3b8a470324e0f881be110a25a3

Observation 16ce0f8b-0daf-4bc3-967d-d938b0dc3d2b · outbound

This paper cites Zorro: the masked multimodal transformer.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer Zorro: the masked multimodal transformer

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T04:09:40.977183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:09:40.977183Z digest=sha256:765cccfd36bcd5ad91e0f9d860f979251ec3bbf23ecca408d07e94f2499e7ffe

Observation d46087ef-f70c-4d11-83e1-910194c6e1e2 · outbound

This paper cites ImageNet-21K Pretraining for the Masses.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer ImageNet-21K Pretraining for the Masses

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T04:09:41.086921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:09:41.086921Z digest=sha256:9afb42dd9566251b781a8b85ebc165f2548678392e99f1d9f09af7ecc70dba01

Observation ea171b79-85fd-47e4-ae91-6a6c7b111834 · outbound

This paper cites and Monro, S.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer and Monro, S

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:48.138516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:09:41.216679Z digest=sha256:2fcd54786ce0e4cbd692745c81ff79bc6e9e1064e89542103420f9a89123bfbc

Observation ebcad2ad-3b0e-4e4e-9aba-961f1d4e01b2 · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T04:09:41.333805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:09:41.333805Z digest=sha256:a8238f617e2416a215c1ead303cf07631300564584ede546b58430ce3771115f

Observation 5ffad669-1854-4da8-9703-6fcf17fd33b9 · outbound

This paper cites L., Tickoo, O., and Huang, J.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer L., Tickoo, O., and Huang, J

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:47.883478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:09:41.440088Z digest=sha256:7c8e01c25a7b1fc3516ee81bccc2ffa920bb6ef2f3f2ebafd52867202302832a

Observation d0858fa7-edf5-4ef4-812c-4c2c3767b01d · outbound

This paper cites H., Bai, S., Liang, P.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer H., Bai, S., Liang, P

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:47.606166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:09:41.552413Z digest=sha256:ef32f8afa1ab0b83f35c56bdb9c2351c7983c40c90d429425554221c299e7e56

Observation 2fe37389-a3c9-4df0-afb3-bb45496bb3d9 · outbound

This paper cites Attention is all you need.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer Attention is all you need

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T04:09:41.645870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:09:41.645870Z digest=sha256:1743657690fe4bed8d2c6e42aef9d293112d03cf3e85e876d9ca317663b100c8

Observation ca52aa6e-4920-443d-9036-e6abdfdecfe2 · outbound

This paper cites H., Zeeshan, M.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer H., Zeeshan, M

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:47.235792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:09:41.762413Z digest=sha256:c7ab04a3d30862cfb16ef059a1136f2f3f99086f6a407ff7290c56327e4584bb

Observation a82048ad-0376-4017-b5ad-0b4f1e841597 · outbound

This paper cites What makes training multi-modal classification networks hard? In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp.\ 12695--12705, 2020 a.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer What makes training multi-modal classification networks hard? In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp.\ 12695--12705, 2020 a

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:47.029696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:09:41.879818Z digest=sha256:4e37b42662e7d6bbdfc797b31304f203b8e19c313d62bcd0d2241de2c1493990

Observation a99a2d56-7b1f-4a7f-a537-54692516e151 · outbound

This paper cites Deep multimodal fusion by channel exchanging.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer Deep multimodal fusion by channel exchanging

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:46.828432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:09:41.996523Z digest=sha256:4ba1122c1e2ef992c0c0e05c9067d9dcbd89c44d233adf19647a23d209a8634b

Observation f21ef0bb-bf4f-461c-800f-d0a1362c8203 · outbound

This paper cites Multimodal token fusion for vision transformers.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer Multimodal token fusion for vision transformers

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:46.490537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:09:42.101140Z digest=sha256:dd17c1d1d2bcbb3a7dcd0212a65b05f5edd360ea7469032ff69f9c12dc0b24c4

Observation 2d4e5a44-063e-4e35-8056-c28ae24e966e · outbound

This paper cites an unresolved cited work.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:09:46.245419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:09:42.170844Z digest=sha256:fc706b8209203794bfc8934bbc0125dc79083ea87e859eb0a2b35eb6e2f50de0

Observation 7323a5ea-4aaa-4e4c-aa74-bd1a0b5851e2 · outbound

This paper cites Enhancing multimodal cooperation via sample-level modality valuation.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer Enhancing multimodal cooperation via sample-level modality valuation

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:46.014133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:09:42.227002Z digest=sha256:6ec64374b3832c13935cdea379282e4aade1a9ab18cb26d1e2daedcdd5ea8a01

Observation a1a04a73-0842-4a2e-beba-ec1c69b862ca · outbound

This paper cites an unresolved cited work.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:09:45.793727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:09:42.283576Z digest=sha256:fb7b3d81390fb2415141986af7c3073d65a1c878d4c9b2755f8cadb5271f2f1b

Observation 9441ca14-8e29-4259-abb4-a2903efd00ea · outbound

This paper cites Multimodal fusion with co-attention networks for fake news detection.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer Multimodal fusion with co-attention networks for fake news detection

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:45.527009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:09:42.335445Z digest=sha256:0f481c781392c4f292e7aedfe0d9478a99dc9c83e7d27584932350dd0709dc96

Observation 049e80b9-65b8-495a-8450-e1a6d4325ab9 · outbound

This paper cites Multimodal multi-loss fusion network for sentiment analysis.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer Multimodal multi-loss fusion network for sentiment analysis

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:45.332401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:09:42.388586Z digest=sha256:1acff9aa61d269c5a5e63e1ccadd1ced83d19b913544f7e0af3bd6e9a8a53109

Observation 2b312de6-de2a-406f-800e-4eb2aee2f3c7 · outbound

This paper cites Balanced Audiovisual Dataset for Imbalance Analysis.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer Balanced Audiovisual Dataset for Imbalance Analysis

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:09:43.246686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:09:42.449291Z digest=sha256:3ad1b4b7d230024c0ca4151dcae8f1409c68c653e326904309abf8235618863a

Observation f9ef6733-d0c6-460b-9bd2-99b47c4c3490 · outbound

This paper cites an unresolved cited work.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:09:45.069587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:09:42.505015Z digest=sha256:3194cc8d9323c276343280d2969ab45ae1f72df35acab807a9f31da411cf21d3

Observation 66ed191c-5d3d-4682-beeb-b3e45e79ae60 · outbound

This paper cites Facilitating multimodal classification via dynamically learning modality gap.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer Facilitating multimodal classification via dynamically learning modality gap

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:44.948613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:09:42.553697Z digest=sha256:cac7d825e97006a146b5b69c3454cba0cfdf19e69797ede95f57927b029721a8

Observation f45c8981-637e-44af-9e1a-b8e6becbbf24 · outbound

This paper cites Learning to rebalance multi-modal optimization by adaptively masking subnetworks.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer Learning to rebalance multi-modal optimization by adaptively masking subnetworks

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:44.656662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:09:42.619478Z digest=sha256:b6815e06adff4ac2b6cdc52c70c44a7da7b724358df007c3578a4a43b56ac093

Observation e4aab786-902f-4950-b114-ce9358f957a4 · outbound

This paper cites Deep modular co-attention networks for visual question answering.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer Deep modular co-attention networks for visual question answering

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:44.419762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:09:42.689333Z digest=sha256:eae5179392987ecb2d8965a31453fe450685778dfa2fb87607e24f8f66c68dc6

Observation 1b29b99a-b9c2-4b6a-87ff-895202aa7c49 · outbound

This paper cites Tensor Fusion Network for Multimodal Sentiment Analysis.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer Tensor Fusion Network for Multimodal Sentiment Analysis

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T04:09:42.779895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:09:42.779895Z digest=sha256:1aabb31f8cf3829257c8e1f2ab8de18fa60d49d4aa260b29f4dae325aefbbafc

Observation 33865e64-506f-4578-9ab0-f0365b44c665 · outbound

This paper cites B., Liang, P.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer B., Liang, P

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:44.221027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:09:42.869769Z digest=sha256:453955d603d302ab06a91ea2a2e754d6c6c9a61957649caf4a388c1f435ace9e

Observation d43883cf-fe5d-4810-887c-33cfd95e3fd6 · outbound

This paper cites T., and Peng, X.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer T., and Peng, X

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:43.959333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:09:42.959924Z digest=sha256:6a2effdd8a3dbb8cee35f2a1836b00921abdf634251e2c046b4d5334aed8e942

Observation c2704c55-8763-4037-aeaa-3d9c784a8eef · outbound

This paper cites Multimodal Fusion on Low-quality Data: A Comprehensive Survey.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer Multimodal Fusion on Low-quality Data: A Comprehensive Survey

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T04:09:43.057849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:09:43.057849Z digest=sha256:28e31d573eec4fb9601dbe672c9a156d14b3cd79c61bbffd82c3b5e125ed103b

Pith citing papers

No inbound Pith citation observations are available.