Pith. sign in

Paper Citation Record · LEDGER

Meta-Transformer: A Unified Framework for Multimodal Learning

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 19 inbound Pith citation observations for arXiv:2307.10802.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2307.10802 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 19 of 19 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T15:45:46.182206Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T13:43:28.559460Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 094c9896-949c-45d8-8035-c54f68d97a4b · inbound

LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment cites this paper.

LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment Meta-Transformer: A Unified Framework for Multimodal Learning

Reference 201

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:27:59.100370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-17T03:27:58.952076Z digest=sha256:410a6d4fda7636cf9d36d23a7a615da1cd08d99363b07a8a213761af993183da

Observation b28a1464-fd51-485c-b3ce-5993f5cd603c · inbound

Fine-tuning Multimodal Transformers on Edge: A Parallel Split Learning Approach cites this paper.

Fine-tuning Multimodal Transformers on Edge: A Parallel Split Learning Approach Meta-Transformer: A Unified Framework for Multimodal Learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T15:45:46.182206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:45:46.182206Z digest=sha256:7c9c2d0ecc389562ffde08ee1eabf106e17aa5bb8910bdc3273284c4508c5cd1

Observation 7d3b72f3-9ac3-47dc-971f-7913d16d98b5 · inbound

VisCon-100K: Leveraging Contextual Web Data for Fine-tuning Vision Language Models cites this paper.

VisCon-100K: Leveraging Contextual Web Data for Fine-tuning Vision Language Models Meta-Transformer: A Unified Framework for Multimodal Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T18:51:10.466981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:51:10.466981Z digest=sha256:9fdb48b09ba511b925f8b1eea6198f9cd7d8281eb91ed18e201a5b438735b4a0

Observation bf6a3967-dc8c-4814-90ad-81b871184152 · inbound

CoNav: Collaborative Cross-Modal Reasoning for Embodied Navigation cites this paper.

CoNav: Collaborative Cross-Modal Reasoning for Embodied Navigation Meta-Transformer: A Unified Framework for Multimodal Learning

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:56.782916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:56.782916Z digest=sha256:7bdb6b5ced26df0d4ac57c3597cfe6a2128f6b81d95f5127dfba87448d2feb13

Observation 5eb4434e-ca3b-48c6-981e-c21030d14ef3 · inbound

CSTrack: Enhancing RGB-X Tracking via Compact Spatiotemporal Features cites this paper.

CSTrack: Enhancing RGB-X Tracking via Compact Spatiotemporal Features Meta-Transformer: A Unified Framework for Multimodal Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:16.348830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:16.348830Z digest=sha256:ff852ee061c1bcbada656529843260509ad39a47a0a1cc6c10a1904e5e1f4757

Observation 03313ad1-34f4-45bd-87f8-5a4fd52734b5 · inbound

Multimodal Representation Alignment for Cross-modal Information Retrieval cites this paper.

Multimodal Representation Alignment for Cross-modal Information Retrieval Meta-Transformer: A Unified Framework for Multimodal Learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T05:08:31.058312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:08:31.058312Z digest=sha256:05b9cd2c87793adb3f641ae5d050c53985ec27e7669093db5330d69292978a77

Observation 073e77fb-ad37-4dd2-97d6-dd1e769597b4 · inbound

UniPre3D: Unified Pre-training of 3D Point Cloud Models with Cross-Modal Gaussian Splatting cites this paper.

UniPre3D: Unified Pre-training of 3D Point Cloud Models with Cross-Modal Gaussian Splatting Meta-Transformer: A Unified Framework for Multimodal Learning

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T04:41:03.961669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:41:03.961669Z digest=sha256:e365a701e7fd3b503304fab43c5ba13f82149d2fb16f6ef98622a77917deeacf

Observation 999da810-be2a-47f6-a707-6f033554932b · inbound

Vision Generalist Model: A Survey cites this paper.

Vision Generalist Model: A Survey Meta-Transformer: A Unified Framework for Multimodal Learning

Reference 204

Resolution
unresolved
no resolver link, observed 2026-08-07T04:44:05.332988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:44:05.332988Z digest=sha256:1fe73219c50419c4b46ccb673d261c8ce0e793549c6774d52c5dd2f53f7004d0

Observation e1b352cc-fa26-4045-88c2-71c9d46df71c · inbound

A Comprehensive Survey on Video Scene Parsing:Advances, Challenges, and Prospects cites this paper.

A Comprehensive Survey on Video Scene Parsing:Advances, Challenges, and Prospects Meta-Transformer: A Unified Framework for Multimodal Learning

Reference 197

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:24.785026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:34:24.785026Z digest=sha256:7d5a0a78fdbee359383a41da05a1cabafbc0e87371ded3214ce494536b1f3040

Observation a5bb0568-c375-4117-9572-1c43c69c6999 · inbound

Grounding Intelligence in Movement cites this paper.

Grounding Intelligence in Movement Meta-Transformer: A Unified Framework for Multimodal Learning

Reference 102

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:32.736057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:32.736057Z digest=sha256:533c417c28dbfb5b74b9d3331c7048418a40de2f3a947d5def56402e5c4f4ed8

Observation af7aa99d-b892-47ac-a506-e0154c640fe8 · inbound

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning cites this paper.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Meta-Transformer: A Unified Framework for Multimodal Learning

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:24.267156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:24.267156Z digest=sha256:ef305178a65d440a97f6ff4533f7f49bad81ad3fbf3cc5a769f9219cc5b58143

Observation bb67f243-f6f9-4431-a34b-1e7889d8cf85 · inbound

SkySense V2: A Unified Foundation Model for Multi-modal Remote Sensing cites this paper.

SkySense V2: A Unified Foundation Model for Multi-modal Remote Sensing Meta-Transformer: A Unified Framework for Multimodal Learning

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-06T16:21:14.242806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:21:14.242806Z digest=sha256:a303be60c8c164782b61f83c501119676f31d177d9e87ab2134711938889f4a8

Observation a81cb44e-c014-4261-ba5b-b64825c1889b · inbound

Toward Long-Tailed Online Anomaly Detection through Class-Agnostic Concepts cites this paper.

Toward Long-Tailed Online Anomaly Detection through Class-Agnostic Concepts Meta-Transformer: A Unified Framework for Multimodal Learning

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-06T15:06:24.737211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:06:24.737211Z digest=sha256:5a24966ed258ac1075a931bf3508890cf28b2714d62a3f5f5b34e9d830af3d32

Observation 8da0476f-1283-4e1f-a0f4-4179457ab785 · inbound

PiERN: Token-Level Routing for Integrating High-Precision Computation and Reasoning cites this paper.

PiERN: Token-Level Routing for Integrating High-Precision Computation and Reasoning Meta-Transformer: A Unified Framework for Multimodal Learning

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T16:06:34.928800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T16:05:57.505355Z digest=sha256:375cca2c78cabf43c8a9b767317947eea1a873e418e442ed31798be97fca7821

Observation 95e42abf-9c27-4b0d-afb5-a5db37cc3d07 · inbound

QA-MoE: Towards a Continuous Reliability Spectrum with Quality-Aware Mixture of Experts for Robust Multimodal Sentiment Analysis cites this paper.

QA-MoE: Towards a Continuous Reliability Spectrum with Quality-Aware Mixture of Experts for Robust Multimodal Sentiment Analysis Meta-Transformer: A Unified Framework for Multimodal Learning

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:30:50.439578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T19:06:47.812339Z digest=sha256:c62ad241d14d704641e54738c968bceef84b564420ba1a3094e46e781987a437

Observation c7895d85-750b-4ae9-8d65-2a04bb2e1486 · inbound

MedMIX: Modality-Internal Expert Fusion for Multimodal Medical Diagnosis cites this paper.

MedMIX: Modality-Internal Expert Fusion for Multimodal Medical Diagnosis Meta-Transformer: A Unified Framework for Multimodal Learning

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:43:44.196528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T19:38:56.999644Z digest=sha256:8a8d4a42c29c450d4b80f56a18c9af3440201fc6054f6a09dbce69e79b27dc0d

Observation 2e4e840f-7192-4aac-80fa-7d2c87bfc363 · inbound

CodeBind: Decoupled Representation Learning for Multimodal Alignment with Unified Compositional Codebook cites this paper.

CodeBind: Decoupled Representation Learning for Multimodal Alignment with Unified Compositional Codebook Meta-Transformer: A Unified Framework for Multimodal Learning

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:43:12.558257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T10:40:56.726020Z digest=sha256:b8f2f032c366d7bd8739ba2eb92ee963cc6b1a02268c8282a7b82570dca1b7c0

Observation 746ca5dc-4c1f-4614-b2aa-0a2feab58d3e · inbound

AREA: Attribute Extraction and Aggregation for CLIP-Based Class-Incremental Learning cites this paper.

AREA: Attribute Extraction and Aggregation for CLIP-Based Class-Incremental Learning Meta-Transformer: A Unified Framework for Multimodal Learning

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:43:28.561098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T13:41:50.250899Z digest=sha256:e1740f8e06ae6cee96eeec163124b72f0752099ffebb1752a37ce2ddbd3ab47c

Observation d9ed1494-8d8a-42e7-9ef6-8ce8174919a0 · inbound

Mixture of Probes: Learning from Privileged Modalities in Multimodal LLMs Through Probing cites this paper.

Mixture of Probes: Learning from Privileged Modalities in Multimodal LLMs Through Probing Meta-Transformer: A Unified Framework for Multimodal Learning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-07-13T06:21:17.706173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T06:21:17.706173Z digest=sha256:13216b793aeab353487f39177f7ba69671b3c191bdd43dbab8b23a5097ef9d5d