Pith. sign in

Paper Citation Record · LEDGER

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework

As of 22 August 2026, this Paper Citation Record lists 66 of 66 outbound references and 0 inbound Pith citation observations for arXiv:2504.12576.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.12576 v1

Coverage vector

measured 66 of 66 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:33:45.234194Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

66 of 66 outbound references displayed

  • verified exact0
  • verified fuzzy34
  • unresolved32
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8ba16635-3d8c-4301-b6a1-6f41a79519c6 · outbound

This paper cites Multimae: Multi-modal multi-task masked autoen- coders.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Multimae: Multi-modal multi-task masked autoen- coders

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:33:46.159804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T12:33:44.920575Z digest=sha256:6e98f3364c45e55ee840c76323f783d247ff08b816073c8b0ee42c5d09f423ed

Observation 5703e613-8c86-4795-86d5-1de470fba762 · outbound

This paper cites Beit: Bert pre-training of image transformers.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Beit: Bert pre-training of image transformers

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:33:46.144879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T12:33:44.925773Z digest=sha256:14b0658628c1a1294e8d745e3723358b89c28645245be7b206c08ad9337d5390

Observation df0011cb-5772-4953-90df-ea39ecaf7b04 · outbound

This paper cites Is Space-Time Attention All You Need for Video Understanding?.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Is Space-Time Attention All You Need for Video Understanding?

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T12:33:44.930487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:33:44.930487Z digest=sha256:f3b1702318ec672d210f508c64a9b382ffe90515816044ac80128a112e224a38

Observation a40b3fa2-2778-4886-8096-0e6779e508ee · outbound

This paper cites DeepSeek LLM: Scaling Open-Source Language Models with Longtermism.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework DeepSeek LLM: Scaling Open-Source Language Models with Longtermism

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T12:33:44.935504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:33:44.935504Z digest=sha256:eb40457584d4aaeb2dcd311c785d6172b477724d879d73f3049060af245c95b5

Observation 9860ca14-a597-4263-b046-560ef9db0abc · outbound

This paper cites Lan- guage models are few-shot learners.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Lan- guage models are few-shot learners

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T12:33:44.940844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:33:44.940844Z digest=sha256:7e6bdc05c731670e4d4ac0ca55f2ce792e98fbf713e11c7f77ece18053a1ab8c

Observation aa93e432-08d9-44cd-b267-138a72124aeb · outbound

This paper cites End-to- end object detection with transformers.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework End-to- end object detection with transformers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T12:33:44.946536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:33:44.946536Z digest=sha256:d95e8a9b09a18bff42f80babe9539676cf98590681a93b5c854865ff48323db9

Observation c08d217c-1bb1-4b4a-8664-9067cb276728 · outbound

This paper cites Weakly misalignment-free adaptive feature alignment for uavs- based multimodal object detection.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Weakly misalignment-free adaptive feature alignment for uavs- based multimodal object detection

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:33:46.110267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T12:33:44.951657Z digest=sha256:5900e951e48627baa89f258ec94f192f3a2868ebfc2e36807dbbb506c06bd277

Observation dfc6f7f1-7290-4ce9-ac27-f489d5702f41 · outbound

This paper cites An empirical study of training self-supervised vision transformers.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework An empirical study of training self-supervised vision transformers

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:33:46.093955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T12:33:44.956835Z digest=sha256:209aa0a1b06a299e5cfee2c62d5da56a50cbc7e536269eb2d381b84408b7c658

Observation 2c7e24f8-b6e3-4255-b13a-092e665e96e1 · outbound

This paper cites Segment any event streams via 11 weighted adaptation of pivotal tokens.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Segment any event streams via 11 weighted adaptation of pivotal tokens

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:33:46.077155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T12:33:44.961205Z digest=sha256:bcdf45240a831e3a842073d4ba3f7031b7e99ac71c525111aa29a95ae55b6fa1

Observation 3439abeb-f364-49b7-b0f9-4d799dd7798b · outbound

This paper cites Unihcp: A unified model for human-centric perceptions.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Unihcp: A unified model for human-centric perceptions

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:33:46.061113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T12:33:44.965721Z digest=sha256:d42a418cfbe615b4c5d871a2d5da5f577893c111fcf33c8fa34d94cc2ac32382

Observation 38d2fb5d-d24e-432b-b290-d20ae97bc3c3 · outbound

This paper cites Satmae: Pre-training transformers for tem- poral and multi-spectral satellite imagery.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Satmae: Pre-training transformers for tem- poral and multi-spectral satellite imagery

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:33:46.046607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T12:33:44.970391Z digest=sha256:70a69f91d1901e89dcf2cce36be53a16e0bb5a314b0b34f2d6487b2b50d8e049

Observation 98ac94c9-fdf6-4ef3-8dc5-d253015a5fa4 · outbound

This paper cites Deepseekmoe: Towards ultimate expert spe- cialization in mixture-of-experts language models.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Deepseekmoe: Towards ultimate expert spe- cialization in mixture-of-experts language models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:33:46.030964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T12:33:44.974774Z digest=sha256:ccc2f9a25055ba58609065dd78e7cea393164fd296aa9ba30ed0b3a282e7c33d

Observation f59cbf1d-53eb-4070-8f65-f78cd35eb765 · outbound

This paper cites Sfod: Spiking fusion object detector.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Sfod: Spiking fusion object detector

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:33:46.015958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T12:33:44.980632Z digest=sha256:c4d29b0ac3e48ac8b83f7f7fbbadf9969d99f04812f79ca9e72062fc522c6573

Observation 9980f297-1682-430c-8ef0-1241ade1e83f · outbound

This paper cites Hypergraph-based multi-view action recognition using event cameras.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Hypergraph-based multi-view action recognition using event cameras

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:33:46.000459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T12:33:44.985177Z digest=sha256:c6ccffb72c863cc6e6e0de768a5dbdaea3d06ef215a5109add178f057350203d

Observation de460341-6127-4e57-8104-119f7d72db4d · outbound

This paper cites Multimodal Masked Autoencoders Learn Transferable Representations.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Multimodal Masked Autoencoders Learn Transferable Representations

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T12:33:44.989651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:33:44.989651Z digest=sha256:7d38bca2405949c80b881c9349314319ca3203f18e35b09daa21fc16fa12b602

Observation fe194490-520d-45cd-81c3-7fae7ea3d348 · outbound

This paper cites The Llama 3 Herd of Models.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework The Llama 3 Herd of Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T12:33:44.995396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:33:44.995396Z digest=sha256:2c062c06b64fb89c3c2d2525bbf9bbaad93df324deaf74e0b9c55fcad5599553

Observation 60b8eedb-36d3-441c-97e2-7c3a88c27f93 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T12:33:45.000275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:33:45.000275Z digest=sha256:3c5143169f0497c59dbaed18910e6a705cfc47417c24f32e036e7188417bd1fd

Observation 4e0369dd-4a64-4df6-b92c-11afd59e42b5 · outbound

This paper cites Momentum contrast for unsupervised visual rep- resentation learning.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Momentum contrast for unsupervised visual rep- resentation learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T12:33:45.006252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:33:45.006252Z digest=sha256:c8df345bc2987c76e822156ae49f305937b3e4d9ee8256863e2e149501a21fc1

Observation 8bd98e9c-45d6-4423-a31f-055b5ce28c91 · outbound

This paper cites Masked autoencoders are scalable vision learners.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Masked autoencoders are scalable vision learners

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:33:45.976142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T12:33:45.011021Z digest=sha256:55dbbee57cff54bccc7e4edcae79b88b92d56f8a812a89c17fa603cfca1a1272

Observation 034413c5-16c7-4253-b51d-d1856d61227a · outbound

This paper cites Data-efficient Event Camera Pre-training via Disentangled Masked Modeling.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Data-efficient Event Camera Pre-training via Disentangled Masked Modeling

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T12:33:45.016104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:33:45.016104Z digest=sha256:d92879691ffe11248ff94d576215f91480b13b11bfd0622a7961d76957b79f8e

Observation ed5fa96b-ccab-45d1-8890-cdbb3cf69b19 · outbound

This paper cites N-imagenet: Towards robust, fine-grained object recognition with event cameras.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework N-imagenet: Towards robust, fine-grained object recognition with event cameras

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:33:45.961258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T12:33:45.021864Z digest=sha256:221932b237f5f984d6fcea33ad7b70955cc70a18b45bd6440550acd5c1b24f9a

Observation fefa5436-bb6e-4ac9-a0fb-110eb4d3a2f5 · outbound

This paper cites Spiking-yolo: spiking neural network for energy- efficient object detection.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Spiking-yolo: spiking neural network for energy- efficient object detection

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T12:33:45.026355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:33:45.026355Z digest=sha256:80fed00b8e58cbc61ca329e2460e41187679c4ec65fe9db4f6b37b5f8d0a3ec2

Observation 800ecb3d-ae9e-4fe9-a87e-f5854ef98554 · outbound

This paper cites Segment any- thing.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Segment any- thing

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T12:33:45.030964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:33:45.030964Z digest=sha256:2c84bb05ad29765e88fe057ea4dce60898944a50aa0dceb1d5e5bdd2ebe65009

Observation c2ff80af-8f69-4d7e-bc4b-cbc8b470212a · outbound

This paper cites Masked event modeling: Self-supervised pretraining for event cameras.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Masked event modeling: Self-supervised pretraining for event cameras

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T12:33:45.035526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:33:45.035526Z digest=sha256:4f80d73f033e416e53b7a950a7f45858680554700af98d17b5bac10fd4aa58c1

Observation de3313fb-2b40-487a-a9d7-94727cda921c · outbound

This paper cites Openess: Event-based semantic scene understanding with open vocabularies.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Openess: Event-based semantic scene understanding with open vocabularies

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:33:45.918387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T12:33:45.040195Z digest=sha256:47ac285d48021ce8bcc7d376adf7bd6326c36a43c4bdadc2d4e2fb03d29ab73a

Observation bd9ab563-4775-4dfa-ba6b-41e4b9ad8aab · outbound

This paper cites Mulfs-cap: Multimodal fusion- supervised cross-modality alignment perception for unreg- istered infrared-visible image fusion.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Mulfs-cap: Multimodal fusion- supervised cross-modality alignment perception for unreg- istered infrared-visible image fusion

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T12:33:45.044826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:33:45.044826Z digest=sha256:5e8c39b4df253c19ea9cce4ece51bf46b105717f89a6d808b0f86db41e45b703

Observation e3869773-b7c9-4d8a-987f-c2c7e5b13175 · outbound

This paper cites Coupled mamba: Enhanced multimodal fusion with coupled state space model.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Coupled mamba: Enhanced multimodal fusion with coupled state space model

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:33:45.893764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T12:33:45.049239Z digest=sha256:8e6a1543ac386169afaebf1bf8f84f9c0f0f152a00a81f0f4149b493d1ea9e96

Observation 97e04996-459f-4629-879e-6143eabe9f48 · outbound

This paper cites Exploring plain vision transformer backbones for object de- tection.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Exploring plain vision transformer backbones for object de- tection

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:33:45.878991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T12:33:45.054058Z digest=sha256:b7f22610c84fcb56006dda0d70c9cd185e585a8b64747c59b8522b2e37fa8ff8

Observation 55f2a1b4-1763-4618-aab0-8d7753b2d2f9 · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T12:33:45.058883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:33:45.058883Z digest=sha256:b2a95b83e2d71d52698c482135bf88cfec70e1b7a8ba6687990fa2ae9089b0d3

Observation 0309dcdf-261d-4b8a-b9b7-223f75749363 · outbound

This paper cites Visual instruction tuning.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Visual instruction tuning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T12:33:45.063794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:33:45.063794Z digest=sha256:5f872c336ea8c42a77a22aec1a9741cb3e98dbe1fea9e212546c3c1a181bcebb

Observation f68ca61a-21d8-465b-9239-dd19f5b72050 · outbound

This paper cites Pixmim: Rethinking pixel reconstruction in masked image modeling.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Pixmim: Rethinking pixel reconstruction in masked image modeling

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:33:45.854175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T12:33:45.068364Z digest=sha256:bca1f800eb5a6cd084ad9d4601ef352a513e77b717f9c2235cb58d1947c3d002

Observation 02ce43ba-edb4-4727-be67-8050ee285cda · outbound

This paper cites Decoupled weight de- cay regularization.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Decoupled weight de- cay regularization

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:33:45.840121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T12:33:45.073164Z digest=sha256:cb0bd0a163a4f0dc58a9b5f8d99e738b4c9685e54cfd7d8a7c9b798f4ae0773f

Observation cfa01b67-b41b-4160-90f4-b4e40bbcf0a2 · outbound

This paper cites Integer-valued training and spike-driven inference spiking neural network for high-performance and energy-efficient object detection.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Integer-valued training and spike-driven inference spiking neural network for high-performance and energy-efficient object detection

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:33:45.824766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T12:33:45.077761Z digest=sha256:0cd96bbb88f3ba580f8af54aabd5cd5ff4070dfb1b4f02dec8eaca6d0764b2d9

Observation c3370488-1ea8-492f-9a40-ae424cf4d65b · outbound

This paper cites Event-based moving object 12 detection and tracking.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Event-based moving object 12 detection and tracking

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:33:45.809465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T12:33:45.082367Z digest=sha256:55c462aaa79d2d184e147bb3a71236a3643daa95247fbed680d601127bb26d67

Observation e3dc42bd-b708-467e-86eb-359e73d63f8a · outbound

This paper cites Rethinking transformers pre-training for multi- spectral satellite imagery.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Rethinking transformers pre-training for multi- spectral satellite imagery

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:33:45.794835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T12:33:45.087166Z digest=sha256:2d5578d783c5830178254968ada1d6762c54d524c5deff1b2f2609561a535c1a

Observation 5b8fe7b3-96cb-40c8-9a09-78a03e83ce5e · outbound

This paper cites Dinov2: Learning robust visual features without supervision.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Dinov2: Learning robust visual features without supervision

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T12:33:45.092015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:33:45.092015Z digest=sha256:243fae5573c6e8c9d52510082412962c447516b714bb663532c0947200d74757

Observation 20971ba1-d35b-46ab-99ed-729cfdd53ed2 · outbound

This paper cites Pytorch: An im- perative style, high-performance deep learning library.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Pytorch: An im- perative style, high-performance deep learning library

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:33:45.770086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T12:33:45.097788Z digest=sha256:26406675fbc1e9f4a4a04a58d24e8fd9aecb945976433bf39d5855edf829db61

Observation f8113807-ab7e-474d-9622-52d29421a468 · outbound

This paper cites BEiT v2: Masked Image Modeling with Vector-Quantized Visual Tokenizers.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework BEiT v2: Masked Image Modeling with Vector-Quantized Visual Tokenizers

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T12:33:45.102760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:33:45.102760Z digest=sha256:42dae82173091820638b25af16d2b59309663f6fdc5efecd66b563f21003b835

Observation 222c146c-c03c-48cb-a699-0c35812d253f · outbound

This paper cites Detectors: Detecting objects with recursive feature pyramid and switch- able atrous convolution.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Detectors: Detecting objects with recursive feature pyramid and switch- able atrous convolution

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T12:33:45.107543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:33:45.107543Z digest=sha256:b91209e42a55754e58b65dc568f85e67900b7e0eaf9151ceb4ae3820239f4876

Observation 0105fd7e-2943-4883-9f02-d8cbedb24a00 · outbound

This paper cites Improving language understanding by gen- erative pre-training.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Improving language understanding by gen- erative pre-training

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T12:33:45.112156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:33:45.112156Z digest=sha256:a322a88a16b037236396fbccf51e981640bc856fe65fcafca493cbe72daf670c

Observation 65c96e97-f43b-4d3d-bd3c-f7c288d743cb · outbound

This paper cites Language models are unsu- pervised multitask learners.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Language models are unsu- pervised multitask learners

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T12:33:45.116617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:33:45.116617Z digest=sha256:56bb05ef67576a1f5ef0281b95f645d8226b9c6045bd1121a2c72196d289fb6d

Observation 1ed4d3bd-0aa9-4d44-b1f5-6f872cb9e01f · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Learning transferable visual models from natural language supervi- sion

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T12:33:45.121084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:33:45.121084Z digest=sha256:c0c3ba6b03c3be4cb4fe4b50f5b60f3ced3c34ee77790d7afc389cf6d23b2a4d

Observation 3674552d-2517-4313-9bdd-5f19b8022cee · outbound

This paper cites Scale-mae: A scale-aware masked autoencoder for multiscale geospatial representation learning.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Scale-mae: A scale-aware masked autoencoder for multiscale geospatial representation learning

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:33:45.717949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T12:33:45.125427Z digest=sha256:7a27c740d00d132c0854de16e4102e5aa33ffc9adb3ac317d93ba8ec1ed9b19d

Observation af078688-5acc-4713-a44c-21535e3343a3 · outbound

This paper cites Revisiting Color-Event based Tracking: A Unified Network, Dataset, and Metric.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Revisiting Color-Event based Tracking: A Unified Network, Dataset, and Metric

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T12:33:45.130092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:33:45.130092Z digest=sha256:9e65b3d1190a1b424e5d07d23e5a245192d3274cac979258db28bb1a9317cc65

Observation e0f732fa-39b3-469c-96f1-09333daf40be · outbound

This paper cites Humanbench: Towards general human- centric perception with projector assisted pretraining.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Humanbench: Towards general human- centric perception with projector assisted pretraining

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:33:45.703106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T12:33:45.135573Z digest=sha256:925e88e4d9fe5ebcf4349116ffd163dfe481d0020a73208e88620233de7467f7

Observation 4ce8beff-01ba-4ca3-abd0-03ee8f555aa6 · outbound

This paper cites VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-16T12:33:45.140311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:33:45.140311Z digest=sha256:0aa096169555d55f895c141216fafe4310fa8400de24e801d62704a5a79d5128

Observation a94de740-76e3-47cc-bd6c-d0cdec7f82f6 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework LLaMA: Open and Efficient Foundation Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-16T12:33:45.145257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:33:45.145257Z digest=sha256:027a1371363c10ee7c93cec145d76ff761c9f697cb3522af0b10af0066cfdc6a

Observation 0b3114b7-928e-4e90-af68-6b680b606807 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-16T12:33:45.149913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:33:45.149913Z digest=sha256:9d0829e1119254287d2ea518cd829357c01da32fd60872a33ddfab14c9737c9c

Observation 114a39bf-a593-480c-8fb7-a6c1635fbd4d · outbound

This paper cites Videomae v2: Scaling video masked autoencoders with dual masking.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Videomae v2: Scaling video masked autoencoders with dual masking

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:33:45.688661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T12:33:45.154973Z digest=sha256:4a227b4efff577485c055524dfe4f6f965ac71e505eb474546005a67e3eefaa1

Observation 2d48d814-d7c5-4c57-ba50-76c3c3e1a92b · outbound

This paper cites Image as a foreign language: Beit pretraining for vision and vision- language tasks.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Image as a foreign language: Beit pretraining for vision and vision- language tasks

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:33:45.673472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T12:33:45.159467Z digest=sha256:462377d764c6a1afcb299aebb9089683bf40c8fc76b7556b65a3b89cf0238f7c

Observation 659f7318-309d-43d2-bc62-70506a10c14d · outbound

This paper cites Vi- sevent: Reliable object tracking via collaboration of frame and event flows.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Vi- sevent: Reliable object tracking via collaboration of frame and event flows

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:33:45.656566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T12:33:45.164092Z digest=sha256:419c427c68bcbcfe9319e9215d181b1ba17da818deae1c7a1868a7983bc7435e

Observation 6cabab66-b653-495d-b316-1007dbcc9457 · outbound

This paper cites Object Detection using Event Camera: A MoE Heat Conduction based Detector and A New Benchmark Dataset.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Object Detection using Event Camera: A MoE Heat Conduction based Detector and A New Benchmark Dataset

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-16T12:33:45.168675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:33:45.168675Z digest=sha256:86ca1f5d55ea9a0e8e00fd8c8f1ead0babcb1cb7e1dd13de3c82682d12c1aef2

Observation d60a012c-cff1-4059-a3cb-6635e5bc8211 · outbound

This paper cites Pre-training on High Definition X-ray Images: An Experimental Study.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Pre-training on High Definition X-ray Images: An Experimental Study

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-16T12:33:45.173911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:33:45.173911Z digest=sha256:c956d9df4904ffc0b7dd51e4663ccbea3ff54b616706b7f979621b48c0cd8fc9

Observation c82d5b01-70e9-4563-b000-c083b4671f72 · outbound

This paper cites Event stream-based visual object tracking: A high-resolution benchmark dataset and a novel baseline.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Event stream-based visual object tracking: A high-resolution benchmark dataset and a novel baseline

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:33:45.641198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T12:33:45.178551Z digest=sha256:3b5bdfdd3689228f2d8bfe1b2879f8e5bdb7e272ffdb0d55699a94b75d735060

Observation 5d19998c-fecb-40ec-a7f8-bb8dfc73f81b · outbound

This paper cites Structural information guided multimodal pre-training for vehicle-centric percep- tion.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Structural information guided multimodal pre-training for vehicle-centric percep- tion

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:33:45.626716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T12:33:45.183525Z digest=sha256:3063ad5effe4cf4cbff5d769c2380276a6d70bf6c2042879eeec0e5de40084a2

Observation 519c5d77-222a-47be-a7be-b6cb6c73b803 · outbound

This paper cites Hardvs: Re- visiting human activity recognition with dynamic vision sen- sors.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Hardvs: Re- visiting human activity recognition with dynamic vision sen- sors

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:33:45.611488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T12:33:45.188277Z digest=sha256:80e084be85784668bbe1a91f91c3e9231a9d3e35bddf8487fec94470bc1ad472

Observation 9f5af015-e868-42e2-9939-4bd60da1cd75 · outbound

This paper cites Multipath event-based network for low-power human action recognition.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Multipath event-based network for low-power human action recognition

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:33:45.596237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T12:33:45.192890Z digest=sha256:64cc4a15c289ccce0add58e349442cbfcdbda27fb084b7f9bf782fa09fe7d559

Observation 6619c4a3-f1c4-4597-b17f-458372397f34 · outbound

This paper cites Event camera data pre-training.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Event camera data pre-training

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-16T12:33:45.197522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:33:45.197522Z digest=sha256:5421cb13be5bdf15a7c305e7e87a0c03f9b063b69c39c76d17eeda55c6cec314

Observation 015ff7b0-d559-4ef7-998a-84cc6be01d65 · outbound

This paper cites Florence: A New Foundation Model for Computer Vision.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Florence: A New Foundation Model for Computer Vision

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-16T12:33:45.202194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:33:45.202194Z digest=sha256:3e3c0929e6cecb0039d74eb1f78229995ded30560a0e8041927e64460d4dca38

Observation f939e99e-2774-471e-8bab-d66d1273ef23 · outbound

This paper cites Dino: Detr with improved denoising anchor boxes for end-to-end object de- tection.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Dino: Detr with improved denoising anchor boxes for end-to-end object de- tection

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:33:45.571713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T12:33:45.207077Z digest=sha256:68b42e96b64a18ccd2cab2b5251a89e7e950cfc04123206cdf6df734ea8a8b23

Observation 33f32570-e839-459e-96b9-e2096d899de8 · outbound

This paper cites Odtrack: Online dense temporal token learning for visual tracking.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Odtrack: Online dense temporal token learning for visual tracking

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-16T12:33:45.211899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:33:45.211899Z digest=sha256:bfa5732442b30ebebbded92062b0da82e1a728496edf4f5ed892485b0f3a0caa

Observation a1c57396-d39c-4ebe-93e8-8329ad2606f8 · outbound

This paper cites Image bert pre-training with online tokenizer.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Image bert pre-training with online tokenizer

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:33:45.547965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T12:33:45.216184Z digest=sha256:c0c42ed387f759a98ce1ff5234b74e7c0c7644368ef5be8d59386caaa02724a9

Observation da8c1892-5d29-40d3-9215-cc601b1c1f3b · outbound

This paper cites Ex- act: Language-guided conceptual reasoning and uncertainty estimation for event-based action recognition and more.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Ex- act: Language-guided conceptual reasoning and uncertainty estimation for event-based action recognition and more

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:33:45.533474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T12:33:45.220773Z digest=sha256:5abf996e04ee9f247f7d7593d7a0562da5ffb74fa2d6a8ca35323e1119cf06b1

Observation b4b144e9-967a-4cd5-9f00-5ef8ae27502a · outbound

This paper cites Event-free moving object segmentation from moving ego vehicle.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Event-free moving object segmentation from moving ego vehicle

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:33:45.518562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T12:33:45.225157Z digest=sha256:5b7324beb9ae24ba925f7906863bc37f8a7f1e02bbc3f4baf07882d0d63bee02

Observation f30c6a45-9e2b-49b5-b6c3-1ba9152998fd · outbound

This paper cites Segment everything everywhere all at once.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework Segment everything everywhere all at once

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-16T12:33:45.229598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:33:45.229598Z digest=sha256:4358c84e371de588409e434954059ff24d5591e78620cae1717e1a31b380e3bd

Observation 4967cac0-fc44-4460-be2f-93c159ec9ba6 · outbound

This paper cites PLIP: Language-Image Pre-training for Person Representation Learning.

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework PLIP: Language-Image Pre-training for Person Representation Learning

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-16T12:33:45.234194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:33:45.234194Z digest=sha256:4d5c1199c738773a3325e09344fddff3aaf48979041191100d9e12955182b956

Pith citing papers

No inbound Pith citation observations are available.