Pith. sign in

Paper Citation Record · LEDGER

SimMIM: A Simple Framework for Masked Image Modeling

As of 19 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2111.09886.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2111.09886 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:41:10.617330Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

15
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0a5c7945-cd9b-41dc-9b6e-6d4085175ca2 · inbound

CoCa: Contrastive Captioners are Image-Text Foundation Models cites this paper.

CoCa: Contrastive Captioners are Image-Text Foundation Models SimMIM: A Simple Framework for Masked Image Modeling

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T10:53:08.409786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T10:53:08.292063Z digest=sha256:8921e2f7655334c3b460495c4955308b724833469eeb893697f333dbb009aba7

Observation acc5e171-64f3-431b-b163-fde77b341ce5 · inbound

Vision Transformers Need Registers cites this paper.

Vision Transformers Need Registers SimMIM: A Simple Framework for Masked Image Modeling

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T09:41:38.236334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-13T09:41:37.937046Z digest=sha256:9fdbce120a6af77e094fe8ca83d452b55094897da55fdb10a29497dd2bdc9df0

Observation bbfc41ed-c0fb-4213-9f46-c6bfd06aad24 · inbound

Revisiting Feature Prediction for Learning Visual Representations from Video cites this paper.

Revisiting Feature Prediction for Learning Visual Representations from Video SimMIM: A Simple Framework for Masked Image Modeling

Reference 290

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T12:40:24.075377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-12T12:40:23.709098Z digest=sha256:5aed6f55d90c6bd2ee9d473c8c9305b3ef3f4f14b6cd16c347e4e8b521e78002

Observation 89988654-f3ab-4129-8c52-d4cec6306de0 · inbound

A Self-supervised Learning Method for Raman Spectroscopy based on Masked Autoencoders cites this paper.

A Self-supervised Learning Method for Raman Spectroscopy based on Masked Autoencoders SimMIM: A Simple Framework for Masked Image Modeling

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T11:41:10.617330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:41:10.617330Z digest=sha256:fbf9540de39a2dd688f7fbe46099aa51a0d1fa6c7b5780fe123d25ea5c08fedb

Observation 4c801cb5-4993-4f15-936b-2da4a8541fb8 · inbound

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation cites this paper.

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation SimMIM: A Simple Framework for Masked Image Modeling

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-16T11:06:08.394168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:06:08.394168Z digest=sha256:c7bc3a0b1dd82dbc41380fcbfb154e003ddcf7e84bd53eceee088bfe4156b973

Observation 14039709-2f8a-4b23-b34f-7ad03ac69994 · inbound

Thoughts on Objectives of Sparse and Hierarchical Masked Image Model cites this paper.

Thoughts on Objectives of Sparse and Hierarchical Masked Image Model SimMIM: A Simple Framework for Masked Image Modeling

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T22:11:15.049428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:11:15.049428Z digest=sha256:99d1e339b833f38d49d7b15604d7c023f68af7cf765458d233640e9a85ae23c7

Observation 6f119842-b305-4082-a2a1-3b1d854aeec2 · inbound

Scaling Up Audio-Synchronized Visual Animation: An Efficient Training Paradigm cites this paper.

Scaling Up Audio-Synchronized Visual Animation: An Efficient Training Paradigm SimMIM: A Simple Framework for Masked Image Modeling

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-06T01:05:34.862848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T01:05:34.862848Z digest=sha256:520bc8802213a6f6739c31c42d1a16cae6ae2c4713369577ce35171052a76d9b

Observation c40d9ca7-a5bf-440c-93e2-cfd801e2816b · inbound

Expectation and Acoustic Neural Network Representations Enhance Music Identification from Brain Activity cites this paper.

Expectation and Acoustic Neural Network Representations Enhance Music Identification from Brain Activity SimMIM: A Simple Framework for Masked Image Modeling

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-21T11:40:03.197599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-21T11:36:06.967549Z digest=sha256:38657ecd12354945fdc0ee44ee448ee23b51eb7985e656c15369a8cc36d478c5

Observation 217eab2f-fac3-4a7d-958c-f66bb3443bf0 · inbound

Transcoda: End-to-End Zero-Shot Optical Music Recognition via Data-Centric Synthetic Training cites this paper.

Transcoda: End-to-End Zero-Shot Optical Music Recognition via Data-Centric Synthetic Training SimMIM: A Simple Framework for Masked Image Modeling

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T05:41:26.797647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-12T05:02:02.367347Z digest=sha256:6a4aa3aaece7f67eb988d1ea56d88563e30e2f091dc003d1535b734b7eb1bd11

Observation cc81ad20-096b-40bf-b0ae-cf70c7e93e2c · inbound

MIMFlow: Integrating Masked Image Modeling with Normalizing Flows for End-to-End Image Generation cites this paper.

MIMFlow: Integrating Masked Image Modeling with Normalizing Flows for End-to-End Image Generation SimMIM: A Simple Framework for Masked Image Modeling

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-07-04T20:50:11.403617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-25T19:34:02.046104Z digest=sha256:ea1798ae4e47348a76c8629b7e4ec085606084e576477617dbaf9ff07cc2da6f