Pith. sign in

Paper Citation Record · LEDGER

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality

As of 20 August 2026, this Paper Citation Record lists 100 of 111 outbound references and 0 inbound Pith citation observations for arXiv:2411.18669.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.18669 v1

Coverage vector

measured 100 of 111 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T11:11:44.115214Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 111 outbound references displayed

  • verified exact0
  • verified fuzzy48
  • unresolved52
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 34ed44b5-4471-49d2-ae87-3d4a94cd94db · outbound

This paper cites Sequential modeling enables scalable learning for large vision models.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Sequential modeling enables scalable learning for large vision models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T11:11:43.637794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:11:43.637794Z digest=sha256:3af12b22af5ad6ee53c9c8d076fef9b5a4d4aef23f0cda4a5fc9d731c15eb64d

Observation 20e2c845-7037-4c58-b29c-45b40b9db29e · outbound

This paper cites Beit: Bert pre-training of image transformers.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Beit: Bert pre-training of image transformers

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T11:11:43.642808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:11:43.642808Z digest=sha256:eefa5b3619f7f6b525d5bdf9bd44f9524b07ec02dfb06a5442eed55438184b73

Observation bb5963a4-2dab-4b91-af83-d7bf7352569c · outbound

This paper cites On the Opportunities and Risks of Foundation Models.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality On the Opportunities and Risks of Foundation Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T11:11:43.647539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:11:43.647539Z digest=sha256:075e80206313a814e2cf0478678fb64029888e6818083a640a2a2678e2eb7e60

Observation c9e98013-8d77-4fa5-bd65-815532dbb8dd · outbound

This paper cites Multi-spectral sift for scene category recognition.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Multi-spectral sift for scene category recognition

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T11:11:43.653165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:11:43.653165Z digest=sha256:685aa6fbc56c27c4f1454483fe20dbacd76e73e53f292a54e4026b421d893d8f

Observation 196e969f-6053-4630-8cc3-6c2273aad0a4 · outbound

This paper cites Language models are few-shot learners.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Language models are few-shot learners

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T11:11:43.658190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:11:43.658190Z digest=sha256:5b3720d05be59b40a2c71936ced4658ef27a2411d5241087b8faf2fe9c73a17d

Observation 5608ce74-84a7-4f4f-af2e-35de469167f0 · outbound

This paper cites Pretrainable geometric graph neural network for antibody affinity maturation.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Pretrainable geometric graph neural network for antibody affinity maturation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T11:11:43.663163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:11:43.663163Z digest=sha256:ecfb7516f75d013d457318d6a9ccd946c64b7ce83b65db163c0147a91ac54184

Observation 16c01c81-68b3-454b-8070-91fff15ebcfa · outbound

This paper cites SAD: Segment Any RGBD.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality SAD: Segment Any RGBD

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T11:11:43.668630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:11:43.668630Z digest=sha256:94fd79658482a9c93471e763743c81a555a16d51fc990fab85b461a0d37bd3e8

Observation af42e5b3-6fee-47b8-ad3f-d326585129a1 · outbound

This paper cites Hts-at: A hierarchical token-semantic audio transformer for sound classification and detection.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Hts-at: A hierarchical token-semantic audio transformer for sound classification and detection

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T11:11:43.673641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:11:43.673641Z digest=sha256:fbeba6b7509ff5ff4b01b18ee0aebac0aee993bce96ef43591e7a8723638d2c4

Observation 7771db4a-236a-438a-9878-2811fc1a62f9 · outbound

This paper cites Domain ad- aptation for semantic segmentation with maximum squares loss.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Domain ad- aptation for semantic segmentation with maximum squares loss

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T11:11:43.678256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:11:43.678256Z digest=sha256:c8e7d864a3425b66e7f0f88ccaa9e31067ee3d385a1d8062862784df1b2fd340

Observation cc6bf476-9401-4b14-af7a-f53bee9d0c0c · outbound

This paper cites Adaptformer: Ad- apting vision transformers for scalable visual recognition.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Adaptformer: Ad- apting vision transformers for scalable visual recognition

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T11:11:43.682832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:11:43.682832Z digest=sha256:ef5739b4d2257f6a4435cf79e8529b4e9a2ea06d3fba0a9328f75afcea3c0484

Observation 6379fb25-0387-4b33-b42b-fc7a63bb6b57 · outbound

This paper cites A simple framework for contrastive learn- ing of visual representations.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality A simple framework for contrastive learn- ing of visual representations

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T11:11:43.687845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:11:43.687845Z digest=sha256:b8df781a7150d96171b149a0c6d3e992d2f34818611efedcf91fc1e224efe6c2

Observation b7a0740e-7a2a-4eb7-86a5-7d84b5581169 · outbound

This paper cites SAM Fails to Segment Anything? -- SAM-Adapter: Adapting SAM in Underperformed Scenes: Camouflage, Shadow, Medical Image Segmentation, and More.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality SAM Fails to Segment Anything? -- SAM-Adapter: Adapting SAM in Underperformed Scenes: Camouflage, Shadow, Medical Image Segmentation, and More

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T11:11:43.692410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:11:43.692410Z digest=sha256:69bf1c485a1e60990fe0bf58c8494204c05a7567727ff724f470a9806cf81de7

Observation 00252cb6-46ce-479e-a8e6-b9e506700168 · outbound

This paper cites Improved Baselines with Momentum Contrastive Learning.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Improved Baselines with Momentum Contrastive Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T11:11:43.697394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:11:43.697394Z digest=sha256:bf624df0c18696a6dd3afe71d1e20c11fac0f4118a73f33b6a5832f1aa93f06a

Observation ff1ea2aa-2ee9-4dce-917f-a334af241bcc · outbound

This paper cites Imagenet: A large-scale hierarchical im- age database.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Imagenet: A large-scale hierarchical im- age database

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T11:11:43.702405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:11:43.702405Z digest=sha256:0cdbb6fab20a5500eb5b41d4280771ea1388edfd427612d2cda04fea685e61fa

Observation 928c5971-fa53-4acb-8471-ba12da6e487c · outbound

This paper cites Lift: Language- interfaced fine-tuning for non-language machine learning tasks.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Lift: Language- interfaced fine-tuning for non-language machine learning tasks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T11:11:43.706955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:11:43.706955Z digest=sha256:550c2acf673bd2da43b3c867c11bbf21b6fbe4494575f0dbfb24df950e084cae

Observation 875e70da-0c68-4c23-8bc3-e5aa735fb8e3 · outbound

This paper cites Hyperspectral image super-resolution via non-negative structured sparse repres- entation.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Hyperspectral image super-resolution via non-negative structured sparse repres- entation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T11:11:43.711124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:11:43.711124Z digest=sha256:f7755da6ab022fb29080715b494b46f1264679967da5d926533610f439d68f5b

Observation bd480787-7f72-4abc-bfd0-cffbc8d25f69 · outbound

This paper cites An image is worth 16x16 words: Trans- formers for image recognition at scale.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality An image is worth 16x16 words: Trans- formers for image recognition at scale

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T11:11:43.715480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:11:43.715480Z digest=sha256:ec02d8c03a720210ebafc30d3cf21ae8c03b215c4e3b1b967a9965d3b126a057

Observation f0b7825d-2414-4013-85c3-3652625f2051 · outbound

This paper cites Multiscale vision transformers.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Multiscale vision transformers

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T11:11:43.719637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:11:43.719637Z digest=sha256:334f4dc2b0bef2e08e0daa465587ed5fcccaded843e9a031a96bd3e4267cc6a6

Observation 8c13dfe7-70b0-4a93-bfc1-3c129c603fde · outbound

This paper cites Event-based vision: A survey.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Event-based vision: A survey

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T11:11:43.724038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:11:43.724038Z digest=sha256:4e15456e6dd0e4a97a9998b3967313d93ca86582425cd697e45870ac2af8f55e

Observation 3ebd8af7-2dc8-4963-9c6c-e7015a5dd3a2 · outbound

This paper cites Low-latency auto- motive vision with event cameras.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Low-latency auto- motive vision with event cameras

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T11:11:43.728339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:11:43.728339Z digest=sha256:01610ecc437f8963ea7a834e073fdb4a6902f5d7432c27f1fdf44674d12360e6

Observation aa4eda63-cb61-43dd-ab1a-2954d1001518 · outbound

This paper cites Imagebind: One embedding space to bind them all.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Imagebind: One embedding space to bind them all

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T11:11:43.733122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:11:43.733122Z digest=sha256:dc7bf51db25aa8a4a74f3746db533b6792f3ab75869d3bad650bcb2387a67759

Observation 0b1aa502-3d72-4f20-8bf4-22a2a8d23daa · outbound

This paper cites AST: Audio Spectrogram Transformer.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality AST: Audio Spectrogram Transformer

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T11:11:43.737620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:11:43.737620Z digest=sha256:6c51607aa88deb5b4df622b4a5a45f10c7366c40bfdbf704814cc45631af9333

Observation f62be8ee-b251-4707-9c55-748b623a5dff · outbound

This paper cites Bootstrap your own latent-a new approach to self-supervised learning.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Bootstrap your own latent-a new approach to self-supervised learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T11:11:43.742245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:11:43.742245Z digest=sha256:db550ac2d6720d6ad18c14d7fd0416130083b8fcc47d48e51318ba7669b76e12

Observation 513cca78-48d8-412c-8db2-0ebe4a256dc9 · outbound

This paper cites Pct: Point cloud transformer.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Pct: Point cloud transformer

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T11:11:43.746222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:11:43.746222Z digest=sha256:1ed4953570ff1e98b5e8ebc6daf223f7ac8a2b56eb526b6fc4c0033f26971099

Observation fb5e766a-4321-4d5e-815a-30dfe248b5c5 · outbound

This paper cites Learning rich features from rgb-d im- ages for object detection and segmentation.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Learning rich features from rgb-d im- ages for object detection and segmentation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T11:11:43.750517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:11:43.750517Z digest=sha256:e1a44dfeeaf570f6f5b192a4c8fbb33b9961bf91c682e391080e65482b5d8535

Observation de1db3b7-f77d-4cac-97d7-79420518623a · outbound

This paper cites Towards a Unified View of Parameter-Efficient Transfer Learning.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Towards a Unified View of Parameter-Efficient Transfer Learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T11:11:43.754991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:11:43.754991Z digest=sha256:381ea1da45b95e175151e36114404ab6f7981091282edc822075ef499748587d

Observation 4cdcacbb-0e1e-4983-8dce-cf8e4c53a219 · outbound

This paper cites Momentum contrast for unsupervised visual rep- resentation learning.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Momentum contrast for unsupervised visual rep- resentation learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T11:11:43.760869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:11:43.760869Z digest=sha256:8650446c1d26926f5546465e9f9ae2869848b29cc19e2dcdae4f21f8892689b3

Observation 7d0f6cc1-663c-40a0-a31b-ff87444dec00 · outbound

This paper cites Masked autoencoders are scal- able vision learners.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Masked autoencoders are scal- able vision learners

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T11:11:43.765839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:11:43.765839Z digest=sha256:b09e0b3a26794392732a8ab5c01adab30d04c1f3cd250703b2f86dc7859ace90

Observation 60bfe44c-850c-4425-bf79-d62a81ec2c09 · outbound

This paper cites Parameter-efficient transfer learning for nlp.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Parameter-efficient transfer learning for nlp

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T11:11:43.771836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:11:43.771836Z digest=sha256:a9e6d9eebfaa4c4c3aeab33c60144bf834f4c7d00cf03a7f627183dc7561f92b

Observation d07be425-4a83-4343-8586-231f47192a01 · outbound

This paper cites Lora: Low- rank adaptation of large language models.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Lora: Low- rank adaptation of large language models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T11:11:43.776897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:11:43.776897Z digest=sha256:5098a7a382481100a59cae80e24e4361c5318e3a9b78665164c685de88560c77

Observation e12e30d7-9be2-46c1-ae23-a30a9a0af1db · outbound

This paper cites Polarization structured light 3d depth image sensor for scenes with reflective sur- faces.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Polarization structured light 3d depth image sensor for scenes with reflective sur- faces

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T11:11:43.782398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:11:43.782398Z digest=sha256:ff04ac3bed5f1bce8d9cd712abf20e0e81653627d8167907e54853e417724c67

Observation 546fd074-1fd8-486d-9478-130b67c5b4fb · outbound

This paper cites Glass segmentation with rgb-thermal image pairs.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Glass segmentation with rgb-thermal image pairs

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T11:11:43.787167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:11:43.787167Z digest=sha256:7881bc74ca0e05df663eb48bcabfcce4b56d7d6f8ff7c968a2d501bdd3614fe2

Observation f3635d6b-34bc-42c4-873c-bb9f7eded2a6 · outbound

This paper cites Perceiver: General perception with iterative attention.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Perceiver: General perception with iterative attention

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T11:11:43.792232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:11:43.792232Z digest=sha256:b5a0c5b2dcda66970fd90b5c0b48c6a6eaff656568655245913a5ace917b6890

Observation 14d56bac-e093-44ad-ad57-cc3f9a47ff27 · outbound

This paper cites Perceiver io: A general architecture for structured inputs & outputs.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Perceiver io: A general architecture for structured inputs & outputs

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T11:11:43.797001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:11:43.797001Z digest=sha256:51d53d696eae6c276b6795002bc5c4d6164c7a14fd9edb2f46feaaf69df6d1e4

Observation 6fb28d69-b0d4-4eee-a5f0-02f02e3bbc74 · outbound

This paper cites SAM Struggles in Concealed Scenes -- Empirical Study on Segment Anything.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality SAM Struggles in Concealed Scenes -- Empirical Study on Segment Anything

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T11:11:43.802397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:11:43.802397Z digest=sha256:741668ef5fda49f6f3f190f57654bf22569d2d1d1bd13e5027574834112ded12

Observation 7c8b73c4-a382-4b95-902b-1030a05a2f36 · outbound

This paper cites Visual prompt tuning.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Visual prompt tuning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T11:11:43.807787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:11:43.807787Z digest=sha256:9d0b384ce180da35cbe74fd17ac1803cbcdc6f74dfc7bbbf431e0fc7264c3079

Observation 0795cdab-86d4-4dee-be71-6065a7f48257 · outbound

This paper cites A multi-modal pre-training transformer for univer- sal transfer learning in metal–organic frameworks.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality A multi-modal pre-training transformer for univer- sal transfer learning in metal–organic frameworks

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T11:11:43.813021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:11:43.813021Z digest=sha256:45c0ee376a03bff11418f1b99b2ea7ccd2a269063ededccdaa4a01dd4192df08

Observation b78f6f17-55fd-49c4-a3f6-9636edc4e4f9 · outbound

This paper cites Bert: Pre-training of deep bidirectional trans- formers for language understanding.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Bert: Pre-training of deep bidirectional trans- formers for language understanding

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T11:11:43.817504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:11:43.817504Z digest=sha256:a6072da324fde9c8ff90d11f133a317d5dbcfd5164db7095273e8f3c073c6253

Observation 0c81c926-f572-471a-b139-222808e45c42 · outbound

This paper cites Segment anything.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Segment anything

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:11:45.500119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:11:43.821737Z digest=sha256:26f502b1191b8abd3fec88d04d37f2bb825751045ed3980ed01e89c905587f12

Observation 941c96fa-a584-48fa-a9af-edf8e2deebe1 · outbound

This paper cites Decou- plenet: Decoupled network for domain adaptive semantic segmentation.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Decou- plenet: Decoupled network for domain adaptive semantic segmentation

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:11:45.484572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:11:43.825736Z digest=sha256:465a7546254fd1932fa9fc2627731c2224f5de56231f317455a6e083c2779cc0

Observation 9dda10d9-39af-4c38-84c0-3d2e4233db57 · outbound

This paper cites Lisa: Reasoning segmentation via large language model.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Lisa: Reasoning segmentation via large language model

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:11:45.467397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:11:43.830207Z digest=sha256:92a262961a1f4bc388101724197993d24354d28fee27ece1bd98350db2784f67

Observation f287ed60-7535-4bfa-af01-3a745882cdfd · outbound

This paper cites Polarized reflection re- moval with perfect alignment in the wild.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Polarized reflection re- moval with perfect alignment in the wild

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:11:45.449280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:11:43.834331Z digest=sha256:163102492ca6568348a7026e8a0feda06f71180c6d37ee6efb86cf46f9cf3959

Observation 4d708cc9-d7d1-4b37-9a77-615e376975d4 · outbound

This paper cites Shape from polarization for com- plex scenes in the wild.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Shape from polarization for com- plex scenes in the wild

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:11:45.433682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:11:43.838920Z digest=sha256:02ec458a070ef5e2b05122499a1f3b5043583bb6569c7c2c30f9201747dc4a45

Observation 25cd7fbd-0fd1-4ae4-8808-49a6b17f93bb · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality LLaVA-OneVision: Easy Visual Task Transfer

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T11:11:43.843444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:11:43.843444Z digest=sha256:e90876b5af772dc407887e6b9cd07204082d4ce7398d7769910d80465809ccd1

Observation 4ad4e292-c23c-45ae-9f99-13789b6532fe · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:11:45.417416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:11:43.849346Z digest=sha256:4ba186030acb405cc1f72634a831a7bf9fe13ea7d1308e931196c51b238491f0

Observation e55e324e-3605-4a85-98e4-ff80721cc9f3 · outbound

This paper cites Prefix-Tuning: Optimizing Continuous Prompts for Generation.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Prefix-Tuning: Optimizing Continuous Prompts for Generation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T11:11:43.854002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:11:43.854002Z digest=sha256:e7a2d47961c885c72f9993e5e8b611a3f9b75730a32a2a84972ff4c5aced5c9c

Observation 10eec4c9-99c6-40dd-8bb8-954998a9e43c · outbound

This paper cites Heterogen- eous domain adaptation: An unsupervised approach.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Heterogen- eous domain adaptation: An unsupervised approach

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:11:45.400239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:11:43.859130Z digest=sha256:abd4cbe5edb7d4aa1f5e4425060c278c90f8ef2e5e6098010260ec02ba424868

Observation 644e45ed-e840-40c7-a6ed-1f5780fa9686 · outbound

This paper cites Visual instruction tuning.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Visual instruction tuning

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:11:45.382106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:11:43.863750Z digest=sha256:13862687875d07d9f8267b3249f2e7e68213f0676326785f3facb907627860db

Observation 4fa89c58-e36e-4b8c-bf18-fe7124be4f49 · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Swin transformer: Hierarchical vision transformer using shifted windows

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T11:11:43.868821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:11:43.868821Z digest=sha256:ed2125ff7739dda0c11d2370de1526c03de5a035e123842b929b3bc0120de568

Observation 7dbbdce5-65bc-4bf2-bfce-02da01f68a64 · outbound

This paper cites Frozen pretrained transformers as universal compu- tation engines.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Frozen pretrained transformers as universal compu- tation engines

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:11:45.354759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:11:43.873580Z digest=sha256:18b44d94c47f665183546821219d900e9dab263fcba34940b8ac21a6712af7c0

Observation d34170dc-33cc-40ab-aaa5-3b0eeb9349e6 · outbound

This paper cites Transferring knowledge fragments for learning dis- tance metric from a heterogeneous domain.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Transferring knowledge fragments for learning dis- tance metric from a heterogeneous domain

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:11:45.337493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:11:43.878532Z digest=sha256:76bde36c422769daa85fbcd495bd0f21596616a461e776e556edc51969253f1d

Observation 9c2216cd-a132-41d5-93c7-02ccf00da84d · outbound

This paper cites Segment Anything in Medical Images.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Segment Anything in Medical Images

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T11:11:43.883286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:11:43.883286Z digest=sha256:a9ad7837183fae8c624153884dcfb355aab5e208bd126f0fddeb6fca468d8c36

Observation 5666c991-91d2-4f8e-8c3b-c8bdbbefdc59 · outbound

This paper cites Segment anything in medical images.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Segment anything in medical images

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T11:11:43.889200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:11:43.889200Z digest=sha256:f0c8199ee60716616fa0c32559a5e73ddb5313d33581c15c844e49e40acddc7a

Observation 46d10936-2142-4fbe-8605-1b83465b5f47 · outbound

This paper cites Mul- timodal tactile sensing fused with vision for dexterous ro- botic housekeeping.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Mul- timodal tactile sensing fused with vision for dexterous ro- botic housekeeping

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:11:45.309705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:11:43.894148Z digest=sha256:1b00e23dda01f5f0c02d3e2633481f9d52bcbd39e6256f4e3e211f183c465372

Observation b55084c3-3729-411f-8ee3-9ba50ca06ed3 · outbound

This paper cites Glass segmentation using intensity and spectral polarization cues.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Glass segmentation using intensity and spectral polarization cues

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:11:45.291183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:11:43.899093Z digest=sha256:97ad2cc163974426845af038d7d53583225058b65372c213428b72e4518a642d

Observation 53e53330-abf0-4386-b828-6d8a80fccbdb · outbound

This paper cites Scaling deep learning for materials discovery.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Scaling deep learning for materials discovery

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:11:45.275679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:11:43.904104Z digest=sha256:710f835c215a1c5a92347d310065d6ce263b45a6e993adfdec6c323812b80c8b

Observation f8043947-5ea6-4ad9-92ff-a2cb7c8b2809 · outbound

This paper cites 4m: Massively multimodal masked modeling.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality 4m: Massively multimodal masked modeling

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:11:45.260078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:11:43.909161Z digest=sha256:075a5580b21ef6ddee76290575073115eb04a5b1a7c562c5b34b24e41b06cb34

Observation accd9ac8-6cc7-4dd2-b6b9-81579f69a7e6 · outbound

This paper cites Indoor segmentation and support inference from rgbd images.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Indoor segmentation and support inference from rgbd images

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:11:45.244205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:11:43.914078Z digest=sha256:312b0b3089c121ab7b84d6b6ddf9878c28138aa542d55f5e1478ab999e1ac3fc

Observation c95163c2-192d-4e88-99c3-61566478a627 · outbound

This paper cites Training lan- guage models to follow instructions with human feedback.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Training lan- guage models to follow instructions with human feedback

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:11:45.227679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:11:43.919217Z digest=sha256:5eeaa33b4722bc8276a7e58362d8e73277fb349463f9773b8569f05127c8f4e5

Observation 73d75103-60e8-4d46-beb7-ef5e3cfbf673 · outbound

This paper cites Foundation model for cancer imaging biomark- ers.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Foundation model for cancer imaging biomark- ers

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:11:45.209297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:11:43.923751Z digest=sha256:fa0d04c16e097f3be458424d3beeae7def53a6d35674c31da9d24c396b27f8cf

Observation 068ef4fc-4dd2-4bf7-b946-93690bb7d9b9 · outbound

This paper cites Unsupervised intra-domain adaptation for semantic segmentation through self-supervision.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Unsupervised intra-domain adaptation for semantic segmentation through self-supervision

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:11:45.192584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:11:43.928030Z digest=sha256:9b716e6a574e06480d993adfe96293fc3f8c4f24a241355cd486dc72d2bac789

Observation f6f15a3b-7835-4e21-91e4-9f9a006069e7 · outbound

This paper cites Transfer learning for metal–organic frameworks.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Transfer learning for metal–organic frameworks

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:11:45.177241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:11:43.933224Z digest=sha256:6279d8b96196fd44385e999daf5dd78c963c4a24accaf2b3d876e7d82051cb05

Observation cff9635a-ef76-4153-b72a-072bfd67f0d2 · outbound

This paper cites A survey on transfer learning.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality A survey on transfer learning

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:11:45.160473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:11:43.939363Z digest=sha256:32fc72c368d4dd4c1b93d30d38b7004ddfc71ea7e14e0c8245b1bdf178f62f7d

Observation 86bad8f6-3810-4f19-acc8-0dc040de502f · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Learn- ing transferable visual models from natural language super- vision

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:11:45.145065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:11:43.944035Z digest=sha256:bb2295cbecaf9a2164fb9b3102c027f2a71ab3aaae952d1047707e60cf134b08

Observation 06b5f924-987f-4566-a010-235b73cf2199 · outbound

This paper cites Transfer learning with kernel methods.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Transfer learning with kernel methods

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:11:45.127675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:11:43.951561Z digest=sha256:ec4baa312f2a0be0b500e1da29b77580df78c0cbd5ef5d44c46a6180d9588d84

Observation 97891fc1-bfeb-41ba-ad2b-052fcd7258f3 · outbound

This paper cites High-resolution image synthesis with latent diffusion models, 2021.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality High-resolution image synthesis with latent diffusion models, 2021

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:11:45.111483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:11:43.956913Z digest=sha256:52671e9c7efa154eba8bddaaf26d2a0573f3113149a7a69eb976c94afc6637c9

Observation c1c4d721-e0e0-44bb-a6bb-5d9b63d3841c · outbound

This paper cites Cross-modal fine-tuning: Align then refine.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Cross-modal fine-tuning: Align then refine

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:11:45.094507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:11:43.961413Z digest=sha256:7699c9e5b81466d35d05384d3148ed49f94365f7b2b767c895e7a317e1118875

Observation 83d2ef80-e836-4ac3-953f-b6132494a1c8 · outbound

This paper cites Depth estimation from camera im- age and mmwave radar point cloud.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Depth estimation from camera im- age and mmwave radar point cloud

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:11:45.075398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:11:43.965948Z digest=sha256:83afe0992068a0a89fb6a4cb718e8a6744294a5049fbb8afa1e4a3716c96e1f4

Observation 7e87b243-10b3-4917-860e-fd4e41809e8f · outbound

This paper cites Rna secondary structure prediction using an en- semble of two-dimensional deep neural networks and trans- fer learning.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Rna secondary structure prediction using an en- semble of two-dimensional deep neural networks and trans- fer learning

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:11:45.052479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:11:43.970400Z digest=sha256:408462d884b758ee40f5b86b08f872767a7f66b0cde2a61ca7b4640a7b3fbafc

Observation 4c2e424d-cf75-40b9-bc24-64034af6a5ab · outbound

This paper cites Rtfnet: Rgb- thermal fusion network for semantic segmentation of urban scenes.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Rtfnet: Rgb- thermal fusion network for semantic segmentation of urban scenes

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:11:45.030411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:11:43.975013Z digest=sha256:162511ad07403b529ee86cb2646e68475bde3d7d38825617c04f04b65c18e8b8

Observation b214f3e3-45fc-48a0-85c1-a2de8414bcdb · outbound

This paper cites Seeing far in the dark with patterned flash.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Seeing far in the dark with patterned flash

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:11:45.012532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:11:43.979622Z digest=sha256:bbf3605d0f376e26fb28262fb05b12a488d33f45ca7dbab93edaa0ef100bffa4

Observation 4777f815-8e25-426e-bc29-61b3e24ab013 · outbound

This paper cites Dabs: A domain-agnostic benchmark for self-supervised learning.Advances in neural information processing systems, 2021.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Dabs: A domain-agnostic benchmark for self-supervised learning.Advances in neural information processing systems, 2021

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:11:44.995428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:11:43.984622Z digest=sha256:f5792c916d809a89bc8fe36a6bf0a10cd676b6050cc6b3c93bb1ba338c893fad

Observation a79ed7bd-6687-42fe-aad5-1f16f47012a6 · outbound

This paper cites Can SAM Segment Anything? When SAM Meets Camouflaged Object Detection.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Can SAM Segment Anything? When SAM Meets Camouflaged Object Detection

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-12T11:11:43.990171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:11:43.990171Z digest=sha256:50edf9669d3f181d780937e987ce3a1659dc48b235b5cde1034c2c16e574aef8

Observation 565f0341-4757-42f8-8218-9e840c729478 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-12T11:11:43.995206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:11:43.995206Z digest=sha256:92fc666cdf922fed4462eeab1e66fd0d0c9012fb966fb1558f2d63692fa096dc

Observation 3c729aae-4a88-4e9d-8863-f4d6003a21a7 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-12T11:11:44.000620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:11:44.000620Z digest=sha256:a48287f85a004a7e6ce07e02e1074adbc3f31a367c5fed32f4b69dd686debda1

Observation cad8bd9b-6993-4ec3-830f-0ad5ba3f59a9 · outbound

This paper cites Neural nano-optics for high-quality thin lens ima- ging.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Neural nano-optics for high-quality thin lens ima- ging

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:11:44.977736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:11:44.005657Z digest=sha256:b0c016bc9103c3b29a279c49848dd43005f88ea3468e22106dc827e3852fdbf9

Observation ca7fac17-64c8-4745-9ace-07f5978b9f2d · outbound

This paper cites Attention is all you need.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Attention is all you need

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-12T11:11:44.010204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:11:44.010204Z digest=sha256:b2c31a3f2bea0cd5d5a34fab569f971fea157d8a5e63d1c88edc3079876aecd6

Observation 3f7daf65-04ca-4676-b784-6acec83ac5b4 · outbound

This paper cites Audio Transformers.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Audio Transformers

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-12T11:11:44.014616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:11:44.014616Z digest=sha256:777b3b1ae0ba2ad57bed20c6e2e9569faf937276f47bc44c3bc6c896d8458d62

Observation d822514d-043d-4ede-969f-f5a5ffdfbc44 · outbound

This paper cites Reprogramming Pretrained Language Models for Protein Sequence Representation Learning.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Reprogramming Pretrained Language Models for Protein Sequence Representation Learning

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-12T11:11:44.019553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:11:44.019553Z digest=sha256:c9ce261b3cbf0610503d1837fdf050df4cbc9213a77ee5e118a4d234e66f2d99

Observation 6823c460-b6af-4460-b198-39ec4edacad4 · outbound

This paper cites Advent: Adversarial entropy min- imization for domain adaptation in semantic segmentation.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Advent: Adversarial entropy min- imization for domain adaptation in semantic segmentation

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:11:44.950380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:11:44.024427Z digest=sha256:2b8dfb5f3256eb717689312baaa1839dc7c3d96dcdc3750db244b0097d71af09

Observation 4f00e47c-49c2-4598-9746-21cd4bd3ebf2 · outbound

This paper cites Predicting fault slip via transfer learning.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Predicting fault slip via transfer learning

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:11:44.934923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:11:44.029156Z digest=sha256:a8c62316f811a833c2b68f30cec05b81d870a6e6b00c93ba4b0cd0c656b441fa

Observation ad138850-2302-4609-b9e5-ccc696837950 · outbound

This paper cites Uncertainty-aware clustering for unsupervised domain adaptive object re- identification.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Uncertainty-aware clustering for unsupervised domain adaptive object re- identification

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:11:44.918456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:11:44.033798Z digest=sha256:7295d260e5c373f74cf23265b8f3f13ead44740b7ff26b5c776fda6000a5a613

Observation 61fa1df1-6362-43c2-864f-0cec33c1367a · outbound

This paper cites Sub-surface thermal measurement in additive man- ufacturing via machine learning-enabled high-resolution fiber optic sensing.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Sub-surface thermal measurement in additive man- ufacturing via machine learning-enabled high-resolution fiber optic sensing

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:11:44.902779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:11:44.038542Z digest=sha256:d8708211f309573ddee416005e6550f7a8962533ae29202e620977b07da4ea2c

Observation 4ecabf92-0f36-4334-a5b7-1c2a354b0b61 · outbound

This paper cites Internimage: Exploring large-scale vision foundation models with deformable convolutions.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Internimage: Exploring large-scale vision foundation models with deformable convolutions

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:11:44.885725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:11:44.043144Z digest=sha256:f03404152214ce53c853015fc92e6866dfdc75abab89c5d03ccaa3a906ce9ab2

Observation f127eddc-b36e-4979-895c-20aee68786c3 · outbound

This paper cites Internimage: Exploring large-scale vision foundation models with deformable convolutions.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Internimage: Exploring large-scale vision foundation models with deformable convolutions

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:11:44.869848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:11:44.047276Z digest=sha256:c9ce3e173a7fbe230819b79fa8c8b67fe311ddb6773615de5e8357aa7a3ce38d

Observation 23c1628a-64b7-4d88-9e73-7b8b630f8893 · outbound

This paper cites Not all images are worth 16x16 words: Dynamic transformers for efficient image recognition.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Not all images are worth 16x16 words: Dynamic transformers for efficient image recognition

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:11:44.853975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:11:44.051394Z digest=sha256:c6340df019c634ef41ce5f27001115ee1c1cd71e5b432f46f64a41d0af89795c

Observation 10f59e66-5ad6-40ac-8963-0d29f393bb8f · outbound

This paper cites Randomized quantization: A generic aug- mentation for data agnostic self-supervised learning.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Randomized quantization: A generic aug- mentation for data agnostic self-supervised learning

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:11:44.837501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:11:44.055767Z digest=sha256:4545af97db3bc27e2a3aeb896643e5dc47de47bb8ea7bc6ccd770c17cc38833a

Observation 066ba7b0-eb8c-4dc6-8b9d-8c6f27240a2a · outbound

This paper cites Medical SAM Adapter: Adapting Segment Anything Model for Medical Image Segmentation.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Medical SAM Adapter: Adapting Segment Anything Model for Medical Image Segmentation

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-12T11:11:44.060072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:11:44.060072Z digest=sha256:59fbdd0288a12192b11a234f102e7fbce50139a6c2766b6d8c0351c3debfc9d9

Observation 7610d777-8719-496f-ac35-6cbbd246f965 · outbound

This paper cites Point transformer v2: Grouped vector attention and partition-based pooling.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Point transformer v2: Grouped vector attention and partition-based pooling

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:11:44.821699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:11:44.064535Z digest=sha256:41ff7c1a39a330eebdb51bc8c4eea3dc9b15681059ffdb7385fe3fc1fc80491f

Observation effa0f06-bd53-440b-ac5a-7ad065461ac3 · outbound

This paper cites Polarization- driven semantic segmentation via efficient attention- bridged fusion.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Polarization- driven semantic segmentation via efficient attention- bridged fusion

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:11:44.804075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:11:44.068979Z digest=sha256:e8c169eccafe0dbfd2bd20b8a8350bd4d87c02bcd6fde7fd16339e29fcbc4998

Observation afeee50c-7441-4767-bb11-0700f8bd5ea3 · outbound

This paper cites Simmim: A simple framework for masked image modeling.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Simmim: A simple framework for masked image modeling

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:11:44.788279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:11:44.073227Z digest=sha256:50e41e8a3b277c7b808dc961dcdd0f54a7375ea714fe04c79daab7f7deeef800

Observation 81c6824e-a3df-404c-9586-bcd5347f2fc7 · outbound

This paper cites Video-rate hyperspectral camera based on a cmos-compatible random array of fabry–p ´erot filters.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Video-rate hyperspectral camera based on a cmos-compatible random array of fabry–p ´erot filters

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:11:44.771216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:11:44.077351Z digest=sha256:a9e7eda4dfb8590852e592039b2acb07c28f353519c4452882d64991aa8838e7

Observation 5dff3ee4-ed84-4a68-a5d6-c8bd09e63425 · outbound

This paper cites A vision chip with complement- ary pathways for open-world sensing.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality A vision chip with complement- ary pathways for open-world sensing

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:11:44.753707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:11:44.082038Z digest=sha256:364e819da725d0f1da7c5d19c5875fe530e99dc28a2eb8c985f954c5565f482a

Observation 446d231f-7ffc-44c6-a56a-d94ac5de401b · outbound

This paper cites Superanimal pretrained pose estimation models for behavioral analysis.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Superanimal pretrained pose estimation models for behavioral analysis

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:11:44.737994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:11:44.086340Z digest=sha256:fd519c8769803bc2498d953a93de190980df90695e1a3af294b9c33cbf328578

Observation 6fed5e57-f360-4e72-8d75-6b9b40c2d8db · outbound

This paper cites Taskonomy: Disentangling task transfer learning.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Taskonomy: Disentangling task transfer learning

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:11:44.720905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:11:44.090316Z digest=sha256:47a02444049cd523bc5a07d916fa8a654c597fc9849ed7b9fdd310137b13e4a1

Observation 2350e08c-a058-4bf7-91aa-4b03f645b773 · outbound

This paper cites Cmx: Cross-modal fusion for rgb-x semantic segmentation with transformers.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Cmx: Cross-modal fusion for rgb-x semantic segmentation with transformers

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:11:44.704284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:11:44.095133Z digest=sha256:cb7099a300f29d8e3c8f2baee533c949d977e8c02088daf480de9d9b18e064cc

Observation 0b4162d3-d1e2-4c53-8912-a03dd8574797 · outbound

This paper cites A generalist vision–language foundation model for diverse biomedical tasks.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality A generalist vision–language foundation model for diverse biomedical tasks

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:11:44.688419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:11:44.100477Z digest=sha256:31c5501966830a0bfd1295b0d58f4af27ae21bdd7db6b2876d21c30f73d7a732

Observation 94ea45ef-deb7-475d-a6db-1711ba12dc73 · outbound

This paper cites Adding conditional control to text-to-image diffusion models.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Adding conditional control to text-to-image diffusion models

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-12T11:11:44.105360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:11:44.105360Z digest=sha256:70238486f8970cbac349d645fa6564b9a0b29c2b41a738713fd91491176d0d6c

Observation 46ade355-5cce-4fbb-bda0-a10bccdad22b · outbound

This paper cites Meta-Transformer: A Unified Framework for Multimodal Learning.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Meta-Transformer: A Unified Framework for Multimodal Learning

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-12T11:11:44.110303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:11:44.110303Z digest=sha256:8f5eaa6221ecfe0b209db6097b74efc0589e3b0ca0cf041a86455c6bc1c1101a

Observation 2f3e07a2-5347-4f40-809f-d94f569c22d7 · outbound

This paper cites Point transformer.

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality Point transformer

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:11:44.660954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:11:44.115214Z digest=sha256:2839841c18dc1fd530a2b57046cf2b8bc3503a1669fedf54df52735b89be712f

Pith citing papers

No inbound Pith citation observations are available.