Pith. sign in

Paper Citation Record · LEDGER

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge

As of 15 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 2 inbound Pith citation observations for arXiv:2506.16673.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.16673 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:41:44.214299Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-11T03:31:33.964471Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T03:40:54.141854Z

Reference resolution

47 of 47 outbound references displayed

  • verified exact3
  • verified fuzzy24
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bb0d8947-a5cd-476a-a08b-6a3b270f23f1 · outbound

This paper cites Layer Normalization.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Layer Normalization

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T23:41:39.653398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:41:39.653398Z digest=sha256:ff26aee948fb9dfcd05a40ba0076ede9cfa0cab0192cb9ef0c73ed67c582efa3

Observation 1210daea-a2a6-4b2a-8591-31dd5b9ec0f7 · outbound

This paper cites Distilling the Knowledge in a Neural Network.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Distilling the Knowledge in a Neural Network

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T23:41:40.603499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:41:40.603499Z digest=sha256:fd4f496a12a1c246b646775ceb36d56e51a5ca82383d80c65e953bff99b9aef6

Observation 8d20089f-b83f-4d4d-af8c-0e25addf1014 · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text su- pervision.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Scaling up visual and vision-language representation learning with noisy text su- pervision

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:41:46.787183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:41:40.818042Z digest=sha256:2c09d9aee3873418373c322f5c025b74f201b1b22b9ed130af5f344344b23d66

Observation e51020db-c6d2-4e42-8726-38decccecdf1 · outbound

This paper cites Learning multiple layers of features from tiny im- ages.Handbook of Systemic Autoimmune Diseases, 1(4),.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Learning multiple layers of features from tiny im- ages.Handbook of Systemic Autoimmune Diseases, 1(4),

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:41:46.777334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:41:40.919321Z digest=sha256:dcd241da82db505c6ad9f5588ffcdbb6f861f6c427af67d2aa6d16d97c68b266

Observation 77274896-7a07-4c53-bc8b-be4c9ac55907 · outbound

This paper cites Clipath: Fine-tune clip with visual fea- ture fusion for pathology image analysis towards min- imizing data collection efforts.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Clipath: Fine-tune clip with visual fea- ture fusion for pathology image analysis towards min- imizing data collection efforts

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:41:46.766897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:41:41.059755Z digest=sha256:8f935b2940c68f989342e199166c43d79315ff9324c402ac89d62fe9353d4204

Observation 9f089dd5-871e-45fd-8274-8df9a9fc5a1f · outbound

This paper cites ALBERT: A Lite BERT for Self-supervised Learning of Language Representations.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge ALBERT: A Lite BERT for Self-supervised Learning of Language Representations

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T23:41:41.156877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:41:41.156877Z digest=sha256:9bb0cf48462746cbfc619ad6bb3220011648dc593ee3cffa5a8c967bc0eddee0

Observation 942f3b57-1b3f-4ff5-aa75-1a36304224d2 · outbound

This paper cites Align before fuse: Vision and language representation learning with momentum distilla- tion.Advances in neural information processing systems, 34:9694–9705,.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Align before fuse: Vision and language representation learning with momentum distilla- tion.Advances in neural information processing systems, 34:9694–9705,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:41:46.756791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:41:41.226742Z digest=sha256:8d12f5589c6220e31824f2eb95fc9b01e3a07b577447568501b58918db4d6313

Observation 34d6268d-ed76-42af-b3aa-77edaa48c71d · outbound

This paper cites Blip-2: Bootstrapping language-image pre- training with frozen image encoders and large language models.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Blip-2: Bootstrapping language-image pre- training with frozen image encoders and large language models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T23:41:41.302093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:41:41.302093Z digest=sha256:23c02d49adffff18c8a32e3d8751e1a2da1cf1b9c04b4e48394cb634ea6481b8

Observation 983c5e6a-388b-4d59-8928-9ce5afa2e772 · outbound

This paper cites Automatic evaluation of machine translation quality using longest common subsequence and skip-bigram statistics.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Automatic evaluation of machine translation quality using longest common subsequence and skip-bigram statistics

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:41:46.738963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:41:41.381057Z digest=sha256:e542566a96fcd54eb65e154115747e75d269cb0417d9a4b1c8f2c372f5d8fc4e

Observation e4239e0a-3066-4fb3-9eaf-840639acf3d5 · outbound

This paper cites FoldGPT: Simple and Effective Large Language Model Compression Scheme.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge FoldGPT: Simple and Effective Large Language Model Compression Scheme

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T23:41:41.576024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:41:41.576024Z digest=sha256:0a81a090b7ba2523bcfa10129f7e4dc582206a172dbf87c5c4d806d616f58c02

Observation 78e710c1-ceb4-43b7-a2bc-32a33ff584d7 · outbound

This paper cites Clip-branches: Interactive fine-tuning for text- image retrieval.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Clip-branches: Interactive fine-tuning for text- image retrieval

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:41:46.721080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:41:41.680090Z digest=sha256:4327804ddce793b7481e84cd35aa6f233f982d1dfef85261152ee5d566c9a131

Observation 59700003-ea25-4573-9220-5fa415e5f49c · outbound

This paper cites ClipCap: CLIP Prefix for Image Captioning.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge ClipCap: CLIP Prefix for Image Captioning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T23:41:41.828926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:41:41.828926Z digest=sha256:5d9224ff534a2e1a3b88eff2eea67f401f649b365fafa3bd1e4ab5099feb36eb

Observation 55ca4b34-9252-4ad3-9b9b-57979f93adda · outbound

This paper cites Compact language models via pruning and knowledge distillation.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Compact language models via pruning and knowledge distillation

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:41:46.710512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:41:41.888375Z digest=sha256:c5a2b749f60f8c782de4498e3a878f1124b11d5841c784e40dbd408b6142e64c

Observation 149bad71-fc54-4a54-929f-b7bfae66a016 · outbound

This paper cites CHiLS: Zero- shot image classification with hierarchical label sets.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge CHiLS: Zero- shot image classification with hierarchical label sets

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:41:46.700607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:41:41.974012Z digest=sha256:6530a44c4d45ffae3e98891b8ba62b12a26fc6eee1e1e1dc4149ec8ca0e0f250

Observation d22b41c4-e8de-4f9d-b8d1-42281a2efdb2 · outbound

This paper cites an unresolved cited work.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:41:46.690119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:41:42.038682Z digest=sha256:04c14cb645f3605953a249bd857227ccb2367a0c52cf6c572348afd6520c6ef9

Observation 1c4ace28-b9e0-45ec-a7f1-0dbdb30aa7ca · outbound

This paper cites Language models are unsupervised multitask learners.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Language models are unsupervised multitask learners

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T23:41:42.417659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:41:42.417659Z digest=sha256:13844bed5698d64971893398be47585a7f37148c2c3d2c7fe825a0179c95d504

Observation e7913e60-7df6-4ba1-b8bb-b6570d741f7e · outbound

This paper cites Learning transferable visual models from nat- ural language supervision.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Learning transferable visual models from nat- ural language supervision

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T23:41:42.525499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:41:42.525499Z digest=sha256:8f6d5f1497722c68a70dce8adbb694389e1af3a316b2146d3f5851ecda19c8f9

Observation a98d3e93-39aa-42b4-b253-5f8146c3135b · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:41:46.648692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:41:42.594229Z digest=sha256:307a1423629b62809948ca15fa7caf19e01167c4b70c0a802da8eb7a6e8d17bf

Observation 87482803-e7a0-4381-952c-bae8b819c481 · outbound

This paper cites Gomez, Lukasz Kaiser, and Illia Polosukhin.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Gomez, Lukasz Kaiser, and Illia Polosukhin

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:41:46.627340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:41:42.730440Z digest=sha256:3488948e428949a8ed5b7101d7709792ae51d9f4366b3031e52165600d038ab0

Observation f62f778a-ab77-42f5-ba9c-37c4ea823d8e · outbound

This paper cites Characterizing and avoid- ing negative transfer.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Characterizing and avoid- ing negative transfer

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:41:46.597299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:41:42.923374Z digest=sha256:7d8dd4b303a59cc039ad3ddfb816c7605cc95a2e829324b2ac97aeed851d7280

Observation 8b136c55-2dc0-455d-87f1-4f9dd86c18f7 · outbound

This paper cites Learngene: From open-world to your learning task.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Learngene: From open-world to your learning task

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:41:46.291286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:41:42.944359Z digest=sha256:e41b943dea1df11c7ffe43695ff23832e34f5e86754185cce4d8e24adf883842

Observation 8ec5d236-9419-43ea-ae21-f9ceb50cc882 · outbound

This paper cites Learngene: Inheriting Condensed Knowledge from the Ancestry Model to Descendant Models.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Learngene: Inheriting Condensed Knowledge from the Ancestry Model to Descendant Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T23:41:43.035082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:41:43.035082Z digest=sha256:958a26ee2a982f0f8bcb6387e9904265de00050c5f13c67646c2b6540d311a57

Observation a355136e-4ae3-4159-beae-2294fd60338a · outbound

This paper cites Vision transformers as probabilistic expan- sion from learngene.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Vision transformers as probabilistic expan- sion from learngene

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:41:46.158677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:41:43.198681Z digest=sha256:21285b18615f655a4c03f3e3ee8f330701a1461f5fc74d96a81b978c5103a88b

Observation a023133f-f8b6-4034-acdd-d7b6fd96b9ca · outbound

This paper cites Exploring Learngene via Stage-wise Weight Sharing for Initializing Variable-sized Models.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Exploring Learngene via Stage-wise Weight Sharing for Initializing Variable-sized Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T23:41:43.321605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:41:43.321605Z digest=sha256:8b45f0708fb8046d673894cb8d04fdd12a2166fe22c9cf4638a9497982ba8483

Observation 996b9d26-f958-4d91-8ac7-493294c8f537 · outbound

This paper cites KIND: Knowledge Integration and Diversion for Training Decomposable Models.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge KIND: Knowledge Integration and Diversion for Training Decomposable Models

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-08-06T23:41:44.756302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:41:43.487929Z digest=sha256:5708cff83a2b11a26fa8ef1196a1dc8f4ac95811d24cc82aa8ee989637e2a7c0

Observation ceb347dc-1b90-4c89-aef2-57be930a8a3b · outbound

This paper cites CLIP-CID: Efficient CLIP Distillation via Cluster-Instance Discrimination.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge CLIP-CID: Efficient CLIP Distillation via Cluster-Instance Discrimination

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T23:41:43.615143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:41:43.615143Z digest=sha256:e89c851556814092d8c4c6d5bdfbb63b5d3b57d4ab0c042d8a977c22fb28222e

Observation 70dfdb34-7ca4-40d8-9c41-b1766c1acbcd · outbound

This paper cites an unresolved cited work.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:41:46.137224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:41:43.735490Z digest=sha256:2408e8e676b9d3f0d55c7c773aa8ff317dab19a1611d55988f4050e3e7c231f6

Observation 2082631f-286b-4761-b5a2-68274aff7dbd · outbound

This paper cites Sigmoid loss for language image pre-training.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Sigmoid loss for language image pre-training

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:41:45.888512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:41:43.840052Z digest=sha256:c102ea7b80e7852cbc5724a9f6fa2c3e113ff270a4ca6f5d6fbdbc38987d87cc

Observation 13984ce2-e02d-4297-8d68-991175959fc3 · outbound

This paper cites Minivit: Compressing vision transformers with weight multiplexing.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Minivit: Compressing vision transformers with weight multiplexing

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:41:45.576440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:41:43.907351Z digest=sha256:90bdb179ce6ddfd57d649199e7e34d01cca89152b411cf77ceb4fcc80f0f344d

Observation db9e8f2c-82da-40c2-9525-65d048c539d0 · outbound

This paper cites CoCoT: Contrastive Chain-of-Thought Prompting for Large Multimodal Models with Multiple Image Inputs.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge CoCoT: Contrastive Chain-of-Thought Prompting for Large Multimodal Models with Multiple Image Inputs

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T23:41:44.018505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:41:44.018505Z digest=sha256:2d26601c93a41b44421553d6a161dc350cc3a95db66f78265e72d46c64930534

Observation 0ee55134-129c-450b-b7d0-5db5854a84ca · outbound

This paper cites Learning clip guided visual- text fusion transformer for video-based pedestrian attribute recognition.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Learning clip guided visual- text fusion transformer for video-based pedestrian attribute recognition

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:41:45.302679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:41:44.116542Z digest=sha256:4c386664e4482e103966daa30016391bca8ff8d12247ec0edad905fb82cd43ed

Observation 64496046-0093-4afe-b0ab-3ea71dd38fb8 · outbound

This paper cites 8 with the loss weightλ set to1.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge 8 with the loss weightλ set to1

Reference 47

Resolution
verified exact
raw_fallback, observed 2026-08-06T23:41:44.546243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:41:44.214299Z digest=sha256:d0aae6f8184f0d92497f86f415165a32aa14e5773ca0e7a00090c1102a780729

Observation 98ba2ed7-0ac0-4d72-8e15-208b5f67e433 · outbound

This paper cites Cats and dogs.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Cats and dogs

Reference 2001

Resolution
unresolved
no resolver link, observed 2026-08-06T23:41:42.159844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:41:42.159844Z digest=sha256:2f8e826ef23a72371e0cfc9046235c09d5f3f0968ed22838a0eec43bec525874

Observation c309051c-b47c-4cb4-aa55-c012e17239d6 · outbound

This paper cites Microsoft coco: Com- mon objects in context.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Microsoft coco: Com- mon objects in context

Reference 2004

Resolution
unresolved
no resolver link, observed 2026-08-06T23:41:41.496223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:41:41.496223Z digest=sha256:9e6ff13a3b43d32170e019934797626de039a6acc1a496d0618d451b0a4a132f

Observation a7a89bf2-2990-40f3-abc5-b4dcd4f3ff9b · outbound

This paper cites Bert: Pre-training of deep bidirectional transformers for language understand- ing.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Bert: Pre-training of deep bidirectional transformers for language understand- ing

Reference 2009

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:41:46.809002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:41:40.172526Z digest=sha256:1d96f56448928753762aeb7156d5640a2fc880ce69f9777b753a6e3113da2a29

Observation 6005b785-7aaf-4522-895b-91e43ca15514 · outbound

This paper cites Clipping: Distilling clip-based models with a stu- dent base for video-language retrieval.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Clipping: Distilling clip-based models with a stu- dent base for video-language retrieval

Reference 2012

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:41:46.673029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:41:42.268837Z digest=sha256:d3afbce371e9651c0e8bb1f6bd67df379963e0759f26c3346b44f0fe6bf25e02

Observation fefd7bed-e213-4290-b97c-7aa5e6336261 · outbound

This paper cites Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts

Reference 2014

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:41:46.839841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:41:39.926867Z digest=sha256:db920155ca99d12b7b6ec0b106f5e6d23648aeee1cc1e67080d48fcb382f3060

Observation 02cce5a3-27e0-4c63-b4aa-0d7b7c05fe3e · outbound

This paper cites Openclip, July.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Openclip, July

Reference 2015

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:41:46.798397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:41:40.705044Z digest=sha256:d4798d852853e45104695252311be9ffd12b3658a788ccfdf1cfae14649d3424

Observation 0f226c26-7bbe-4ff2-a102-1d8c3ee06c94 · outbound

This paper cites Vlmo: Unified vision-language pre-training with mixture-of-modality- experts.Advances in Neural Information Processing Sys- tems, 35:32897–32912,.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Vlmo: Unified vision-language pre-training with mixture-of-modality- experts.Advances in Neural Information Processing Sys- tems, 35:32897–32912,

Reference 2016

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:41:46.856955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:41:39.722053Z digest=sha256:c98405019e5b48e033dd7e74b175e51699ff89db1b3a348a3036ca6e75a0e241

Observation 8661aa16-e518-4425-bef4-13ac93d6980a · outbound

This paper cites Lawrence Zitnick, and Devi Parikh.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Lawrence Zitnick, and Devi Parikh

Reference 2017

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:41:46.616326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:41:42.843705Z digest=sha256:d0a203763f354435a9e13620c1c7ce6f0ffd92a6e7b3348de120009c649d3bee

Observation 0da41141-771f-40f8-aeda-fcc84f6ac6df · outbound

This paper cites Building variable- sized models via learngene pool.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Building variable- sized models via learngene pool

Reference 2018

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:41:46.637519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:41:42.654304Z digest=sha256:4431f2c7494213fcb2290cd529a276450e5646379fb8309412e484dffd03034f

Observation 73c67a20-8c1a-40ae-bdf5-6971a45b3686 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-06T23:41:40.275604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:41:40.275604Z digest=sha256:03c14f7ee2ffd9887255fbf4c3e1ea303c42af45b56602787a2f56bae1627278

Observation f24fbc33-3743-436b-89f8-99ac3074417b · outbound

This paper cites Transferring Core Knowledge via Learngenes.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Transferring Core Knowledge via Learngenes

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T23:41:40.379368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:41:40.379368Z digest=sha256:28fcd32851b9f33ee678ba445c2058671aad91368629635c4d96d36aa41ea761

Observation 59e47a0f-3e1a-43cd-b200-b7144e282dc5 · outbound

This paper cites Reproducible scaling laws for contrastive language-image learning.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Reproducible scaling laws for contrastive language-image learning

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:41:46.829621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:41:40.005207Z digest=sha256:2d3c1dcb065e3ee42403fdcfd8974c5ac937f6e4391a0c1e0ff2e9c1ff1d5002

Observation 2a080e28-5759-4fe3-ad22-af9afa26175d · outbound

This paper cites Food-101–mining discriminative com- ponents with random forests.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Food-101–mining discriminative com- ponents with random forests

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T23:41:39.837013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:41:39.837013Z digest=sha256:6d431b0ea903d1c6ac1dd106422e1130a2ef982c1b258e3af1d0c70fc3c77d84

Observation 9612f030-0cc3-4c8e-b862-f57471afba35 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Imagenet: A large-scale hierarchical image database

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T23:41:40.089980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:41:40.089980Z digest=sha256:2997196aea713da8d15faa281a8448bc773e327a7af29e2b898cc0e35b6c9000

Observation 1645c199-146a-4897-910e-bac736c52528 · outbound

This paper cites WAVE: Weight Templates for Adaptive Initialization of Variable-sized Models.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge WAVE: Weight Templates for Adaptive Initialization of Variable-sized Models

Reference 2024

Resolution
verified exact
local_arxiv, observed 2026-08-06T23:41:45.019297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:41:40.522210Z digest=sha256:1ac8193f8c0c30c4d8511c5878bda7d168da466f49b9b8effc24ab2dac94c960

Pith citing papers

Observation cddb2aaa-c23f-4a4c-ae77-8b645044def7 · inbound

Understanding Performance Collapse in Layer-Pruned Large Language Models via Decision Representation Transitions cites this paper.

Understanding Performance Collapse in Layer-Pruned Large Language Models via Decision Representation Transitions Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-11T02:25:53.878606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-11T02:23:52.589354Z digest=sha256:7733c1b6f61e70049a872ec6b2d63a95ef7c5d98b33dbf084a7af80074465d49

Observation ce73f04e-e77b-4e33-a86e-6f8d47c2ade2 · inbound

Chain-based Distillation for Effective Initialization of Variable-Sized Small Language Models cites this paper.

Chain-based Distillation for Effective Initialization of Variable-Sized Small Language Models Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:40:54.143600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-11T03:31:33.964471Z digest=sha256:f9030f318223ec1660eb75dc3c69d56c7bfa83ea25b4247a5479b8abc6877dc2