Pith. sign in

Paper Citation Record · LEDGER

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers

As of 16 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 0 inbound Pith citation observations for arXiv:2411.14789.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.14789 v2

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T14:59:52.281615Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

46 of 46 outbound references displayed

  • verified exact0
  • verified fuzzy18
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3667f728-116f-4d70-8df0-b86563989a6b · outbound

This paper cites Learning transferable visual models from natural language supervision.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Learning transferable visual models from natural language supervision

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.114205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.114205Z digest=sha256:8f17bb14a0b6e3d03c1f444f4edd94c29dceeafaaae1b763cca326ec72072368

Observation 7f3b4057-6abf-4e5c-ba0f-6635cd45248c · outbound

This paper cites Mobileclip: Fast image-text models through multi-modal reinforced training.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Mobileclip: Fast image-text models through multi-modal reinforced training

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:59:52.788275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:59:52.119201Z digest=sha256:06df24131878dbbaa0b17f4a019610aeef3099d9ecb952822cfce2668ba76483

Observation bd51ce5b-e1d3-47b9-939d-baad2c37c84a · outbound

This paper cites Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.123160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.123160Z digest=sha256:6e1a1bf6fa697311e374ee890da989231e74de47a2e6717f71b84621ed22decb

Observation 95ba7f40-9837-4a6a-96f4-48f3644cb71e · outbound

This paper cites FILIP: Fine-grained Interactive Language-Image Pre-Training.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.127302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.127302Z digest=sha256:7f47651cc525f3683fad620ffeee088ff76096be4dabcc75c7bda301ade66db6

Observation 772f5775-3f49-4cc9-96ca-86ee78d549a0 · outbound

This paper cites Slip: Self-supervision meets language-image pre-training.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Slip: Self-supervision meets language-image pre-training

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:59:52.768761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:59:52.131546Z digest=sha256:42ddca0417b099fed8661f107a6f07c1137ac0a6153cfbc91958c7e2ecb6cd87

Observation 9e9d7ef6-4019-4fdb-ad6a-63b163c52140 · outbound

This paper cites Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.135248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.135248Z digest=sha256:d6f750b6fb65d044920f2ca362def3dbd2befd7fab350d21fc02dedef50f3cbd

Observation deefb527-540a-4655-a01e-204eb4de7628 · outbound

This paper cites Unified contrastive learning in image-text-label space.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Unified contrastive learning in image-text-label space

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:59:52.757540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:59:52.139565Z digest=sha256:bb59d8055c85319bf239af094e7b222a7dba9f70d4550c21211cc58174513536

Observation befa4a88-7088-4800-977f-9293a0e45170 · outbound

This paper cites Sigmoid loss for language image pre- training.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Sigmoid loss for language image pre- training

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:59:52.745608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:59:52.142687Z digest=sha256:f3b061ade79074a93f1ec49076fc4d3e84d72b095a53c09e3a026d47c24cd6b4

Observation d224ba9a-6194-4892-b457-f70d39b92c84 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:59:52.733608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:59:52.145599Z digest=sha256:8b246fd27112dbd11b49ed85adeb2bec5080c25c89e4526e05768d5995abca80

Observation d2113434-c7b6-416a-b61d-d75c170e37ad · outbound

This paper cites Distilling the Knowledge in a Neural Network.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Distilling the Knowledge in a Neural Network

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.148493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.148493Z digest=sha256:41fe7683ed163f9e095a202a3f2ce4f516a92c751aa5502b3c184e738afc2d40

Observation 0e814fee-c16d-467d-a4cc-56dc16025f96 · outbound

This paper cites Tinyclip: Clip distillation via affinity mimicking and weight inheritance.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Tinyclip: Clip distillation via affinity mimicking and weight inheritance

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.151714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.151714Z digest=sha256:70708aa88600397e42893c19b5ad225d504b5f77339353741318cba0075ed905

Observation d09a2960-9757-4414-b699-6877ec18ba32 · outbound

This paper cites Clip-kd: An empirical study of clip model distillation.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Clip-kd: An empirical study of clip model distillation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.154940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.154940Z digest=sha256:1dedb784e980fe3e4756cde828eb3124a14eb029e363ebdd08a27052aca2d62c

Observation 06934296-ab32-4eb8-921d-1b70242d165a · outbound

This paper cites Data Filtering Networks.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Data Filtering Networks

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.157799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.157799Z digest=sha256:e2a0822968025d9a85eaec61048c416245e87a865925705b53f8175848291e73

Observation 6da76b83-de18-419b-b00f-2480161d7073 · outbound

This paper cites Datacomp: In search of the next generation of multimodal datasets.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Datacomp: In search of the next generation of multimodal datasets

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.161256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.161256Z digest=sha256:5bb355bb02649e3912a43f80ba7e912de475fbd9cabe761fcb43364b481b7dd7

Observation 0046f2d9-6ffa-40b8-8898-93617acc3956 · outbound

This paper cites Alip: Adaptive language-image pre-training with synthetic caption.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Alip: Adaptive language-image pre-training with synthetic caption

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:59:52.689960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:59:52.164341Z digest=sha256:de35fed1a9461fdf107582ff9ba0d263b46ec21ae9fe04c9a69a16f481a65b66

Observation 99f775da-c713-4249-abd7-54cd86e75a8d · outbound

This paper cites VeCLIP: Improving CLIP Training via Visual-enriched Captions.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers VeCLIP: Improving CLIP Training via Visual-enriched Captions

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.167679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.167679Z digest=sha256:340d15ac945d530a7674ebf07b3bb9aea78298b464abf5ea85359fb6ce961431

Observation 36e15436-1244-43c3-8391-60b30af1edca · outbound

This paper cites Metaformer is actually what you need for vision.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Metaformer is actually what you need for vision

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:59:52.679633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:59:52.171120Z digest=sha256:5fb5a9fea72eb720edd0cdf6927523066bd1e3cd3eaa5649dabd9fb47e6804ed

Observation c8118bb7-c484-474a-83da-f2b90090f27f · outbound

This paper cites Deeply supervised salient object detection with short connections.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Deeply supervised salient object detection with short connections

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:59:52.668728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:59:52.174562Z digest=sha256:2debdb91c11eec69993bb1468caf9edbd03cac94653a91e14c483aec9d2919f7

Observation ecfb34f2-c296-4364-86e2-a67e3bae9621 · outbound

This paper cites Cascaded partial decoder for fast and accurate salient object detection.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Cascaded partial decoder for fast and accurate salient object detection

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:59:52.657079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:59:52.178714Z digest=sha256:e97ddb3bf678890521f4277ada462ced3749bf96e247ba2d8d1fae9957690a23

Observation ec24a1d5-3abc-4585-ae60-f988163ab6d0 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.182930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.182930Z digest=sha256:333b00f893c4f58ff56895c1e27ee9707aa24c977a7384212b8e3dd0d78a8810

Observation fd94dcaf-8fd9-41bc-acb9-11a1dab9fa50 · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Swin transformer: Hierarchical vision transformer using shifted windows

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.186821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.186821Z digest=sha256:10ff63c0ba844a8b0e467d7184fc7255cb2785578784f8aa5d06338ef4bd3983

Observation abbf80c6-af9b-4da1-b80a-8f4a4d48a0d3 · outbound

This paper cites Less is more: Pay less attention in vision transformers.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Less is more: Pay less attention in vision transformers

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:59:52.637839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:59:52.190581Z digest=sha256:6222db21cc7bf362153270d25fee79dd9691ccbefa55f9abab7e256faa62531f

Observation fd86234b-74ed-4a34-821a-ad2ef9365662 · outbound

This paper cites Cmt: Convolutional neural networks meet vision transformers.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Cmt: Convolutional neural networks meet vision transformers

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:59:52.625709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:59:52.194481Z digest=sha256:5519b5a0ca9004bc3a6d568767b509ef092dd1ac6c3c007f8d4323bac67ce5d1

Observation ddca9952-036c-4d8c-8fc8-ac33cd6bcebb · outbound

This paper cites Fastvit: A fast hybrid vision transformer using structural reparameterization.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Fastvit: A fast hybrid vision transformer using structural reparameterization

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:59:52.613709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:59:52.200277Z digest=sha256:2f92fd7f5452fd497c8e0cf09e385936bce846e18bc38d3517d1d56a3e71660b

Observation 97a4c6e0-62c2-4894-9c15-4ad64edd5671 · outbound

This paper cites Universal Transformers.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Universal Transformers

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.205965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.205965Z digest=sha256:4cb935ea30cadb9d832c82e81cf581429968707103675b4f1b63b0e958e159c6

Observation e5f31a3b-c752-4246-8d4a-651867b554eb · outbound

This paper cites Perceiver: General perception with iterative attention.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Perceiver: General perception with iterative attention

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.210190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.210190Z digest=sha256:9b5109c78dd1a12bb1e45b7431c90c5407dc3e6ad85e75c59e3705e1c1a44412

Observation 645b9ae5-35a5-4397-b4ef-4028d0cb4d43 · outbound

This paper cites Sharing low rank conformer weights for tiny always-on ambient speech recognition models.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Sharing low rank conformer weights for tiny always-on ambient speech recognition models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:59:52.593450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:59:52.214395Z digest=sha256:936b3323d889b5888a3e583c7b51c5c307fdbb97b985f7b120118da63af548ab

Observation de702da3-e804-42f5-9342-b6696ed0b4fb · outbound

This paper cites Conformer: Convolution-augmented Transformer for Speech Recognition.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.219465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.219465Z digest=sha256:04a59eb13569f072d213d45829d5731fa65f40c25b0e1e95799edd6d17334f0e

Observation 19bc353d-c96a-4333-9622-b5e048a40494 · outbound

This paper cites Simplifying transformer blocks.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Simplifying transformer blocks

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:59:52.581833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:59:52.223873Z digest=sha256:6a24193c2424ac3b08b2bd1bcd8c9dbc58e36708f34582554b80061bf0eaca34

Observation 26350de0-be99-4fa9-aadd-01befde4bec1 · outbound

This paper cites Attention is all you need.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Attention is all you need

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.227589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.227589Z digest=sha256:4857541d00fbd0ae45f4cd63e502c0b894ed5db7eea564b3e885704d30ab3347

Observation 011605a2-487f-4b23-815c-7495e8f36552 · outbound

This paper cites The shaped transformer: Attention models in the infinite depth-and-width limit.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers The shaped transformer: Attention models in the infinite depth-and-width limit

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:59:52.564010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:59:52.232264Z digest=sha256:dd7f7d30bc97381aec80ab59310c4ce016846744401d7d10e5a8c7aabdcec73d

Observation 5d5940fc-69ea-4bae-852d-4586dd2e2c14 · outbound

This paper cites Augmenting Self-attention with Persistent Memory.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Augmenting Self-attention with Persistent Memory

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.236035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.236035Z digest=sha256:f84a1c8eb909ec9f5bd1a58fe7bcdf9e951223f442251a014bdccdc97c1c980e

Observation d4744f33-043d-4ad0-b297-e92d696d6a8e · outbound

This paper cites Beyond Scaling Laws: Understanding Transformer Performance with Associative Memory.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Beyond Scaling Laws: Understanding Transformer Performance with Associative Memory

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.240108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.240108Z digest=sha256:a4f5ec39eb4d4cb922fe3709fde5de9e9123345397fab87b3193f93639bf791b

Observation 4a2d84b0-0353-4449-84eb-a5c79032403e · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers LoRA: Low-Rank Adaptation of Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.244204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.244204Z digest=sha256:ccfbfc2fd3766cb246ab8f8406fce1006da7e520fdab993480da703944cfadbf

Observation a43af1cc-e3a0-4340-9a0b-9d5d8037fd69 · outbound

This paper cites CoCa: Contrastive Captioners are Image-Text Foundation Models.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers CoCa: Contrastive Captioners are Image-Text Foundation Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.247115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.247115Z digest=sha256:0d62a1f92e8a9147aa31d5ba263677f75f9ee25d253e3d7dff914cbd5fa3f1a9

Observation 57d172d9-f745-4b17-acf7-c7d30bbe706b · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models.Advances in Neural Information Processing Systems, 35:25278–25294, 2022.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Laion-5b: An open large-scale dataset for training next generation image-text models.Advances in Neural Information Processing Systems, 35:25278–25294, 2022

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.250581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.250581Z digest=sha256:d11323852bf74280ee10781a0c2681e89060c85e10bbd64eda8743bea16dd52b

Observation 1ed292a2-cff2-4386-9a2c-90537d6deea3 · outbound

This paper cites Rils: Masked visual reconstruction in language semantic space.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Rils: Masked visual reconstruction in language semantic space

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:59:52.543998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:59:52.253198Z digest=sha256:33165f72c7185da49b4b5e23bce399e68a11e53ddb28732bbbf91eba6e73d446

Observation 8e900ed7-dc7f-4e8d-b091-74c8b1ba186c · outbound

This paper cites Extract free dense labels from clip.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Extract free dense labels from clip

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.255907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.255907Z digest=sha256:29a536a65725bc3aa8c4c3177de7d8aafa9728330466f17b8656b6b4434ef9f6

Observation a2a50556-d133-47dc-bc8d-39e181098183 · outbound

This paper cites Openclip repository, 2021.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Openclip repository, 2021

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:59:52.520710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:59:52.258639Z digest=sha256:00c0d601a41d1e39301d4233939de1f195fee1918b0d00beec9f363e3f66cec2

Observation f7f7c6fd-a8e6-4aab-911b-4e83350e8020 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Imagenet: A large-scale hierarchical image database

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.261429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.261429Z digest=sha256:e1b34db473f2f83cdb21777b26769121afe9d193583b155b1315cfd402ef8560

Observation 4241ff8b-9f35-40f9-bf3d-38bb0188fd6d · outbound

This paper cites Do imagenet classifiers generalize to imagenet? In International conference on machine learning, pages 5389–5400.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Do imagenet classifiers generalize to imagenet? In International conference on machine learning, pages 5389–5400

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.264271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.264271Z digest=sha256:fbb10daa88373d76ae4ad640f49aff4bdc5b71cfc88a7a353ec781ca792c570c

Observation f0d35b14-1eaa-45ee-b25b-470128a85c44 · outbound

This paper cites The many faces of robustness: A critical analysis of out-of-distribution generalization.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers The many faces of robustness: A critical analysis of out-of-distribution generalization

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.267287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.267287Z digest=sha256:ff86d4151e2bf49ae7bfb74f709a9924904012fb24297eaaeaef67332e80974d

Observation c38f53c5-6c36-4c81-917b-9db95a4c2a22 · outbound

This paper cites Learning robust global representations by penalizing local predictive power.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Learning robust global representations by penalizing local predictive power

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.270399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.270399Z digest=sha256:b53a1f8363274765dc774dd7955fb7ebabbe1fae30f209e8a0a31d0604aba293

Observation 7d573126-e6cb-48ac-b674-e921a159f385 · outbound

This paper cites Microsoft coco: Common objects in context.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Microsoft coco: Common objects in context

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.273910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.273910Z digest=sha256:86a955f4efc598e239c52b7199b6feffb2710450391ad8fa4ac2931994ef4de2

Observation 84f5fab0-faaf-4a37-ba18-c40afd5f6426 · outbound

This paper cites Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.277763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.277763Z digest=sha256:baac91acbd6d36635d16145b3b7be1b17eb08207b4b497f54f4a2b28ecc439c7

Observation d1580e7f-5c9a-4e48-959a-6db2de7f8d86 · outbound

This paper cites Randaugment: Practical automated data augmentation with a reduced search space.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Randaugment: Practical automated data augmentation with a reduced search space

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:59:52.452649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:59:52.281615Z digest=sha256:c8c3bebbe4e55da81088466bcb742cd207c7b70c0f8927205213f49d4f47ac79

Pith citing papers

No inbound Pith citation observations are available.