Pith. sign in

Paper Citation Record · LEDGER

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers

As of 18 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 0 inbound Pith citation observations for arXiv:2411.14789.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.14789 v2

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T14:59:52.281615Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

46 of 46 outbound references displayed

  • verified exact0
  • verified fuzzy18
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3667f728-116f-4d70-8df0-b86563989a6b · outbound

This paper cites Learning transferable visual models from natural language supervision.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Learning transferable visual models from natural language supervision

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.114205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.114205Z digest=sha256:cae0598f7f4b9b99c563309271a29990b514aaadc8c1bc46894c2a31e2fddb6a

Observation 7f3b4057-6abf-4e5c-ba0f-6635cd45248c · outbound

This paper cites Mobileclip: Fast image-text models through multi-modal reinforced training.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Mobileclip: Fast image-text models through multi-modal reinforced training

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:59:52.788275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:59:52.119201Z digest=sha256:472823d53b35e8680140b41b3b896f21612f979838d67eeffdc68afc9a51cdf8

Observation bd51ce5b-e1d3-47b9-939d-baad2c37c84a · outbound

This paper cites Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.123160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.123160Z digest=sha256:800187571612609bcd1b53dcce805ef5f782c920ab51c03028b901c49f37f1b0

Observation 95ba7f40-9837-4a6a-96f4-48f3644cb71e · outbound

This paper cites FILIP: Fine-grained Interactive Language-Image Pre-Training.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.127302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.127302Z digest=sha256:218237a7ce6173304d5524badc993ab51ba420a92d682737cf3e246526c5d6e5

Observation 772f5775-3f49-4cc9-96ca-86ee78d549a0 · outbound

This paper cites Slip: Self-supervision meets language-image pre-training.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Slip: Self-supervision meets language-image pre-training

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:59:52.768761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:59:52.131546Z digest=sha256:66dd58aec94b10b679b68029bca0f6bc1fce8e90b9572050957c23d8fd8d8f25

Observation 9e9d7ef6-4019-4fdb-ad6a-63b163c52140 · outbound

This paper cites Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.135248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.135248Z digest=sha256:8d36f903aeaf340906756c064af376918e1c25c4c4ba8ad143818acfd103bb7d

Observation deefb527-540a-4655-a01e-204eb4de7628 · outbound

This paper cites Unified contrastive learning in image-text-label space.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Unified contrastive learning in image-text-label space

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:59:52.757540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:59:52.139565Z digest=sha256:daddc1709186b0f7b5b3f1ab990dbe7e0c2e7e9eedd9f2065de21d296c909231

Observation befa4a88-7088-4800-977f-9293a0e45170 · outbound

This paper cites Sigmoid loss for language image pre- training.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Sigmoid loss for language image pre- training

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:59:52.745608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:59:52.142687Z digest=sha256:3eb158b7761d964925c2f9819f2d563ab54134abd16f38153833d7ce7175c8c0

Observation d224ba9a-6194-4892-b457-f70d39b92c84 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:59:52.733608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:59:52.145599Z digest=sha256:8b1245422fd2da210e66eda05018c71ad0e4f602f5a10f6ddb05446b0e866883

Observation d2113434-c7b6-416a-b61d-d75c170e37ad · outbound

This paper cites Distilling the Knowledge in a Neural Network.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Distilling the Knowledge in a Neural Network

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.148493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.148493Z digest=sha256:658501f07721f5c52f76f34fa41af5097243b968a0279120575d5cf885db5082

Observation 0e814fee-c16d-467d-a4cc-56dc16025f96 · outbound

This paper cites Tinyclip: Clip distillation via affinity mimicking and weight inheritance.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Tinyclip: Clip distillation via affinity mimicking and weight inheritance

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.151714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.151714Z digest=sha256:2c3ad1a48c2e9c9405ee352a740f196dbd991c803663a31bfe8f6489ddc435fb

Observation d09a2960-9757-4414-b699-6877ec18ba32 · outbound

This paper cites Clip-kd: An empirical study of clip model distillation.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Clip-kd: An empirical study of clip model distillation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.154940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.154940Z digest=sha256:66434115092ff792d5d540e57b77ec185b4f0bf0d1d2d3caf214cd5fa03b6a15

Observation 06934296-ab32-4eb8-921d-1b70242d165a · outbound

This paper cites Data Filtering Networks.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Data Filtering Networks

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.157799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.157799Z digest=sha256:4fd501327f35bc15d630ee84382db8ae73293948c918e5b282677cc4dbe88098

Observation 6da76b83-de18-419b-b00f-2480161d7073 · outbound

This paper cites Datacomp: In search of the next generation of multimodal datasets.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Datacomp: In search of the next generation of multimodal datasets

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.161256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.161256Z digest=sha256:95e42ae8ba5a4e21a0f98a2a046dba4502ed904b5d3d52609dc479f58f7d9484

Observation 0046f2d9-6ffa-40b8-8898-93617acc3956 · outbound

This paper cites Alip: Adaptive language-image pre-training with synthetic caption.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Alip: Adaptive language-image pre-training with synthetic caption

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:59:52.689960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:59:52.164341Z digest=sha256:63a64433a944f52922b644e911e1b0fffbf54876d668b6b7f7717a57350e06a9

Observation 99f775da-c713-4249-abd7-54cd86e75a8d · outbound

This paper cites VeCLIP: Improving CLIP Training via Visual-enriched Captions.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers VeCLIP: Improving CLIP Training via Visual-enriched Captions

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.167679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.167679Z digest=sha256:49dc1823b2fde230c02e6e20d762b489a769825dcf5453217fe22973a9458bd7

Observation 36e15436-1244-43c3-8391-60b30af1edca · outbound

This paper cites Metaformer is actually what you need for vision.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Metaformer is actually what you need for vision

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:59:52.679633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:59:52.171120Z digest=sha256:be685891c5aa5ce67d06eaccf72c42cf95a83bf7afd141173c4bfc8dcb149044

Observation c8118bb7-c484-474a-83da-f2b90090f27f · outbound

This paper cites Deeply supervised salient object detection with short connections.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Deeply supervised salient object detection with short connections

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:59:52.668728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:59:52.174562Z digest=sha256:cde3436bfd6d96d5a1bee12372f5f3b024f0c6459755d408ad7e3bbe5a20da57

Observation ecfb34f2-c296-4364-86e2-a67e3bae9621 · outbound

This paper cites Cascaded partial decoder for fast and accurate salient object detection.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Cascaded partial decoder for fast and accurate salient object detection

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:59:52.657079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:59:52.178714Z digest=sha256:7d57b6a4fa4fba4516694d9784c8554316c5d6824f515483c0f587d2bbadef1f

Observation ec24a1d5-3abc-4585-ae60-f988163ab6d0 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.182930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.182930Z digest=sha256:7ad4fc5a56e3bcb7657cb1b8dffe756d34fbdc43a476204d388493101c013485

Observation fd94dcaf-8fd9-41bc-acb9-11a1dab9fa50 · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Swin transformer: Hierarchical vision transformer using shifted windows

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.186821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.186821Z digest=sha256:112adb1308aa5a80ae8147bf2a050d484c3a1d1508d4ce08a8b627e8e26f2a32

Observation abbf80c6-af9b-4da1-b80a-8f4a4d48a0d3 · outbound

This paper cites Less is more: Pay less attention in vision transformers.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Less is more: Pay less attention in vision transformers

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:59:52.637839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:59:52.190581Z digest=sha256:000185324d77cf6bbd69f74d5af1ed536026a48bcc87200ffb98e2e3be206bcf

Observation fd86234b-74ed-4a34-821a-ad2ef9365662 · outbound

This paper cites Cmt: Convolutional neural networks meet vision transformers.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Cmt: Convolutional neural networks meet vision transformers

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:59:52.625709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:59:52.194481Z digest=sha256:6f21304681bc5adc2c9b6499cdf6ef74e6c3a3cc3a7c6f50e27556abf86cd567

Observation ddca9952-036c-4d8c-8fc8-ac33cd6bcebb · outbound

This paper cites Fastvit: A fast hybrid vision transformer using structural reparameterization.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Fastvit: A fast hybrid vision transformer using structural reparameterization

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:59:52.613709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:59:52.200277Z digest=sha256:f8770f536d8999b05b672124eced9b0a95c00044d71129d3023e27416e138bba

Observation 97a4c6e0-62c2-4894-9c15-4ad64edd5671 · outbound

This paper cites Universal Transformers.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Universal Transformers

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.205965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.205965Z digest=sha256:d6cf40fed0e11fd4a27031ad0994aaf08ecfcc4ad767be0d8bce1ef231b114d1

Observation e5f31a3b-c752-4246-8d4a-651867b554eb · outbound

This paper cites Perceiver: General perception with iterative attention.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Perceiver: General perception with iterative attention

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.210190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.210190Z digest=sha256:53acb40686de45ea30cb25f43ad3727d03d7dc579f9cb0be3df57efcded54d4d

Observation 645b9ae5-35a5-4397-b4ef-4028d0cb4d43 · outbound

This paper cites Sharing low rank conformer weights for tiny always-on ambient speech recognition models.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Sharing low rank conformer weights for tiny always-on ambient speech recognition models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:59:52.593450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:59:52.214395Z digest=sha256:2698e6f40c8ac3b7754343d876a5c0f97c210b3410f068056640f7678cede1fe

Observation de702da3-e804-42f5-9342-b6696ed0b4fb · outbound

This paper cites Conformer: Convolution-augmented Transformer for Speech Recognition.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.219465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.219465Z digest=sha256:d20610b9c69144e09f1cb10b284a34e9e35c64946e7fdb1735831b3878d914d3

Observation 19bc353d-c96a-4333-9622-b5e048a40494 · outbound

This paper cites Simplifying transformer blocks.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Simplifying transformer blocks

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:59:52.581833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:59:52.223873Z digest=sha256:2a7850b126c1a40f571198b432d2b05281b81767a21443dab85878fb20abdb7b

Observation 26350de0-be99-4fa9-aadd-01befde4bec1 · outbound

This paper cites Attention is all you need.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Attention is all you need

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.227589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.227589Z digest=sha256:3d3449221668736bead39d1467ef43fb22d03a1f2d845e3fbbad5c703122f4cc

Observation 011605a2-487f-4b23-815c-7495e8f36552 · outbound

This paper cites The shaped transformer: Attention models in the infinite depth-and-width limit.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers The shaped transformer: Attention models in the infinite depth-and-width limit

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:59:52.564010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:59:52.232264Z digest=sha256:b30988158978e2384fb831d8e7e835ac83ce4c2a51d4a425674bbf0f348bf280

Observation 5d5940fc-69ea-4bae-852d-4586dd2e2c14 · outbound

This paper cites Augmenting Self-attention with Persistent Memory.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Augmenting Self-attention with Persistent Memory

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.236035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.236035Z digest=sha256:a7010026e55732467d2a5d7f86fbba1c8891efd386f4360080fc18840c2abaa4

Observation d4744f33-043d-4ad0-b297-e92d696d6a8e · outbound

This paper cites Beyond Scaling Laws: Understanding Transformer Performance with Associative Memory.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Beyond Scaling Laws: Understanding Transformer Performance with Associative Memory

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.240108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.240108Z digest=sha256:3191c351256f833e627fbcec7af93de5bc6a313cced90fb66efde84500f40403

Observation 4a2d84b0-0353-4449-84eb-a5c79032403e · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers LoRA: Low-Rank Adaptation of Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.244204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.244204Z digest=sha256:bc64c0e15a762c7b39d7595b52f9d698dad6b96e834b4977e8d66d3c5c024d8e

Observation a43af1cc-e3a0-4340-9a0b-9d5d8037fd69 · outbound

This paper cites CoCa: Contrastive Captioners are Image-Text Foundation Models.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers CoCa: Contrastive Captioners are Image-Text Foundation Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.247115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.247115Z digest=sha256:9b596d77573131a64baf27beb85fe687e02227602eefe8606c412a3214eb91af

Observation 57d172d9-f745-4b17-acf7-c7d30bbe706b · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models.Advances in Neural Information Processing Systems, 35:25278–25294, 2022.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Laion-5b: An open large-scale dataset for training next generation image-text models.Advances in Neural Information Processing Systems, 35:25278–25294, 2022

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.250581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.250581Z digest=sha256:ec0f02db9b2b7c048d86e98bace3b21fc5b5a54c02fe5abc2eba42babe478c26

Observation 1ed292a2-cff2-4386-9a2c-90537d6deea3 · outbound

This paper cites Rils: Masked visual reconstruction in language semantic space.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Rils: Masked visual reconstruction in language semantic space

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:59:52.543998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:59:52.253198Z digest=sha256:a181ef7f11d82c2ca340c1935637a9c122b0a8f4c4d451a762202048425867ca

Observation 8e900ed7-dc7f-4e8d-b091-74c8b1ba186c · outbound

This paper cites Extract free dense labels from clip.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Extract free dense labels from clip

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.255907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.255907Z digest=sha256:9575402dfd910fb9d732fe6aedb9ee2695e3f266a32e6835fc06eaa90216ced2

Observation a2a50556-d133-47dc-bc8d-39e181098183 · outbound

This paper cites Openclip repository, 2021.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Openclip repository, 2021

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:59:52.520710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:59:52.258639Z digest=sha256:8ebd41691414f23306555ea1e29f42ce65e59a94aaf311ea7f4a64e8e72ba0b4

Observation f7f7c6fd-a8e6-4aab-911b-4e83350e8020 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Imagenet: A large-scale hierarchical image database

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.261429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.261429Z digest=sha256:9c8f960a626102a71647faf9a03b9df452ed8867fc5d79135b65b8a22df35d15

Observation 4241ff8b-9f35-40f9-bf3d-38bb0188fd6d · outbound

This paper cites Do imagenet classifiers generalize to imagenet? In International conference on machine learning, pages 5389–5400.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Do imagenet classifiers generalize to imagenet? In International conference on machine learning, pages 5389–5400

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.264271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.264271Z digest=sha256:69ec4c2501714a7a95070905d0855748b203b448fdb75522b57c56a390f1fa8b

Observation f0d35b14-1eaa-45ee-b25b-470128a85c44 · outbound

This paper cites The many faces of robustness: A critical analysis of out-of-distribution generalization.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers The many faces of robustness: A critical analysis of out-of-distribution generalization

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.267287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.267287Z digest=sha256:28cfadbceedac6c6f96c17fbfad41fd72e9c9025b0e82d8e505dff2a1fb1a5a6

Observation c38f53c5-6c36-4c81-917b-9db95a4c2a22 · outbound

This paper cites Learning robust global representations by penalizing local predictive power.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Learning robust global representations by penalizing local predictive power

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.270399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.270399Z digest=sha256:2ec13cdccd71b0c90d9c6fad9bbf1687a60b295c8b2b91a25d114f003576e44e

Observation 7d573126-e6cb-48ac-b674-e921a159f385 · outbound

This paper cites Microsoft coco: Common objects in context.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Microsoft coco: Common objects in context

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.273910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.273910Z digest=sha256:304394b10d4a8e12fea26adc4fc8a8978e4fcfed61ac9de5300c871ed377663c

Observation 84f5fab0-faaf-4a37-ba18-c40afd5f6426 · outbound

This paper cites Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.277763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.277763Z digest=sha256:0e5b102140372611aff6e3e322f4f939e3705304f12448576497868b2de194a6

Observation d1580e7f-5c9a-4e48-959a-6db2de7f8d86 · outbound

This paper cites Randaugment: Practical automated data augmentation with a reduced search space.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Randaugment: Practical automated data augmentation with a reduced search space

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:59:52.452649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:59:52.281615Z digest=sha256:319048c60d8aab9b65f0006d451d7a119a8232f2f8d7863f76c327005b64e6c9

Pith citing papers

No inbound Pith citation observations are available.