Pith. sign in

REVIEW 1 cited by

Towards flexible perception with visual memory

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.08172 v3 pith:UC3TBDE7 submitted 2024-08-15 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords knowledgememorydatanetworkvisualabilityacrosscapabilities
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Training a neural network is a monolithic endeavor, akin to carving knowledge into stone: once the process is completed, editing the knowledge in a network is hard, since all information is distributed across the network's weights. We here explore a simple, compelling alternative by marrying the representational power of deep neural networks with the flexibility of a database. Decomposing the task of image classification into image similarity (from a pre-trained embedding) and search (via fast nearest neighbor retrieval from a knowledge database), we build on well-established components to construct a simple and flexible visual memory that has the following key capabilities: (1.) The ability to flexibly add data across scales: from individual samples all the way to entire classes and billion-scale data; (2.) The ability to remove data through unlearning and memory pruning; (3.) An interpretable decision-mechanism on which we can intervene to control its behavior. Taken together, these capabilities comprehensively demonstrate the benefits of an explicit visual memory. We hope that it might contribute to a conversation on how knowledge should be represented in deep vision models -- beyond carving it in "stone" weights.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion

    cs.CV 2025-02 conditional novelty 6.0 of 10

    CLIP's intra-modal embeddings are miscalibrated: converting one side of an image-image or text-text comparison into the other modality improves retrieval accuracy on 15+ datasets.

Pith tools