REVIEW 4 cited by
Learning to Count Anything: Reference-less Class-agnostic Counting with Weak Supervision
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Current class-agnostic counting methods can generalise to unseen classes but usually require reference images to define the type of object to be counted, as well as instance annotations during training. Reference-less class-agnostic counting is an emerging field that identifies counting as, at its core, a repetition-recognition task. Such methods facilitate counting on a changing set composition. We show that a general feature space with global context can enumerate instances in an image without a prior on the object type present. Specifically, we demonstrate that regression from vision transformer features without point-level supervision or reference images is superior to other reference-less methods and is competitive with methods that use reference images. We show this on the current standard few-shot counting dataset FSC-147. We also propose an improved dataset, FSC-133, which removes errors, ambiguities, and repeated images from FSC-147 and demonstrate similar performance on it. To the best of our knowledge, we are the first weakly-supervised reference-less class-agnostic counting method.
Forward citations
Cited by 4 Pith papers
-
SFOOD: A Multimodal Benchmark for Comprehensive Food Attribute Analysis Beyond RGB with Spectral Insights
SFOOD combines existing food datasets with self-collected hyperspectral images to create a six-task benchmark, and its evaluations suggest spectral bands improve sweetness and herbal classification while current model...
-
ISAC: Training-Free Instance-to-Semantic Attention Control for Multi-Instance Generation
ISAC improves multi-instance image generation by carving out instance regions from self-attention first and then assigning semantics to those regions.
-
Single Domain Generalization for Few-Shot Counting via Universal Representation Matching
URM distills CLIP vision-language representations into learnable prototypes for few-shot counting, improving single-domain generalization on unseen datasets.
-
Expanding Zero-Shot Object Counting with Rich Prompts
RichCount improves zero-shot object counting by enriching text prompts with MLLM-generated descriptions and aligning them to CLIP visual features, achieving state-of-the-art mean absolute error on three counting benchmarks.
Discussion (0). Continue with ORCID to comment.