Text distillation from BioCLIP-2 into BioLingual creates audio-image alignment for bird species retrieval without any audio-image training pairs.
Bioclip 2: Emergent properties from scaling hierarchical contrastive learning
7 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
roles
background 1polarities
background 1representative citing papers
AVA-Bench evaluates vision foundation models by disentangling 14 atomic visual abilities with aligned training-test distributions to reveal precise ability fingerprints.
Cross-modal agreement between zero-shot visual detection and acoustic classification recovers known Milu deer activity patterns consistent with published behavioral priors, enabling annotation-light validation.
PRIMA boosts 3D quadruped mesh recovery by injecting BioCLIP biological priors and using test-time adaptation with 2D constraints to build the Quadruped3D pseudo-3D dataset and reach SOTA on imbalanced animal benchmarks.
Compact binary hypercube embeddings enable efficient text-to-image and text-to-audio retrieval in wildlife databases with performance competitive to continuous embeddings but far lower memory and search costs.
CropVLM is a domain-adapted vision-language model that achieves 72.51% zero-shot crop classification accuracy and superior open-set detection performance on novel species without retraining.
Applying multi-object tracking to fuse softmax probabilities across frames in camera trap data yields weighted F1-score gains of 5.1%, 3.1%, and 2.0% over standalone classifiers on three datasets.
citing papers explorer
-
AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models
AVA-Bench evaluates vision foundation models by disentangling 14 atomic visual abilities with aligned training-test distributions to reveal precise ability fingerprints.
-
Cross-Modal Corroboration for Annotation-Free Wildlife Monitoring
Cross-modal agreement between zero-shot visual detection and acoustic classification recovers known Milu deer activity patterns consistent with published behavioral priors, enabling annotation-light validation.
-
PRIMA: Boosting Animal Mesh Recovery with Biological Priors and Test-Time Adaptation
PRIMA boosts 3D quadruped mesh recovery by injecting BioCLIP biological priors and using test-time adaptation with 2D constraints to build the Quadruped3D pseudo-3D dataset and reach SOTA on imbalanced animal benchmarks.
-
Compact Hypercube Embeddings for Fast Text-based Wildlife Observation Retrieval
Compact binary hypercube embeddings enable efficient text-to-image and text-to-audio retrieval in wildlife databases with performance competitive to continuous embeddings but far lower memory and search costs.
-
CropVLM: A Domain-Adapted Vision-Language Model for Open-Set Crop Analysis
CropVLM is a domain-adapted vision-language model that achieves 72.51% zero-shot crop classification accuracy and superior open-set detection performance on novel species without retraining.
-
Multi-Object Tracking Consistently Improves Wildlife Inference
Applying multi-object tracking to fuse softmax probabilities across frames in camera trap data yields weighted F1-score gains of 5.1%, 3.1%, and 2.0% over standalone classifiers on three datasets.