{"work":{"id":"effdb28b-742e-4840-b3ca-d89502a6cd4d","openalex_id":"https://openalex.org/W2108598243","doi":"10.1109/cvpr.2009.5206848","arxiv_id":"2009.520684","raw_key":null,"title":"ImageNet: A large- scale hierarchical image database","authors":null,"authors_text":"Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei","year":2009,"venue":"2009 IEEE Conference on Computer Vision and Pattern Recognition","abstract":null,"external_url":"https://arxiv.org/abs/2009.520684","cited_by_count":61841,"metadata_source":"arxiv_reference","metadata_fetched_at":"2026-08-05T02:28:24.338817+00:00","pith_arxiv_id":null,"created_at":"2026-05-09T19:05:10.388036+00:00","updated_at":"2026-08-05T02:28:24.338817+00:00","title_quality_ok":false,"display_title":"author Dong, W","render_title":"author Dong, W"},"hub":{"state":{"work_id":"effdb28b-742e-4840-b3ca-d89502a6cd4d","tier":"super_hub","tier_reason":"100+ Pith inbound or 10,000+ external citations","pith_inbound_count":112,"external_cited_by_count":61841,"distinct_field_count":21,"first_pith_cited_at":"2019-07-09T21:39:23+00:00","last_pith_cited_at":"2026-07-09T17:51:50+00:00","author_build_status":"needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"needed","reader_status":"not_needed","recognition_status":"not_needed","updated_at":"2026-08-21T15:09:24.005877+00:00","tier_text":"super_hub"},"tier":"super_hub","role_counts":[{"context_role":"background","n":9},{"context_role":"dataset","n":4},{"context_role":"method","n":2}],"polarity_counts":[{"context_polarity":"background","n":8},{"context_polarity":"use_dataset","n":4},{"context_polarity":"use_method","n":2},{"context_polarity":"unclear","n":1}],"runs":{"ask_index":{"job_type":"ask_index","status":"succeeded","result":{"title":"author Dong, W","claims":[{"claim_text":"ratio as baseline first@τdivided by ours first@τ. A speedup ratio greater than 1.0×means ours reaches the same target earlier with fewer epochs or steps. For higher-is-better metrics (Top-1, AP50), first@τis the first epoch with metric at or aboveτ. For lower-is-better metrics (FID), first@τis the first step at or belowτ. Gate and Hyperparameter Selection.For ImageNet classification [7], we useτ= 65for ResNet-50 [14] andτ= 50for ViT-S/16 [8]. For CIFAR early-stage classification [26], we use fix","claim_type":"dataset","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"ϕ(cchild,c parent)< η text(∥˜ cparent∥)·ω(˜ cparent).(8) This allows users to prune entire branches of spurious concepts with a single interaction, substantially reducing the number of interventions required to correct a prediction. 4 Experiments We evaluate HypCBM across three domains:CIFAR-100[ 20] for general object classification, SUN397[ 51] for (hierarchical) scene understanding, andImageNet[ 6] to assess scalability to real- world complexity. Additional results onCUB-200[ 50] are provided","claim_type":"dataset","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"URL https://www.datanami.com/2020/07/06/ data-prep-still-dominates-data-scientists-time-survey-finds/. [8] Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large- scale hierarchical image database. In2009 IEEE Conference on Computer Vision and Pattern Recognition, pages 248-255, 2009. doi: 10.1109/CVPR.2009.5206848. [9] Robert Dorfman. A formula for the gini coefficient.The Review of Economics and Statistics, 61 (1):146-49, 1979. URL https://EconPapers.repec.org","claim_type":"background","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"The inverse-rendering model, invRend-BFM, was trained to infer the BFM generative parameters of a 2D face image, including identity-related shape and texture latents as well as expres- sion, pose, light direction, and light intensity. The object-categorization model, objCat-ImageNet, was trained to classify natural images into ImageNet object categories [46]. Details of the training objective, architectural modi- fications, and training dataset for each model are provided in Methods 4.1. For cop","claim_type":"dataset","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"by shared tasks, common data, and open leaderboards, was the engine behind transformative progress 2 Figure 1:MC 2 pipeline.A low-budget Monte Carlo WoS estimate is corrected in a single forward pass by a learned operator, yielding an improved solution for the PDE. in NLP and computer vision, where benchmarks like GLUE [35], SuperGLUE [36], and ImageNet [8] created a culture of head-to-head comparison on identical inputs. PDE solving has no analog. Existing benchmarks each occupy narrow regimes:","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"Instance Segmentation.Cityscapes [ 6], ADE20K [42], LVIS [12], and Mapillary Vistas [28] cover outdoor driving and general scenes but apply no domain-specific vocabulary tailored to commercial spaces-the escalators, retail shelves, display cases, hotel beds, and food presentations that define the majority of Urban-ImageNet's images. Scaling Behaviour.ImageNet [ 7] established scale as a performance driver; GPT-3 [4] and scaling laws [18] showed predictable growth; LAION-5B [35] demonstrated bill","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks author Dong, W because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (9 contexts).","role_counts":[{"n":9,"context_role":"background"},{"n":4,"context_role":"dataset"},{"n":2,"context_role":"method"}]},"error":null,"updated_at":"2026-06-30T12:19:51.769495+00:00"},"author_expand":{"job_type":"author_expand","status":"succeeded","result":{"authors_linked":[{"id":"178ae3c3-0ee2-4711-b5da-46f62752e506","orcid":null,"display_name":"bchapter Deng"}]},"error":null,"updated_at":"2026-06-30T12:19:52.316064+00:00"},"context_extract":{"job_type":"context_extract","status":"succeeded","result":{"enqueued_papers":25},"error":null,"updated_at":"2026-06-26T09:44:57.092101+00:00"},"graph_features":{"job_type":"graph_features","status":"succeeded","result":{"co_cited":[{"title":"Deep residual learning for image recognition","work_id":"b353bda2-591d-479a-9c8b-22dfcba12431","shared_citers":15},{"title":"Walk in the cloud: Learning curves for point clouds shape analysis, pp","work_id":"3820f598-11b0-45c3-8c99-0079181ac0a7","shared_citers":10},{"title":"A ConvNet for the 2020s","work_id":"0a23d1b7-bd56-43cc-8a80-7c43ce994e1e","shared_citers":9},{"title":"Emogen: Emotional image content generation with text-to-image diffusion models","work_id":"7efbc2dd-b0f2-4f71-bb1c-d2fcf110d805","shared_citers":8},{"title":"In: 2023 IEEE/CVF International Conference on Computer Vision (ICCV)","work_id":"b8a8bb9e-1d31-40e2-9cab-ae21e338dde6","shared_citers":8},{"title":"In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp","work_id":"b9701eca-d05e-4d2e-9045-6761df4ba175","shared_citers":8},{"title":"An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale","work_id":"e96730e3-129b-4db6-b981-15ab7932e297","shared_citers":6},{"title":"BERT : Pre-training of Deep Bidirectional Transformers for Language Understanding","work_id":"3e3c8ac8-b858-4b22-af32-393d98c883e0","shared_citers":6},{"title":"Dickerson","work_id":"5c2060c6-427c-4321-be22-49ccae439d80","shared_citers":6},{"title":"In: Proceedings of the IEEE/CVF Conference on Computer 25 Vision and Pattern Recognition, pp","work_id":"9da51225-b7bd-4032-b7db-ca577971dafe","shared_citers":6},{"title":"Derf: Decomposed radiance fields","work_id":"7083a41e-5666-435b-ab26-c753f6490b9a","shared_citers":5},{"title":"DINOv3","work_id":"c8b07deb-8fe7-4e18-9620-f3569d3529ce","shared_citers":5},{"title":"LLaMA: Open and Efficient Foundation Language Models","work_id":"c018fc23-6f3f-4035-9d02-28a2173b2b9d","shared_citers":5},{"title":"PoseNet: A convolutional network for real-time 6-dof camera relocalization","work_id":"135418b1-cafd-49fd-803d-1ca6433d4b1b","shared_citers":5},{"title":"Very Deep Convolutional Networks for Large-Scale Image Recognition","work_id":"1c4b4409-c14b-488b-a086-c57a5aab8a29","shared_citers":5},{"title":"author Han, D","work_id":"b6acb671-02be-4b1f-ad2a-31ae513ee884","shared_citers":4},{"title":"Deep Residual Learning for Image Recognition","work_id":"ae9e5671-23e8-4853-82a4-699b5b8dd639","shared_citers":4},{"title":"Densely connected convolutional networks","work_id":"2199d436-33c2-4b30-9d6f-ce9b8904101e","shared_citers":4},{"title":"DINOv2: Learning Robust Visual Features without Supervision","work_id":"26b304e5-b54a-4f26-be7e-83299eca52e4","shared_citers":4},{"title":"Hierarchical Text-Conditional Image Generation with CLIP Latents","work_id":"0c6a768b-70b8-4242-bb0e-459f1008c9fc","shared_citers":4},{"title":"Masset, R","work_id":"238df2e4-a3e5-46f3-860e-3ae2b0094b97","shared_citers":4},{"title":"Microsoft COCO: common objects in context","work_id":"9a41dfbd-d7c8-4b32-943f-59c08fdd6db7","shared_citers":4},{"title":"The Llama 3 Herd of Models","work_id":"1549a635-88af-4ac1-acfe-51ae7bb53345","shared_citers":4},{"title":"2016.280","work_id":"45b0bfd8-65dc-4252-b2ab-2f6b411d04d0","shared_citers":3}],"time_series":[{"n":2,"year":2019},{"n":1,"year":2022},{"n":7,"year":2024},{"n":15,"year":2025},{"n":50,"year":2026}],"dependency_candidates":[{"n":1,"role":"dataset","polarity":"use_dataset","paper_title":"Human face perception reflects inverse-generative and naturalistic discriminative objectives","primary_cat":"q-bio.NC","context_text":"The inverse-rendering model, invRend-BFM, was trained to infer the BFM generative parameters of a 2D face image, including identity-related shape and texture latents as well as expres- sion, pose, light direction, and light intensity. The object-categorization model, objCat-ImageNet, was trained to classify natural images into ImageNet object categories [46]. Details of the training objective, architectural modi- fications, and training dataset for each model are provided in Methods 4.1. For copyright reasons, the VGGFace2 [43] images in this figure were replaced with synthetic faces generated using https://thispersondoesnotexist.com/; the ImageNet [46] illustration was obtained from https://commons.wikimedia.","citing_arxiv_id":"2605.12619"},{"n":1,"role":"dataset","polarity":"use_dataset","paper_title":"Removing the Watermark Is Not Enough: Forensic Stealth in Generative-AI Watermark Removal","primary_cat":"cs.CR","context_text":"A separate foren- sic literature studies the traces left by learned image generators and manipulators, and provides much of the methodological foundation we build on. Wang et al. [ 37] show that CNN-generated images contain characteristic fingerprints that simple detectors can exploit, with transfer across unseen generators. Frank et al. [12] and Durall et al. [8] connect these traces to spectral irregularities introduced by upsampling. Marra et al. [ 28] show that GANs leave model- specific fingerprints analogous to camera sensor noise, and Ojha et al. [31] leverage pretrained representations for more universal fake detection across architectures. This literature shows that learned image pipelines leave statistical traces.","citing_arxiv_id":"2605.09203"},{"n":1,"role":"dataset","polarity":"use_dataset","paper_title":"Hyperbolic Concept Bottleneck Models","primary_cat":"cs.LG","context_text":"ϕ(cchild,c parent)< η text(∥˜ cparent∥)·ω(˜ cparent).(8) This allows users to prune entire branches of spurious concepts with a single interaction, substantially reducing the number of interventions required to correct a prediction. 4 Experiments We evaluate HypCBM across three domains:CIFAR-100[ 20] for general object classification, SUN397[ 51] for (hierarchical) scene understanding, andImageNet[ 6] to assess scalability to real- world complexity. Additional results onCUB-200[ 50] are provided in Appendix A. To isolate the impact of geometry, we control for training data by utilizing the pre-trained checkpoints from Pal et al. [33], ensuring both the hyperbolic HyCoCLIP and Euclidean CLIP [38] are trained on the same GrIT- 20M dataset [35]. We also report results for CLIP [38], trained on 400M samples, to contextualize","citing_arxiv_id":"2605.06440"},{"n":1,"role":"method","polarity":"use_method","paper_title":"StomaD2: An All-in-One System for Intelligent Stomatal Phenotype Analysis via Diffusion-Based Restoration Detection Network","primary_cat":"cs.CV","context_text":"Stage Two: Although the regression stage can recover the major structure of the stoma, residual degradation may still remain, especially in regions with severe blur or noise contamination. Therefore, a diffusion-based refinement stage is further introduced to suppress residual degradation while preserving local stomatal structures. We use the fine-tuned Stable Diffusion [26] to reconstruct the image with the obtained (𝐼𝐼𝑟𝑟𝑟𝑟𝑟𝑟− 𝐼𝐼𝐻𝐻𝐿𝐿) . The VAE encoder [27] maps (𝐼𝐼𝑟𝑟𝑟𝑟𝑟𝑟) into the latent space, resulting in the conditional latent (ε�𝐼𝐼𝑟𝑟𝑟𝑟𝑟𝑟�). Objective function is as follow: ℒ𝐷𝐷𝐷𝐷𝐷𝐷𝐷𝐷= 𝔼𝔼𝑧𝑧𝑡𝑡,𝑐𝑐,𝑡𝑡,𝜀𝜀,𝜀𝜀�𝐼𝐼𝑟𝑟𝑟𝑟𝑟𝑟� ��𝜀𝜀− 𝜀𝜀𝜃𝜃�𝑧𝑧𝑡𝑡, 𝑐𝑐, 𝑡𝑡, 𝜀𝜀, 𝜀𝜀�𝐼𝐼𝑟𝑟𝑟𝑟𝑟𝑟��� 2 2 �, (2) where ε ∼ 𝒩𝒩(0, 𝐼𝐼), 𝑧𝑧𝑡𝑡 is the latent code, 𝑐𝑐 is the sample point, and 𝑡𝑡 represents the time step, which is","citing_arxiv_id":"2604.18632"},{"n":1,"role":"dataset","polarity":"use_dataset","paper_title":"Weak-to-Strong Knowledge Distillation Accelerates Visual Learning","primary_cat":"cs.CV","context_text":"ratio as baseline first@τdivided by ours first@τ. A speedup ratio greater than 1.0×means ours reaches the same target earlier with fewer epochs or steps. For higher-is-better metrics (Top-1, AP50), first@τis the first epoch with metric at or aboveτ. For lower-is-better metrics (FID), first@τis the first step at or belowτ. Gate and Hyperparameter Selection.For ImageNet classification [7], we useτ= 65for ResNet-50 [14] andτ= 50for ViT-S/16 [8]. For CIFAR early-stage classification [26], we use fixed dataset-level gates:τ= 75for CIFAR-10 and τ= 60for CIFAR-100. For object detection on the COCO dataset [28], we use a fixed task-level AP50 target (τ= 20%). For diffusion generation on the CIFAR-10 dataset [18,26,32], we use a fixed task-level FID target (τ= 60),","citing_arxiv_id":"2604.15451"},{"n":1,"role":"method","polarity":"use_method","paper_title":"Zero-shot World Models Are Developmentally Efficient Learners","primary_cat":"cs.AI","context_text":"task-performant models for visual tasks as well as the most accurate models of neural responses across the visual cortex [9, 10, 11, 12, 13] and for human-like error patterns [14]. Initially, DNNs required supervision on large labeled datasets [15, 16] and did not transfer broadly to downstream tasks. These issues motivated a shift to self-supervised models, which learn representations by grouping similar or temporally-proximate images [ 17, 18, 19, 20, 21, 22, 23, 24]. Such self- supervised models also turn out to provide a description of neural and cognitive patterns in the primate visual system, rivaling or exceeding the accuracy of the earlier supervised systems [25, 26]. However, from a developmental point of view, modern self-supervised learning methods are a \"glass half full\". When trained on the visual data diet of real infants and children, they achieve","citing_arxiv_id":"2604.10333"}]},"error":null,"updated_at":"2026-06-26T09:44:52.321480+00:00"},"identity_refresh":{"job_type":"identity_refresh","status":"succeeded","result":{"items":[{"title":"Qwen3 Technical Report","outcome":"unchanged","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","resolver":"local_arxiv","confidence":0.98,"old_work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e"}],"counts":{"fixed":0,"merged":0,"unchanged":1,"quarantined":0,"needs_external_resolution":0},"errors":[],"attempted":1},"error":null,"updated_at":"2026-06-26T09:45:08.709104+00:00"},"role_polarity":{"job_type":"role_polarity","status":"succeeded","result":{"title":"author Dong, W","claims":[{"claim_text":"ratio as baseline first@τdivided by ours first@τ. A speedup ratio greater than 1.0×means ours reaches the same target earlier with fewer epochs or steps. For higher-is-better metrics (Top-1, AP50), first@τis the first epoch with metric at or aboveτ. For lower-is-better metrics (FID), first@τis the first step at or belowτ. Gate and Hyperparameter Selection.For ImageNet classification [7], we useτ= 65for ResNet-50 [14] andτ= 50for ViT-S/16 [8]. For CIFAR early-stage classification [26], we use fix","claim_type":"dataset","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"ϕ(cchild,c parent)< η text(∥˜ cparent∥)·ω(˜ cparent).(8) This allows users to prune entire branches of spurious concepts with a single interaction, substantially reducing the number of interventions required to correct a prediction. 4 Experiments We evaluate HypCBM across three domains:CIFAR-100[ 20] for general object classification, SUN397[ 51] for (hierarchical) scene understanding, andImageNet[ 6] to assess scalability to real- world complexity. Additional results onCUB-200[ 50] are provided","claim_type":"dataset","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"URL https://www.datanami.com/2020/07/06/ data-prep-still-dominates-data-scientists-time-survey-finds/. [8] Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large- scale hierarchical image database. In2009 IEEE Conference on Computer Vision and Pattern Recognition, pages 248-255, 2009. doi: 10.1109/CVPR.2009.5206848. [9] Robert Dorfman. A formula for the gini coefficient.The Review of Economics and Statistics, 61 (1):146-49, 1979. URL https://EconPapers.repec.org","claim_type":"background","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"The inverse-rendering model, invRend-BFM, was trained to infer the BFM generative parameters of a 2D face image, including identity-related shape and texture latents as well as expres- sion, pose, light direction, and light intensity. The object-categorization model, objCat-ImageNet, was trained to classify natural images into ImageNet object categories [46]. Details of the training objective, architectural modi- fications, and training dataset for each model are provided in Methods 4.1. For cop","claim_type":"dataset","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"by shared tasks, common data, and open leaderboards, was the engine behind transformative progress 2 Figure 1:MC 2 pipeline.A low-budget Monte Carlo WoS estimate is corrected in a single forward pass by a learned operator, yielding an improved solution for the PDE. in NLP and computer vision, where benchmarks like GLUE [35], SuperGLUE [36], and ImageNet [8] created a culture of head-to-head comparison on identical inputs. PDE solving has no analog. Existing benchmarks each occupy narrow regimes:","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"Instance Segmentation.Cityscapes [ 6], ADE20K [42], LVIS [12], and Mapillary Vistas [28] cover outdoor driving and general scenes but apply no domain-specific vocabulary tailored to commercial spaces-the escalators, retail shelves, display cases, hotel beds, and food presentations that define the majority of Urban-ImageNet's images. Scaling Behaviour.ImageNet [ 7] established scale as a performance driver; GPT-3 [4] and scaling laws [18] showed predictable growth; LAION-5B [35] demonstrated bill","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks author Dong, W because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (9 contexts).","role_counts":[{"n":9,"context_role":"background"},{"n":4,"context_role":"dataset"},{"n":2,"context_role":"method"}]},"error":null,"updated_at":"2026-06-30T12:20:09.306269+00:00"},"summary_claims":{"job_type":"summary_claims","status":"succeeded","result":{"title":"author Dong, W","claims":[{"claim_text":"ratio as baseline first@τdivided by ours first@τ. A speedup ratio greater than 1.0×means ours reaches the same target earlier with fewer epochs or steps. For higher-is-better metrics (Top-1, AP50), first@τis the first epoch with metric at or aboveτ. For lower-is-better metrics (FID), first@τis the first step at or belowτ. Gate and Hyperparameter Selection.For ImageNet classification [7], we useτ= 65for ResNet-50 [14] andτ= 50for ViT-S/16 [8]. For CIFAR early-stage classification [26], we use fix","claim_type":"dataset","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"ϕ(cchild,c parent)< η text(∥˜ cparent∥)·ω(˜ cparent).(8) This allows users to prune entire branches of spurious concepts with a single interaction, substantially reducing the number of interventions required to correct a prediction. 4 Experiments We evaluate HypCBM across three domains:CIFAR-100[ 20] for general object classification, SUN397[ 51] for (hierarchical) scene understanding, andImageNet[ 6] to assess scalability to real- world complexity. Additional results onCUB-200[ 50] are provided","claim_type":"dataset","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"URL https://www.datanami.com/2020/07/06/ data-prep-still-dominates-data-scientists-time-survey-finds/. [8] Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large- scale hierarchical image database. In2009 IEEE Conference on Computer Vision and Pattern Recognition, pages 248-255, 2009. doi: 10.1109/CVPR.2009.5206848. [9] Robert Dorfman. A formula for the gini coefficient.The Review of Economics and Statistics, 61 (1):146-49, 1979. URL https://EconPapers.repec.org","claim_type":"background","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"The inverse-rendering model, invRend-BFM, was trained to infer the BFM generative parameters of a 2D face image, including identity-related shape and texture latents as well as expres- sion, pose, light direction, and light intensity. The object-categorization model, objCat-ImageNet, was trained to classify natural images into ImageNet object categories [46]. Details of the training objective, architectural modi- fications, and training dataset for each model are provided in Methods 4.1. For cop","claim_type":"dataset","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"by shared tasks, common data, and open leaderboards, was the engine behind transformative progress 2 Figure 1:MC 2 pipeline.A low-budget Monte Carlo WoS estimate is corrected in a single forward pass by a learned operator, yielding an improved solution for the PDE. in NLP and computer vision, where benchmarks like GLUE [35], SuperGLUE [36], and ImageNet [8] created a culture of head-to-head comparison on identical inputs. PDE solving has no analog. Existing benchmarks each occupy narrow regimes:","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"Instance Segmentation.Cityscapes [ 6], ADE20K [42], LVIS [12], and Mapillary Vistas [28] cover outdoor driving and general scenes but apply no domain-specific vocabulary tailored to commercial spaces-the escalators, retail shelves, display cases, hotel beds, and food presentations that define the majority of Urban-ImageNet's images. Scaling Behaviour.ImageNet [ 7] established scale as a performance driver; GPT-3 [4] and scaling laws [18] showed predictable growth; LAION-5B [35] demonstrated bill","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks author Dong, W because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (9 contexts).","role_counts":[{"n":9,"context_role":"background"},{"n":4,"context_role":"dataset"},{"n":2,"context_role":"method"}]},"error":null,"updated_at":"2026-06-26T09:44:57.094675+00:00"}},"summary":{"title":"author Dong, W","claims":[{"claim_text":"ratio as baseline first@τdivided by ours first@τ. A speedup ratio greater than 1.0×means ours reaches the same target earlier with fewer epochs or steps. For higher-is-better metrics (Top-1, AP50), first@τis the first epoch with metric at or aboveτ. For lower-is-better metrics (FID), first@τis the first step at or belowτ. Gate and Hyperparameter Selection.For ImageNet classification [7], we useτ= 65for ResNet-50 [14] andτ= 50for ViT-S/16 [8]. For CIFAR early-stage classification [26], we use fix","claim_type":"dataset","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"ϕ(cchild,c parent)< η text(∥˜ cparent∥)·ω(˜ cparent).(8) This allows users to prune entire branches of spurious concepts with a single interaction, substantially reducing the number of interventions required to correct a prediction. 4 Experiments We evaluate HypCBM across three domains:CIFAR-100[ 20] for general object classification, SUN397[ 51] for (hierarchical) scene understanding, andImageNet[ 6] to assess scalability to real- world complexity. Additional results onCUB-200[ 50] are provided","claim_type":"dataset","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"URL https://www.datanami.com/2020/07/06/ data-prep-still-dominates-data-scientists-time-survey-finds/. [8] Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large- scale hierarchical image database. In2009 IEEE Conference on Computer Vision and Pattern Recognition, pages 248-255, 2009. doi: 10.1109/CVPR.2009.5206848. [9] Robert Dorfman. A formula for the gini coefficient.The Review of Economics and Statistics, 61 (1):146-49, 1979. URL https://EconPapers.repec.org","claim_type":"background","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"The inverse-rendering model, invRend-BFM, was trained to infer the BFM generative parameters of a 2D face image, including identity-related shape and texture latents as well as expres- sion, pose, light direction, and light intensity. The object-categorization model, objCat-ImageNet, was trained to classify natural images into ImageNet object categories [46]. Details of the training objective, architectural modi- fications, and training dataset for each model are provided in Methods 4.1. For cop","claim_type":"dataset","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"by shared tasks, common data, and open leaderboards, was the engine behind transformative progress 2 Figure 1:MC 2 pipeline.A low-budget Monte Carlo WoS estimate is corrected in a single forward pass by a learned operator, yielding an improved solution for the PDE. in NLP and computer vision, where benchmarks like GLUE [35], SuperGLUE [36], and ImageNet [8] created a culture of head-to-head comparison on identical inputs. PDE solving has no analog. Existing benchmarks each occupy narrow regimes:","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"Instance Segmentation.Cityscapes [ 6], ADE20K [42], LVIS [12], and Mapillary Vistas [28] cover outdoor driving and general scenes but apply no domain-specific vocabulary tailored to commercial spaces-the escalators, retail shelves, display cases, hotel beds, and food presentations that define the majority of Urban-ImageNet's images. Scaling Behaviour.ImageNet [ 7] established scale as a performance driver; GPT-3 [4] and scaling laws [18] showed predictable growth; LAION-5B [35] demonstrated bill","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks author Dong, W because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (9 contexts).","role_counts":[{"n":9,"context_role":"background"},{"n":4,"context_role":"dataset"},{"n":2,"context_role":"method"}]},"graph":{"co_cited":[{"title":"Deep residual learning for image recognition","work_id":"b353bda2-591d-479a-9c8b-22dfcba12431","shared_citers":15},{"title":"Walk in the cloud: Learning curves for point clouds shape analysis, pp","work_id":"3820f598-11b0-45c3-8c99-0079181ac0a7","shared_citers":10},{"title":"A ConvNet for the 2020s","work_id":"0a23d1b7-bd56-43cc-8a80-7c43ce994e1e","shared_citers":9},{"title":"Emogen: Emotional image content generation with text-to-image diffusion models","work_id":"7efbc2dd-b0f2-4f71-bb1c-d2fcf110d805","shared_citers":8},{"title":"In: 2023 IEEE/CVF International Conference on Computer Vision (ICCV)","work_id":"b8a8bb9e-1d31-40e2-9cab-ae21e338dde6","shared_citers":8},{"title":"In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp","work_id":"b9701eca-d05e-4d2e-9045-6761df4ba175","shared_citers":8},{"title":"An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale","work_id":"e96730e3-129b-4db6-b981-15ab7932e297","shared_citers":6},{"title":"BERT : Pre-training of Deep Bidirectional Transformers for Language Understanding","work_id":"3e3c8ac8-b858-4b22-af32-393d98c883e0","shared_citers":6},{"title":"Dickerson","work_id":"5c2060c6-427c-4321-be22-49ccae439d80","shared_citers":6},{"title":"In: Proceedings of the IEEE/CVF Conference on Computer 25 Vision and Pattern Recognition, pp","work_id":"9da51225-b7bd-4032-b7db-ca577971dafe","shared_citers":6},{"title":"Derf: Decomposed radiance fields","work_id":"7083a41e-5666-435b-ab26-c753f6490b9a","shared_citers":5},{"title":"DINOv3","work_id":"c8b07deb-8fe7-4e18-9620-f3569d3529ce","shared_citers":5},{"title":"LLaMA: Open and Efficient Foundation Language Models","work_id":"c018fc23-6f3f-4035-9d02-28a2173b2b9d","shared_citers":5},{"title":"PoseNet: A convolutional network for real-time 6-dof camera relocalization","work_id":"135418b1-cafd-49fd-803d-1ca6433d4b1b","shared_citers":5},{"title":"Very Deep Convolutional Networks for Large-Scale Image Recognition","work_id":"1c4b4409-c14b-488b-a086-c57a5aab8a29","shared_citers":5},{"title":"author Han, D","work_id":"b6acb671-02be-4b1f-ad2a-31ae513ee884","shared_citers":4},{"title":"Deep Residual Learning for Image Recognition","work_id":"ae9e5671-23e8-4853-82a4-699b5b8dd639","shared_citers":4},{"title":"Densely connected convolutional networks","work_id":"2199d436-33c2-4b30-9d6f-ce9b8904101e","shared_citers":4},{"title":"DINOv2: Learning Robust Visual Features without Supervision","work_id":"26b304e5-b54a-4f26-be7e-83299eca52e4","shared_citers":4},{"title":"Hierarchical Text-Conditional Image Generation with CLIP Latents","work_id":"0c6a768b-70b8-4242-bb0e-459f1008c9fc","shared_citers":4},{"title":"Masset, R","work_id":"238df2e4-a3e5-46f3-860e-3ae2b0094b97","shared_citers":4},{"title":"Microsoft COCO: common objects in context","work_id":"9a41dfbd-d7c8-4b32-943f-59c08fdd6db7","shared_citers":4},{"title":"The Llama 3 Herd of Models","work_id":"1549a635-88af-4ac1-acfe-51ae7bb53345","shared_citers":4},{"title":"2016.280","work_id":"45b0bfd8-65dc-4252-b2ab-2f6b411d04d0","shared_citers":3}],"time_series":[{"n":2,"year":2019},{"n":1,"year":2022},{"n":7,"year":2024},{"n":15,"year":2025},{"n":50,"year":2026}],"dependency_candidates":[{"n":1,"role":"dataset","polarity":"use_dataset","paper_title":"Human face perception reflects inverse-generative and naturalistic discriminative objectives","primary_cat":"q-bio.NC","context_text":"The inverse-rendering model, invRend-BFM, was trained to infer the BFM generative parameters of a 2D face image, including identity-related shape and texture latents as well as expres- sion, pose, light direction, and light intensity. The object-categorization model, objCat-ImageNet, was trained to classify natural images into ImageNet object categories [46]. Details of the training objective, architectural modi- fications, and training dataset for each model are provided in Methods 4.1. For copyright reasons, the VGGFace2 [43] images in this figure were replaced with synthetic faces generated using https://thispersondoesnotexist.com/; the ImageNet [46] illustration was obtained from https://commons.wikimedia.","citing_arxiv_id":"2605.12619"},{"n":1,"role":"dataset","polarity":"use_dataset","paper_title":"Removing the Watermark Is Not Enough: Forensic Stealth in Generative-AI Watermark Removal","primary_cat":"cs.CR","context_text":"A separate foren- sic literature studies the traces left by learned image generators and manipulators, and provides much of the methodological foundation we build on. Wang et al. [ 37] show that CNN-generated images contain characteristic fingerprints that simple detectors can exploit, with transfer across unseen generators. Frank et al. [12] and Durall et al. [8] connect these traces to spectral irregularities introduced by upsampling. Marra et al. [ 28] show that GANs leave model- specific fingerprints analogous to camera sensor noise, and Ojha et al. [31] leverage pretrained representations for more universal fake detection across architectures. This literature shows that learned image pipelines leave statistical traces.","citing_arxiv_id":"2605.09203"},{"n":1,"role":"dataset","polarity":"use_dataset","paper_title":"Hyperbolic Concept Bottleneck Models","primary_cat":"cs.LG","context_text":"ϕ(cchild,c parent)< η text(∥˜ cparent∥)·ω(˜ cparent).(8) This allows users to prune entire branches of spurious concepts with a single interaction, substantially reducing the number of interventions required to correct a prediction. 4 Experiments We evaluate HypCBM across three domains:CIFAR-100[ 20] for general object classification, SUN397[ 51] for (hierarchical) scene understanding, andImageNet[ 6] to assess scalability to real- world complexity. Additional results onCUB-200[ 50] are provided in Appendix A. To isolate the impact of geometry, we control for training data by utilizing the pre-trained checkpoints from Pal et al. [33], ensuring both the hyperbolic HyCoCLIP and Euclidean CLIP [38] are trained on the same GrIT- 20M dataset [35]. We also report results for CLIP [38], trained on 400M samples, to contextualize","citing_arxiv_id":"2605.06440"},{"n":1,"role":"method","polarity":"use_method","paper_title":"StomaD2: An All-in-One System for Intelligent Stomatal Phenotype Analysis via Diffusion-Based Restoration Detection Network","primary_cat":"cs.CV","context_text":"Stage Two: Although the regression stage can recover the major structure of the stoma, residual degradation may still remain, especially in regions with severe blur or noise contamination. Therefore, a diffusion-based refinement stage is further introduced to suppress residual degradation while preserving local stomatal structures. We use the fine-tuned Stable Diffusion [26] to reconstruct the image with the obtained (𝐼𝐼𝑟𝑟𝑟𝑟𝑟𝑟− 𝐼𝐼𝐻𝐻𝐿𝐿) . The VAE encoder [27] maps (𝐼𝐼𝑟𝑟𝑟𝑟𝑟𝑟) into the latent space, resulting in the conditional latent (ε�𝐼𝐼𝑟𝑟𝑟𝑟𝑟𝑟�). Objective function is as follow: ℒ𝐷𝐷𝐷𝐷𝐷𝐷𝐷𝐷= 𝔼𝔼𝑧𝑧𝑡𝑡,𝑐𝑐,𝑡𝑡,𝜀𝜀,𝜀𝜀�𝐼𝐼𝑟𝑟𝑟𝑟𝑟𝑟� ��𝜀𝜀− 𝜀𝜀𝜃𝜃�𝑧𝑧𝑡𝑡, 𝑐𝑐, 𝑡𝑡, 𝜀𝜀, 𝜀𝜀�𝐼𝐼𝑟𝑟𝑟𝑟𝑟𝑟��� 2 2 �, (2) where ε ∼ 𝒩𝒩(0, 𝐼𝐼), 𝑧𝑧𝑡𝑡 is the latent code, 𝑐𝑐 is the sample point, and 𝑡𝑡 represents the time step, which is","citing_arxiv_id":"2604.18632"},{"n":1,"role":"dataset","polarity":"use_dataset","paper_title":"Weak-to-Strong Knowledge Distillation Accelerates Visual Learning","primary_cat":"cs.CV","context_text":"ratio as baseline first@τdivided by ours first@τ. A speedup ratio greater than 1.0×means ours reaches the same target earlier with fewer epochs or steps. For higher-is-better metrics (Top-1, AP50), first@τis the first epoch with metric at or aboveτ. For lower-is-better metrics (FID), first@τis the first step at or belowτ. Gate and Hyperparameter Selection.For ImageNet classification [7], we useτ= 65for ResNet-50 [14] andτ= 50for ViT-S/16 [8]. For CIFAR early-stage classification [26], we use fixed dataset-level gates:τ= 75for CIFAR-10 and τ= 60for CIFAR-100. For object detection on the COCO dataset [28], we use a fixed task-level AP50 target (τ= 20%). For diffusion generation on the CIFAR-10 dataset [18,26,32], we use a fixed task-level FID target (τ= 60),","citing_arxiv_id":"2604.15451"},{"n":1,"role":"method","polarity":"use_method","paper_title":"Zero-shot World Models Are Developmentally Efficient Learners","primary_cat":"cs.AI","context_text":"task-performant models for visual tasks as well as the most accurate models of neural responses across the visual cortex [9, 10, 11, 12, 13] and for human-like error patterns [14]. Initially, DNNs required supervision on large labeled datasets [15, 16] and did not transfer broadly to downstream tasks. These issues motivated a shift to self-supervised models, which learn representations by grouping similar or temporally-proximate images [ 17, 18, 19, 20, 21, 22, 23, 24]. Such self- supervised models also turn out to provide a description of neural and cognitive patterns in the primate visual system, rivaling or exceeding the accuracy of the earlier supervised systems [25, 26]. However, from a developmental point of view, modern self-supervised learning methods are a \"glass half full\". When trained on the visual data diet of real infants and children, they achieve","citing_arxiv_id":"2604.10333"}]},"authors":[{"id":"178ae3c3-0ee2-4711-b5da-46f62752e506","orcid":null,"display_name":"bchapter Deng","source":"manual","import_confidence":0.72}]}}