REVIEW 4 major objections 5 minor 2 cited by
Towards RAW Object Detection in Diverse Conditions
T0 review · 4 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read Detectors pre-trained on synthetic RAW sensor data outperform sRGB-pretrained models on a new 62-category adverse-weather benchmark, with the largest gains in rain.
desk verdict Solid new RAW detection benchmark; the pre-training gain is confounded with brightness/noise augmentation, so the causal claim needs a control. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is a paired recipe: synthetic ImageNet-RAW pre-training plus cross-domain distillation. Synthetic ImageNet-RAW is produced by the unprocessing method of [2], which reverses an image signal processor to convert sRGB images back to 16-bit RAW-like data and simulates camera noise; because the unprocessing runs inside the data-augmentation pipeline, brightness and noise are randomized each iteration. Cross-domain distillation then trains the RAW-pretrained student with logit-based Kullback-Leibler divergence and feature-based L1 loss against an off-the-shelf sRGB-pretrained teacher of the same architecture, giving the student stable semantic targets that do not vary with the synthesized noise. Together these two pieces let the backbone learn representations that are invariant to brightness and noise before it is fine-tuned on real RAW detection data.
What would settle it
Train the identical detector with sRGB ImageNet pre-training while applying the exact same random brightness and noise augmentation that ImageNet-RAW uses, then fine-tune on AODRaw RAW images; if that control reaches or exceeds 34.8% AP, the claimed RAW-domain advantage is an artifact of the augmentation. A complementary check is to pre-train on a comparably sized set of real RAW images and see whether the gap over sRGB pre-training persists.
Extended reading notes
Core claim
On its own terms, the paper claims two things. First, the domain gap between sRGB pre-training and RAW fine-tuning is real and costly: a Cascade R-CNN trained on sRGB drops from 34.0% to 28.0% AP when evaluated on RAW, and models pre-trained on sRGB then fine-tuned on RAW underperform models trained and tested entirely in sRGB. Second, pre-training directly in the RAW domain closes most of that gap. Because no large real RAW pre-training set exists, the authors synthesize ImageNet-RAW by unprocessing ImageNet-1K images, inverting the camera pipeline and adding random brightness and shot noise inside the augmentation loop. To help the student cope with noise, they distill logit and feature knowledge from an off-the-shelf sRGB-pretrained teacher. The result is 34.8% AP on AODRaw with Cascade R-CNN and ConvNeXt-T, 1.1 points above sRGB pre-training and 0.8 points above sRGB-based detection, with a 4.8-point gain in rain, achieved without any neural ISP.
Load-bearing premise
The load-bearing premise is that the synthetic ImageNet-RAW images created by unprocessing, with random brightness and simulated shot noise, faithfully represent real camera RAW data well enough that pre-training on them transfers to real AODRaw images; the paper does not include an sRGB pre-training control with the same augmentation schedule, so if that premise fails the reported gains could be an augmentation effect rather than a RAW-domain effect.
Editorial extensions
If this is right
- A single RAW-pretrained model can serve all nine light and weather conditions at once, eliminating the need for per-condition models or trainable ISP adapters.
- The benefit of RAW pre-training is concentrated where it matters most: +0.8 AP in low light, +4.8 AP in rain, and +1.2 AP in fog, compared with +0.9 AP in normal conditions.
- The gains are not tied to one architecture: they appear with Cascade R-CNN and ConvNeXt-T in both down-sampled and sliced-image settings, and the raw data supports real-time YOLO-scale detectors without destroying frame rate.
- Distillation makes the pre-trained representation measurably more robust to brightness and noise shifts, so downstream fine-tuning starts from a sturdier feature space.
Reading between the lines
- An implication the paper leaves implicit is that the same unprocessing-plus-distillation recipe should transfer to other sensor modalities that lack large labeled datasets, such as multispectral, high-dynamic-range, or polarization imaging; the sRGB teacher supplies semantics while the synthetic sensor data supplies the input distribution.
- The paper does not report an sRGB pre-training control that applies the same random brightness and noise augmentation to ordinary ImageNet-1K; adding that control would decisively separate the benefit of the RAW input distribution from the benefit of the augmentation schedule.
- Because the largest gain appears in rain, a natural next experiment is to push the simulated noise and brightness ranges further and test whether detector AP keeps climbing under increasingly extreme synthetic conditions or saturates.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces AODRaw, a RAW-image object detection dataset with 7,785 high-resolution images, 135,601 annotated instances across 62 categories, and 9 combinations of light and weather conditions. It benchmarks multiple detection architectures under sRGB and RAW inputs, and proposes pre-training detectors on synthetic RAW images generated from ImageNet via the unprocessing method, augmented with cross-domain knowledge distillation from an off-the-shelf sRGB pre-trained teacher. The central claim is that RAW pre-training improves detection on real RAW images, especially under adverse conditions, without needing neural ISP adapters; the headline result is 34.8% AP for Cascade RCNN/ConvNeXt-T on RAW, versus 33.7% AP with sRGB pre-training.
Significance. The AODRaw dataset is a valuable new resource: it is substantially larger and more diverse than existing RAW detection datasets, with 62 categories and nine condition combinations, and the paper provides a careful statistical analysis of the data. The benchmark results across diverse backbones and detector families are extensive and internally consistent. The proposed pre-training recipe is practical and removes the need for learnable ISP modules at inference time. However, the paper's central causal claim—that the RAW input domain, rather than the accompanying brightness/noise augmentation or the distillation procedure, is responsible for the observed gains—is not yet established by the experiments. The dataset and benchmark alone would support a useful paper; the method claim requires additional controlled comparisons.
major comments (4)
- [Section 5.1, 5.2; Table 3] The comparison between RAW pre-training and sRGB pre-training is confounded by the data augmentation schedule. Section 5.1 states that in ImageNet-RAW synthesis, 'the unprocessed operation is inserted into the pipeline of data augmentations. Thus, we can randomly adjust the average brightness and simulated noise in each iteration.' The sRGB pre-training baseline in Table 3 (rows 'sRGB RAW') uses standard ImageNet augmentation without such brightness/noise randomization. Since Figures 5 and 6 show sRGB pre-trained models degrade by 14.0% and 13.1% under these shifts, a robustness gain from the augmentation alone is plausible. The paper needs a control that pre-trains on sRGB ImageNet with the same random brightness/noise augmentation (e.g., applied to the sRGB images) to attribute the reported 1.1% AP improvement to the RAW domain rather than to the augmentation schedule.
- [Section 5.2; Table 8] The headline result is also confounded by the cross-domain distillation. The 'RAW RAW' rows in Table 3 correspond to the full proposed method that includes distillation, while the 'sRGB RAW' rows use plain sRGB pre-training without distillation. Table 8 shows that distillation contributes 0.5% AP (34.1% without distillation vs. 34.8% with distillation) within RAW pre-training. To isolate the effect of the RAW domain, the paper should report: (a) RAW pre-training without distillation versus sRGB pre-training with the same augmentation, and (b) sRGB pre-training with the same augmentation and with an analogous distillation from the same teacher. Without these controls, the 1.1% AP gain cannot be attributed to the RAW domain.
- [Section 4.1 vs. Section 5.1] There is a mismatch between the pre-training and fine-tuning input domains. Section 4.1 says that real RAW images are 'further processed through gamma correction' before fine-tuning, while Section 5.1 describes synthesizing ImageNet-RAW using the unprocessing method, which produces linear RAW images with simulated noise. The paper does not state whether the synthetic pre-training data is gamma-corrected or linear. If the pre-training uses linear RAW and the fine-tuning uses gamma-corrected RAW, then the two domains differ by more than just the RAW-vs-sRGB distinction, and the term 'RAW pre-training' is ambiguous. The authors should either apply the same gamma correction to the synthetic pre-training data or explicitly justify the mismatch and demonstrate that it does not affect the main comparison.
- [Section 3.1] The dataset annotation section provides no information about annotator training, quality control, or inter-annotator agreement. Since AODRaw is a core contribution and is used to evaluate the pre-training method, the reliability of the annotations is essential. The authors should report the annotation protocol, the number of annotators, and a quality metric such as inter-annotator agreement or a manual verification rate on a subset, so that the benchmark results can be trusted as ground truth.
minor comments (5)
- [Throughout] The paper contains a recurring typo where 'RAW' appears as 'RA W' (e.g., in the title and several headings); this should be fixed.
- [Section 5.1; Figures 5 and 6] The text says 'when reducing the brightness from 791 to 80', but the caption of Figure 5 states the maximum average brightness is 216. These numbers are inconsistent; please clarify the brightness scale and correct the text or figure.
- [Table 5] The 'Baseline' row reports 33.4% AP for ConvNeXt-T with sRGB pre-training and RAW fine-tuning, while the analogous row in Table 3 (Cascade RCNN/ConvNeXt-T, 'sRGB RAW') reports 33.7% AP. The discrepancy should be explained, since the reader cannot tell whether the baseline in Table 5 uses a different training recipe or a different definition.
- [Section 4.1] The text says 'too tiny objects with an area of less than 32^2 are ignored because they will disappear after down-sampling,' but the object-area thresholds for down-sampling are defined as [0,128^2), [128^2,320^2), and [320^2,+∞). The 32^2 cutoff appears inconsistent with these ranges; please verify the intended threshold.
- [Supplementary, Eq. (1)] The KL divergence formula in the supplementary material is typeset as 'yt log yt / ys', which is missing parentheses and the summation over classes; it should be written as sum_i y_t(i) log(y_t(i)/y_s(i)).
Circularity Check
No circular derivation chain: the RAW pre-training gains are empirical measurements, and the only self-citation (YOLO-MS) is a non-load-bearing baseline reference.
full rationale
The paper contains no formal derivation chain whose output is equivalent to an input. The central quantitative claims—RAW pre-training improving Cascade R-CNN + ConvNeXt-T from 34.0/33.7 to 34.8 AP on AODRaw, and the per-condition gains (0.8 APlow, 4.8 APrain, 1.2 APfog)—are measured results on a fixed test set, not constants fitted to that set. The pipeline in Section 5.1 synthesizes ImageNet-RAW using unprocessing [2] and distills from an off-the-shelf sRGB teacher; neither step defines the reported AP in terms of the method's own choices. The dataset is constructed and annotated by the authors and used both to motivate and to evaluate the approach, but that is a benchmark-creation pattern, not a circular reduction. The only author-overlap citation is [6] (YOLO-MS, including author M.-M. Cheng), used as a real-time baseline; it does not carry a load-bearing premise for the RAW-pretraining claim. The main weakness is experimental-design-related: the sRGB pre-training control does not include the same random brightness/noise augmentation as ImageNet-RAW (Section 5.1), so part of the 1.1% AP gain could in principle be augmentation robustness; however, this is a confounding-variable concern, not a circularity. I therefore assign 2 for the benign self-citation, with no circular steps identified.
Assumptions & free parameters
free parameters (3)
- Random brightness and noise augmentation ranges in ImageNet-RAW synthesis
- Down-sampling resolution and slicing settings =
2000x1333; 1280x1280 patches, overlap 300, IoU threshold 0.4
- Small, medium, and large object area thresholds =
[0,128^2), [128^2,320^2), [320^2,...) for downsampling; [0,64^2), [64^2,160^2), [160^2,...) for slicing
assumptions (4)
- domain assumption RAW images preserve information that is lost in sRGB conversion and useful for detection under adverse conditions.
- domain assumption Synthetic RAW from unprocessing [2] with simulated noise and brightness variation is a sufficient proxy for real RAW in pre-training.
- domain assumption MMDetection implementations and the chosen hyperparameters for all detectors are correct and comparable.
- domain assumption The AODRaw annotations are complete and consistent across 62 categories.
Cite this review
Pith. "Pith review of Towards RAW Object Detection in Diverse Conditions." pith.science (2026). https://pith.science/paper/HPW4T26G
@misc{pith2026241115678,
author = {Pith},
title = {Pith review of: Towards RAW Object Detection in Diverse Conditions},
year = {2026},
howpublished = {\url{https://pith.science/paper/HPW4T26G}},
note = {Machine review of arXiv:2411.15678}
}
read the original abstract
Existing object detection methods often consider sRGB input, which was compressed from RAW data using ISP originally designed for visualization. However, such compression might lose crucial information for detection, especially under complex light and weather conditions. We introduce the AODRaw dataset, which offers 7,785 high-resolution real RAW images with 135,601 annotated instances spanning 62 categories, capturing a broad range of indoor and outdoor scenes under 9 distinct light and weather conditions. Based on AODRaw that supports RAW and sRGB object detection, we provide a comprehensive benchmark for evaluating current detection methods. We find that sRGB pre-training constrains the potential of RAW object detection due to the domain gap between sRGB and RAW, prompting us to directly pre-train on the RAW domain. However, it is harder for RAW pre-training to learn rich representations than sRGB pre-training due to the camera noise. To assist RAW pre-training, we distill the knowledge from an off-the-shelf model pre-trained on the sRGB domain. As a result, we achieve substantial improvements under diverse and adverse conditions without relying on extra pre-processing modules. Code and dataset are available at https://github.com/lzyhha/AODRaw.
Figures
Figures from the paper (7 more)
Forward citations
Cited by 2 Pith papers
-
UNICE: Training A Universal Image Contrast Enhancer
UNICE trains a two-stage model to generate and fuse a pseudo multi-exposure sequence from one image, generalizing across four contrast-enhancement tasks without human labels.
-
Depth Anything at Any Condition
A fine-tuned Depth Anything V2 model using perturbation consistency and spatial distance constraints improves monocular depth estimation under adverse conditions without any labeled data.
Reference graph
Works this paper leans on
-
[1]
Seeing through fog without seeing fog: Deep multimodal sensor fu- sion in unseen adverse weather
Mario Bijelic, Tobias Gruber, Fahim Mannan, Florian Kraus, Werner Ritter, Klaus Dietmayer, and Felix Heide. Seeing through fog without seeing fog: Deep multimodal sensor fu- sion in unseen adverse weather. In CVPR, 2020. 3
work page 2020
-
[2]
Unprocessing im- ages for learned raw denoising
Tim Brooks, Ben Mildenhall, Tianfan Xue, Jiawen Chen, Dillon Sharlet, and Jonathan T Barron. Unprocessing im- ages for learned raw denoising. In CVPR, 2019. 6
work page 2019
-
[3]
Cascade r-cnn: High quality object detection and instance segmentation
Zhaowei Cai and Nuno Vasconcelos. Cascade r-cnn: High quality object detection and instance segmentation. IEEE TPAMI, 2019. 2, 5, 6
work page 2019
-
[4]
MMDetection: Open mmlab detection toolbox and benchmark
Kai Chen, Jiaqi Wang, Jiangmiao Pang, Yuhang Cao, Yu Xiong, Xiaoxiao Li, Shuyang Sun, Wansen Feng, Ziwei Liu, Jiarui Xu, Zheng Zhang, Dazhi Cheng, Chenchen Zhu, Tian- heng Cheng, Qijie Zhao, Buyu Li, Xin Lu, Rui Zhu, Yue Wu, Jifeng Dai, Jingdong Wang, Jianping Shi, Wanli Ouyang, Chen Change Loy, and Dahua Lin. MMDetection: Open mmlab detection toolbox and...
arXiv 1906
-
[5]
Instance segmentation in the dark
Linwei Chen, Ying Fu, Kaixuan Wei, Dezhi Zheng, and Fe- lix Heide. Instance segmentation in the dark. International Journal of Computer Vision, 131(8):2198–2218, 2023. 2
work page 2023
-
[6]
Yolo-ms: Rethinking multi- scale representation learning for real-time object detection
Yuming Chen, Xinbin Yuan, Ruiqi Wu, Jiabao Wang, Qibin Hou, and Ming-Ming Cheng. Yolo-ms: Rethinking multi- scale representation learning for real-time object detection. arXiv preprint arXiv:2308.05480, 2023. 8
arXiv 2023
-
[7]
Towards large-scale small object detection: Survey and benchmarks
Gong Cheng, Xiang Yuan, Xiwen Yao, Kebing Yan, Qinghua Zeng, Xingxing Xie, and Junwei Han. Towards large-scale small object detection: Survey and benchmarks. IEEE TPAMI, 45(11):13467–13488, 2023. 5
work page 2023
-
[8]
Raw-adapter: Adapting pre- trained visual model to camera raw images
Ziteng Cui and Tatsuya Harada. Raw-adapter: Adapting pre- trained visual model to camera raw images. In ECCV, 2024. 1, 2, 3, 6, 7
work page 2024
Show all 42 references
-
[9]
Dirty pixels: Towards end-to-end image processing and percep- tion
Steven Diamond, Vincent Sitzmann, Frank Julca-Aguilar, Stephen Boyd, Gordon Wetzstein, and Felix Heide. Dirty pixels: Towards end-to-end image processing and percep- tion. ACM Trans. Graph., 40(3), 2021. 2
2021
-
[10]
Everingham, L
M. Everingham, L. Gool, Christopher K. I. Williams, J. Winn, and Andrew Zisserman. The pascal visual object classes (voc) challenge. IJCV, 88:303–338, 2009. 1, 3
2009
-
[11]
YOLOX: Exceeding yolo series in 2021.arXiv preprint arXiv:2107.08430, 2021
Zheng Ge, Songtao Liu, Feng Wang, Zeming Li, and Jian Sun. YOLOX: Exceeding yolo series in 2021.arXiv preprint arXiv:2107.08430, 2021. 2, 8
2021 arXiv
-
[12]
Gamma cor- rection for digital fringe projection profilometry
Hongwei Guo, Haitao He, and Mingyi Chen. Gamma cor- rection for digital fringe projection profilometry. Appl. Opt., 43(14):2906–2914, 2004. 7
2004
-
[13]
Learn- ing degradation-independent representations for camera isp pipelines
Yanhui Guo, Fangzhou Luo, and Xiaolin Wu. Learn- ing degradation-independent representations for camera isp pipelines. In CVPR, 2024. 2
2024
-
[14]
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR,
-
[15]
Masked autoencoders are scalable vision learners
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Doll´ar, and Ross Girshick. Masked autoencoders are scalable vision learners. In CVPR, 2022. 6
2022
-
[16]
Distilling the knowledge in a neural net- work
Geoffrey Hinton. Distilling the knowledge in a neural net- work. arXiv preprint arXiv:1503.02531, 2015. 7
2015 arXiv
-
[17]
Craft- ing object detection in very low light
Yang Hong, Kaixuan Wei, Linwei Chen, and Ying Fu. Craft- ing object detection in very low light. In BMVC, 2021. 1, 3, 4
2021
-
[18]
YOLO by Ultralytics, 2023
Glenn Jocher, Ayush Chaurasia, and Jing Qiu. YOLO by Ultralytics, 2023. 8
2023
-
[19]
Generalized focal loss: Learning qualified and distributed bounding boxes for dense object detection
Xiang Li, Wenhai Wang, Lijun Wu, Shuo Chen, Xiaolin Hu, Jun Li, Jinhui Tang, and Jian Yang. Generalized focal loss: Learning qualified and distributed bounding boxes for dense object detection. In NeurIPS, 2020. 5, 6
2020
-
[20]
Salman Asif, and Zhan Ma
Zhihao Li, Ming Lu, Xu Zhang, Xin Feng, M. Salman Asif, and Zhan Ma. Efficient visual computing with camera raw snapshots. IEEE TPAMI, 46(7):4684–4701, 2024. 2, 4
2024
-
[21]
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll´ar, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In ECCV, 2014. 1, 2, 3, 5
2014
-
[22]
Focal loss for dense object detection
Tsung-Yi Lin, Priya Goyal, Ross Girshick, Kaiming He, and Piotr Doll´ar. Focal loss for dense object detection. In ICCV,
-
[23]
Swin transformer: Hierarchical vision transformer using shifted windows
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. In ICCV, 2021. 6
2021
-
[24]
A convnet for the 2020s
Zhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feicht- enhofer, Trevor Darrell, and Saining Xie. A convnet for the 2020s. In CVPR, 2022. 2, 6, 7
2022
-
[25]
Hardware-in- the-loop end-to-end optimization of camera image process- ing pipelines
Ali Mosleh, Avinash Sharma, Emmanuel Onzon, Fahim Mannan, Nicolas Robidoux, and Felix Heide. Hardware-in- the-loop end-to-end optimization of camera image process- ing pipelines. In CVPR, 2020. 1, 2, 3
2020
-
[26]
Pas- calraw: raw image database for object detection
Alex Omid-Zohoor, David Ta, and Boris Murmann. Pas- calraw: raw image database for object detection. Stanford Digital Repository, 2014. 1, 3, 4
2014
-
[27]
Attention-aware learning for hyperparameter prediction in image processing pipelines
Haina Qin, Longfei Han, Juan Wang, Congxuan Zhang, Yan- wei Li, Bing Li, and Weiming Hu. Attention-aware learning for hyperparameter prediction in image processing pipelines. In ECCV, 2022. 2
2022
-
[28]
Faster r-cnn: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. Faster r-cnn: Towards real-time object detection with region proposal networks. IEEE TPAMI, 2017. 2, 5, 6
2017
-
[29]
Imagenet large scale visual recognition challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, San- jeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al. Imagenet large scale visual recognition challenge. IJCV, 2015. 6
2015
-
[30]
Sparse r-cnn: End-to-end object detection with learnable proposals
Peize Sun, Rufeng Zhang, Yi Jiang, Tao Kong, Chenfeng Xu, Wei Zhan, Masayoshi Tomizuka, Lei Li, Zehuan Yuan, Changhu Wang, and Ping Luo. Sparse r-cnn: End-to-end object detection with learnable proposals. In CVPR, 2021. 2, 5, 6
2021
-
[31]
Adaptiveisp: Learning an adaptive image signal proces- sor for object detection
Yujin Wang, Tianyi Xu, Fan Zhang, Tianfan Xue, and Jinwei Gu. Adaptiveisp: Learning an adaptive image signal proces- sor for object detection. In NeurIPS, 2024. 2
2024
-
[32]
Contrastive learning rivals masked image modeling in fine-tuning via feature distillation
Yixuan Wei, Han Hu, Zhenda Xie, Zheng Zhang, Yue Cao, Jianmin Bao, Dong Chen, and Baining Guo. Contrastive learning rivals masked image modeling in fine-tuning via feature distillation. arXiv preprint arXiv:2205.14141, 2022. 7 9
2022 arXiv
-
[33]
Isikdogan, Sushma Rao, Bhavin Nayak, Timo Gerasimow, Aleksandar Sutic, Liron Ain- kedem, and Gilad Michael
Chyuan-Tyng Wu, Leo F. Isikdogan, Sushma Rao, Bhavin Nayak, Timo Gerasimow, Aleksandar Sutic, Liron Ain- kedem, and Gilad Michael. Visionisp: Repurposing the im- age signal processor for computer vision applications. In ICIP, 2019. 2
2019
-
[34]
Toward raw object detection: A new benchmark and a new model
Ruikang Xu, Chang Chen, Jingyang Peng, Cheng Li, Yibin Huang, Fenglong Song, Youliang Yan, and Zhiwei Xiong. Toward raw object detection: A new benchmark and a new model. In CVPR, 2023. 1, 2, 3, 4, 6, 7, 8
2023
-
[35]
Dynamicisp: Dynamically controlled im- age signal processor for image recognition
Masakazu Yoshimura, Junji Otsuka, Atsushi Irie, and Takeshi Ohashi. Dynamicisp: Dynamically controlled im- age signal processor for image recognition. In ICCV, 2023. 2
2023
-
[36]
Reconfigisp: Reconfigurable camera image processing pipeline
Ke Yu, Zexian Li, Yue Peng, Chen Change Loy, and Jinwei Gu. Reconfigisp: Reconfigurable camera image processing pipeline. In ICCV, 2021. 2, 4
2021
-
[37]
Darkvision: a bench- mark for low-light image/video perception
Bo Zhang, Yuchen Guo, Runzhao Yang, Zhihong Zhang, Ji- ayi Xie, Jinli Suo, and Qionghai Dai. Darkvision: a bench- mark for low-light image/video perception. arXiv preprint arXiv:2301.06269, 2023. 3
2023 arXiv
-
[38]
Deformable detr: Deformable transformers for end-to-end object detection
Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang, and Jifeng Dai. Deformable detr: Deformable transformers for end-to-end object detection. In ICLR, 2021. 2, 4, 5, 6 10 Towards RA W Object Detection in Diverse Conditions Supplementary Material
2021
-
[39]
We collect images under 9 conditions, as shown in Tab
Statistics of AODRaw More examples from AODRaw. We collect images under 9 conditions, as shown in Tab. 2 of the main paper. Tab. 9 shows a specific example for each condition for a bet- ter understanding. Furthermore, Fig. 7 shows more exam- ples of our AODRaw dataset and the ...
-
[41]
For data augmentations, the images are re- sized between 800 and 1024 along the shorter side, while the longer side is no larger than 2048
Experiments Settings Most hyperparameters follow the COCO dataset in the mmdetection. For data augmentations, the images are re- sized between 800 and 1024 along the shorter side, while the longer side is no larger than 2048. And the RandomFlip is used to augment images. For d...
-
[42]
Besides the supervised classification loss func- tion, we use logit-based and feature-based distillation for cross-domain distillation
Distillation Implementation Method. Besides the supervised classification loss func- tion, we use logit-based and feature-based distillation for cross-domain distillation. For logit-based distillation, we denote zs and zt as the output of student and teacher, re- spectively. z...
1993
-
[100]
For cases exceeding 100, since there are fewer images in this range, there is some deviation between several con- ditions and the whole, e.g., the condition of low-light and fog in outdoor scenes, as shown in Fig. 10h. Bounding box size. Fig. 11 shows the distribution of the b...
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.