Res2Net: A New Multi-scale Backbone Architecture

2 Apr 2019  ·  Shang-Hua Gao, Ming-Ming Cheng, Kai Zhao, Xin-Yu Zhang, Ming-Hsuan Yang, Philip Torr ·

Representing features at multiple scales is of great importance for numerous vision tasks. Recent advances in backbone convolutional neural networks (CNNs) continually demonstrate stronger multi-scale representation ability, leading to consistent performance gains on a wide range of applications. However, most existing methods represent the multi-scale features in a layer-wise manner. In this paper, we propose a novel building block for CNNs, namely Res2Net, by constructing hierarchical residual-like connections within one single residual block. The Res2Net represents multi-scale features at a granular level and increases the range of receptive fields for each network layer. The proposed Res2Net block can be plugged into the state-of-the-art backbone CNN models, e.g., ResNet, ResNeXt, and DLA. We evaluate the Res2Net block on all these models and demonstrate consistent performance gains over baseline models on widely-used datasets, e.g., CIFAR-100 and ImageNet. Further ablation studies and experimental results on representative computer vision tasks, i.e., object detection, class activation mapping, and salient object detection, further verify the superiority of the Res2Net over the state-of-the-art baseline methods. The source code and trained models are available on https://mmcheng.net/res2net/.

PDF Abstract
Task Dataset Model Metric Name Metric Value Global Rank Result Benchmark
Image Classification CIFAR-100 Res2NeXt-29 Percentage correct 83.44 # 73
Instance Segmentation COCO minival Res2Net-101+HTC mask AP 41.3 # 39
Object Detection COCO minival Res2Net101+HTC box AP 47.5 # 52
AP50 66.5 # 24
AP75 51.3 # 20
APS 28.6 # 19
APM 51.6 # 15
APL 62.1 # 17
Instance Segmentation COCO minival Faster R-CNN (Res2Net-50) mask AP 35.6 # 62
AP50 57.6 # 13
APL 53.7 # 6
APM 37.9 # 13
APS 15.7 # 13
Object Detection COCO minival Faster R-CNN (Res2Net-50) box AP 33.7 # 158
AP50 53.6 # 93
APS 14 # 78
APM 38.3 # 76
APL 51.1 # 64
RGB Salient Object Detection DUT-OMRON DSS (Res2Net-50) MAE 0.071 # 9
F-measure 0.800 # 3
RGB Salient Object Detection ECSSD DSS (Res2Net-50) MAE 0.056 # 8
F-measure 0.926 # 3
RGB Salient Object Detection HKU-IS DSS (Res2Net-50) MAE 0.05 # 9
F-measure 0.905 # 3
Image Classification ImageNet Res2Net-101 Top 1 Accuracy 81.23% # 339
Top 5 Accuracy 94.43% # 143
Image Classification ImageNet Res2Net-50-299 Top 1 Accuracy 78.59% # 445
Top 5 Accuracy 94.12% # 151
RGB Salient Object Detection PASCAL-S DSS (Res2Net-50) MAE 0.099 # 7
F-measure 0.841 # 3
Semantic Segmentation PASCAL VOC 2012 val Deeplab v3+ (Res2Net-101) mIoU 79.3% # 11

Methods