A multiple-instance densely-connected ConvNet for aerial scene classification

In contrast with nature scenes, aerial scenes are often composed of many objects crowdedly distributed on the surface in bird’s view, the description of which usually demands more discriminative features as well as local semantics. However, when applied to scene classification, most of the existing convolution neural networks (ConvNets) tend to depict global semantics of images, and the loss of low- and mid-level features can hardly be avoided, especially when the model goes deeper. To tackle these challenges, in this paper, we propose a multiple-instance densely-connected ConvNet (MIDC-Net) for aerial scene classification. It regards aerial scene classification as a multiple-instance learning problem so that local semantics can be further investigated. Our classification model consists of an instance-level classifier, a multiple instance pooling and followed by a bag-level classification layer. In the instance-level classifier, we propose a simplified dense connection structure to effectively preserve features from different levels. The extracted convolution features are further converted into instance feature vectors. Then, we propose a trainable attention-based multiple instance pooling. It highlights the local semantics relevant to the scene label and outputs the bag-level probability directly. Finally, with our bag-level classification layer, this multiple instance learning framework is under the direct supervision of bag labels. Experiments on three widely-utilized aerial scene benchmarks demonstrate that our proposed method outperforms many state-of-the-art methods by a large margin with much fewer parameters.

PDF

Results from the Paper


Task Dataset Model Metric Name Metric Value Global Rank Benchmark
Scene Recognition AID MIDC-Net Accuracy 92.95 # 3
Aerial Scene Classification AID (20% as trainset) MIDC-Net Accuracy 88.51 # 10
Aerial Scene Classification AID (50% as trainset) MIDC-Net Accuracy 92.95 # 9
Aerial Scene Classification NWPU (10% as trainset) MIDC-Net Accuracy 86.12 # 9
Aerial Scene Classification NWPU (20% as trainset) MIDC-Net Accuracy 87.99 # 13
Image Classification RESISC45 MIDC-Net Top 1 Accuracy 87.99 # 13
Aerial Scene Classification UCM (50% as trainset) MIDC-Net Accuracy 95.41 # 6
Aerial Scene Classification UCM (80% as trainset) MIDC-Net Accuracy 97.40 # 7
Scene Classification UC Merced Land Use Dataset MIDC-Net Accuracy (%) 97.40 # 6

Methods


No methods listed for this paper. Add relevant methods here