Volume 10,Issue 7
In autonomous driving perception, point cloud-based 3D object detection plays an important role. This task still faces two challenges in long-range and small-object detection: loss of fine details and weak context modeling. To solve these problems, this paper proposes HFA-RCNN based on PV-RCNN. The method adds an encoder-decoder structure to the 3D sparse convolution backbone. This design improves multi-scale context modeling and preserves more detailed features. In the BEV feature generation stage, the method also designs a spatial-frequency aggregation network. This network combines complementary information from the spatial domain and the frequency domain. This design improves feature representation. Results on the KITTI dataset show that the proposed method preserves strong detection performance for the Car category and further improves detection accuracy for the Pedestrian and Cyclist categories. These results confirm the effectiveness of the method in long-range and small-object detection.