Posture-driven multistage region proposal and fast R-CNN architecture for advanced real-time autonomous vehicle object detection
Abstract
Accurate and real-time detection of objects is important in the field of autonomous vehicle (AV) safety, but the current detection systems have numerous issues. Conventional Region Proposal Networks (RPNs) do not typically succeed in the technical urban setting with tiny or hidden features, whereas Fast R-CNN-based techniques do not possess the capability to use the posture of vulnerable road individuals, like pedestrians and cyclists. The research gaps exist in the ability to incorporate the posture information directly into RPNs, real-time performance on edge devices, and performance in various conditions of the environment. The paper will deal with these drawbacks by proposing a posture-sensitive object detection model that combines Automated Posture Indexing (API) with a RPN and Fast R-CNN with HRNet. Posture and orientation gross will contribute to the following objectives: Ensuring posture and orientation characteristics are incorporated into region proposals; Keeping the feature extraction resolution up while assuring good performance for detecting small areas and occlusions and optimizing the inference operation for edge devices; Furthermore, the kitchen team is testing the human-robot interaction in real space and getting better detection results of the animals. Its methodology consisted of seven execution steps, namely, data acquisition and preprocessing; API; posture-aware RPN; HRNet feature extraction; Fast R-CNN classification and localization; real-time optimization; and decision-making with navigation support. It was tested with each stage by means of properly defined test cases and real-life situations, such as the pedestrian running, the cyclist turning, and multi-object interactions. An urban intersection with a running pedestrian, a leaning-to-the-left cyclist, and a distant, obscured traffic sign was used to validate a sample dataset. Sample performance indicated that sensor fusion could achieve sensor fusion latency of 8 ms, posture key points detection of 96 percent, and posture-aware RPN of 12.4 percent higher recall than baselines. HRNet achieved an improvement of 9.1% mean average precision (mAP) for small objects with only 12% latency overhead, and Fast R-CNN achieved 91.7% classification accuracy and a 0.87 posture F1-score. On an NVIDIA Jetson AGX Xavier, inference operated at 31 FPS with 46 ms latency to frame firmware according to real-time constraints. In the scenarios of navigation, situations of early braking between running pedestrians were activated in 428 ms with a minimum time to collision of 2.3 s, whereas a cyclist turning intention with an F1 of 0.89 was predicted, while a 100% collision avoidance rate in a single scenario was achieved as of the sample run. The results prove that the proposed posture-conscious framework not only addresses the existing research gaps but also guarantees the robust perception, effective edge deployment, and predictive navigation safety. This work provides an authenticated methodology for improving the safety and reliability of autonomous vehicles in dynamic city settings by consistently attaining more than 90% detection accuracy, more than 95% collision avoidance, and real-time performance.
// Source
Authors: Kalvacherla Kiran, T. Sampath Kumar, Santosh Kumar Henge
Institutions: Artificial Intelligence in Medicine (Canada)