VISTA-3D : Training-free unfolding for vision-based 3D object detection.

Journal: Neural networks : the official journal of the International Neural Network Society
Published Date:

Abstract

3D object detection from vision inputs powers autonomous driving and embodied AI but remains compute- and energy-intensive at inference. While neuromorphic (spike-driven) computation promises event-driven sparsity and efficiency, prior training-free conversions have largely focused on 2D tasks. They are not directly applicable to modern 3D detectors due to normalization-dependent inconsistencies. We present VISTA-3D, a training-free unfolding method for vision-based 3D object detection. The approach replaces all LayerNorm blocks with a calibrated Exponential Normalization (ExpNorm) and emits incremental temporal updates whose sum matches the one-shot ANN output, producing a temporally unfolded representation suitable for spike-style execution. On the KITTI benchmark, VISTA-3D preserves the original detector's accuracy, achieving the same 3D Average Precision under the standard R40 evaluation protocol for both the monocular and depth-augmented variants. Experiments on nuScenes further show that the unfolding mechanism generalizes to larger transformer-based detectors, maintaining competitive accuracy while enabling sparse spike-style computation. VISTA-3D reduces analytic latency and achieves a normalized SOP-based energy proxy of approximately 0.17, without introducing new parameters. A controlled ablation confirms that directly enabling test-time spiking without replacing LayerNorm severely degrades performance, whereas our calibrated unfolding preserves accuracy robustly. VISTA-3D provides a principled, plug-and-play route toward neuromorphic-ready 3D perception, offering a stable temporal representation that may serve as a basis for future fully spiking 3D detectors.

Authors

Keywords

No keywords available for this article.