An efficient and low-latency attention model for event denoising.

Journal: Neural networks : the official journal of the International Neural Network Society
Published Date:

Abstract

Event-based cameras are bio-inspired vision sensors that capture brightness changes at each pixel independently, generating sparse events with ultra-low latency. Unlike frame-based cameras, they provide highly efficient, temporally precise scene information, making them ideal for real-time processing. However, the output data from event-based cameras is often contaminated by noise, particularly Background Activity (BA) noise, which degrades data quality and affects downstream tasks. In this paper, we introduce StatFormer, a novel event denoising method designed to enhance temporal feature modeling while maintaining computational efficiency. Specifically, StatFormer incorporates statistical information into the event embedding process, allowing temporal dynamics to be directly modeled at the input level. To further enhance denoising performance, we adopt a two-stage transformer architecture focused on local attention. In the first stage, local attention is extracted within short sub-sequences to capture fine-grained spatial-temporal dependencies. In the second stage, local attention is strengthened over the entire event sequence. Compared to previous attention-based model, our approach significantly reduces both model size and inference time. Experimental results demonstrate that StatFormer achieves a 13.41 ×  speedup in inference, with an average processing time of just 1.38 μs per event, while delivering best denoising performance across several publicly available datasets.

Authors

Keywords

No keywords available for this article.