1.3DCJun 16
From GPU to Microcontroller: Online Ridge Regression for Edge-Deployable Traffic PredictionSuresh Purini, Archit Narwadkar, Deepak Gangadharan
State-of-the-art traffic flow forecasting models, including Graph Convolutional Networks and graph-less MLPs, require centralized GPU training across all sensors, making them impractical for resource-constrained intelligent transportation deployments. We show that much of this complexity is unnecessary. A parametric analysis of the recent graph-less model GLMST reveals that reducing its internal embedding dimension from 64 to 4 degrades MAPE by less than one percentage point, suggesting that the model's effective capacity far exceeds what the task requires. Motivated by this finding, we replace the neural architecture entirely with per-sensor Ridge regression using horizon-aligned periodic features, combined with Recursive Least Squares (RLS) for online adaptation. With only 444 parameters per sensor (80x fewer than GLMST) and test-time online adaptation, our method achieves the best MAPE on three of four PEMS benchmarks, and remains within one percentage point on the fourth. Because each sensor's model is self-contained and involves only elementary linear algebra, the entire pipeline (training, inference, and online adaptation) runs on edge hardware without a GPU. An ESP32 microcontroller (160 MHz, 520 KB SRAM) completes cold-start training in 7.4s and each predict-and-update in under 2ms with zero heap allocation; a single Raspberry Pi 5 core completes cold-start training in 0.21s and each predict-and-update in 0.26ms.
3.6IVFeb 18, 2024
Accelerating local laplacian filters on FPGAsShashwat Khandelwal, Ziaul Choudhury, Shashwat Shrivastava et al.
Images when processed using various enhancement techniques often lead to edge degradation and other unwanted artifacts such as halos. These artifacts pose a major problem for photographic applications where they can denude the quality of an image. There is a plethora of edge-aware techniques proposed in the field of image processing. However, these require the application of complex optimization or post-processing methods. Local Laplacian Filtering is an edge-aware image processing technique that involves the construction of simple Gaussian and Laplacian pyramids. This technique can be successfully applied for detail smoothing, detail enhancement, tone mapping and inverse tone mapping of an image while keeping it artifact-free. The problem though with this approach is that it is computationally expensive. Hence, parallelization schemes using multi-core CPUs and GPUs have been proposed. As is well known, they are not power-efficient, and a well-designed hardware architecture on an FPGA can do better on the performance per watt metric. In this paper, we propose a hardware accelerator, which exploits fully the available parallelism in the Local Laplacian Filtering algorithm, while minimizing the utilization of on-chip FPGA resources. On Virtex-7 FPGA, we obtain a 7.5x speed-up to process a 1 MB image when compared to an optimized baseline CPU implementation. To the best of our knowledge, we are not aware of any other hardware accelerators proposed in the research literature for the Local Laplacian Filtering problem.