Julian D. Schiller

LGFeb 11

Tuning the burn-in phase in training recurrent neural networks improves their performance

Julian D. Schiller, Malte Heinrich, Victor G. Lopez et al.

Training recurrent neural networks (RNNs) with standard backpropagation through time (BPTT) can be challenging, especially in the presence of long input sequences. A practical alternative to reduce computational and memory overhead is to perform BPTT repeatedly over shorter segments of the training data set, corresponding to truncated BPTT. In this paper, we examine the training of RNNs when using such a truncated learning approach for time series tasks. Specifically, we establish theoretical bounds on the accuracy and performance loss when optimizing over subsequences instead of the full data sequence. This reveals that the burn-in phase of the RNN is an important tuning knob in its training, with significant impact on the performance guarantees. We validate our theoretical results through experiments on standard benchmarks from the fields of system identification and time series forecasting. In all experiments, we observe a strong influence of the burn-in phase on the training process, and proper tuning can lead to a reduction of the prediction error on the training and test data of more than 60% in some cases.

30.2SYApr 8

Small-gain analysis of exponential incremental input/output-to-state stability for large-scale distributed systems

Christian Gatke, Julian D. Schiller, Matthias A. Müller

We provide a detectability analysis for nonlinear large-scale distributed systems in the sense of exponential incremental input/output-to-state stability (i-IOSS). In particular, we prove that the overall system is exponentially i-IOSS if each subsystem is i-IOSS, with interconnections treated as external inputs, and a suitable small-gain condition holds. The analysis is extended to a Lyapunov characterization, resulting in a different quantitative outcome regarding the small-gain condition, which is further analyzed within this work. Moreover, we derive linear matrix inequality conditions posed solely on the local subsystems and their interconnections, which guarantee exponential i-IOSS of the overall distributed system. The results are illustrated on a numerical example.

Julian D. Schiller

2 Papers