Elasticity in Parallel Sparse Triangular Solve
It improves parallel sparse triangular solve performance for high-performance computing users, but the gains are incremental.
The paper introduces stale synchronous parallel execution for sparse triangular solve, achieving 7-30% geometric-mean speedup over GrowLocal on ARM and 19-60% over SpMP on x86 with 48 cores.
We introduce stale synchronous parallel as a mode of execution in parallel sparse triangular linear system solve and present a general directed-acyclic-graph scheduler capable of producing such schedules. Stale-synchronous-parallel schedules allow the overlap of synchronisation and compute which results in a geometric-mean speed-up of $7$-$30\%$ of our scheduler, ElasticDivide, over state-of-the-art synchronous scheduler GrowLocal on an ARM machine using 48 cores. On an x86 machine using 48 cores, we report geometric-mean speed-ups of $19$-$60\%$ over SpMP.