7.2LGApr 23
Insect-inspired modular architectures as inductive biases for reinforcement learningAnne E. Staples
Most reinforcement-learning (RL) controllers used in continuous control are architecturally centralized: observations are compressed into a single latent state from which both value estimates and actions are produced. Biological control systems are often organized differently. Insects, in particular, coordinate navigation, heading stabilization, memory, and context-dependent action selection through distributed circuits rather than a single monolithic controller. Motivated by this contrast, we study an RL policy architecture that decomposes control into interacting modules for sensory encoding, heading representation, sparse associative memory, recurrent command generation, and local motor control, with a learned arbitration mechanism that allocates motor authority across modules. The model is evaluated on a two-dimensional navigation task that require simultaneous food seeking, obstacle avoidance, and predator escape. In a six-seed predator-navigation experiment trained with Proximal Policy Optimization (PPO) for 75 updates, the modular policy achieves the strongest final mean performance among the tested controllers, with final episodic return $-2798.8\pm964.4$ versus $-3778.0\pm628.1$ for a centralized gated recurrent unit (GRU) and $-4727.5\pm772.5$ for a centralized multilayer perceptron (MLP). The modular policy also attains the lowest final value loss and stable PPO optimization statistics while driving module-assignment entropy to $0.0457\pm0.0244$, indicating highly selective control allocation. These results suggest that distributed control can serve as a useful inductive bias for RL problems involving dynamically competing behavioral objectives.
1.2COMP-PHOct 22, 2016
A Finite-Element Coarse-Grid Projection Method: A Dual Acceleration/Mesh Refinement Tool for Incompressible FlowsA. Kashefi, A. E. Staples
Coarse grid projection (CGP) methodology is a novel multigrid method for systems involving decoupled nonlinear evolution equations and linear elliptic Poisson equations. The nonlinear equations are solved on a fine grid and the linear equations are solved on a corresponding coarsened grid. Mapping operators execute data transfer between the grids. The CGP framework is constructed upon spatial and temporal discretization schemes. This framework has been established for finite volume/difference discretizations as well as explicit time integration methods. In this article we present for the first time a version of CGP for finite element discretizations, which uses a semi-implicit time integration scheme. The mapping functions correspond to the finite-element shape functions. With the novel data structure introduced, the mapping computational cost becomes insignificant. We apply CGP to pressure-correction schemes used for the incompressible Navier-Stokes flow computations. This version is validated on standard test cases with realistic boundary conditions using unstructured triangular meshes. We also pioneer investigations of the effects of CGP on the accuracy of the pressure field. It is found that although CGP reduces the pressure field accuracy, it preserves the accuracy of the pressure gradient and thus the velocity field, while achieving speedup factors ranging from approximately 2 to 30. Exploring the influence of boundary conditions on CGP, the minimum speedup occurs for velocity Dirichlet boundary conditions, while the maximum speedup occurs for open boundary conditions. We discuss the CGP method as a guide for partial mesh refinement of incompressible flow computations and show its application for simulations of flow over a backward-facing step and flow past a cylinder.