Shiheng Zhang

NA
h-index9
8papers
309citations
Novelty48%
AI Score48

8 Papers

16.8MLJun 9, 2023
Energy-Dissipative Evolutionary Deep Operator Neural Networks

Jiahao Zhang, Shiheng Zhang, Jie Shen et al.

Energy-Dissipative Evolutionary Deep Operator Neural Network is an operator learning neural network. It is designed to seed numerical solutions for a class of partial differential equations instead of a single partial differential equation, such as partial differential equations with different parameters or different initial conditions. The network consists of two sub-networks, the Branch net and the Trunk net. For an objective operator G, the Branch net encodes different input functions u at the same number of sensors, and the Trunk net evaluates the output function at any location. By minimizing the error between the evaluated output q and the expected output G(u)(y), DeepONet generates a good approximation of the operator G. In order to preserve essential physical properties of PDEs, such as the Energy Dissipation Law, we adopt a scalar auxiliary variable approach to generate the minimization problem. It introduces a modified energy and enables unconditional energy dissipation law at the discrete level. By taking the parameter as a function of time t, this network can predict the accurate solution at any further time with feeding data only at the initial state. The data needed can be generated by the initial conditions, which are readily available. In order to validate the accuracy and efficiency of our neural networks, we provide numerical simulations of several partial differential equations, including heat equations, parametric heat equations and Allen-Cahn equations.

6.9NAApr 12
On the stability of the low-rank projector-splitting integrators for hyperbolic and parabolic equations

Shiheng Zhang, Jingwei Hu

We study the stability of a class of dynamical low-rank methods--the projector-splitting integrator (PSI)--applied to linear hyperbolic and parabolic equations. Using a von Neumann-type analysis, we investigate the stability of such low-rank time integrator coupled with standard spatial discretizations, including upwind and central finite difference schemes, under two commonly used formulations: discretize-then-project (DtP) and project-then-discretize (PtD). For hyperbolic equations, we show that the stability conditions for DtP and PtD are the same under Lie-Trotter splitting, and that the stability region can be significantly enlarged by using Strang splitting. For parabolic equations, despite the presence of a negative S-step, unconditional stability can still be achieved by employing Crank-Nicolson or a hybrid forward-backward Euler scheme in time stepping. While our analysis focuses on simplified model problems, it offers insight into the stability behavior of PSI for more complex systems, such as those arising in kinetic theory.

8.5NAMay 27
A separable and asymptotic-preserving dynamical low-rank method for the Vlasov-Poisson-Fokker-Planck system

Shiheng Zhang, Jingwei Hu

We present a dynamical low-rank (DLR) method for the Vlasov-Poisson-Fokker-Planck (VPFP) system. Our main contributions are two-fold: (i) a conservative spatial discretization of the Fokker-Planck operator that factors into velocity-only and space-only components, enabling efficient low-rank projection, and (ii) a time discretization within the DLR framework that properly handles stiff collisions. We propose both first-order and second-order low-rank IMEX schemes. For the first-order scheme, we prove an asymptotic-preserving (AP) property when the field fluctuation is small. Numerical experiments demonstrate accuracy, robustness, and AP property at modest ranks.

5.1LGMay 18
Planner-Admissible Graph-PDE Value Extensions for Sparse Goal-Conditioned Planning

Shiheng Zhang

Sparse goal-conditioned planning with few cost-to-go labels can be viewed as a graph-PDE Dirichlet extension problem: extend sparse labels on a goal-dependent boundary to unlabelled graph vertices so that greedy rollouts reach the goal. We study which graph value extensions are planner-admissible under the operational argmin-Q planner. Our main result is a local action-gap certificate: if the surrogate value error along the rollout stays below half the true action gap, then the greedy rollout reaches the goal. Absolutely Minimal Lipschitz Extension (AMLE), the p=infinity endpoint of the graph p-Laplacian family, instantiates this certificate through a comparison-principle fill-distance bound. Harmonic extension, by contrast, can mis-rank local actions because its values reflect boundary hitting probabilities rather than shortest-path greedy order. On 120 AntMaze layout-derived graph configurations, harmonic extension achieves 0.584 aggregate rollout success, while AMLE reaches 0.970. Finite high-p methods also enter a high-success regime, with success 0.903 for p=4, 0.973 for p=8, and 0.982 for a fixed-budget p=16 solver, though the p=16 row is not used as a converged endpoint ranking due to incomplete solver certification. Mechanism audits show that many rollout decisions occur in AMLE-compatible but harmonic-incompatible local geometry, and that AMLE corrects most harmonic inversions on the rollout-weighted decision scope.

12.1LGJul 8
Guidance Breaks the Fitted Operator: A Terminal-Fitted Repair for Classifier-Free Guidance

Shiheng Zhang

Classifier-free guidance (CFG) is the standard way to strengthen class-conditioning in diffusion and flow-matching samplers, yet at large guidance it oversaturates and destabilizes, symptoms practitioners suppress with more steps or limited-interval schedules. We analyze CFG through an asymptotic-preserving, numerical-analysis lens. Building on a recent result that the deterministic DDIM step is the unique fitted operator for the unguided terminal layer, exact on the final small-sigma stretch of sampling, we show that guidance re-stiffens exactly the discriminative subspace to an anomalous exponent 1+w. DDIM is therefore no longer fitted there, and on coarse meshes its guided residual diverges as sigma_min goes to zero. We prove a guided clock barrier with three ordered step-size thresholds, and read one-step oversaturation as its endpoint: a solver artifact on the calibration model rather than the continuous guided law. The same analysis yields a one-coefficient, zero-extra-NFE repair: replace CFG's w(r-1) by r^(1+w)-r on the guidance direction. On the calibration model's discriminative crossover, this removes CFG's sigma_min-divergent blow-up and is first-order accurate against the exact guided flow as sigma_min goes to zero. On learned CIFAR-10 checkpoints, and as a cross-domain smoke test on Stable Diffusion 1.5 DDIM, it acts as a high-guidance stabilizer at no extra cost rather than a universal quality knob: it cuts residual amplification and saturation, gives 9/9 point-FID wins over CFG on the tested grid, and preserves classifier-proxy target accuracy in the hard-cell blocks. We report the limits alongside: it is not a universal image-quality win, and against a dense vanilla-CFG reference it is not a uniformly better integrator of that field.

2.4OCSep 7, 2023
An Element-wise RSAV Algorithm for Unconstrained Optimization Problems

Shiheng Zhang, Jiahao Zhang, Jie Shen et al.

We present a novel optimization algorithm, element-wise relaxed scalar auxiliary variable (E-RSAV), that satisfies an unconditional energy dissipation law and exhibits improved alignment between the modified and the original energy. Our algorithm features rigorous proofs of linear convergence in the convex setting. Furthermore, we present a simple accelerated algorithm that improves the linear convergence rate to super-linear in the univariate case. We also propose an adaptive version of E-RSAV with Steffensen step size. We validate the robustness and fast convergence of our algorithm through ample numerical experiments.

3.1NAJun 17
Scalar-Tracking SAV Schemes with Pullback Corrections for Gradient Flows

Shiheng Zhang, Jie Shen

The scalar auxiliary variable (SAV) method constructs linear, unconditionally energy-stable time discretizations of gradient flows. In a first-order SAV step, eliminating the auxiliary variable shows that the state equation is a semi-implicit update augmented by a rank-one positive semidefinite correction from the previous nonlinear force. The multiple-SAV (MSAV) method produces this correction componentwise, yielding a correction of rank up to the number of energy components. This separates two mechanisms usually coupled in MSAV: the number of scalar variables tracking the nonlinear energy and the rank of the correction applied to the state equation. We introduce a pullback-corrected SAV (PB-SAV) family that keeps a single scalar auxiliary variable but replaces the rank-one SAV correction by the pullback correction induced by an admissible component decomposition. The correction remains positive semidefinite, has rank at most the number of components, and may change from step to step without changing the scalar energy tracker. We prove modified-energy dissipation laws for fixed and step-dependent decompositions, derive a refinement identity whose gain is an explicit weighted variance, and give a Sherman-Morrison-Woodbury implementation of the low-rank perturbation of the standard semi-implicit solve. We also show, in finite dimensions, that the pullback correction is the Gauss-Newton matrix of a least-squares representation of the nonlinear energy. Numerical experiments on finite-dimensional gradient flows, Allen-Cahn dynamics, and nonlocal Cahn-Hilliard models illustrate regimes in which PB-SAV mainly changes the first-order error constant and regimes in which it substantially improves trajectory accuracy.

7.8MLSep 19, 2025
Low-Rank Adaptation of Evolutionary Deep Neural Networks for Efficient Learning of Time-Dependent PDEs

Jiahao Zhang, Shiheng Zhang, Guang Lin

We study the Evolutionary Deep Neural Network (EDNN) framework for accelerating numerical solvers of time-dependent partial differential equations (PDEs). We introduce a Low-Rank Evolutionary Deep Neural Network (LR-EDNN), which constrains parameter evolution to a low-rank subspace, thereby reducing the effective dimensionality of training while preserving solution accuracy. The low-rank tangent subspace is defined layer-wise by the singular value decomposition (SVD) of the current network weights, and the resulting update is obtained by solving a well-posed, tractable linear system within this subspace. This design augments the underlying numerical solver with a parameter efficient EDNN component without requiring full fine-tuning of all network weights. We evaluate LR-EDNN on representative PDE problems and compare it against corresponding baselines. Across cases, LR-EDNN achieves comparable accuracy with substantially fewer trainable parameters and reduced computational cost. These results indicate that low-rank constraints on parameter velocities, rather than full-space updates, provide a practical path toward scalable, efficient, and reproducible scientific machine learning for PDEs.