Graphical conditional generative modeling for digital twin modeling

arXiv:2606.162197.3
Predicted impact top 56% in CE · last 90 daysOriginality Incremental advance
AI Analysis

This work addresses the fidelity problem in digital twin modeling by providing a principled way to build simpler, interpretable stochastic surrogates, which is important for maintaining, validating, and stress-testing complex systems.

The paper introduces a framework for discovering parsimonious stochastic surrogate models from observational data by identifying which candidate inputs influence the full conditional law of a target quantity, rather than just its mean. The method couples conditional generative modeling with Gaussian-process-based analysis of variance to prune non-influential inputs, yielding interpretable surrogates that perform comparably to full-variable models across diverse examples.

Digital twin modeling, including control and data assimilation under model uncertainty, often faces an open-ended fidelity problem: adding variables, data streams, and time scales can indefinitely increase model complexity, ultimately producing systems that are difficult to maintain, validate, interpret, and use for stress or safety testing. As an alternative, one can seek parsimonious stochastic surrogate models built only on the variables needed to describe the relevant quantities of interest. We introduce a framework for discovering such variables from observational data by identifying which candidate inputs influence the full conditional law of a target quantity, rather than only its conditional mean. This distinction is essential in stochastic, coarse-grained, or partially observed systems, where dependencies may appear through changes in variability, tail behavior, multimodality, or uncertainty rather than through deterministic functional relationships. The framework couples conditional generative modeling, which learns the conditional distribution of the target given candidate inputs, with Gaussian-process-based analysis of variance (through kernel mode decomposition), which enables iterative pruning of non-influential inputs and interpretable structure discovery. In control settings, the resulting surrogate can be interpreted as a learned Markov decision process: the method identifies not only a transition model, but also the state, action, and memory variables needed to make the learned dynamics effectively Markovian. Across examples involving stochastic dynamical systems, missing variables, PDE control, reinforcement learning, and economic data, the discovered structures yield interpretable stochastic surrogates whose downstream performance is comparable to models trained on the full variable set.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes