21.1CVMar 14, 2023
I$^2$-SDF: Intrinsic Indoor Scene Reconstruction and Editing via Raytracing in Neural SDFsJingsen Zhu, Yuchi Huo, Qi Ye et al.
In this work, we present I$^2$-SDF, a new method for intrinsic indoor scene reconstruction and editing using differentiable Monte Carlo raytracing on neural signed distance fields (SDFs). Our holistic neural SDF-based framework jointly recovers the underlying shapes, incident radiance and materials from multi-view images. We introduce a novel bubble loss for fine-grained small objects and error-guided adaptive sampling scheme to largely improve the reconstruction quality on large-scale indoor scenes. Further, we propose to decompose the neural radiance field into spatially-varying material of the scene as a neural field through surface-based, differentiable Monte Carlo raytracing and emitter semantic segmentations, which enables physically based and photorealistic scene relighting and editing applications. Through a number of qualitative and quantitative experiments, we demonstrate the superior quality of our method on indoor scene reconstruction, novel view synthesis, and scene editing compared to state-of-the-art baselines.
22.5CVJul 22, 2022
NeurAR: Neural Uncertainty for Autonomous 3D Reconstruction with Implicit Neural RepresentationsYunlong Ran, Jing Zeng, Shibo He et al.
Implicit neural representations have shown compelling results in offline 3D reconstruction and also recently demonstrated the potential for online SLAM systems. However, applying them to autonomous 3D reconstruction, where a robot is required to explore a scene and plan a view path for the reconstruction, has not been studied. In this paper, we explore for the first time the possibility of using implicit neural representations for autonomous 3D scene reconstruction by addressing two key challenges: 1) seeking a criterion to measure the quality of the candidate viewpoints for the view planning based on the new representations, and 2) learning the criterion from data that can generalize to different scenes instead of a hand-crafting one. To solve the challenges, firstly, a proxy of Peak Signal-to-Noise Ratio (PSNR) is proposed to quantify a viewpoint quality; secondly, the proxy is optimized jointly with the parameters of an implicit neural network for the scene. With the proposed view quality criterion from neural networks (termed as Neural Uncertainty), we can then apply implicit representations to autonomous 3D reconstruction. Our method demonstrates significant improvements on various metrics for the rendered image quality and the geometry quality of the reconstructed 3D models when compared with variants using TSDF or reconstruction without view planning. Project webpage https://kingteeloki-ran.github.io/NeurAR/
14.5CVOct 4, 2022
ImmFusion: Robust mmWave-RGB Fusion for 3D Human Body Reconstruction in All Weather ConditionsAnjun Chen, Xiangyu Wang, Kun Shi et al.
3D human reconstruction from RGB images achieves decent results in good weather conditions but degrades dramatically in rough weather. Complementary, mmWave radars have been employed to reconstruct 3D human joints and meshes in rough weather. However, combining RGB and mmWave signals for robust all-weather 3D human reconstruction is still an open challenge, given the sparse nature of mmWave and the vulnerability of RGB images. In this paper, we present ImmFusion, the first mmWave-RGB fusion solution to reconstruct 3D human bodies in all weather conditions robustly. Specifically, our ImmFusion consists of image and point backbones for token feature extraction and a Transformer module for token fusion. The image and point backbones refine global and local features from original data, and the Fusion Transformer Module aims for effective information fusion of two modalities by dynamically selecting informative tokens. Extensive experiments on a large-scale dataset, mmBody, captured in various environments demonstrate that ImmFusion can efficiently utilize the information of two modalities to achieve a robust 3D human body reconstruction in all weather conditions. In addition, our method's accuracy is significantly superior to that of state-of-the-art Transformer-based LiDAR-camera fusion methods.
17.3CVSep 12, 2022
mmBody Benchmark: 3D Body Reconstruction Dataset and Analysis for Millimeter Wave RadarAnjun Chen, Xiangyu Wang, Shaohao Zhu et al.
Millimeter Wave (mmWave) Radar is gaining popularity as it can work in adverse environments like smoke, rain, snow, poor lighting, etc. Prior work has explored the possibility of reconstructing 3D skeletons or meshes from the noisy and sparse mmWave Radar signals. However, it is unclear how accurately we can reconstruct the 3D body from the mmWave signals across scenes and how it performs compared with cameras, which are important aspects needed to be considered when either using mmWave radars alone or combining them with cameras. To answer these questions, an automatic 3D body annotation system is first designed and built up with multiple sensors to collect a large-scale dataset. The dataset consists of synchronized and calibrated mmWave radar point clouds and RGB(D) images in different scenes and skeleton/mesh annotations for humans in the scenes. With this dataset, we train state-of-the-art methods with inputs from different sensors and test them in various scenarios. The results demonstrate that 1) despite the noise and sparsity of the generated point clouds, the mmWave radar can achieve better reconstruction accuracy than the RGB camera but worse than the depth camera; 2) the reconstruction from the mmWave radar is affected by adverse weather conditions moderately while the RGB(D) camera is severely affected. Further, analysis of the dataset and the results shadow insights on improving the reconstruction from the mmWave radar and the combination of signals from different sensors.
Seal-3D: Interactive Pixel-Level Editing for Neural Radiance FieldsXiangyu Wang, Jingsen Zhu, Qi Ye et al.
With the popularity of implicit neural representations, or neural radiance fields (NeRF), there is a pressing need for editing methods to interact with the implicit 3D models for tasks like post-processing reconstructed scenes and 3D content creation. While previous works have explored NeRF editing from various perspectives, they are restricted in editing flexibility, quality, and speed, failing to offer direct editing response and instant preview. The key challenge is to conceive a locally editable neural representation that can directly reflect the editing instructions and update instantly. To bridge the gap, we propose a new interactive editing method and system for implicit representations, called Seal-3D, which allows users to edit NeRF models in a pixel-level and free manner with a wide range of NeRF-like backbone and preview the editing effects instantly. To achieve the effects, the challenges are addressed by our proposed proxy function mapping the editing instructions to the original space of NeRF models in the teacher model and a two-stage training strategy for the student model with local pretraining and global finetuning. A NeRF editing system is built to showcase various editing types. Our system can achieve compelling editing effects with an interactive speed of about 1 second.
1.4CVOct 17, 2022
Pixel-Aligned Non-parametric Hand Mesh ReconstructionShijian Jiang, Guwen Han, Danhang Tang et al.
Non-parametric mesh reconstruction has recently shown significant progress in 3D hand and body applications. In these methods, mesh vertices and edges are visible to neural networks, enabling the possibility to establish a direct mapping between 2D image pixels and 3D mesh vertices. In this paper, we seek to establish and exploit this mapping with a simple and compact architecture. The network is designed with these considerations: 1) aggregating both local 2D image features from the encoder and 3D geometric features captured in the mesh decoder; 2) decoding coarse-to-fine meshes along the decoding layers to make the best use of the hierarchical multi-scale information. Specifically, we propose an end-to-end pipeline for hand mesh recovery tasks which consists of three phases: a 2D feature extractor constructing multi-scale feature maps, a feature mapping module transforming local 2D image features to 3D vertex features via 3D-to-2D projection, and a mesh decoder combining the graph convolution and self-attention to reconstruct mesh. The decoder aggregate both local image features in pixels and geometric features in vertices. It also regresses the mesh vertices in a coarse-to-fine manner, which can leverage multi-scale information. By exploiting the local connection and designing the mesh decoder, Our approach achieves state-of-the-art for hand mesh reconstruction on the public FreiHAND dataset.
3.3NASep 1, 2011
Reproducing Kernels of Generalized Sobolev Spaces via a Green Function Approach with Differential OperatorsQi Ye
In this paper we introduce a generalization of the classical $\Leb_2(\Rd)$-based Sobolev spaces with the help of a vector differential operator $\mathbf{P}$ which consists of finitely or countably many differential operators $P_n$ which themselves are linear combinations of distributional derivatives. We find that certain proper full-space Green functions $G$ with respect to $L=\mathbf{P}^{\ast T}\mathbf{P}$ are positive definite functions. Here we ensure that the vector distributional adjoint operator $\mathbf{P}^{\ast}$ of $\mathbf{P}$ is well-defined in the distributional sense. We then provide sufficient conditions under which our generalized Sobolev space will become a reproducing-kernel Hilbert space whose reproducing kernel can be computed via the associated Green function $G$. As an application of this theoretical framework we use $G$ to construct multivariate minimum-norm interpolants $s_{f,X}$ to data sampled from a generalized Sobolev function $f$ on $X$. Among other examples we show the reproducing-kernel Hilbert space of the Gaussian function is equivalent to a generalized Sobolev space.
1.2NAFeb 19, 2015
Kernel-based Methods for Stochastic Partial Differential EquationsQi Ye
This article gives a new insight of kernel-based (approximation) methods to solve the high-dimensional stochastic partial differential equations. We will combine the techniques of meshfree approximation and kriging interpolation to extend the kernel-based methods for the deterministic data to the stochastic data. The main idea is to endow the Sobolev spaces with the probability measures induced by the positive definite kernels such that the Gaussian random variables can be well-defined on the Sobolev spaces. The constructions of these Gaussian random variables provide the kernel-based approximate solutions of the stochastic models. In the numerical examples of the stochastic Poisson and heat equations, we show that the approximate probability distributions are well-posed for various kinds of kernels such as the compactly supported kernels (Wendland functions) and the Sobolev-spline kernels (Matérn functions).
2.4OCMar 1, 2023
Composite Optimization Algorithms for Sigmoid NetworksHuixiong Chen, Qi Ye
In this paper, we use composite optimization algorithms to solve sigmoid networks. We equivalently transfer the sigmoid networks to a convex composite optimization and propose the composite optimization algorithms based on the linearized proximal algorithms and the alternating direction method of multipliers. Under the assumptions of the weak sharp minima and the regularity condition, the algorithm is guaranteed to converge to a globally optimal solution of the objective function even in the case of non-convex and non-smooth problems. Furthermore, the convergence results can be directly related to the amount of training data and provide a general guide for setting the size of sigmoid networks. Numerical experiments on Franke's function fitting and handwritten digit recognition show that the proposed algorithms perform satisfactorily and robustly.
1.2NAOct 14, 2017
Kernel-based Approximation Methods for Generalized Interpolations: A Deterministic or Stochastic Problem?Qi Ye
In this article, we solve a deterministically generalized interpolation problem by a stochastic approach. We introduce a kernel-based probability measure on a Banach space by a covariance kernel which is defined on the dual space of the Banach space. The kernel-based probability measure provides a numerical tool to construct and analyze the kernel-based estimators conditioned on non-noise data or noisy data including algorithms and error analysis. Same as meshfree methods, we can also obtain the kernel-based approximate solutions of elliptic partial differential equations by the kernel-based probability measure.
5.9CVDec 27, 2023
In-Hand 3D Object Reconstruction from a Monocular RGB VideoShijian Jiang, Qi Ye, Rengan Xie et al.
Our work aims to reconstruct a 3D object that is held and rotated by a hand in front of a static RGB camera. Previous methods that use implicit neural representations to recover the geometry of a generic hand-held object from multi-view images achieved compelling results in the visible part of the object. However, these methods falter in accurately capturing the shape within the hand-object contact region due to occlusion. In this paper, we propose a novel method that deals with surface reconstruction under occlusion by incorporating priors of 2D occlusion elucidation and physical contact constraints. For the former, we introduce an object amodal completion network to infer the 2D complete mask of objects under occlusion. To ensure the accuracy and view consistency of the predicted 2D amodal masks, we devise a joint optimization method for both amodal mask refinement and 3D reconstruction. For the latter, we impose penetration and attraction constraints on the local geometry in contact regions. We evaluate our approach on HO3D and HOD datasets and demonstrate that it outperforms the state-of-the-art methods in terms of reconstruction surface quality, with an improvement of $52\%$ on HO3D and $20\%$ on HOD. Project webpage: https://east-j.github.io/ihor.
1.2QUANT-PHFeb 7, 2025
Quantum automated learning with provable and explainable trainabilityQi Ye, Shuangyue Geng, Zizhao Han et al. · tsinghua
Machine learning is widely believed to be one of the most promising practical applications of quantum computing. Existing quantum machine learning schemes typically employ a quantum-classical hybrid approach that relies crucially on gradients of model parameters. Such an approach lacks provable convergence to global minima and will become infeasible as quantum learning models scale up. Here, we introduce quantum automated learning, where no variational parameter is involved and the training process is converted to quantum state preparation. In particular, we encode training data into unitary operations and iteratively evolve a random initial state under these unitaries and their inverses, with a target-oriented perturbation towards higher prediction accuracy sandwiched in between. Under reasonable assumptions, we rigorously prove that the evolution converges exponentially to the desired state corresponding to the global minimum of the loss function. We show that such a training process can be understood from the perspective of preparing quantum states by imaginary time evolution, where the data-encoded unitaries together with target-oriented perturbations would train the quantum learning model in an automated fashion. We further prove that the quantum automated learning paradigm features good generalization ability with the generalization error upper bounded by the ratio between a logarithmic function of the Hilbert space dimension and the number of training samples. In addition, we carry out extensive numerical simulations on real-life images and quantum data to demonstrate the effectiveness of our approach and validate the assumptions. Our results establish an unconventional quantum learning strategy that is gradient-free with provable and explainable trainability, which would be crucial for large-scale practical applications of quantum computing in machine learning scenarios.
1.2QUANT-PHDec 7, 2024
No-Free-Lunch Theories for Tensor-Network Machine Learning ModelsJing-Chuan Wu, Qi Ye, Dong-Ling Deng et al.
Tensor network machine learning models have shown remarkable versatility in tackling complex data-driven tasks, ranging from quantum many-body problems to classical pattern recognitions. Despite their promising performance, a comprehensive understanding of the underlying assumptions and limitations of these models is still lacking. In this work, we focus on the rigorous formulation of their no-free-lunch theorem -- essential yet notoriously challenging to formalize for specific tensor network machine learning models. In particular, we rigorously analyze the generalization risks of learning target output functions from input data encoded in tensor network states. We first prove a no-free-lunch theorem for machine learning models based on matrix product states, i.e., the one-dimensional tensor network states. Furthermore, we circumvent the challenging issue of calculating the partition function for two-dimensional Ising model, and prove the no-free-lunch theorem for the case of two-dimensional projected entangled-pair state, by introducing the combinatorial method associated to the "puzzle of polyominoes". Our findings reveal the intrinsic limitations of tensor network-based learning models in a rigorous fashion, and open up an avenue for future analytical exploration of both the strengths and limitations of quantum-inspired machine learning frameworks.
1.6LGSep 7, 2021
Analysis of Regularized Learning in Banach Spaces for Linear-functional DataQi Ye
This article delves into the study of the theory of regularized learning in Banach spaces for linear-functional data. It encompasses discussions on representer theorems, pseudo-approximation theorems, and convergence theorems. Regularized learning is designed to minimize regularized empirical risks over a Banach space. The empirical risks are calculated by utilizing training data and multi-loss functions. The input training data are composed of linear functionals in a predual space of the Banach space to capture discrete local information from multimodal data and multiscale models. Through the regularized learning, approximations of the exact solution to an unidentified or uncertain original problem are globally achieved. In the convergence theorems, the convergence of the approximate solutions to the exact solution is established through the utilization of the weak* topology of the Banach space. The theorems of regularized learning are utilized in the interpretation of classical machine learning, such as support vector machines and artificial neural networks.
1.2NASep 27, 2011
Reproducing Kernels of Sobolev Spaces via a Green Kernel Approach with Differential Operators and Boundary OperatorsGregory E. Fasshauer, Qi Ye
We introduce a vector differential operator $\mathbf{P}$ and a vector boundary operator $\mathbf{B}$ to derive a reproducing kernel along with its associated Hilbert space which is shown to be embedded in a classical Sobolev space. This reproducing kernel is a Green kernel of differential operator $L:=\mathbf{P}^{\ast T}\mathbf{P}$ with homogeneous or nonhomogeneous boundary conditions given by $\mathbf{B}$, where we ensure that the distributional adjoint operator $\mathbf{P}^{\ast}$ of $\mathbf{P}$ is well-defined in the distributional sense. We represent the inner product of the reproducing-kernel Hilbert space in terms of the operators $\mathbf{P}$ and $\mathbf{B}$. In addition, we find relationships for the eigenfunctions and eigenvalues of the reproducing kernel and the operators with homogeneous or nonhomogeneous boundary conditions. These eigenfunctions and eigenvalues are used to compute a series expansion of the reproducing kernel and an orthonormal basis of the reproducing-kernel Hilbert space. Our theoretical results provide perhaps a more intuitive way of understanding what kind of functions are well approximated by the reproducing kernel-based interpolant to a given multivariate data sample.