Learning to Generate Semantic Layouts for Higher Text-Image Correspondence in Text-to-Image SynthesisMinho Park, Jooyeol Yun, Seunghwan Choi et al.
Existing text-to-image generation approaches have set high standards for photorealism and text-image correspondence, largely benefiting from web-scale text-image datasets, which can include up to 5~billion pairs. However, text-to-image generation models trained on domain-specific datasets, such as urban scenes, medical images, and faces, still suffer from low text-image correspondence due to the lack of text-image pairs. Additionally, collecting billions of text-image pairs for a specific domain can be time-consuming and costly. Thus, ensuring high text-image correspondence without relying on web-scale text-image datasets remains a challenging task. In this paper, we present a novel approach for enhancing text-image correspondence by leveraging available semantic layouts. Specifically, we propose a Gaussian-categorical diffusion process that simultaneously generates both images and corresponding layout pairs. Our experiments reveal that we can guide text-to-image generation models to be aware of the semantics of different image regions, by training the model to generate semantic labels for each pixel. We demonstrate that our approach achieves higher text-image correspondence compared to existing text-to-image generation approaches in the Multi-Modal CelebA-HQ and the Cityscapes dataset, where text-image pairs are scarce. Codes are available in this https://pmh9960.github.io/research/GCDP
0.9CVJul 2, 2019
Generative Guiding Block: Synthesizing Realistic Looking Variants Capable of Even Large Change DemandsMinho Park, Hak Gu Kim, Yong Man Ro
Realistic image synthesis is to generate an image that is perceptually indistinguishable from an actual image. Generating realistic looking images with large variations (e.g., large spatial deformations and large pose change), however, is very challenging. Handing large variations as well as preserving appearance needs to be taken into account in the realistic looking image generation. In this paper, we propose a novel realistic looking image synthesis method, especially in large change demands. To do that, we devise generative guiding blocks. The proposed generative guiding block includes realistic appearance preserving discriminator and naturalistic variation transforming discriminator. By taking the proposed generative guiding blocks into generative model, the latent features at the layer of generative model are enhanced to synthesize both realistic looking- and target variation- image. With qualitative and quantitative evaluation in experiments, we demonstrated the effectiveness of the proposed generative guiding blocks, compared to the state-of-the-arts.
1.2NAJun 15, 2015
Improved Multilevel Monte Carlo Methods for Finite Volume Discretisations of Darcy Flow in Randomly Layered MediaMinho Park, Aretha Teckentrup
We consider the application of multilevel Monte Carlo methods to steady state Darcy flow in a random porous medium, described mathematically by elliptic partial differential equations with random coefficients. The levels in the multilevel estimator are defined by finite volume discretisations of the governing equations with different mesh parameters. To simulate different layers in the subsurface, the permeability is modelled as a piecewise constant or piecewise spatially correlated random field, including the possibility of piecewise log-normal random fields. The location of the layers is assumed unknown, and modelled by a random process. We prove new convergence results of the spatial discretisation error required to quantify the mean square error of the multilevel estimator, and provide an optimal implementation of the method based on algebraic multigrid methods and a novel variance reduction technique termed Coarse Grid Variates.