Jiyoung Park

CV
h-index26
3papers
17citations
Novelty32%
AI Score31

3 Papers

3.6CVNov 28, 2025
Leveraging Textual Compositional Reasoning for Robust Change Captioning

Kyu Ri Park, Jiyoung Park, Seong Tae Kim et al.

Change captioning aims to describe changes between a pair of images. However, existing works rely on visual features alone, which often fail to capture subtle but meaningful changes because they lack the ability to represent explicitly structured information such as object relationships and compositional semantics. To alleviate this, we present CORTEX (COmpositional Reasoning-aware TEXt-guided), a novel framework that integrates complementary textual cues to enhance change understanding. In addition to capturing cues from pixel-level differences, CORTEX utilizes scene-level textual knowledge provided by Vision Language Models (VLMs) to extract richer image text signals that reveal underlying compositional reasoning. CORTEX consists of three key modules: (i) an Image-level Change Detector that identifies low-level visual differences between paired images, (ii) a Reasoning-aware Text Extraction (RTE) module that use VLMs to generate compositional reasoning descriptions implicit in visual features, and (iii) an Image-Text Dual Alignment (ITDA) module that aligns visual and textual features for fine-grained relational reasoning. This enables CORTEX to reason over visual and textual features and capture changes that are otherwise ambiguous in visual features alone.

7.5IRJun 27, 2019
Representation Learning of Music Using Artist, Album, and Track Information

Jongpil Lee, Jiyoung Park, Juhan Nam

Supervised music representation learning has been performed mainly using semantic labels such as music genres. However, annotating music with semantic labels requires time and cost. In this work, we investigate the use of factual metadata such as artist, album, and track information, which are naturally annotated to songs, for supervised music representation learning. The results show that each of the metadata has individual concept characteristics, and using them jointly improves overall performance.

4.1SDJul 24, 2018
A Hybrid of Deep Audio Feature and i-vector for Artist Recognition

Jiyoung Park, Donghyun Kim, Jongpil Lee et al.

Artist recognition is a task of modeling the artist's musical style. This problem is challenging because there is no clear standard. We propose a hybrid method of the generative model i-vector and the discriminative model deep convolutional neural network. We show that this approach achieves state-of-the-art performance by complementing each other. In addition, we briefly explain the advantages and disadvantages of each approach.