Yue Yang

CV
h-index12
5papers
721citations
Novelty53%
AI Score37

5 Papers

20.5CVNov 27, 2024
OpenING: A Comprehensive Benchmark for Judging Open-ended Interleaved Image-Text Generation

Pengfei Zhou, Xiaopeng Peng, Jiajun Song et al.

Multimodal Large Language Models (MLLMs) have made significant strides in visual understanding and generation tasks. However, generating interleaved image-text content remains a challenge, which requires integrated multimodal understanding and generation abilities. While the progress in unified models offers new solutions, existing benchmarks are insufficient for evaluating these methods due to limitations in data size and diversity. To bridge this gap, we introduce OpenING, a comprehensive benchmark comprising 5,400 high-quality human-annotated instances across 56 real-world tasks. OpenING covers diverse daily scenarios such as travel guide, design, and brainstorming, offering a robust platform for challenging interleaved generation methods. In addition, we present IntJudge, a judge model for evaluating open-ended multimodal generation methods. Trained with a novel data pipeline, our IntJudge achieves an agreement rate of 82.42% with human judgments, outperforming GPT-based evaluators by 11.34%. Extensive experiments on OpenING reveal that current interleaved generation methods still have substantial room for improvement. Key findings on interleaved image-text generation are further presented to guide the development of next-generation models.

11.1CVNov 17, 2021
Induce, Edit, Retrieve: Language Grounded Multimodal Schema for Instructional Video Retrieval

Yue Yang, Joongwon Kim, Artemis Panagopoulou et al.

Schemata are structured representations of complex tasks that can aid artificial intelligence by allowing models to break down complex tasks into intermediate steps. We propose a novel system that induces schemata from web videos and generalizes them to capture unseen tasks with the goal of improving video retrieval performance. Our system proceeds in three major phases: (1) Given a task with related videos, we construct an initial schema for a task using a joint video-text model to match video segments with text representing steps from wikiHow; (2) We generalize schemata to unseen tasks by leveraging language models to edit the text within existing schemata. Through generalization, we can allow our schemata to cover a more extensive range of tasks with a small amount of learning data; (3) We conduct zero-shot instructional video retrieval with the unseen task names as the queries. Our schema-guided approach outperforms existing methods for video retrieval, and we demonstrate that the schemata induced by our system are better than those generated by other models.

50.4CVApr 12, 2021Code
Visual Goal-Step Inference using wikiHow

Yue Yang, Artemis Panagopoulou, Qing Lyu et al.

Understanding what sequence of steps are needed to complete a goal can help artificial intelligence systems reason about human activities. Past work in NLP has examined the task of goal-step inference for text. We introduce the visual analogue. We propose the Visual Goal-Step Inference (VGSI) task, where a model is given a textual goal and must choose which of four images represents a plausible step towards that goal. With a new dataset harvested from wikiHow consisting of 772,277 images representing human actions, we show that our task is challenging for state-of-the-art multimodal models. Moreover, the multimodal representation learned from our data can be effectively transferred to other datasets like HowTo100m, increasing the VGSI accuracy by 15 - 20%. Our task will facilitate multimodal reasoning about procedural events.

2.0IVOct 7, 2020
Discriminative Cross-Modal Data Augmentation for Medical Imaging Applications

Yue Yang, Pengtao Xie

While deep learning methods have shown great success in medical image analysis, they require a number of medical images to train. Due to data privacy concerns and unavailability of medical annotators, it is oftentimes very difficult to obtain a lot of labeled medical images for model training. In this paper, we study cross-modality data augmentation to mitigate the data deficiency issue in the medical imaging domain. We propose a discriminative unpaired image-to-image translation model which translates images in source modality into images in target modality where the translation task is conducted jointly with the downstream prediction task and the translation is guided by the prediction. Experiments on two applications demonstrate the effectiveness of our method.

1.2ITMar 4, 2016
OFDM demodulation using virtual time reversal processing in underwater acoustic communication

Yanling Yin, Songzuo Liu, Gang Qiao et al.

The extremely long underwater channel delay spread causes severe inter-symbol interference (ISI) for underwater acoustic communications. Passive time reversal processing (PTRP) can effectively reduce the channel time dispersion in a simple way via convolving the received packet with a time reversed probe signal. However the probe signal itself may introduce extra noise and interference (self-correlation of the probe signal). In this paper, we propose a virtual time reversal processing (VTRP) for single input single output (SISO) Orthogonal Frequency Division Multiplexing (OFDM) systems. It convolves the received packet with the reversed estimated channel, instead of the probe signal to reduce the interference. Two sparse channel estimation methods, matching pursuit (MP), and basis pursuit de-noising (BPDN), are adopted to estimate the channel impulse response (CIR). We compare the performance of VTRP with the PTRP and without any time reversal processing through MATLAB simulations and the pool experiments. The results reveal that VTRP has outstanding performance over time-invariant channels.