Ning Wang

CL
h-index36
3papers
94citations
Novelty37%
AI Score34

3 Papers

1.2DBDec 2, 2025
QJoin: Transformation-aware Joinable Data Discovery Using Reinforcement Learning

Ning Wang, Sainyam Galhotra

Discovering which tables in large, heterogeneous repositories can be joined and by what transformations is a central challenge in data integration and data discovery. Traditional join discovery methods are largely designed for equi-joins, which assume that join keys match exactly or nearly so. These techniques, while efficient in clean, well-normalized databases, fail in open or federated settings where identifiers are inconsistently formatted, embedded, or split across multiple columns. Approximate or fuzzy joins alleviate minor string variations but cannot capture systematic transformations. We introduce QJoin, a reinforcement-learning framework that learns and reuses transformation strategies across join tasks. QJoin trains an agent under a uniqueness-aware reward that balances similarity with key distinctiveness, enabling it to explore concise, high-value transformation chains. To accelerate new joins, we introduce two reuse mechanisms: (i) agent transfer, which initializes new policies from pretrained agents, and (ii) transformation reuse, which caches successful operator sequences for similar column clusters. On the AutoJoin Web benchmark (31 table pairs), QJoin achieves an average F1-score of 91.0%. For 19,990 join tasks in NYC+Chicago open datasets, Qjoin reduces runtime by up to 7.4% (13,747 s) by using reusing. These results demonstrate that transformation learning and reuse can make join discovery both more accurate and more efficient.

3.4CLJun 26, 2024
Methodology of Adapting Large English Language Models for Specific Cultural Contexts

Wenjing Zhang, Siqi Xiao, Xuejiao Lei et al.

The rapid growth of large language models(LLMs) has emerged as a prominent trend in the field of artificial intelligence. However, current state-of-the-art LLMs are predominantly based on English. They encounter limitations when directly applied to tasks in specific cultural domains, due to deficiencies in domain-specific knowledge and misunderstandings caused by differences in cultural values. To address this challenge, our paper proposes a rapid adaptation method for large models in specific cultural contexts, which leverages instruction-tuning based on specific cultural knowledge and safety values data. Taking Chinese as the specific cultural context and utilizing the LLaMA3-8B as the experimental English LLM, the evaluation results demonstrate that the adapted LLM significantly enhances its capabilities in domain-specific knowledge and adaptability to safety values, while maintaining its original expertise advantages.

4.3SPAug 27, 2019
Physical Layer Key Generation in 5G Wireless Networks

Long Jiao, Ning Wang, Pu Wang et al.

The bloom of the fifth generation (5G) communication and beyond serves as a catalyst for physical layer key generation techniques. In 5G communications systems, many challenges in traditional physical layer key generation schemes, such as co-located eavesdroppers, the high bit disagreement ratio, and high temporal correlation, could be overcome. This paper lists the key-enabler techniques in 5G wireless networks, which offer opportunities to address existing issues in physical layer key generation. We survey the existing key generation methods and introduce possible solutions for the existing issues. The new solutions include applying the high signal directionality in beamforming to resist co-located eavesdroppers, utilizing the sparsity of millimeter wave (mmWave) channel to achieve a low bit disagreement ratio under low signal-to-noise-ratio (SNR), and exploiting hybrid precoding to reduce the temporal correlation among measured samples. Finally, the future trends of physical layer key generation in 5G and beyond communications are discussed.