Jie Tang

IR
h-index19
4papers
149citations
Novelty29%
AI Score25

4 Papers

17.0LGFeb 27, 2024
Does Negative Sampling Matter? A Review with Insights into its Theory and Applications

Zhen Yang, Ming Ding, Tinglin Huang et al. · tsinghua

Negative sampling has swiftly risen to prominence as a focal point of research, with wide-ranging applications spanning machine learning, computer vision, natural language processing, data mining, and recommender systems. This growing interest raises several critical questions: Does negative sampling really matter? Is there a general framework that can incorporate all existing negative sampling methods? In what fields is it applied? Addressing these questions, we propose a general framework that leverages negative sampling. Delving into the history of negative sampling, we trace the development of negative sampling through five evolutionary paths. We dissect and categorize the strategies used to select negative sample candidates, detailing global, local, mini-batch, hop, and memory-based approaches. Our review categorizes current negative sampling methods into five types: static, hard, GAN-based, Auxiliary-based, and In-batch methods, providing a clear structure for understanding negative sampling. Beyond detailed categorization, we highlight the application of negative sampling in various areas, offering insights into its practical benefits. Finally, we briefly discuss open problems and future directions for negative sampling.

12.1IRApr 21, 2018
Expert Finding in Community Question Answering: A Review

Sha Yuan, Yu Zhang, Jie Tang et al.

The rapid development recently of Community Question Answering (CQA) satisfies users quest for professional and personal knowledge about anything. In CQA, one central issue is to find users with expertise and willingness to answer the given questions. Expert finding in CQA often exhibits very different challenges compared to traditional methods. Sparse data and new features violate fundamental assumptions of traditional recommendation systems. This paper focuses on reviewing and categorizing the current progress on expert finding in CQA. We classify all the existing solutions into four different categories: matrix factorization based models (MF-based models), gradient boosting tree based models (GBT-based models), deep learning based models (DL-based models) and ranking based models (R-based models). We find that MF-based models outperform other categories of models in the field of expert finding in CQA. Moreover, we use innovative diagrams to clarify several important concepts of ensemble learning, and find that ensemble models with several specific single models can further boosting the performance. Further, we compare the performance of different models on different types of matching tasks, including text vs. text, graph vs. text, audio vs. text and video vs. text. The results can help the model selection of expert finding in practice. Finally, we explore some potential future issues in expert finding research in CQA.

1.7AIOct 13, 2017
Fast Top-k Area Topics Extraction with Knowledge Base

Fang Zhang, Xiaochen Wang, Jingfei Han et al.

What are the most popular research topics in Artificial Intelligence (AI)? We formulate the problem as extracting top-$k$ topics that can best represent a given area with the help of knowledge base. We theoretically prove that the problem is NP-hard and propose an optimization model, FastKATE, to address this problem by combining both explicit and latent representations for each topic. We leverage a large-scale knowledge base (Wikipedia) to generate topic embeddings using neural networks and use this kind of representations to help capture the representativeness of topics for given areas. We develop a fast heuristic algorithm to efficiently solve the problem with a provable error bound. We evaluate the proposed model on three real-world datasets. Experimental results demonstrate our model's effectiveness, robustness, real-timeness (return results in $<1$s), and its superiority over several alternative methods.

2.2IRMar 13, 2017
Multiple User Context Inference by Fusing Data Sources

Jinliang Xu, Shangguang Wang, Fangchun Yang et al.

Inference of user context information, including user's gender, age, marital status, location and so on, has been proven to be valuable for building context aware recommender system. However, prevalent existing studies on user context inference have two shortcommings: 1. focusing on only a single data source (e.g. Internet browsing logs, or mobile call records), and 2. ignoring the interdependence of multiple user contexts (e.g. interdependence between age and marital status), which have led to poor inference performance. To solve this problem, in this paper, we first exploit tensor outer product to fuse multiple data sources in the feature space to obtain an extensional user feature representation. Following this, by taking this extensional user feature representation as input, we propose a multiple attribute probabilistic model called MulAProM to infer user contexts that can take advantage of the interdependence between them. Our study is based on large telecommunication datasets from the local mobile operator of Shanghai, China, and consists of two data sources, 4.6 million call detail records and 7.5 million data traffic records of 8,000 mobile users, collected in the course of six months. The experimental results show that our model can outperform other models in terms of \emph{recall}, \emph{precision}, and the \emph{F1-measure}.