CLJun 10

Detecting Sensitive Personal Information in Japanese Pre-Training Corpora for Large Language Models

arXiv:2606.12114v115.2h-index: 5
Predicted impact top 66% in CL · last 90 daysOriginality Synthesis-oriented
AI Analysis

For Japanese LLM developers and privacy regulators, this work provides a first-step tool to detect sensitive personal information in pre-training data, though it is incremental as it applies existing methods to a new language.

This study addresses the lack of research on detecting sensitive personal information in Japanese pre-training corpora for LLMs, constructing a dataset and training a classifier that effectively identifies special care-required personal information (SCPI) under Japan's APPI. The classifier achieves effective detection, marking the first exploration of SCPI detection in Japanese text.

Sensitive personal information can appear in large-scale pre-training corpora for large language models (LLMs). Detecting and filtering such information is therefore essential to ensure compliance with privacy regulations and prevent unintended information leakage. However, in contrast to English and other languages, research into sensitive personal information has been limited in the Japanese language. In this study, we focus on sensitive personal data defined as special care-required personal information (SCPI) under Japan's Act on the Protection of Personal Information (APPI). We construct an SCPI dataset using LLM-based annotation and train machine learning models to rapidly detect SCPI in text. As a result, our SCPI classifier can effectively identify information related to SCPI. This study is the first to explore SCPI detection in Japanese text corpora, highlighting the challenges of accurate detection.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes