CLMar 5

VietJobs: A Vietnamese Job Advertisement Dataset

arXiv:2603.05262v1Has Code
Originality Incremental advance
AI Analysis

This dataset addresses the lack of a large-scale, publicly available Vietnamese job advertisement corpus for NLP researchers and labor market analysts, providing a new benchmark for Vietnamese NLP.

This paper introduces VietJobs, the first large-scale, publicly available dataset of Vietnamese job advertisements, containing 48,092 postings and over 15 million words. It provides extensive linguistic and structured information, including job titles, categories, salaries, skills, and employment conditions across 16 occupational domains. The dataset was used to benchmark generative LLMs on job category classification and salary estimation, with instruction-tuned models like Qwen2.5-7B-Instruct and Llama-SEA-LION-v3-8B-IT showing notable gains.

VietJobs is the first large-scale, publicly available corpus of Vietnamese job advertisements, comprising 48,092 postings and over 15 million words collected from all 34 provinces and municipalities across Vietnam. The dataset provides extensive linguistic and structured information, including job titles, categories, salaries, skills, and employment conditions, covering 16 occupational domains and multiple employment types (full-time, part-time, and internship). Designed to support research in natural language processing and labour market analytics, VietJobs captures substantial linguistic, regional, and socio-economic diversity. We benchmark several generative large language models (LLMs) on two core tasks: job category classification and salary estimation. Instruction-tuned models such as Qwen2.5-7B-Instruct and Llama-SEA-LION-v3-8B-IT demonstrate notable gains under few-shot and fine-tuned settings, while highlighting challenges in multilingual and Vietnamese-specific modelling for structured labour market prediction. VietJobs establishes a new benchmark for Vietnamese NLP and offers a valuable foundation for future research on recruitment language, socio-economic representation, and AI-driven labour market analysis. All code and resources are available at: https://github.com/VinNLP/VietJobs.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes