SDASJun 18

PolSeT: Polish Semantics of Timbre Dataset

arXiv:2606.199877.1
Predicted impact top 65% in SD · last 90 daysOriginality Synthesis-oriented
AI Analysis

This dataset fills a gap in open timbre research data for Polish, supporting multilingual semantic embedding models and psychoacoustic studies.

PolSeT provides a Polish-language dataset for timbre semantics, including a lexicon of 701 unique descriptors from free-verbalization and ratings of 18 instrument sounds on 8 bipolar scales, enabling cross-cultural psychoacoustic and MIR research.

This data report introduces PolSeT (Polish Semantic Timbre), a dataset designed to facilitate research in psychoacoustics and Music Information Retrieval (MIR) in Polish and cross-cultural contexts. The dataset contains data from two sequential experiments. Experiment 1 (N=60) was a free-verbalization task aimed at creating a lexicon of Polish semantic descriptors. Using 11 stimuli, a total of 1901 descriptors (701 unique) were gathered. Experiment 2 (N=105) utilized this lexicon to conduct a semantic differential study, where participants rated 18 instrument sounds on 8 bipolar scales, with repeated trials for reliability analysis. The released dataset includes raw listener responses, comprehensive demographics (experience, gender, age), audio stimuli, and extracted acoustic features with Python extraction code. This dataset addresses a gap in open timbre research data, providing both the qualitative linguistic groundwork and the quantitative ratings necessary for psychoacoustic research and the training of multilingual semantic embedding models.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes