Cross-dataset generalization of EEG-based classification under weak, proxy-derived labels remains an open problem for altered-states research. We present a reproducible eight-dataset alignment pipeline that maps eight heterogeneous EEG corpora (712,832 windows; 697,906 with valid labels) to a common 14-channel EPOC+ montage with 63-dimensional spectral features, and we recover the real 1-9 arousal self-assessments for MAHNOB-HCI from session.xml metadata. As a benchmark, Random Forest classifiers are trained on seven source domains and evaluated on the held-out target under both zero-shot and 20%-participant few-shot calibration. The benchmark exposes two concrete methodological pitfalls rather than a performance result: (i) per-class recall shows every target collapsing to a single majority class, and (ii) a within-dataset upper-bound experiment (Table 3) shows that six of eight proxy label sets sit at or below three-class chance even when trained and tested on the same dataset, so the cross-dataset failure is a label-validity problem rather than a transfer-method problem. Across the eight targets (20 seeds, 8,000 evaluation windows per target), zero-shot accuracy averages 36.85% (95% CI 34.40-39.30) and calibrated 43.76% (41.77-45.75), but balanced accuracy stays at 33.01-35.62% (Cohen's kappa <= 0.068), i.e. at chance. The +6.91pp mean change is driven almost entirely by a single target, ds006437 (6.31% -> 60.60%): the median paired change across all 160 seed-pairs is 0.00pp, and after Holm-Bonferroni correction only ds006437 and ds004572 remain significant, the latter with a practically null effect (+0.39pp). Balanced accuracy stays between 33.01% and 35.62% and Cohen's kappa at 0.009 +/- 0.032, i.e. at or barely above three-class chance, while per-class recall shows six of eight targets collapsing to Deep (96.7-100% recall) and two to Light (68.5-99.1%). The collapse persists under SMOTE oversampling, under an EEGNet-v4 deep-learning baseline, and under CORAL and AdaBN feature alignment, which locates the bottleneck in proxy-label validity and class overlap in the feature space rather than in classifier capacity. We position this work as a preliminary methodological study: its contribution is a reproducible eight-dataset alignment pipeline, recovered MAHNOB-HCI arousal self-assessments, a quantitative estimate of split-leakage inflation, and a transparently reported negative result rather than a performance claim.
Weng, Z., Jung, M.
Advertisement
Stats
- Recommendations n/a n/a positive of 0 vote(s)
- Views 21
- Comments 0
