Class Modelling

Domain: Sensori dan Riset Konsumen · SQalytics · SIMCA one-class classification + Q-residual + Hotelling T² + Coomans plot untuk authentication dan adulteration detection

1 Introduksi

1.1 Latar Belakang

Class modelling atau klasifikasi satu-kelas berbeda dari analisis diskriminan konvensional. Alih-alih memaksa setiap sampel masuk ke salah satu dari K kelas, class modelling membangun model untuk setiap kelas lalu mengevaluasi apakah sampel baru masih cocok dengan kelas tersebut atau menjadi outlier. Konsekuensinya, satu sampel dapat masuk ke beberapa kelas, tidak masuk kelas mana pun, atau masuk ke semua kelas; sifat ini sangat berguna untuk autentikasi, deteksi outlier, dan screening adulteran (Wold & Sjöström, 1977; De Maesschalck et al., 2000). SIMCA (Soft Independent Modelling of Class Analogy) (Wold, 1976) adalah metode klasik class modelling berbasis PCA per kelas dan metrik jarak seperti Q-residual serta T²-Hotelling. Aplikasinya mencakup autentikasi asal geografis produk, deteksi pemalsuan, atau klasifikasi batch ideal dibanding batch anomali.

1.2 Tujuan Modul

1.3 Posisi di Antara Alternatif

Pilih Class Modelling untuk klasifikasi satu-kelas atau klasifikasi lunak. Untuk diskriminasi keras berbasis LDA, gunakan E-nose / E-tongue Bridge. Untuk pendekatan berbasis regresi, gunakan PLSR Studio. Untuk clustering tidak-terawasi, gunakan Cluster Analysis Explorer (segera hadir).

2 Metode

2.1 Dasar Teoretis

SIMCA algorithm (Wold, 1976; Wold & Sjöström, 1977):

Per class $c$:

  1. Subset training data from class $c$: $X_c$ (matrix $n_c \times p$).
  2. Autoscale $X_c$ (mean-center + unit variance) — note: scaling per-class, not global.
  3. PCA: $X_c = T_c P_c^T + E_c$ with $A_c$ optimal components (via cross-validation $Q^2$).
  4. Calculate threshold: Q-residual threshold (orthogonal distance): $Q_{\alpha, c}$ from $F$-distribution. $T^2$-Hotelling threshold (Mahalanobis-like): $$T^2_{\alpha, c} = \frac{A_c (n_c^2 - 1)}{n_c (n_c - A_c)} F_{\alpha, A_c, n_c - A_c}$$

Prediction for new sample $x_{\text{new}}$:

  1. Project to PCA of each class: $t_{\text{new}} = x_{\text{new}} P_c$.
  2. Calculate Q-residual: $Q_{\text{new}} = \|x_{\text{new}} - t_{\text{new}} P_c^T\|^2$.
  3. Calculate $T^2$: $T^2_{\text{new}} = t_{\text{new}}^T S_t^{-1} t_{\text{new}}$ with $S_t$ scores covariance.
  4. Decision: if $Q_{\text{new}} \leq Q_{\alpha}$ AND $T^2_{\text{new}} \leq T^2_{\alpha}$ → belong to class $c$.

Coomans plot: scatter $T^2$ vs $Q$ with thresholds drawn — visualization of acceptance region.

Sensitivity and Specificity (model validation):

$$\text{Sensitivity} = \frac{\text{true positives}}{\text{actual positives}}$$ $$\text{Specificity} = \frac{\text{true negatives}}{\text{actual negatives}}$$

Target for authentication: ≥ 90% both.

2.2 Persamaan Inti

2.3 Asumsi & Batas Validitas

AsumsiKonsekuensi jika dilanggarCara cek di SQalytics
Training set representative classMisclassificationCek CV accuracy
$n_c \geq 20$ per classPCA tidak stableModul flag
Number PC dipilih via CVOverfitPakai $Q^2$ max
Variables relevant ($p$ tidak terlalu banyak)Curse of dimensionalityVariable selection
Class within-variance homogeneousThreshold biasCek Levene test
Outlier training data identifiedSkew PCAPre-screen outlier

3 Cara Kerja

3.1 Step-by-Step di SQalytics

  1. Buka Class Modelling dari domain Sensori dan Riset Konsumen.
  2. Muat training set: kolom Class + features (mis. dari e-nose, NIR, QDA).
  3. Muat test set (opsional, tanpa class label).
  4. Pilih method: SIMCA (default).
  5. Pre-process: autoscaling per class.
  6. Atur CV strategy: leave-one-out, 5-fold, atau 10-fold.
  7. Klik Run Class Modelling.
  8. Tinjau hasil: Tab Per-Class PCA Summary, Tab Coomans Plot, Tab Sensitivity / Specificity, Tab Test Set Predictions.

3.2 Template Tabel Input + Contoh Data Sintetis

Training set (kolom: Sample, Class, S1–S8):

SampleClassS1S2S3S4S5S6S7S8
K001Aceh250380145520290410180320
K002Aceh245385142525285415178318
K003Toraja290320180480340380210290
… (20 Aceh, 18 Toraja, 22 Gayo training samples)

Test set (kolom: Sample, S1–S8, tanpa Class):

SampleS1S2S3S4S5S6S7S8
U001248382143522287412179319
U002292322182482338382212288
U003238408133552272438167348
U004310290210450390350250260
SYNTHETIC SIMCA class modelling untuk authentication 3 origin kopi Indonesia (Aceh, Toraja, Gayo) dari profil e-nose 8-sensor. Training: 60 samples (20 per class). Test: 4 samples termasuk 1 outlier potensial. CSV setara: docs/assets/example-data/id/sensory/template_sensory_class_simca.csv.

3.3 Contoh Luaran

Per-class SIMCA summary:

Class$n_c$Optimal $A_c$Explained variance$Q_\alpha$$T^2_\alpha$
Aceh20288%2508.5
Toraja18291%1808.8
Gayo22393%2209.2

Test set predictions:

Sample$Q$ Aceh$T^2$ Aceh$Q$ Toraja$T^2$ TorajaClass membership
U0011203.285025.1Aceh (within)
U00275022.4954.1Toraja
U00358018.292028.5Gayo
U00462019.871022.0OUTLIER (not in any class)

Validation metrics:

ClassSensitivitySpecificity
Aceh95%98%
Toraja92%96%
Gayo100%95%
Coomans plot SIMCA class modelling 3 origin kopi Indonesia dengan validation metrics
Gambar 1. (a) Coomans plot untuk model kelas Aceh: titik training (biru=Aceh, teal=Toraja, amber=Gayo) dan test samples (diamond). Garis putus-putus = threshold $Q_\alpha$ (250) dan $T^2_\alpha$ (8.5). U001 berada di dalam region Aceh; U004 di luar semua threshold (outlier). (b) Validation: sensitivity 92–100% dan specificity 95–98% untuk 3 kelas — excellent untuk authentication task.
SIMCA class modelling untuk 3 origin kopi mencapai sensitivity 92–100% dan specificity 95–98% — excellent untuk authentication. Test set: U001 = Aceh, U002 = Toraja, U003 = Gayo (semua confident). U004 OUTLIER — $Q$ dan $T^2$ di atas threshold semua kelas — kemungkinan adulteration, unknown origin, atau batch anomaly. Action: investigate U004 dengan analytical lanjutan (GC-MS, isotope). Untuk QC routine, U004 reject dari shipment. Lanjut ke PLSR Studio untuk quantification adulterant fraction bila pattern match adulterant known.

4 Kesimpulan

4.1 Relevansi Real-World

4.2 Where to Go from Here

Troubleshooting Cepat

Sensitivity < 80%. Tambah training samples; tune $A_c$.
Specificity low (banyak false positives). Threshold $Q$ atau $T^2$ terlalu loose; pakai $\alpha = 0.01$.
Outlier semua sample. Possible scale mismatch — re-autoscale.

i Riwayat Revisi

TanggalRevisiPenulis
2026-05-12Draft v2 publikasi (KaTeX SIMCA + Q-residual + $T^2$ + Coomans + APA Wold/De Maesschalck/Brereton)Claude

4 Referensi

  • Wold, S. (1976). Pattern recognition by means of disjoint principal components models. Pattern Recognition, 8(3), 127–139. https://doi.org/10.1016/0031-3203(76)90014-5
  • Wold, S., & Sjöström, M. (1977). SIMCA: A method for analyzing chemical data. In Chemometrics: Theory and application (Vol. 52, pp. 243–282). American Chemical Society.
  • De Maesschalck, R., Jouan-Rimbaud, D., & Massart, D. L. (2000). The Mahalanobis distance. Chemometrics and Intelligent Laboratory Systems, 50(1), 1–18. https://doi.org/10.1016/S0169-7439(99)00047-7
  • Brereton, R. G. (2009). Chemometrics for pattern recognition. John Wiley & Sons.
  • Vanden Branden, K., & Hubert, M. (2005). Robust classification in high dimensions. Chemometrics and Intelligent Laboratory Systems, 79(1–2), 10–21.
  • Forina, M., Casale, M., & Oliveri, P. (2009). Application of chemometrics to food chemistry. In Comprehensive chemometrics (Vol. 4, pp. 75–128). Elsevier.