1 Introduksi
1.1 Latar Belakang
Class modelling atau klasifikasi satu-kelas berbeda dari analisis diskriminan konvensional. Alih-alih memaksa setiap sampel masuk ke salah satu dari K kelas, class modelling membangun model untuk setiap kelas lalu mengevaluasi apakah sampel baru masih cocok dengan kelas tersebut atau menjadi outlier. Konsekuensinya, satu sampel dapat masuk ke beberapa kelas, tidak masuk kelas mana pun, atau masuk ke semua kelas; sifat ini sangat berguna untuk autentikasi, deteksi outlier, dan screening adulteran (Wold & Sjöström, 1977; De Maesschalck et al., 2000). SIMCA (Soft Independent Modelling of Class Analogy) (Wold, 1976) adalah metode klasik class modelling berbasis PCA per kelas dan metrik jarak seperti Q-residual serta T²-Hotelling. Aplikasinya mencakup autentikasi asal geografis produk, deteksi pemalsuan, atau klasifikasi batch ideal dibanding batch anomali.
1.2 Tujuan Modul
- Menerima data latih dengan label kelas dan data uji tanpa label.
- Membangun model SIMCA: PCA per kelas dengan jumlah komponen optimal melalui cross-validation.
- Mengukur Q-residual (jarak ortogonal) dan Hotelling $T^2$ (jarak di dalam model).
- Menampilkan Coomans plot ($T^2$ vs $Q$) untuk visualisasi keanggotaan kelas.
- Mendukung deteksi adulterasi, yaitu sampel yang menjadi outlier dari semua kelas yang dikenal.
- Audiens: peneliti autentikasi pangan, tim QC industri, dan peneliti kemometrika.
1.3 Posisi di Antara Alternatif
Pilih Class Modelling untuk klasifikasi satu-kelas atau klasifikasi lunak. Untuk diskriminasi keras berbasis LDA, gunakan E-nose / E-tongue Bridge. Untuk pendekatan berbasis regresi, gunakan PLSR Studio. Untuk clustering tidak-terawasi, gunakan Cluster Analysis Explorer (segera hadir).
2 Metode
2.1 Dasar Teoretis
SIMCA algorithm (Wold, 1976; Wold & Sjöström, 1977):
Per class $c$:
- Subset training data from class $c$: $X_c$ (matrix $n_c \times p$).
- Autoscale $X_c$ (mean-center + unit variance) — note: scaling per-class, not global.
- PCA: $X_c = T_c P_c^T + E_c$ with $A_c$ optimal components (via cross-validation $Q^2$).
- Calculate threshold: Q-residual threshold (orthogonal distance): $Q_{\alpha, c}$ from $F$-distribution. $T^2$-Hotelling threshold (Mahalanobis-like): $$T^2_{\alpha, c} = \frac{A_c (n_c^2 - 1)}{n_c (n_c - A_c)} F_{\alpha, A_c, n_c - A_c}$$
Prediction for new sample $x_{\text{new}}$:
- Project to PCA of each class: $t_{\text{new}} = x_{\text{new}} P_c$.
- Calculate Q-residual: $Q_{\text{new}} = \|x_{\text{new}} - t_{\text{new}} P_c^T\|^2$.
- Calculate $T^2$: $T^2_{\text{new}} = t_{\text{new}}^T S_t^{-1} t_{\text{new}}$ with $S_t$ scores covariance.
- Decision: if $Q_{\text{new}} \leq Q_{\alpha}$ AND $T^2_{\text{new}} \leq T^2_{\alpha}$ → belong to class $c$.
Coomans plot: scatter $T^2$ vs $Q$ with thresholds drawn — visualization of acceptance region.
Sensitivity and Specificity (model validation):
$$\text{Sensitivity} = \frac{\text{true positives}}{\text{actual positives}}$$ $$\text{Specificity} = \frac{\text{true negatives}}{\text{actual negatives}}$$Target for authentication: ≥ 90% both.
2.2 Persamaan Inti
- Per-class PCA: $X_c = T_c P_c^T + E_c$
- $T^2$ threshold: $T^2_{\alpha} = \frac{A(n^2-1)}{n(n-A)} F_{\alpha, A, n-A}$
- Q-residual: $Q = \|x - t P^T\|^2$
- Decision rule: $Q \leq Q_\alpha$ AND $T^2 \leq T^2_\alpha$ → belong
- Sensitivity / Specificity for validation
2.3 Asumsi & Batas Validitas
| Asumsi | Konsekuensi jika dilanggar | Cara cek di SQalytics |
|---|---|---|
| Training set representative class | Misclassification | Cek CV accuracy |
| $n_c \geq 20$ per class | PCA tidak stable | Modul flag |
| Number PC dipilih via CV | Overfit | Pakai $Q^2$ max |
| Variables relevant ($p$ tidak terlalu banyak) | Curse of dimensionality | Variable selection |
| Class within-variance homogeneous | Threshold bias | Cek Levene test |
| Outlier training data identified | Skew PCA | Pre-screen outlier |
3 Cara Kerja
3.1 Step-by-Step di SQalytics
- Buka
Class Modellingdari domain Sensori dan Riset Konsumen. - Muat training set: kolom
Class+ features (mis. dari e-nose, NIR, QDA). - Muat test set (opsional, tanpa class label).
- Pilih method: SIMCA (default).
- Pre-process: autoscaling per class.
- Atur CV strategy: leave-one-out, 5-fold, atau 10-fold.
- Klik Run Class Modelling.
- Tinjau hasil: Tab
Per-Class PCA Summary, TabCoomans Plot, TabSensitivity / Specificity, TabTest Set Predictions.
3.2 Template Tabel Input + Contoh Data Sintetis
Training set (kolom: Sample, Class, S1–S8):
| Sample | Class | S1 | S2 | S3 | S4 | S5 | S6 | S7 | S8 |
|---|---|---|---|---|---|---|---|---|---|
| K001 | Aceh | 250 | 380 | 145 | 520 | 290 | 410 | 180 | 320 |
| K002 | Aceh | 245 | 385 | 142 | 525 | 285 | 415 | 178 | 318 |
| K003 | Toraja | 290 | 320 | 180 | 480 | 340 | 380 | 210 | 290 |
| … (20 Aceh, 18 Toraja, 22 Gayo training samples) | |||||||||
Test set (kolom: Sample, S1–S8, tanpa Class):
| Sample | S1 | S2 | S3 | S4 | S5 | S6 | S7 | S8 |
|---|---|---|---|---|---|---|---|---|
| U001 | 248 | 382 | 143 | 522 | 287 | 412 | 179 | 319 |
| U002 | 292 | 322 | 182 | 482 | 338 | 382 | 212 | 288 |
| U003 | 238 | 408 | 133 | 552 | 272 | 438 | 167 | 348 |
| U004 | 310 | 290 | 210 | 450 | 390 | 350 | 250 | 260 |
docs/assets/example-data/id/sensory/template_sensory_class_simca.csv.
3.3 Contoh Luaran
Per-class SIMCA summary:
| Class | $n_c$ | Optimal $A_c$ | Explained variance | $Q_\alpha$ | $T^2_\alpha$ |
|---|---|---|---|---|---|
| Aceh | 20 | 2 | 88% | 250 | 8.5 |
| Toraja | 18 | 2 | 91% | 180 | 8.8 |
| Gayo | 22 | 3 | 93% | 220 | 9.2 |
Test set predictions:
| Sample | $Q$ Aceh | $T^2$ Aceh | $Q$ Toraja | $T^2$ Toraja | Class membership |
|---|---|---|---|---|---|
| U001 | 120 | 3.2 | 850 | 25.1 | Aceh (within) |
| U002 | 750 | 22.4 | 95 | 4.1 | Toraja |
| U003 | 580 | 18.2 | 920 | 28.5 | Gayo |
| U004 | 620 | 19.8 | 710 | 22.0 | OUTLIER (not in any class) |
Validation metrics:
| Class | Sensitivity | Specificity |
|---|---|---|
| Aceh | 95% | 98% |
| Toraja | 92% | 96% |
| Gayo | 100% | 95% |
PLSR Studio untuk quantification adulterant fraction bila pattern match adulterant known.4 Kesimpulan
4.1 Relevansi Real-World
- Food authentication — olive oil, honey, coffee origin.
- Pharmaceutical lot release — classify batch to specification model.
- Adulteration detection — milk powder, spice.
- Counterfeit detection — dengan spectroscopy.
- Sensory panel quality — model ideal panelist, flag bad panelist.
4.2 Where to Go from Here
⚙ Troubleshooting Cepat
i Riwayat Revisi
| Tanggal | Revisi | Penulis |
|---|---|---|
| 2026-05-12 | Draft v2 publikasi (KaTeX SIMCA + Q-residual + $T^2$ + Coomans + APA Wold/De Maesschalck/Brereton) | Claude |
4 Referensi
- Wold, S. (1976). Pattern recognition by means of disjoint principal components models. Pattern Recognition, 8(3), 127–139. https://doi.org/10.1016/0031-3203(76)90014-5
- Wold, S., & Sjöström, M. (1977). SIMCA: A method for analyzing chemical data. In Chemometrics: Theory and application (Vol. 52, pp. 243–282). American Chemical Society.
- De Maesschalck, R., Jouan-Rimbaud, D., & Massart, D. L. (2000). The Mahalanobis distance. Chemometrics and Intelligent Laboratory Systems, 50(1), 1–18. https://doi.org/10.1016/S0169-7439(99)00047-7
- Brereton, R. G. (2009). Chemometrics for pattern recognition. John Wiley & Sons.
- Vanden Branden, K., & Hubert, M. (2005). Robust classification in high dimensions. Chemometrics and Intelligent Laboratory Systems, 79(1–2), 10–21.
- Forina, M., Casale, M., & Oliveri, P. (2009). Application of chemometrics to food chemistry. In Comprehensive chemometrics (Vol. 4, pp. 75–128). Elsevier.