1 Introduksi
1.1 Latar Belakang
Partial Least Squares Regression (PLSR) adalah teknik multivariate yang memodelkan hubungan antara matriks prediktor $X$ (banyak kolom) dan respons $y$ dengan memproyeksikan keduanya ke ruang laten yang memaksimumkan kovariansi $X$-$y$. Diperkenalkan oleh Herman Wold (1966) dan dipopulerkan oleh Geladi & Kowalski (1986) untuk chemometrics, PLSR telah menjadi standar di spektroskopi NIR/UV-Vis, sensoris-instrumental linkage, omics, dan process analytical technology (PAT) (Martens & Næs, 1989; Eriksson, Byrne, Johansson, Trygg, & Wikström, 2013).
Keunggulan PLSR atas OLS multiple regression: (i) menangani multicollinearity secara elegant, (ii) menangani $p \gg n$ (lebih banyak prediktor daripada observasi), (iii) interpretasi via VIP scores untuk identifikasi prediktor kunci, dan (iv) biplot loadings untuk visualisasi struktur hubungan $X$-$Y$ dalam 2D atau 3D.
PLSR Studio adalah desk expert yang mengintegrasikan semua aspek PLSR di satu workspace: pilihan jumlah komponen, cross-validation, VIP ranking, dan correlation loading biplot. Modul ini melengkapi Regression Studio (single-X) dan Predict One Result from Many Variables (multi-X OLS) untuk kasus dengan prediktor banyak + multicollinearity.
1.2 Tujuan Modul
- Memodelkan satu respons $Y$ numerik terhadap banyak prediktor $X$ (≥ 2) dengan PLS.
- Memilih jumlah komponen laten optimal via cross-validation (LOOCV atau 10-fold).
- Menghitung VIP scores untuk ranking variable importance (Wold, Sjöström, & Eriksson, 2001).
- Menampilkan correlation loading biplot 2D atau 3D.
- Menampilkan metrik model: $R^2$, $Q^2$ (predictive $R^2$ dari CV), RMSEC, RMSECV.
- Multi-mode preset: Screening view, Focus top features, Explore in 3D.
- Audiens: mahasiswa S2/S3 chemometrics, peneliti NIR/UV-Vis quantification, R&D preference mapping.
1.3 Posisi di Antara Alternatif
Pilih PLSR Studio ketika prediktor $X$ banyak (≥ 5) + multicollinearity tinggi. Untuk satu prediktor, pakai Regression Studio. Untuk multi-prediktor tanpa multicollinearity, pakai Predict One Result from Many Variables. Untuk eksplorasi unsupervised, gunakan PCA Explorer (akan datang). Untuk uji statistik perbedaan kelompok, gunakan ANOVA Studio.
2 Metode
2.1 Dasar Teoretis
Setup PLSR — $X$ ukuran $n \times p$ (autoscaled) dan $y$ ukuran $n \times 1$. PLSR mendekomposisi:
$$ X = T P^T + E, \quad y = T q + f $$$T$ = scores ($n \times A$, $A$ komponen), $P$ = X-loadings ($p \times A$), $q$ = Y-loadings, $E$ dan $f$ residual. Krusial: $T$ dipilih untuk memaksimumkan kovariansi $\text{cov}(X w, y)$, bukan hanya variansi $X$ (seperti PCR/PCA).
Algoritma NIPALS (Wold, 1966) — iterative deflation per komponen $a$:
- Initialize $u_a$ random.
- $w_a = X^T u_a / \|X^T u_a\|$ (weights).
- $t_a = X w_a$ (scores).
- $q_a = y^T t_a / (t_a^T t_a)$ (Y-loading).
- $u_a = y q_a / (q_a^T q_a)$.
- Iterasi 2–5 hingga konvergen.
- $p_a = X^T t_a / (t_a^T t_a)$ (X-loading).
- Deflasi: $X \leftarrow X - t_a p_a^T$, $y \leftarrow y - t_a q_a$.
- Ulangi untuk komponen $a+1$.
Predicted Y: $\hat{y} = X B_{\text{PLS}}$ dengan $B_{\text{PLS}} = W (P^T W)^{-1} q$.
Cross-validation metrik — k-fold CV (default 10-fold):
$\hat{y}_{i,-i}$ adalah prediksi $y_i$ tanpa observasi $i$. Komponen optimal dipilih untuk memaksimumkan $Q^2$ atau minimumkan RMSECV:
$$ \text{RMSECV} = \sqrt{\sum (y_i - \hat{y}_{i,-i})^2 / n} $$Rule of thumb (Eriksson et al., 2013): $Q^2 > 0.5$ acceptable, $> 0.7$ good, $> 0.9$ excellent.
VIP score (Wold, Sjöström, & Eriksson, 2001):
$$ \text{VIP}_j = \sqrt{\frac{p \sum_{a=1}^{A} \text{SS}_a (w_{aj} / \|w_a\|)^2}{\sum_{a=1}^{A} \text{SS}_a}} $$Threshold: VIP ≥ 1 signifikan; VIP < 1 kontribusi marginal.
Correlation loading biplot — plot $X$-loadings dan $y$-loading di ruang komponen 1 vs 2. Posisi loading pada inner circle (radius 0.5) berarti variabel poorly explained; outer circle (radius 1.0) berarti variabel well explained. Sudut antar loading menunjukkan korelasi: $0°$ positif kuat, $90°$ tidak berkorelasi, $180°$ negatif kuat.
Bahaya overfitting: terlalu banyak komponen menghasilkan $R^2$ tinggi tetapi $Q^2$ rendah. Decision rule (Eriksson et al., 2013): stop ketika $\Delta Q^2 < 0.05$ atau $Q^2$ mulai menurun.
2.2 Persamaan Inti
Dekomposisi: $X = T P^T + E$, $y = T q + f$
Koefisien PLS: $B_{\text{PLS}} = W (P^T W)^{-1} q$
$Q^2$ predictive: $1 - \text{PRESS} / \text{SS}_y$
RMSECV: $\sqrt{\sum (y_i - \hat{y}_{i,-i})^2 / n}$
VIP score: $\sqrt{p \cdot \sum_a \text{SS}_a (w_{aj}/\|w_a\|)^2 / \sum_a \text{SS}_a}$
Thresholds: $Q^2 \geq 0.5$ acceptable, VIP $\geq 1$ signifikan.
2.3 Asumsi & Batas Validitas
| Asumsi | Konsekuensi jika dilanggar | Cara cek di SQalytics |
|---|---|---|
| Hubungan linier $X$-$Y$ (ruang laten) | Bias prediksi | Cek residual plot |
| Autoscaling sebelum NIPALS | Bias ke variabel skala besar | Modul auto-center dan scale |
| Tidak ada nilai konstan di $X$ | Variabel tidak kontribusi | Modul auto-drop konstan |
| $n \geq 3$ observasi lengkap | Algoritma fail | Modul flag warning |
| Komponen optimal via CV | Overfit | Modul default: pilih $A$ maks $Q^2$ |
| Missing value minimal | Bias | Imputation atau drop kolom |
| Y tidak konstan | $Q^2$ undefined | Modul flag varian Y |
| VIP threshold = 1 (standar) | Tidak universal | Sesuaikan dengan domain |
3 Cara Kerja
3.1 Step-by-Step di SQalytics
- Buka
PLSR Studiodari domain Statistika Terapan. - (Opsional) Muat seed
stats_group_compare. - Pada
DJ Entry Panel, pilih preset:Screening view,Focus top features, atauExplore in 3D. - Pada panel
PLSR Correlation Analysis, sesuaikan: Response (Y), Predictor Variables (X) (minimal 2), Components (auto-pilih via $Q^2$). - (Opsional) Buka
Visualization Settings: Plot Dimensions (2D/3D), Top N Features (by VIP). - Klik Run from DJ Entry Panel atau Run Analysis.
- Tinjau hasil: conclusion → Correlation Loading Biplot → Top Features → Detailed Feature Importance → Model Summary.
- Klik Download Report.
3.2 Template Tabel Input + Contoh Data Sintetis
| Kolom | Tipe | Wajib | Catatan |
|---|---|---|---|
<Y> | numeric | ✓ | Response variable (1 kolom) |
<X1>, <X2>, ... | numeric | ✓ (≥2) | Predictor variables |
Contoh data sintetis (15 baris — Yield target vs 5 prediktor proses biskuit):
| Batch | Moisture | Protein | Texture | ColorScore | Sugar | Yield |
|---|---|---|---|---|---|---|
| B01 | 12.5 | 11.2 | 7.4 | 8.5 | 18.0 | 85.2 |
| B02 | 12.8 | 11.4 | 7.6 | 8.6 | 18.3 | 86.1 |
| B03 | 12.4 | 11.1 | 7.2 | 8.4 | 17.8 | 84.8 |
| B04 | 12.7 | 11.3 | 7.5 | 8.7 | 18.2 | 85.9 |
| B05 | 12.6 | 11.5 | 7.3 | 8.5 | 18.0 | 85.5 |
| B06 | 12.9 | 11.6 | 7.8 | 8.8 | 18.5 | 86.4 |
| B07 | 12.3 | 11.0 | 7.1 | 8.3 | 17.7 | 84.5 |
| … (15 batch total) | ||||||
docs/assets/example-data/id/statistics/template_stats_groups.csv.
3.3 Contoh Luaran
Tabel Model Summary (PLSR untuk Yield):
| Komponen | $R^2_X$ cum. | $R^2_Y$ cum. | $Q^2$ cum. | RMSEC | RMSECV |
|---|---|---|---|---|---|
| 1 | 0.892 | 0.946 | 0.871 | 0.148 | 0.231 |
| 2 | 0.965 | 0.972 | 0.882 | 0.107 | 0.222 |
| 3 | 0.991 | 0.978 | 0.876 (↓) | 0.094 | 0.227 |
Decision: Komponen optimal = 1 komponen ($\Delta Q^2 < 0.05$ untuk komponen 2; komponen 3 menurunkan $Q^2$ → overfit).
Tabel Top Features (by VIP) dengan 1 komponen:
| Variable | VIP | Coefficient $B_{\text{PLS}}$ | Direction | Rank |
|---|---|---|---|---|
| Texture | 1.082 | +0.282 | Positive | 1 |
| Protein | 1.071 | +0.279 | Positive | 2 |
| Moisture | 1.053 | +0.275 | Positive | 3 |
| Sugar | 1.039 | +0.271 | Positive | 4 |
| ColorScore | 1.025 | +0.267 | Positive | 5 |
Semua VIP > 1 — semua prediktor signifikan. Koefisien hampir identik (0.27–0.28) — konsisten dengan multicollinearity ekstrem, PLS mendistribusikan kontribusi merata.
Regression Studio untuk uji single-predictor cross-check."Grafik utama: dua panel — (a) correlation loading biplot 2D: 5 X-loadings + 1 Y-loading dalam ruang LV1 vs LV2; (b) bar chart VIP score 5 prediktor dengan garis ambang VIP = 1.
4 Kesimpulan
4.1 Relevansi Real-World
- NIR/UV-Vis quantification — kalibrasi multi-wavelength → konsentrasi (protein, moisture, lemak).
- Sensometrics — mapping profil instrumental ke acceptance sensoris (external preference mapping).
- Omics integration — gen expression / metabolomics → fenotipe.
- Process Analytical Technology (PAT) — real-time monitoring multi-sensor → quality.
- HPLC fingerprinting — peak profile multi-wavelength → species identification atau quantification.
- Spectroscopy + chromatography fusion — multi-block PLS untuk integrasi instrumen.
- Survey research multivariate — banyak item Likert → composite score.
Pada Rencana Publikasi Singkil v5, modul ini dipakai pada T3 Tahap 3 untuk hubungan profil HPLC fingerprint → activity antimikroba, dan T5 Tahap 3 untuk preference mapping (instrumental → sensoris).
4.2 Where to Go from Here
Pembacaan lanjutan:
- Geladi & Kowalski (1986) — paper tutorial klasik PLSR.
- Wold, Sjöström, & Eriksson (2001) — review PLS-regression untuk chemometrics.
- Eriksson et al. (2013) — buku rujukan modern multivariate analysis (SIMCA-style).
- Martens & Næs (1989) — multivariate calibration.
⚙ Troubleshooting Cepat
i Riwayat Revisi
| Tanggal | Revisi | Penulis |
|---|---|---|
| 2026-05-12 | Migrasi MD v2 → HTML final dengan figure dual-panel biplot + VIP barplot + caption Elsevier-style | Claude |
| 2026-05-12 | Migrasi v1 → v2 (template publikasi + KaTeX NIPALS/PLS coefficient/VIP/Q² + biplot interpretation + APA Wold/Geladi/Martens/Eriksson) | Claude |
| 2026-05-09 | Draft awal v1 | Tim docs |
4 Referensi
- Wold, H. (1966). Estimation of principal components and related models by iterative least squares. In P. R. Krishnaiah (Ed.), Multivariate analysis (pp. 391–420). Academic Press.
- Geladi, P., & Kowalski, B. R. (1986). Partial least-squares regression: A tutorial. Analytica Chimica Acta, 185, 1–17. https://doi.org/10.1016/0003-2670(86)80028-9
- Martens, H., & Næs, T. (1989). Multivariate calibration. John Wiley & Sons.
- Wold, S., Sjöström, M., & Eriksson, L. (2001). PLS-regression: A basic tool of chemometrics. Chemometrics and Intelligent Laboratory Systems, 58(2), 109–130. https://doi.org/10.1016/S0169-7439(01)00155-1
- Eriksson, L., Byrne, T., Johansson, E., Trygg, J., & Wikström, C. (2013). Multi- and megavariate data analysis: Basic principles and applications (3rd ed.). Umetrics Academy.
- Mevik, B.-H., & Wehrens, R. (2007). The pls package: Principal component and partial least squares regression in R. Journal of Statistical Software, 18(2), 1–23. https://doi.org/10.18637/jss.v018.i02
- Tenenhaus, M. (1998). La régression PLS: Théorie et pratique. Editions Technip.