1 Introduksi
1.1 Latar Belakang
Near-Infrared Spectroscopy (NIR) (780–2500 nm) memberikan fingerprinting cepat non-destruktif untuk identifikasi dan kuantifikasi composition pangan (protein, moisture, lemak, gula) tanpa sample preparation. NIR mengukur overtone dan combination bands dari ikatan C–H, O–H, N–H — sehingga sensitif terhadap matriks pangan. Sayangnya, spektrum NIR mentah penuh noise: baseline drift, scattering effect dari partikel, dan overlapping bands membuat preprocessing wajib sebelum modeling kuantitatif via PLSR atau PCA (Næs, Isaksson, Fearn, & Davies, 2002; Rinnan, van den Berg, & Engelsen, 2009).
Empat teknik preprocessing paling umum: Standard Normal Variate (SNV) untuk koreksi multiplicative scatter, Multiplicative Scatter Correction (MSC) untuk koreksi mean-shift + scaling, Savitzky-Golay derivative (1st atau 2nd order) untuk reduksi baseline + sharpening peak, dan Savitzky-Golay smoothing untuk reduksi noise. Pemilihan kombinasi yang tepat dapat menggandakan $R^2$ kalibrasi dan menurunkan RMSECV 50% (Rinnan et al., 2009; Pasquini, 2003).
1.2 Tujuan Modul
Modul NIR Preprocessing Explorer di SQalytics ditujukan untuk:
- Menerapkan 4 metode preprocessing standar: SNV, MSC, Savitzky-Golay derivative (1st/2nd), Savitzky-Golay smoothing.
- Memungkinkan pipeline (chained preprocessing, mis. Smooth → 1st derivative → SNV).
- Mendukung before/after visualization overlay multi-spektrum.
- Menampilkan summary: mean spectrum, SD bands, dan effect size preprocessing.
- Ekspor preprocessed spectra ke CSV untuk pipeline modeling downstream (PLSR Studio).
- Audiens: mahasiswa S2/S3 yang menjalankan NIR pertama kali, R&D yang membangun kalibrasi NIR rapid quality, dan QC industri pangan.
1.3 Posisi di Antara Alternatif
Pilih NIR Preprocessing Explorer untuk pre-modeling NIR atau spektrum panjang gelombang banyak (Raman, MIR). Untuk overlay UV-Vis sederhana, pakai UV-Vis Spectra. Untuk identifikasi gugus fungsi FTIR, gunakan FTIR Peak Identification. Untuk modeling kuantitatif setelah preprocessing, lanjut ke PLSR Studio. Untuk eksplorasi unsupervised, gunakan PCA Explorer (akan datang).
2 Metode
2.1 Dasar Teoretis
Standard Normal Variate (SNV) (Barnes, Dhanoa, & Lister, 1989) — koreksi scaling per spektrum:
dengan $\bar{x}_i$ mean dan $s_i$ SD spektrum sample $i$ across all wavelengths. SNV menghilangkan additive (baseline offset) dan multiplicative (slope) scatter dalam satu langkah.
Multiplicative Scatter Correction (MSC) (Geladi, MacDougall, & Martens, 1985) — regress setiap spektrum terhadap mean spectrum:
$$ x_i = a_i \cdot \bar{x}_{\text{ref}} + b_i + \varepsilon_i $$ $$ x_i^{\text{MSC}} = (x_i - b_i) / a_i $$MSC mathematically equivalent ke SNV untuk pre-processing tetapi membutuhkan reference spectrum (biasanya mean of calibration set).
Savitzky-Golay derivative (Savitzky & Golay, 1964) — fit polynomial $d$ pada window $w$ titik, ambil derivatif analitis:
$$ \frac{dy_i}{dx} = \sum_{j=-(w-1)/2}^{(w-1)/2} c_j^{(1)} \cdot y_{i+j} $$dengan $c_j^{(1)}$ koefisien Savitzky-Golay untuk 1st derivative. Untuk 2nd derivative pakai $c_j^{(2)}$. Window $w$ tipikal 11–21 titik, polynomial $d$ = 2 atau 3.
Savitzky-Golay smoothing (zero derivative) — least-squares smoothing tanpa mengubah orde:
$$ \hat{y}_i = \sum_{j=-(w-1)/2}^{(w-1)/2} c_j^{(0)} \cdot y_{i+j} $$Detrend (Barnes et al., 1989) — fit polynomial ke spectrum lalu subtract:
$$ x_i^{\text{detrend}} = x_i - P_d(\lambda) $$dengan $P_d$ polynomial fit orde $d$ (biasanya 2).
Pipeline preprocessing — chained operations (Rinnan et al., 2009):
- Smooth → Derivative → SNV — rekomendasi default untuk NIR pangan.
- MSC → 2nd derivative — untuk dilute samples dengan strong scatter.
- SNV → Detrend — untuk powder dengan baseline curvature.
2.2 Persamaan Inti
SNV: $x^{\text{SNV}} = (x - \bar{x}) / s$
MSC: $x^{\text{MSC}} = (x - b) / a$ dari regression $x = a \bar{x}_{\text{ref}} + b$
Savitzky-Golay derivative orde-$n$: $d^n y / dx^n = \sum_j c_j^{(n)} y_{i+j}$
Detrend orde-$d$: $x^{\text{detrend}} = x - P_d(\lambda)$
2.3 Asumsi & Batas Validitas
| Asumsi | Konsekuensi jika dilanggar | Cara cek di SQalytics |
|---|---|---|
| Resolusi spektrum cukup (≥ 5 nm) | Smoothing/derivative artefact | Cek raw spectral resolution |
| Range wavelength konsisten antar sampel | Preprocessing inconsistent | Modul require fixed grid |
| Outlier spectra diidentifikasi sebelum mean (MSC) | Mean reference bias | Visual pre-scan |
| Window Savitzky-Golay sesuai (small enough untuk preserve peak, big enough untuk noise) | Over-smoothing atau noise persist | Tune window 11–21 |
| Tidak ada bad pixel atau saturated regions | Artefact derivative | Mask region sebelum SG derivative |
| Replikasi $\geq$ 3 per sample untuk averaging | Noise tinggi | Average sebelum preprocessing |
3 Cara Kerja
3.1 Step-by-Step di SQalytics
- Buka
NIR Preprocessing Explorerdari domain Kimia. - Muat tabel: kolom
Wavelength_nm+ multi-kolom sampel spektrum (mis.Sample_001,Sample_002, ...). - Pilih preprocessing pipeline:
-Raw(no preprocessing, baseline).
-SNV.
-MSC.
-Savitzky-Golay smoothing($w$, $d$).
-Savitzky-Golay 1st derivative($w$, $d$).
-Savitzky-Golay 2nd derivative.
-Detrend(polynomial order).
- Custom pipeline: chain ≤ 3 operasi. - Atur parameter Savitzky-Golay: window $w$ (default 15), polynomial $d$ (default 2).
- Klik
Run Preprocessing. - Tinjau hasil:
- TabRaw Spectra— overlay original.
- TabProcessed Spectra— setelah preprocessing.
- TabMean ± SD— mean spectrum + SD bands.
- TabEffect Comparison— side-by-side raw vs processed (statistik: total variance reduced).
- Ekspor preprocessed CSV untuk pipeline modeling.
Setelah hasil utama tampil di Step 4, gunakan Step 5 opsional untuk mengirim Raw spectra atau Preprocessed comparison ke Publication Graph Studio (PGS). Di sana Anda bisa lanjut mengatur typography, legend, preview Print, lalu ekspor SVG/PNG/PDF untuk kebutuhan publikasi.
3.2 Template Tabel Input + Contoh Data Sintetis
Skema input minimum:
| Kolom | Tipe | Wajib | Catatan |
|---|---|---|---|
Wavelength_nm |
numeric | ✓ | Sumbu X (1000–2500 nm tipikal) |
Sample_ |
numeric | ✓ (≥1) | Spektrum reflectance/absorbance |
Contoh data sintetis (3 spektrum NIR biji kopi pada 5 wavelength representatif, 10 baris ringkas):
| Wavelength_nm | Sample_001 | Sample_002 | Sample_003 |
|---|---|---|---|
| 1100 | 0.245 | 0.252 | 0.238 |
| 1450 | 0.380 (water OH 1st overtone) | 0.395 | 0.372 |
| 1700 | 0.298 | 0.310 | 0.291 |
| 1940 | 0.520 (water OH combination) | 0.535 | 0.510 |
| 2100 | 0.412 | 0.425 | 0.408 |
| 2300 (CH 1st overtone, oil/lipid) | 0.385 | 0.400 | 0.380 |
3.3 Contoh Luaran
Pipeline yang dipilih: Savitzky-Golay smoothing ($w=15$, $d=2$) → 1st derivative ($w=15$, $d=2$) → SNV.
Summary statistik:
| Metric | Raw | After Smooth | After 1st Deriv | After SNV |
|---|---|---|---|---|
| Mean reflectance | 0.385 | 0.385 | 0.000 | 0.000 |
| SD across samples | 0.012 | 0.011 | 0.0024 | 1.000 |
| Total variance | 1.45e-4 | 1.21e-4 | 5.76e-6 | 1.000 |
| Baseline drift | High | Reduced | Removed | Removed |
| Scatter effect | High | High | Reduced | Removed |
PLSR Studio dengan preprocessed spectra (X-matrix) dan measured moisture (Y-vector) untuk kalibrasi rapid moisture NIR — target $R^2$ kalibrasi > 0.95 dan RMSECV < 0.3% moisture."4 Kesimpulan
4.1 Relevansi Real-World
- Kalibrasi rapid moisture/protein/lemak — NIR online untuk grain quality, milk powder, daging.
- Pharmaceutical NIR PAT — content uniformity API rapid testing.
- Quality grading — kopi, kakao, gandum, palm oil.
- Process Analytical Technology (PAT) — inline monitoring kelembaban granulation.
- Authentication — fingerprinting madu, olive oil origin.
- Compositional mapping — NIR hyperspectral imaging untuk distribusi spasial.
4.2 Where to Go from Here
⚙ Troubleshooting Cepat
i Riwayat Revisi
| Tanggal | Revisi | Penulis |
|---|---|---|
| 2026-05-12 | Migrasi MD v2 → HTML final dengan figure publikasi + caption Elsevier-style (W1 batch malam 12 Mei) | Claude |
| 2026-05-12 | Draft v2 publikasi (KaTeX SNV/MSC/Savitzky-Golay + APA Næs/Rinnan/Barnes/Geladi) | Claude |
4 Referensi
- Næs, T., Isaksson, T., Fearn, T., & Davies, T. (2002). A user-friendly guide to multivariate calibration and classification. NIR Publications.
- Rinnan, Å., van den Berg, F., & Engelsen, S. B. (2009). Review of the most common pre-processing techniques for near-infrared spectra. TrAC Trends in Analytical Chemistry, 28(10), 1201–1222. https://doi.org/10.1016/j.trac.2009.07.007
- Barnes, R. J., Dhanoa, M. S., & Lister, S. J. (1989). Standard normal variate transformation and de-trending of near-infrared diffuse reflectance spectra. Applied Spectroscopy, 43(5), 772–777. https://doi.org/10.1366/0003702894202201
- Geladi, P., MacDougall, D., & Martens, H. (1985). Linearization and scatter-correction for near-infrared reflectance spectra of meat. Applied Spectroscopy, 39(3), 491–500. https://doi.org/10.1366/0003702854248656
- Savitzky, A., & Golay, M. J. E. (1964). Smoothing and differentiation of data by simplified least squares procedures. Analytical Chemistry, 36(8), 1627–1639. https://doi.org/10.1021/ac60214a047
- Pasquini, C. (2003). Near infrared spectroscopy: Fundamentals, practical aspects and analytical applications. Journal of the Brazilian Chemical Society, 14(2), 198–219. https://doi.org/10.1590/S0103-50532003000200006