1 Introduksi
1.1 Latar Belakang
Analisis deskriptif adalah langkah pertama dalam Exploratory Data Analysis (Tukey, 1977) — tahap "look at the data" sebelum analisis inferensial. Peneliti perlu menjawab: "Berapa central tendency? Berapa dispersi? Apakah distribusinya simetris? Adakah outlier? Bagaimana hubungan antar variabel?". Descriptive Analysis Studio adalah desk expert yang menjawab semua pertanyaan ini dalam satu workspace, dengan kontrol lebih dalam daripada Quick Stats Check.
Modul ini menggabungkan statistik standar (mean, SD, median, IQR) dengan statistik robust (trimmed mean, MAD, Huber location) yang tahan terhadap outlier (Wilcox, 2017). Ditambah uji normalitas (Shapiro-Wilk, Anderson-Darling) dan analisis korelasi multi-metode (Pearson, Spearman, Kendall). Modul ini menjembatani eksplorasi data dengan publication-quality Table 1.
1.2 Tujuan Modul
- Statistik standar: mean, SD, median, IQR, range, skewness, kurtosis.
- Statistik robust: trimmed mean (20%), MAD, winsorized mean, Huber location.
- Uji normalitas: Shapiro-Wilk dan Anderson-Darling.
- Korelasi multi-metode: Pearson (linear), Spearman (rank monotonic), Kendall (tau ordinal).
- Group by opsional: ringkasan terpisah per kategori.
- Box plot multi-seri + opsional 3D scatter.
- Audiens: mahasiswa S2/S3, R&D yang membuat Table 1, analis QC pre-uji formal.
1.3 Posisi di Antara Alternatif
Pilih Descriptive Analysis Studio untuk eksplorasi numerik dalam dengan output publikasi. Untuk review cepat satu kolom, pakai Quick Stats Check. Untuk uji statistik perbandingan, lanjut ke ANOVA Studio. Untuk regresi, pakai Regression Studio. Untuk heatmap korelasi kustom, pakai Correlation Explorer. Untuk multivariate latent variable, gunakan PLSR Studio.
2 Metode
2.1 Dasar Teoretis
Central tendency: Mean ($\bar{x} = \frac{1}{n}\sum x_i$, sensitif outlier), Median (robust 50% breakdown), Trimmed mean (Tukey, default $\alpha = 20\%$):
$$ \bar{x}_{\alpha} = \frac{1}{n - 2g} \sum_{i=g+1}^{n-g} x_{(i)}, \quad g = \lfloor \alpha n \rfloor $$Huber M-estimator (Huber, 1964) — iterative reweighted dengan $\psi(u) = u$ bila $|u| \leq k$, $k \cdot \text{sign}(u)$ bila $|u| > k$, dengan $k = 1.345$ untuk 95% efisiensi pada normal.
Dispersi:
- SD: $s = \sqrt{\frac{1}{n-1} \sum (x_i - \bar{x})^2}$.
- IQR: $Q_3 - Q_1$ (robust 25% breakdown).
- MAD: $\text{MAD} = 1.4826 \cdot \text{median}(|x_i - \tilde{x}|)$ (robust 50% breakdown).
Shape:
- Skewness: $g_1 = \frac{1}{n} \sum z_i^3$. Interpretasi: $|g_1| < 0.5$ symmetric; 0.5–1 moderate; $\geq 1$ highly skewed.
- Kurtosis (excess): $g_2 = \frac{1}{n} \sum z_i^4 - 3$. Interpretasi: 0 mesokurtic; > 0 leptokurtic; < 0 platykurtic.
Uji normalitas:
• Shapiro-Wilk (paling powerful $n \leq 50$):
$$ W = \frac{(\sum a_i x_{(i)})^2}{\sum (x_i - \bar{x})^2} $$• Anderson-Darling (robust untuk $n$ besar, weight pada tail):
$$ A^2 = -n - \frac{1}{n} \sum_i (2i-1)[\ln F(x_{(i)}) + \ln(1 - F(x_{(n-i+1)}))] $$Korelasi:
Interpretasi Cohen (1988): $|r| < 0.3$ lemah, 0.3–0.7 sedang, $\geq 0.7$ kuat. Inferensi: $t = r\sqrt{(n-2)/(1-r^2)} \sim t_{n-2}$.
2.2 Persamaan Inti
Standard: $\bar{x} = \frac{1}{n}\sum x_i$, $s = \sqrt{\frac{1}{n-1}\sum (x_i - \bar{x})^2}$
Trimmed mean: $\bar{x}_{0.20} = \frac{1}{n-2g} \sum_{i=g+1}^{n-g} x_{(i)}$
MAD: $1.4826 \cdot \text{median}(|x_i - \tilde{x}|)$
Skew / Kurt: $g_1 = \frac{1}{n}\sum z_i^3$, $g_2 = \frac{1}{n}\sum z_i^4 - 3$
Shapiro-Wilk W: $(\sum a_i x_{(i)})^2 / \sum (x_i - \bar{x})^2$
Pearson: $r = \sum (x_i - \bar{x})(y_i - \bar{y}) / \sqrt{\sum (x_i - \bar{x})^2 \sum (y_i - \bar{y})^2}$
Spearman: $1 - 6 \sum d_i^2 / [n(n^2-1)]$
Kendall: $(n_{\text{conc}} - n_{\text{disc}}) / \binom{n}{2}$
2.3 Asumsi & Batas Validitas
| Asumsi | Konsekuensi jika dilanggar | Cara cek di SQalytics |
|---|---|---|
| Data numerik kontinu | Statistik invalid | Modul validate tipe |
| Mean/SD assume normal distribusi | Bias bila skew/outlier | Bandingkan dengan trimmed mean / MAD |
| Pearson assume bivariate normal | Bias bila outlier | Auto-toggle ke Spearman bila non-normal |
| Shapiro-Wilk paling power $n \leq 50$ | Kurang sensitif untuk $n$ besar | Anderson-Darling untuk $n > 50$ |
| Outlier dideteksi via IQR rule | Bias pada distribusi non-normal | Visual box plot |
| Korelasi tidak menangkap non-linier | Korelasi rendah padahal ada pola | 3D scatter; polinomial di Regression Studio |
| Replikasi $n \geq 8$ untuk Shapiro reliable | Power rendah | Modul flag warning $n < 8$ |
3 Cara Kerja
3.1 Step-by-Step di SQalytics
- Buka
Descriptive Analysis Studiodari domain Statistika Terapan. - (Opsional) Muat seed
stats_group_compare. - Pada
Analysis Function, pilih:Basic Statistics,Correlation Matrix, atau3D Scatter. - Pada
Analyze Num. Vars., pilih satu atau lebih kolom numerik. - Isi
Group by (Optional)bila ingin ringkasan per kategori. - Klik Run.
- Tinjau hasil: Standard → Robust → Shapiro-Wilk → Box Plot → (untuk korelasi) Heatmap.
- Klik Save Results to TXT.
3.2 Template Tabel Input + Contoh Data Sintetis
| Kolom | Tipe | Wajib | Catatan |
|---|---|---|---|
<var1>, <var2>, ... | numeric | ✓ (≥1) | Kolom analisis |
Group | category | ◯ | Group by opsional |
Contoh data sintetis (15 baris — 5 atribut quality multi-batch):
| Batch | Moisture | Protein | Texture | ColorScore | Yield |
|---|---|---|---|---|---|
| B01 | 12.5 | 11.2 | 7.4 | 8.5 | 85.2 |
| B02 | 12.8 | 11.4 | 7.6 | 8.6 | 86.1 |
| B03 | 12.4 | 11.1 | 7.2 | 8.4 | 84.8 |
| B04 | 12.7 | 11.3 | 7.5 | 8.7 | 85.9 |
| B05 | 12.6 | 11.5 | 7.3 | 8.5 | 85.5 |
| B06 | 12.9 | 11.6 | 7.8 | 8.8 | 86.4 |
| B07 | 12.3 | 11.0 | 7.1 | 8.3 | 84.5 |
| B08 | 12.8 | 11.4 | 7.5 | 8.6 | 85.8 |
| B09 (outlier) | 13.4 | 11.7 | 7.9 | 8.9 | 86.8 |
| B10–B15 | … (6 batch lainnya) | ||||
docs/assets/example-data/id/statistics/template_stats_groups.csv.
3.3 Contoh Luaran
Tabel Standard Descriptive Statistics:
| Variable | $n$ | Mean | SD | Median | IQR | Skewness | Kurtosis |
|---|---|---|---|---|---|---|---|
| Moisture | 15 | 12.67 | 0.27 | 12.6 | 0.30 | 1.21 | 1.74 |
| Protein | 15 | 11.33 | 0.20 | 11.3 | 0.30 | 0.18 | -0.71 |
| Texture | 15 | 7.47 | 0.25 | 7.5 | 0.30 | 0.20 | -0.93 |
| ColorScore | 15 | 8.56 | 0.17 | 8.5 | 0.30 | 0.42 | -0.45 |
| Yield | 15 | 85.57 | 0.66 | 85.5 | 1.10 | 0.41 | -0.49 |
Tabel Robust Descriptive Statistics:
| Variable | Trimmed Mean (20%) | MAD | Huber Location | $\Delta$(Mean - Trim) |
|---|---|---|---|---|
| Moisture | 12.64 | 0.30 | 12.65 | +0.03 (slight skew) |
| Protein | 11.33 | 0.30 | 11.33 | 0.00 (symmetric) |
| Texture | 7.47 | 0.30 | 7.47 | 0.00 (symmetric) |
| ColorScore | 8.55 | 0.30 | 8.56 | +0.01 (symmetric) |
| Yield | 85.55 | 0.74 | 85.56 | +0.02 (symmetric) |
Tabel Shapiro-Wilk Normality Test:
| Variable | $W$ statistic | p-value | Decision (α = 0.05) |
|---|---|---|---|
| Moisture | 0.864 | 0.028 | ⚠ Non-normal (skewed by B09) |
| Protein | 0.954 | 0.586 | ✓ Normal |
| Texture | 0.953 | 0.572 | ✓ Normal |
| ColorScore | 0.945 | 0.456 | ✓ Normal |
| Yield | 0.969 | 0.840 | ✓ Normal |
Tabel Correlation Matrix (Pearson):
| Moisture | Protein | Texture | ColorScore | Yield | |
|---|---|---|---|---|---|
| Moisture | 1.00 | — | — | — | — |
| Protein | 0.82*** | 1.00 | — | — | — |
| Texture | 0.87*** | 0.84*** | 1.00 | — | — |
| ColorScore | 0.79*** | 0.82*** | 0.86*** | 1.00 | — |
| Yield | 0.89*** | 0.91*** | 0.92*** | 0.88*** | 1.00 |
*** = $p < 0.001$ — semua korelasi strong positive
Grafik utama: dua panel — (a) box plot multi-seri z-scored dengan jitter + B09 outlier highlighted; (b) correlation heatmap Pearson.
4 Kesimpulan
4.1 Relevansi Real-World
- Table 1: Descriptive characteristics — tabel pertama hampir semua artikel kuantitatif.
- QC release reporting — ringkasan batch quality dengan target spec.
- Pre-modeling diagnostik — cek asumsi sebelum ANOVA / regresi / PLSR.
- Outlier detection — IQR rule, robust statistics, visual box plot.
- Korelasi exploration — Pearson (linear), Spearman (monotonic), Kendall (ordinal).
- Sensory descriptive — ringkasan multi-atribut hedonic per produk.
- Process capability index — $C_p$, $C_{pk}$ dari mean + SD dibanding spec.
Pada Rencana Publikasi Singkil v5, modul ini dipakai pada semua tahap sebagai langkah pertama eksplorasi data.
4.2 Where to Go from Here
Pembacaan lanjutan:
- Tukey (1977) — buku klasik EDA.
- Wilcox (2017) — robust methods komprehensif.
- Shapiro & Wilk (1965) — paper sumber uji normalitas.
- Cohen (1988) — interpretasi correlation strength.
⚙ Troubleshooting Cepat
PLSR Studio.i Riwayat Revisi
| Tanggal | Revisi | Penulis |
|---|---|---|
| 2026-05-12 | Migrasi MD v2 → HTML final dengan figure dual-panel box plot + correlation heatmap + caption Elsevier-style | Claude |
| 2026-05-12 | Migrasi v1 → v2 (template publikasi + KaTeX descriptive/robust/normality/correlation + APA Tukey/Wilcox/Shapiro) | Claude |
| 2026-05-09 | Draft awal v1 | Tim docs |
4 Referensi
- Tukey, J. W. (1977). Exploratory data analysis. Addison-Wesley.
- Wilcox, R. R. (2017). Introduction to robust estimation and hypothesis testing (4th ed.). Academic Press. https://doi.org/10.1016/C2015-0-01807-3
- Shapiro, S. S., & Wilk, M. B. (1965). An analysis of variance test for normality (complete samples). Biometrika, 52(3/4), 591–611. https://doi.org/10.2307/2333709
- Anderson, T. W., & Darling, D. A. (1954). A test of goodness of fit. Journal of the American Statistical Association, 49(268), 765–769. https://doi.org/10.1080/01621459.1954.10501232
- Huber, P. J. (1964). Robust estimation of a location parameter. The Annals of Mathematical Statistics, 35(1), 73–101. https://doi.org/10.1214/aoms/1177703732
- Spearman, C. (1904). The proof and measurement of association between two things. The American Journal of Psychology, 15(1), 72–101. https://doi.org/10.2307/1412159
- Kendall, M. G. (1938). A new measure of rank correlation. Biometrika, 30(1/2), 81–93. https://doi.org/10.2307/2332226
- Cohen, J. (1988). Statistical power analysis for the behavioral sciences (2nd ed.). Lawrence Erlbaum Associates.