Differential Expression Analysis Quality Control
This report checks whether quantified protein abundances are suitable for differential-expression interpretation. It examines missing values, within-group variability, abundance distributions, and sample structure before the differential-expression diagnostics are read.
Input: quantified protein abundances from 12 samples, analysed with DIANN.
Missing values can reveal potential biases or technical problems in the data. Figure 1 summarizes missing protein abundance estimates per group. Panel A shows how many proteins have \(0-N\) missing values. Ideally, most proteins are quantified in all samples within a group. Panel B shows the distribution of mean protein intensity by the number of missing values. Proteins without missing values usually have higher average abundance than proteins with one or more missing values because low-abundance proteins are more likely to remain undetected. Strongly overlapping distributions can point to other sources of missingness, such as large sample heterogeneity or technical problems.
Panel A in Figure 2 shows the coefficients of variation (CVs) for all proteins computed from non-normalized data. Ideally, the within-group CV should be smaller than the CV across all samples. Panel B shows the standard deviation distribution for log2-transformed data, while panel C shows the corresponding distribution after normalization. Normalization should reduce within-group variance compared with the overall variance. If normalization increases within-group variance relative to the overall variance, the selected normalization method may not be compatible with the data.
Table 1 shows the median CV and SD values for all groups and across all samples (all).
| what | A | B | Ctrl | All |
|---|---|---|---|---|
| CV | 3.47 | 3.37 | 3.64 | 9.32 |
| sd_log2 | 0.05 | 0.05 | 0.05 | 0.14 |
| sd | 0.06 | 0.05 | 0.06 | 0.14 |
Most proteins in a dataset are usually not differentially expressed; therefore, differences between the two groups should be centered close to zero. Figure 3 shows the distribution of group differences for all proteins. Ideally, the median of this distribution (red line) should be close to zero (green line). If the median and mode of the difference distribution are non-zero, this should be considered when interpreting the differential expression results.
Panel B in Figure 3 shows the distribution of p-values for all proteins. If the null hypothesis is true, p-values should be approximately uniformly distributed. A subset of differentially expressed proteins produces a higher frequency of small p-values. A higher frequency of large p-values close to 1 can indicate that the linear model does not describe the variance of the data well, for example because of outliers or an unmodeled source of variability.
The MA plot in Figure 4 helps identify whether large fold changes are concentrated among high- or low-abundance proteins. Panel A shows the group difference (y-axis) as a function of average protein abundance (x-axis). The observed fold change should not depend on protein abundance. Panel B shows the same group difference against the rank of average protein abundance.
| Field | Value |
|---|---|
| Workunit ID | 23000 |
| Order ID | 6200 |
| Project ID | 3000 |
| Project name | n/a |
| Creator | runner |
| Created at | 2026-07-29 11:12:58 UTC |
| Input data | https://fgcz-bfabric.uzh.ch/bfabric/ |
| Quantification software | DIANN |
| Model | lm_impute |
| prolfquapp version | 2.6.1 |
R version 4.6.1 (2026-06-24)
Platform: x86_64-pc-linux-gnu
Running under: Ubuntu 24.04.4 LTS
Matrix products: default
BLAS: /usr/lib/x86_64-linux-gnu/openblas-pthread/libblas.so.3
LAPACK: /usr/lib/x86_64-linux-gnu/openblas-pthread/libopenblasp-r0.3.26.so; LAPACK version 3.12.0
locale:
[1] LC_CTYPE=C.UTF-8 LC_NUMERIC=C LC_TIME=C.UTF-8
[4] LC_COLLATE=C.UTF-8 LC_MONETARY=C.UTF-8 LC_MESSAGES=C.UTF-8
[7] LC_PAPER=C.UTF-8 LC_NAME=C LC_ADDRESS=C
[10] LC_TELEPHONE=C LC_MEASUREMENT=C.UTF-8 LC_IDENTIFICATION=C
time zone: UTC
tzcode source: system (glibc)
attached base packages:
[1] stats graphics grDevices utils datasets methods base
loaded via a namespace (and not attached):
[1] RColorBrewer_1.1-3 jsonlite_2.0.0
[3] shape_1.4.6.1 magrittr_2.0.5
[5] jomo_2.7-6 farver_2.1.2
[7] logistf_1.26.1 nloptr_2.2.1
[9] rmarkdown_2.31 GlobalOptions_0.1.4
[11] vctrs_0.7.3 minqa_1.2.8
[13] progress_1.2.3 htmltools_0.5.9
[15] S4Arrays_1.12.0 forcats_1.0.1
[17] broom_1.0.13 cellranger_1.1.0
[19] SparseArray_1.12.2 mitml_0.4-5
[21] htmlwidgets_1.6.4 plyr_1.8.9
[23] plotly_4.12.1 mime_0.13
[25] lifecycle_1.0.5 iterators_1.0.14
[27] pkgconfig_2.0.3 Matrix_1.7-5
[29] R6_2.6.1 fastmap_1.2.0
[31] rbibutils_2.4.1 MatrixGenerics_1.24.0
[33] clue_0.3-68 digest_0.6.39
[35] dtplyr_1.3.3 colorspace_2.1-3
[37] lobstr_1.2.1 S4Vectors_0.50.1
[39] crosstalk_1.2.2 GenomicRanges_1.64.0
[41] labeling_0.4.3 httr_1.4.8
[43] abind_1.4-8 mgcv_1.9-4
[45] compiler_4.6.1 withr_3.0.3
[47] bit64_4.8.2 doParallel_1.0.17
[49] pander_0.6.6 S7_0.2.2
[51] backports_1.5.1 logger_0.4.2
[53] UpSetR_1.4.1 prolfquasaint_0.1.5
[55] pan_2.0 MASS_7.3-65
[57] DelayedArray_0.38.2 rjson_0.2.23
[59] optparse_1.8.2 tools_4.6.1
[61] otel_0.2.0 nnet_7.3-20
[63] glue_1.8.1 nlme_3.1-169
[65] grid_4.6.1 cluster_2.1.8.2
[67] generics_0.1.4 operator.tools_1.6.3.1
[69] gtable_0.3.6 tzdb_0.5.0
[71] formula.tools_1.7.1 preprocessCore_1.74.0
[73] tidyr_1.3.2 data.table_1.18.4
[75] hms_1.1.4 XVector_0.52.0
[77] BiocGenerics_0.58.1 ggrepel_0.9.8
[79] foreach_1.5.2 pillar_1.11.1
[81] stringr_1.6.0 limma_3.68.4
[83] circlize_0.4.18 splines_4.6.1
[85] dplyr_1.2.1 lattice_0.22-9
[87] survival_3.8-6 bit_4.6.0
[89] tidyselect_1.2.1 ComplexHeatmap_2.28.0
[91] knitr_1.51 reformulas_0.4.4
[93] gridExtra_2.3.1 prolfquapp_2.6.1
[95] bookdown_0.47 IRanges_2.46.0
[97] Seqinfo_1.2.0 SummarizedExperiment_1.42.0
[99] stats4_4.6.1 xfun_0.60
[101] prolfqua_1.7.0 Biobase_2.72.0
[103] statmod_1.5.2 matrixStats_1.5.0
[105] stringi_1.8.7 yaml_2.3.12
[107] boot_1.3-32 evaluate_1.0.5
[109] codetools_0.2-20 tibble_3.3.1
[111] BiocManager_1.30.27 cli_3.6.6
[113] affyio_1.82.0 rpart_4.1.27
[115] arrow_25.0.0 Rdpack_2.6.6
[117] Rcpp_1.1.2 readxl_1.5.0
[119] png_0.1-9 parallel_4.6.1
[121] ggplot2_4.0.3 readr_2.2.0
[123] assertthat_0.2.1 prettyunits_1.2.0
[125] lme4_2.0-6 glmnet_5.0
[127] viridisLite_0.4.3 scales_1.4.0
[129] affy_1.90.0 purrr_1.2.2
[131] crayon_1.5.3 writexl_1.5.4
[133] GetoptLong_1.1.1 rlang_1.3.0
[135] vsn_3.80.0 mice_3.19.0