Preprocessing of gene-expression data related to breast cancer diagnosis

Publication details

The work is performed in close cooperation with the University of Tromsø and professor Eiliv Lund and is financed by the ERC TICE project. This note describes the preprocessing steps of gene expression data and focuses particularly on the filtering and normalization steps as the choices made here greatly affects the set of probes used in later analyses. In the filtering step, two parameters are set. Firstly, a cut-off for the detection p-value for each probe is set, and a probe is present in a given sample if its detection p-value is smaller than this cut-off. Secondly, the present limit is set. It is used to decide in how many samples a probe has to be present in order to be included in the dataset. The results show that a p-value cut-off at 0.01 and a present limit at 0.01 are reasonable choices. After filtering, the data can be normalized. Four different approaches are evaluated, and for the available dataset, quantile normalization of the data on original scale gives the most stable results.