Scientific Publications by FDA Staff
PLoS One 2013 Aug 5;8(8):e71680
Identification of bicluster regions in a binary matrix and its applications.
Chen HC, Zou W, Tien YJ, Chen JJ
Biclustering has emerged as an important approach to the analysis of large-scale datasets. A biclustering technique identifies a subset of rows that exhibit similar patterns on a subset of columns in a data matrix. Many biclustering methods have been proposed, and most, if not all, algorithms are developed to detect regions of "coherence" patterns. These methods perform unsatisfactorily if the purpose is to identify biclusters of a constant level. This paper presents a two-step biclustering method to identify constant level biclusters for binary or quantitative data. This algorithm identifies the maximal dimensional submatrix such that the proportion of non-signals is less than a pre-specified tolerance delta. The proposed method has much higher sensitivity and slightly lower specificity than several prominent biclustering methods from the analysis of two synthetic datasets. It was further compared with the Bimax method for two real datasets. The proposed method was shown to perform the most robust in terms of sensitivity, number of biclusters and number of serotype-specific biclusters identified. However, dichotomization using different signal level thresholds usually leads to different sets of biclusters; this also occurs in the present analysis.
|Category: Journal Article|
|PubMed ID: #23940779||DOI: 10.1371/journal.pone.0071680|
|PubMed Central ID: #PMC3733970|
|Includes FDA Authors from Scientific Area(s): Toxicological Research|
|Entry Created: 2013-08-14||Entry Last Modified: 2013-10-26|