ProgPerm: Progressive permutation for a dynamic representation of the robustness of microbiome discoveries.

Overview

abstract

BACKGROUND: Identification of features is a critical task in microbiome studies that is complicated by the fact that microbial data are high dimensional and heterogeneous. Masked by the complexity of the data, the problem of separating signals (differential features between groups) from noise (features that are not differential between groups) becomes challenging and troublesome. For instance, when performing differential abundance tests, multiple testing adjustments tend to be overconservative, as the probability of a type I error (false positive) increases dramatically with the large numbers of hypotheses. Moreover, the grouping effect of interest can be obscured by heterogeneity. These factors can incorrectly lead to the conclusion that there are no differences in the microbiome compositions. RESULTS: We translate and represent the problem of identifying differential features, which are differential in two-group comparisons (e.g., treatment versus control), as a dynamic layout of separating the signal from its random background. More specifically, we progressively permute the grouping factor labels of the microbiome samples and perform multiple differential abundance tests in each scenario. We then compare the signal strength of the most differential features from the original data with their performance in permutations, and will observe a visually apparent decreasing trend if these features are true positives identified from the data. Simulations and applications on real data show that the proposed method creates a U-curve when plotting the number of significant features versus the proportion of mixing. The shape of the U-Curve can convey the strength of the overall association between the microbiome and the grouping factor. We also define a fragility index to measure the robustness of the discoveries. Finally, we recommend the identified features by comparing p-values in the observed data with p-values in the fully mixed data. CONCLUSIONS: We have developed this into a user-friendly and efficient R-shiny tool with visualizations. By default, we use the Wilcoxon rank sum test to compute the p-values, since it is a robust nonparametric test. Our proposed method can also utilize p-values obtained from other testing methods, such as DESeq. This demonstrates the potential of the progressive permutation method to be extended to new settings.

authors

Zhang, Liangliang

Shi, Yushu
Do, Kim-Anh
Peterson, Christine B
Jenq, Robert R

publication date

March 17, 2021

published in

BMC bioinformatics Journal

Research

keywords

Microbiota
Statistics, Nonparametric

Identity

PubMed Central ID

PMC7972227

Scopus Document Identifier

85102703969

Digital Object Identifier (DOI)

10.1186/s12859-021-04061-3

PubMed ID

33731016

Additional Document Info

has global citation frequency

3

volume

22

issue

1

VIVO Weill Cornell Medical College

ProgPerm: Progressive permutation for a dynamic representation of the robustness of microbiome discoveries. Academic Article

Overview

abstract

authors

publication date

published in

Research

keywords

Identity

PubMed Central ID

Scopus Document Identifier

Digital Object Identifier (DOI)

PubMed ID

Additional Document Info

has global citation frequency

volume

issue