Identifying personal microbiomes using metagenomic codes.

Proc Natl Acad Sci U S A
Authors
Keywords
Abstract

Community composition within the human microbiome varies across individuals, but it remains unknown if this variation is sufficient to uniquely identify individuals within large populations or stable enough to identify them over time. We investigated this by developing a hitting set-based coding algorithm and applying it to the Human Microbiome Project population. Our approach defined body site-specific metagenomic codes: sets of microbial taxa or genes prioritized to uniquely and stably identify individuals. Codes capturing strain variation in clade-specific marker genes were able to distinguish among 100s of individuals at an initial sampling time point. In comparisons with follow-up samples collected 30-300 d later, ∼30% of individuals could still be uniquely pinpointed using metagenomic codes from a typical body site; coincidental (false positive) matches were rare. Codes based on the gut microbiome were exceptionally stable and pinpointed >80% of individuals. The failure of a code to match its owner at a later time point was largely explained by the loss of specific microbial strains (at current limits of detection) and was only weakly associated with the length of the sampling interval. In addition to highlighting patterns of temporal variation in the ecology of the human microbiome, this work demonstrates the feasibility of microbiome-based identifiability-a result with important ethical implications for microbiome study design. The datasets and code used in this work are available for download from huttenhower.sph.harvard.edu/idability.

Year of Publication
2015
Journal
Proc Natl Acad Sci U S A
Volume
112
Issue
22
Pages
E2930-8
Date Published
2015 Jun 02
ISSN
1091-6490
URL
DOI
10.1073/pnas.1423854112
PubMed ID
25964341
PubMed Central ID
PMC4460507
Links
Grant list
R01 AI101018 / AI / NIAID NIH HHS / United States
U54HG004969 / HG / NHGRI NIH HHS / United States
HHSN272200900018C / AI / NIAID NIH HHS / United States
P50 GM098911 / GM / NIGMS NIH HHS / United States
HHSN272200900018C / PHS HHS / United States
R01 HG005969 / HG / NHGRI NIH HHS / United States
U54 HG004969 / HG / NHGRI NIH HHS / United States
R01HG005969 / HG / NHGRI NIH HHS / United States
P50GM098911 / GM / NIGMS NIH HHS / United States