Global search

Find data and evidence

Type at least 2 characters. Use arrow keys to review and Enter to open.

Peer reviewedOpen accessSARSCOVID-19

The SARS-CoV-2 Integrated Genomic Epidemiology Database (IGED): Linking viral genomes with patient-level metadata to advance statewide genomic surveillance in California

PLOS Digital Health·

Jesse Elder, Rahil Ryder, Mayuri Panditrao, Kaitlin Grosgebauer, Rebecca Katz, Lawrence Tello, Ellaison Carroll, Deva Borthwick, Chaman Kaur, Romario Smith, Victor Shiau, Will Wheeler, Emilia Reilly, Jennifer Myers, Lauren Nelson, Esther Lim, Phacharee Arunleung, Elizabeth Baylis, Sabrina Gilliam, Tamara Hennessy-Burt, Brooke Bregman, Elana Silver, Curtis Kapsak, Sage Wright, Tomas Leon, John Bell, Christina Morales, Debra A. Wadford

DOI
10.1371/journal.pdig.0001473
PMID
PMCID
OpenAlex
Study type
Genomic study
Publisher
Public Library of Science (PLoS)
Article type
journal-article
Integrity
current

Why this research matters now

Centralizing genomic and clinical data streams allows continuous assessment of viral behavior and transmission networks across a large jurisdiction. The established pipeline offers a replicable template for health departments aiming to strengthen routine variant tracking and resource allocation.

01

Structured evidence summary

Research question

Establish a centralized infrastructure to aggregate viral genomic sequencing outputs alongside associated patient information for improved epidemiological tracking.

Study design

Descriptive database development and genomic surveillance implementation study utilizing cloud-based computing and standardized data processing workflows.

Population and setting

Diagnostic facilities conducting SARS-CoV-2 whole genome sequencing across California, with reporting activity predominantly concentrated in Southern California and Los Angeles County.

Main findings

The integrated platform successfully connected over eight hundred thousand viral sequences with epidemiological records, capturing the majority of state submissions. Variant distributions aligned closely with international repositories, though average processing delays surpassed thirty days. The architecture facilitates predictive modeling of viral evolutionary patterns and demonstrates scalability for additional pathogens.

Public-health relevance

Centralizing genomic and clinical data streams allows continuous assessment of viral behavior and transmission networks across a large jurisdiction. The established pipeline offers a replicable template for health departments aiming to strengthen routine variant tracking and resource allocation.

Important limitations

Processing timelines extended beyond one month, and sequencing activity remained disproportionately localized to specific regions. Sustained operational success depends on consistent laboratory participation, dedicated financial resources, and cross-disciplinary coordination.

GIDS interpretation

This publication outlines a technical blueprint for harmonizing decentralized sequencing outputs into a unified analytical repository. It contextualizes how standardized data ingestion and computational management can improve dataset accessibility for future epidemiological investigations.

02

Related GIDS surveillance

Literature context does not validate, explain, or change a surveillance signal. Exact and contextual relationships are shown separately.

03

Evidence relationships

This article has 11 auditable classifier relationships to diseases, places, topics, and study design.

about diseaseabout diseaseaddresses topicaddresses topicaddresses topicaddresses topicevaluates interventionhas pathogen typestudied population settingstudies pathogenuses study design