The SARS-CoV-2 Integrated Genomic Epidemiology Database (IGED): Linking viral genomes with patient-level metadata to advance statewide genomic surveillance in California
PLOS Digital Health·
- DOI
- 10.1371/journal.pdig.0001473
- PMID
- —
- PMCID
- —
- OpenAlex
- —
- Study type
- Genomic study
- Publisher
- Public Library of Science (PLoS)
- Article type
- journal-article
- Integrity
- current
Why this research matters now
Centralizing genomic and clinical data streams allows continuous assessment of viral behavior and transmission networks across a large jurisdiction. The established pipeline offers a replicable template for health departments aiming to strengthen routine variant tracking and resource allocation.
Structured evidence summary
Research question
Establish a centralized infrastructure to aggregate viral genomic sequencing outputs alongside associated patient information for improved epidemiological tracking.
Study design
Descriptive database development and genomic surveillance implementation study utilizing cloud-based computing and standardized data processing workflows.
Population and setting
Diagnostic facilities conducting SARS-CoV-2 whole genome sequencing across California, with reporting activity predominantly concentrated in Southern California and Los Angeles County.
Main findings
The integrated platform successfully connected over eight hundred thousand viral sequences with epidemiological records, capturing the majority of state submissions. Variant distributions aligned closely with international repositories, though average processing delays surpassed thirty days. The architecture facilitates predictive modeling of viral evolutionary patterns and demonstrates scalability for additional pathogens.
Public-health relevance
Centralizing genomic and clinical data streams allows continuous assessment of viral behavior and transmission networks across a large jurisdiction. The established pipeline offers a replicable template for health departments aiming to strengthen routine variant tracking and resource allocation.
Important limitations
Processing timelines extended beyond one month, and sequencing activity remained disproportionately localized to specific regions. Sustained operational success depends on consistent laboratory participation, dedicated financial resources, and cross-disciplinary coordination.
GIDS interpretation
This publication outlines a technical blueprint for harmonizing decentralized sequencing outputs into a unified analytical repository. It contextualizes how standardized data ingestion and computational management can improve dataset accessibility for future epidemiological investigations.
Related GIDS surveillance
Literature context does not validate, explain, or change a surveillance signal. Exact and contextual relationships are shown separately.
Evidence relationships
This article has 11 auditable classifier relationships to diseases, places, topics, and study design.