Video walkthrough

Implementation of a high-throughput whole genome sequencing approach with the goal of maximizing efficiency and cost effectiveness to improve public health

Student 17:16 CC AI

paperi.ai
0:00 / 0:00

Michelle C. Dickinson, Samantha E. Wirth, Deborah J. Baker, Anna Kidney, Kara Mitchell, Elizabeth Nazarian, Matthew Shudt, Lisa M. Thompson, Sai Laxmi Gubbala Venkata, Kimberlee A. Musser, Lisa Mingle

What if a public health laboratory could sequence thousands of bacterial genomes faster, while cutting the workflow cost to under eighty dollars per sample? This study shows how batching, automation, and quarter-volume library preparation made that possible.

Abstract

This manuscript describes the development of a streamlined, cost-effective laboratory workflow to meet the demands of increased whole genome sequence (WGS) capacity while achieving mandated quality metrics. From 2020 to 2021, the Wadsworth Center Bacteriology Laboratory (WCBL) used a streamlined workflow to sequence 5,743 genomes that contributed sequence data to nine different projects. The combined use of the QIAcube HT, Illumina DNA Prep using quarter volume reactions, and the NextSeq allowed the WCBL to process all samples that required WGS while also achieving a median turn-around time of 7 days (range 4 to 10 days) and meeting minimum sequence quality requirements. Public Health Laboratories should consider implementing these methods to aid in meeting testing requirements within budgetary restrictions.

Transcript

What if a public health laboratory could sequence thousands of bacterial genomes faster, while cutting the workflow cost to under eighty dollars per sample? This study shows how batching, automation, and quarter-volume library preparation made that possible.

Public Health Laboratories that implement whole genome sequencing technologies may struggle to find the balance between sample volume and cost effectiveness. That tension is the starting point for this workflow. The method allows sequencing of a variety of bacterial isolates in a cost-effective manner, including an innovative, low-cost solution using a novel quarter volume sequencing library preparation.

The methods support high-throughput DNA extraction and whole genome sequencing within budgetary constraints, strengthening public health responses to outbreaks and disease surveillance. Public Health Laboratories are continually challenged to meet increased testing capacity, often without increased funding.

The response here was a cost-effective, streamlined laboratory-wide workflow. The workflow combines a high-throughput DNA extraction platform, refined library preparation methods, and the high-throughput NextSeq sequencer to reduce costs and hands-on time.

Its backbone is coordination across multiple Wadsworth Center Bacteriology Laboratory units, allowing large numbers of bacterial isolates requiring whole genome sequencing to be batched together. That design mattered during the COVID-19 pandemic, when up to ninety percent of laboratory staff were reassigned, yet all Wadsworth Center whole genome sequencing continued with little disruption.

The workload grew as sequencing moved into more public health programs. By the end of 2019, the Wadsworth Center Bacteriology Laboratory was sequencing isolates for five programs, increasing to nine in 2021. PulseNet’s 2019 transition from Pulsed-Field Gel Electrophoresis as the primary surveillance method to whole genome sequencing created a significant increase in workload, with Legionella and hospital-associated isolates also transitioning.

From 2020 to 2021, a total of 5,743 genomes were sequenced, making a scalable workflow necessary rather than optional. Samples came to the reference laboratory from local public health departments, hospitals, clinics, health departments in other states, and other healthcare facilities requesting testing services.

Pure bacterial isolates were submitted weekly to the Wadsworth Center whole genome sequencing unit from the Bacteriology Laboratory’s units, except the Mycobacteriology Unit. Metadata such as a unique sample identifier, genus, species, Gram stain result, genome size, and project name were entered into Excel and sorted first by Gram stain, then project, then sample identifier.

After sorting, an extraction worksheet was created to guide the workflow. Two custom off-board lysis protocols pre-processed both Gram-negative and Gram-positive bacteria before they were combined in one high-throughput QIAcube HT DNA extraction run.

Gram-negative bacteria used buffer ATL and Proteinase K, while Gram-positive bacteria used enzymatic lysis buffer and lysozyme before Proteinase K was added for the remaining incubation. The lysis solutions were dispensed into a ninety-six-well S-block with an eight-channel pipette, using the extraction worksheet to determine which solution belonged in each well.

Two wells were reserved for reagent-only extraction controls. A one-microliter inoculating loop collected bacterial growth from each culture sample and transferred it into the appropriate S-block well. The loop needed only enough growth to fill it, without a dome of growth.

Excessive bacterial cells increased the risk of clogging the QIAamp ninety-six filter plate during extraction. For samples that were difficult to collect or resuspend, an inoculation needle was used to collect growth and resuspend it in the lysis solution.

DNA was submitted in a ninety-six-well plate, and next-generation sequencing libraries were prepared with the Illumina DNA Prep kit using a modified quarter-scale protocol. The method could be performed manually or with automated liquid handlers.

The Wadsworth Center Advanced Genomic Technologies Cluster used Sciclone and Zephyr workstations for library preparation, cleanup, and pooling. The quarter-volume programs were developed in-house with vendor support and were based on a full-volume Illumina DNA Prep program provided by PerkinElmer.

Table three compares per-sample costs for two library-preparation approaches across MiSeq and NextSeq sequencing platforms. Nextera XT is listed at one hundred ninety-two dollars with MiSeq-v2-500 cycle and one hundred twelve dollars with NextSeq-Mid Output v2.5–300 cycle, while the AGTC-developed quarter-volume Nextera Flex protocol ranges from one hundred forty-five to sixty-eight dollars per sample.

These flat rates include reagents, consumables, TapeStation analysis, control and repeat-library preparation, and a fifteen percent administrative charge. Libraries were pooled with calculations based on genome sizes, molar concentration, and desired sequencing coverage.

The pool was then denatured and diluted using the standard Illumina method for NextSeq five-hundred or five-fifty. One microliter of twenty picomole PhiX was added for run quality control, resulting in roughly one percent PhiX reads aligned. The typical loading concentration was between one and one point three picomolar.

Whole genome sequencing used the Illumina NextSeq five-hundred or five-fifty with a Mid Output version two point five kit for three hundred cycles. A genome load of approximately four hundred ninety-nine megabases was chosen to provide sufficient, but not overly excessive, average read depth.

After sequencing, staff verified basic run quality standards: Q30 of at least seventy-five percent, clusters passing filter of at least seventy-five percent, and cluster density from one hundred forty to two hundred forty thousand clusters per square millimeter. A quality-control report included coverage and the percentage of reads passing filter for each sample, along with total reads passing filter for the run.

Before bioinformatic analysis, eight BCL files per sample were demultiplexed and converted into one R one FASTQ file and one R two FASTQ file using bcl2fastq2 Conversion Software on a local Galaxy instance. Each Wadsworth Center unit then performed project-specific analyses, additional quality control, data sharing when appropriate, and required reporting.

The cost evaluation included DNA extraction, library preparation, and sequencing. Extraction cost included supplies, DNA quantitation, hands-on technician time, and instrument service contract cost per sample. Instrument service cost per sample used the average number of samples sequenced from 2020 to 2021, which was 2,861.50, divided by the annual service contract cost at 2021 pricing.

Instrument costs themselves were not included in the calculation. Sample run sizes of twelve, twenty-four, thirty-six, and ninety-six were assessed to provide different price points. The QIAcube Classic, QIAcube HT, and MagNAPure 24 were compared. For ninety-six samples, the QIAcube HT required five point two five hours and cost ten dollars and thirty-nine cents per sample.

The QIAcube Classic had a similar cost per sample, but processed fewer samples per run. The MagNAPure 24 was the most expensive at twenty dollars and thirty-three cents per sample for a run of twenty-four. The least expensive library preparation and sequencing option was quarter-volume Illumina DNA Prep with NextSeq Mid Output, at sixty-eight dollars per sample.

The QIAcube Classic and QIAcube HT showed no significant differences in Q30 scores, cluster densities, or coverage. The selected combination was therefore QIAcube HT, quarter-volume DNA Prep, and NextSeq, totaling seventy-eight dollars and thirty-nine cents per sample.

Table Two compares automated and manual DNA extraction platforms across batch sizes, run time, and cost per sample. For ninety-six samples, the QIAcube HT is reported at five point two five hours and ten point thirty-nine dollars per sample, while the QIAcube Classic is twelve hours and eleven point zero six dollars, and the MagNA Pure twenty-four is eight point five hours and nineteen point zero four dollars.

The table also highlights practical pros and cons, showing why throughput, consumables, hands-on work, and instrument availability matter alongside price. Figure one shows how the WCBL’s five thousand seven hundred forty-three sequenced genomes were distributed across nine projects from twenty twenty to twenty twenty-one.

The pie chart indicates that PulseNet accounted for forty-eight percent, while AR Lab Network and GenomeTrakr each accounted for eighteen percent; the remaining projects made up smaller shares. This matters because it shows how sequencing capacity supported multiple public-health and research programs, rather than a single surveillance effort.

From 2020 to 2021, the Wadsworth Center Bacteriology Laboratory sequenced 5,743 genomes contributing data to nine projects. PulseNet received the largest share, with 2,741 genomes, or forty-seven percent, followed by GenomeTrakr with 1,051, or eighteen point three percent, and the Antimicrobial Resistance Laboratory Network with 1,045, or eighteen point two percent.

Salmonella enterica was the most sequenced organism, with 2,336 genomes, or forty-one percent, followed by Listeria monocytogenes with 926, or sixteen percent, and Escherichia coli with 630, or eleven percent. Table 4 breaks down the twenty-twenty to twenty-twenty-one sequencing workload by organism and project, with a grand total of five thousand seven hundred forty-three genomes.

In this section, Streptococcus pneumoniae accounts for two hundred forty-six genomes, including one hundred ninety-seven from the Emerging Infections Program, while Staphylococcus aureus accounts for fifty-two across three projects. The authors use this project-level detail to show how whole-genome sequencing supported distinct surveillance, research, and public-health objectives.

Table 4 shows that whole genome sequencing data were analyzed for each bacterial species according to unique objectives for different projects. Among 630 Escherichia coli isolates, 456 genomes, or seventy-two point four percent, contributed to PulseNet; 96 to the Antimicrobial Resistance Laboratory Network; 53 to GenomeTrakr; and 25 to private and public partnerships.

These data supported purposes ranging from detection of novel antimicrobial resistance mechanisms to assessing genetic relatedness for foodborne disease surveillance. For 246 Streptococcus pneumoniae isolates, 197 genomes, or eighty point one percent, supported the Emerging Infections Program, while 45 supported the Antimicrobial Resistance Laboratory Network.

Figure two breaks the WCBL sequencing workload into monthly, color-coded project contributions across twenty twenty and twenty twenty-one. The chart shows two thousand forty-five genomes sequenced in twenty twenty and three thousand six hundred ninety-eight in twenty twenty-one, with monthly workload increasing markedly in the latter year.

This matters because it makes the scale and changing composition of the laboratory’s sequencing demand visible, in the context of pandemic-related staff reassignment and suspended research projects. Monthly workload showed a marked increase from 2020 to 2021. The laboratory sequenced 2,045 genomes in 2020 and 3,698 in 2021.

The difference between samples received in 2020 and 2021 was fifty-seven point six percent and was likely a result of the COVID-19 pandemic. During 2020 and 2021, 5,743 bacterial genomes moved through one workflow because custom off-board heat and chemical lysis enabled Gram-negative and Gram-positive isolates to be processed side by side in the same deep-well lysis plate.

Each QIAcube HT run used two reagent-only extraction controls: one with Gram-negative lysis solution and one with Gram-positive lysis solution. These controls went through the entire workflow and were quantified with the samples.

The controls were not submitted to the Wadsworth Center Advanced Genomic Technologies Cluster for sequencing. A flat one-microliter loop of fresh bacterial growth provided adequate DNA yield for Illumina DNA Prep without clogging the QIAamp ninety-six filter plate.

The median turnaround time from scheduling a pure bacterial isolate for DNA extraction to having whole genome sequencing data available for analysis was seven days, with a range of four to ten days. That represented an approximate two-day improvement compared with the prior workflow.

Low-throughput MiSeq runs required more time than NextSeq runs. The high-throughput workflow also used less technician time because one dedicated person processed all samples rather than one person from each unit processing that unit’s samples. When sample volume was high, two runs could be performed to meet turnaround time, while urgent or high-public-health-importance samples could still use manual extraction and MiSeq sequencing.

The workflow’s success depended on library preparation and the choice of sequencing instrument. A centralized sequencing facility used dedicated staff, specialized instruments, next-generation sequencers, and automated high-throughput liquid handlers.

Illumina DNA Prep was chosen instead of Nextera XT because input DNA did not have to be standardized to two nanograms per microliter, and coverage uniformity and sequence-data quality were more uniform. The protocol was modified for quarter-volume reactions, reducing required input DNA from one hundred to five hundred nanograms to twenty-five to one hundred twenty-five nanograms per sample.

Inputs as low as seven point five nanograms produced adequate library yields and sequence data that passed established quality-control thresholds, although lower inputs produced less library yield and required quantification before pooling. Public Health Laboratories should research DNA extraction platforms and next-generation sequencers according to expected whole genome sequencing sample volume.

Collaboration between units can batch larger numbers of samples, facilitating high-throughput workflows and reducing costs. Although this workflow relied on weekly batching for extraction and sequencing, whole genome sequencing data were available for analysis within four to ten days.

Future improvements will focus on flexibility, including fully automated next-generation sequencing solutions such as Clear Labs Clear Dx, to reduce hands-on time and provide timely analysis for pathogens of public health significance. The central takeaway is that a coordinated workflow combining QIAcube HT extraction, quarter-volume Illumina DNA Prep, and NextSeq sequencing processed 5,743 genomes with a total cost of 78.39 dollars per sample and a seven-day median turnaround, supporting public health surveillance within budgetary constraints.

A derivative work by Paperi · AI-generated script, voice and captions · pages and figures unaltered

Made with Paperi.

Drop in a research PDF — get a narrated video walkthrough like this one, with highlights that follow the narration. Free to start.

Try it with your paper →

More in Biochemistry, Genetics and Molecular Biology

ULTRA-effective labeling of tandem repeats in genomic sequence 4:16

ULTRA-effective labeling of tandem repeats in genomic sequence

Some of the most important parts of a genome are made from repeated DNA—but once those repeats become damaged and irregular, computer programs can mistake them for something else entirely. This paper introduces a way to keep finding them.

Combined SERS-Raman screening of HER2-overexpressing or silenced breast cancer cell lines 4:31

Combined SERS-Raman screening of HER2-overexpressing or silenced breast cancer cell lines

Breast cancers that look alike can behave very differently. This study uses light to find a key cancer marker on individual cells—and then reveals a surprising change that remains even after that marker is silenced.

Alteration of the N6-methyladenosine methylation landscape in a mouse model of polycystic ovary syndrome 4:26

Alteration of the N6-methyladenosine methylation landscape in a mouse model of polycystic ovary syndrome

Polycystic ovary syndrome can affect periods, fertility, and long-term health, yet there is no single treatment for its underlying cause. This study points to a tiny chemical mark on genetic messages as one possible piece of the puzzle.

All 20 papers in Biochemistry, Genetics and Molecular Biology →