By: Noelymar Gonzalez-Maldonado and Amisha Poret-Peterson | 08/12/2026 Crops Pathology and Genetics Research Unit, U.S. Department of Agriculture – ARS, Davis, CA, USA
Although winegrapes are a highly valuable crop produced in the USA, they are seriously threatened by stressors like severe droughts, heavy rainfall, extreme temperatures, salinity, and erosion. To manage these challenges, growers have shown strong interest in understanding the potential of microbial communities to support soil health and vineyard resiliency for sustainable winegrape production. However, winegrapes have distinctive production and management goals compared to other perennial cropping systems. Instead of yield maximization, winegrowers focus on achieving certain grape quality parameters. These quality goals can influence growers’ views on soil health and their decisions for soil management practices (https://doi.org/10.1016/j.jrurstud.2024.103373). Hence, it is important to understand how soil microbial diversity and activity correlates with soil health metrics, including the most relevant soil conditions for winegrape growers, to enhance sustainable winegrape production.
In collaboration with winegrowers of California, Drs. Gonzalez-Maldonado and Poret-Peterson and collaborators from the University of California (UC) - Davis and USDA-ARS in Davis conducted a landscape-level soil health and microbiome assessment to better understand how soil conditions influence winegrape production (Figure 1). They collected soil samples from the Napa Valley, Paso Robles, and Lodi wine regions, and measured soil health indicators for each sample, including carbon and nitrogen concentrations, pH, salinity, and bulk density, and soil microbial diversity and their potential activity using shotgun metagenomics. Metagenomics is a sequencing tool that allows for the study of all DNA present in a sample and provides information on which microorganisms are present and their potential functions. The DNA extracted from the soil samples was sequenced by the UC Davis DNA Technologies Core Facility and yielded a total of 1.32 TB of raw genomic sequence fragments or “reads”. A series of processing steps are required to go from raw reads to interpretable information. One of the most important and computationally demanding steps early in the pipeline is co-assembly of reads.
The co-assembly of reads is the process in which reads from multiple samples are combined into a single set of contiguous sequences using metagenome assembler programs such as MEGAHIT. Terabyte-sized sequence datasets pose a challenge for co-assembly due to memory requirements. To overcome high memory requirements, large datasets are typically split into smaller ones, co-assembled separately, and sometimes combined later into a single assembly, but this requires significant time and memory use (https://doi.org/10.1038/sdata.2017.203). To use MEGAHIT, the dataset had to be split by region and sampling location within a vineyard, generating 12 assemblies. To enhance efficiency, the research team decided to try an assembler called MetaHipMer2 that handles terabyte-sized datasets and can co-assemble the entire dataset into one assembly (https://doi.org/10.1038/s41598-020-67416-5). All data processing was conducted on SCINet’s Ceres supercomputer. Co-assembly using MetaHipMer2 was performed with the support of Yasasvy Nanyam from the SCINet Virtual Research Support Core (VRSC).
MetaHipMer2 requires significant memory; however, it is extremely fast. It required 35.7 TB of memory and took 3 hours and 24 minutes to assemble the terabyte-sized dataset. In contrast, MEGAHIT failed to assemble the full dataset, which necessitated splitting the data into 12 sets. The largest MEGAHIT assembly set required 446 GB of memory and 117 hours to complete. MetaHipMer2 yielded a higher number of high-quality genomes (n=395) compared to the longest MEGAHIT assembly (n=350).
Overall, these results demonstrate that MetaHipMer2 successfully assembled this project’s terabyte-sized metagenomic dataset and more effectively used supercomputer-scale resources than alternative assemblers such as MEGAHIT. This represents a meaningful capability gain for this research since it provides opportunities to work with large metagenomic data without the memory bottlenecks that would otherwise force subsampling or compromise assembly contiguity. This is especially relevant to current research needs to address challenges and answer biological research questions across multiple regions of the country.
Preliminary analyses of the MetaHipMer2 assembly suggest that soil properties covary with differences in microbial community composition across geographic sampling locations. The next steps of the study include in-depth analysis of functional annotation metagenomic data, incorporation of primary metabolites data, and multi-omics integration of soil, metagenomic, and metabolomic datasets to better understand the relationships between soil health and microbial diversity and activity to enhance sustainable viticulture.
Figure 1. Summary of the study including the regions (Napa Valley, Lodi, and Paso Robles) in California, USA; grower-led site selection based on their perspectives on soil conditions (ideal and challenging) for achieving winegrape production goals; and soil analyses evaluated (soil health indicators and soil DNA metagenomics and metabolomics). Adapted from https://doi.org/10.1111/ejss.70265.