Research Spotlight
Terabyte-scale soil metagenomics assessment to understand potential microbial functions across grower-defined soil health for winegrape production
Noelymar Gonzalez-Maldonado and Amisha Poret-Peterson
Crops Pathology and Genetics Research Unit, U.S. Department of Agriculture – ARS, Davis, CA, USA
Although winegrapes are a highly valuable crop produced in the USA, they are seriously threatened by stressors like severe droughts, heavy rainfall, extreme temperatures, salinity, and erosion. To manage these challenges, growers have shown strong interest in understanding the potential of microbial communities to support soil health and vineyard resiliency for sustainable winegrape production. However, winegrapes have distinctive production and management goals compared to other perennial cropping systems. Instead of yield maximization, winegrowers focus on achieving certain grape quality parameters. These quality goals can influence growers’ views on soil health and their decisions for soil management practices (https://doi.org/10.1016/j.jrurstud.2024.103373). Hence, it is important to understand how soil microbial diversity and activity correlates with soil health metrics, including the most relevant soil conditions for winegrape growers, to enhance sustainable winegrape production.
In collaboration with winegrowers of California, Drs. Gonzalez-Maldonado and Poret-Peterson and collaborators from the University of California (UC) - Davis and USDA-ARS in Davis conducted a landscape-level soil health and microbiome assessment to better understand how soil conditions influence winegrape production (Figure 1). They collected soil samples from the Napa Valley, Paso Robles, and Lodi wine regions, and measured soil health indicators for each sample, including carbon and nitrogen concentrations, pH, salinity, and bulk density, and soil microbial diversity and their potential activity using shotgun metagenomics. Metagenomics is a sequencing tool that allows for the study of all DNA present in a sample and provides information on which microorganisms are present and their potential functions. The DNA extracted from the soil samples was sequenced by the UC Davis DNA Technologies Core Facility and yielded a total of 1.32 TB of raw genomic sequence fragments or “reads”. A series of processing steps are required to go from raw reads to interpretable information. One of the most important and computationally demanding steps early in the pipeline is co-assembly of reads.
The co-assembly of reads is the process in which reads from multiple samples are combined into a single set of contiguous sequences using metagenome assembler programs such as MEGAHIT. Terabyte-sized sequence datasets pose a challenge for co-assembly due to memory requirements. To overcome high memory requirements, large datasets are typically split into smaller ones, co-assembled separately, and sometimes combined later into a single assembly, but this requires significant time and memory use (https://doi.org/10.1038/sdata.2017.203). To use MEGAHIT, the dataset had to be split by region and sampling location within a vineyard, generating 12 assemblies. To enhance efficiency, the research team decided to try an assembler called MetaHipMer2 that handles terabyte-sized datasets and can co-assemble the entire dataset into one assembly (https://doi.org/10.1038/s41598-020-67416-5). All data processing was conducted on SCINet’s Ceres supercomputer. Co-assembly using MetaHipMer2 was performed with the support of Yasasvy Nanyam from the SCINet Virtual Research Support Core (VRSC).
MetaHipMer2 requires significant memory; however, it is extremely fast. It required 35.7 TB of memory and took 3 hours and 24 minutes to assemble the terabyte-sized dataset. In contrast, MEGAHIT failed to assemble the full dataset, which necessitated splitting the data into 12 sets. The largest MEGAHIT assembly set required 446 GB of memory and 117 hours to complete. MetaHipMer2 yielded a higher number of high-quality genomes (n=395) compared to the longest MEGAHIT assembly (n=350).
Overall, these results demonstrate that MetaHipMer2 successfully assembled this project’s terabyte-sized metagenomic dataset and more effectively used supercomputer-scale resources than alternative assemblers such as MEGAHIT. This represents a meaningful capability gain for this research since it provides opportunities to work with large metagenomic data without the memory bottlenecks that would otherwise force subsampling or compromise assembly contiguity. This is especially relevant to current research needs to address challenges and answer biological research questions across multiple regions of the country.
Preliminary analyses of the MetaHipMer2 assembly suggest that soil properties covary with differences in microbial community composition across geographic sampling locations. The next steps of the study include in-depth analysis of functional annotation metagenomic data, incorporation of primary metabolites data, and multi-omics integration of soil, metagenomic, and metabolomic datasets to better understand the relationships between soil health and microbial diversity and activity to enhance sustainable viticulture.
Figure 1. Summary of the study including the regions (Napa Valley, Lodi, and Paso Robles) in California, USA; grower-led site selection based on their perspectives on soil conditions (ideal and challenging) for achieving winegrape production goals; and soil analyses evaluated (soil health indicators and soil DNA metagenomics and metabolomics). Adapted from https://doi.org/10.1111/ejss.70265.
News
Ceres upgrades
SCINet’s Ceres supercomputer underwent routine maintenance in June that included a variety of upgrades and improvements. Ceres users might have noticed the updated appearance of Ceres Open OnDemand (OOD), which has been upgraded to version 4.2. The CLC software was updated to the latest version during the maintenance period, and Galaxy was updated shortly after (the Galaxy upgrade was delayed due to a training workshop). The system’s software across multiple components was also updated: the operating systems on compute and service nodes, multiple virtual machines, the Slurm resource manager, the storage system, the identity management system, and the software hosting the SCINet Forum.
In addition, the VRSC team conducted a stress test to determine the maximum cluster load that could be sustained by the data center’s cooling and power infrastructure and to validate automated response processes for potential cooling system failures. The results of this test will help us ensure that we are able to respond to any unexpected power or cooling issues with minimal disruption.
Ceres partition changes
During Ceres’ June maintenance, the priority partitions were retired, and all compute nodes are now accessible through the “ceres” partition. The “scavenger” partition remains available, allowing users to use resources beyond the standard allocation limits on the “ceres” partition when capacity is available. The “ceres” and “scavenger” partitions share the same hardware and both have the same maximum resource allocation limits per partition. This means that if you have jobs that require large amounts of resources, by running your jobs in both the “ceres” and “scavenger” partitions, your jobs could be allocated up to twice as many resources as if you only used the “ceres” partition.
However, a job running on the “scavenger” partition can be canceled if the “ceres” partition needs access to the same hardware allocated to the “scavenger” job. Submitting to the “scavenger” partition is therefore advantageous when you are reaching the maximum resource allocation limits on the “ceres” partition and your job either is short running or saves checkpoints or interim results and can be later resumed if canceled. This partition scheme improves overall cluster utilization while maintaining fair access and preventing resource monopolization.
As part of this partition scheme update, maximum resource allocation limits were also increased. The per-user limits per partition are now:
- CPUs: 4,000 (~15% of the cluster)
- Nodes: 32 (~20% of the cluster)
- Memory: 36 TB (~15% of the cluster)
AI-COE/SCINet Graduate Student Internships update and symposium
Our 2026 AI-COE/SCINet Graduate Student Internships Program summer interns are busy wrapping up their research projects with their ARS mentors!
We have had 21 graduate students participate in summer or spring internships with ARS researchers this year. This year’s interns were affiliated with the New College of Florida, the University of Florida, and North Carolina State University. Participants were selected by their host institutions in part because of their expertise in computer science, machine learning (ML), and data science, and were paired with ARS researchers for projects that could benefit from the intern’s computational skills.
To recognize the accomplishments of these students, we held our annual internships research symposium on Thursday, August 13, 2026. The symposium highlighted many inspiring applications of data science, AI, and ML methods to ARS research projects.
Many thanks to the students who dedicated their time and hard work to these internships, the ARS scientists who volunteered to serve as internship mentors, and the universities who have partnered with us for this internships program.
FY26 AI Innovation Fund awardees
Congratulations to the following ARS scientists for earning FY2026 Artificial Intelligence Innovation Fund awards! There were many excellent proposals submitted this year (far more than previous years!), making for a very competitive field and a difficult review process. (Awardees are listed in alphabetical order by lead PI’s last name.)
- Xianran Li: Rapid Assessment of Root Rot Severity Using Multimodal Imaging and Deep Learning
- Timothy Porch: Accelerating Common Bean Seed Yield and Quality Assessment Using an AI-Enabled Edge Vision System with Descriptive Machine Reasoning
- Lester Pordesimo, Ronnie Serfa Juan, and Alison Gerken: Advancing a Scalable, AI-Based Automated Surveillance of Stored Product Insect Pests in Food Facilities
- Jixiang Wu, Yanbo Huang, and Haibo Yao: Scalable Deep Learning for Cotton Bloom Mapping and Yield Analysis Using UAV Imagery
- Heping Zhu, Anna Testen, and Hongyoung Jeon: Development of an Automated Continuous Machine Learning System for Precision Plant Disease Detection in Controlled Environment Agriculture
We look forward to sharing the results from these exciting research projects with the SCINet community!
FY26 SCINet/AI-COE Fellowship mentors
Congratulations to the following ARS scientists for earning funding to host FY2026 SCINet/AI-COE postdoctoral fellows! This funding opportunity was extremely competitive this year - we received more proposals than ever before! (Awardees are listed in alphabetical order by lead mentor’s last name.)
- Carson Andorf and Hye-Seon Kim: FungalCAD: A DNA Foundation Model for Fungal Pathogen Risk, Adaptation, and Surveillance
- Jeremy Edwards: AI-Driven Discovery of High-Impact Genetic Variants for Rice Grain Quality Improvement
- Elizabeth Ann French: Employing Machine Learning to Optimize Productivity on Dairy Farms with Automated Milking Systems
- Alexander Hernandez: Floral resources mapping via transparent upscaling – from the field to the drone to the satellite
- George Liu: Bovine-NT: Bridging the Genotype-Phenotype Gap with Deep Genomic Learning
- Peter Olsoy and Stella Copeland: AI Transfer Model of Rangeland Plant Species
- Joseph Purswell: Forecasting highly pathogenic avian influenza risk using a spatiotemporal deep-learning-based model on the SCINet high performance computing cluster
- Yelena Sapozhnikova: Artificial Intelligence for Food Safety: Machine Learning Solutions for Forever Chemical Challenges
- Jacob Washburn: To Predict AND to Explain: Developing a Biologically Informed Neural Network Library for ARS, SCINet and their Stakeholders
- Jayne Wiarda: Applying bulk and single-cell transcriptomics to identify predictors and facilitators of immune protection against highly virulent porcine reproductive and respiratory syndrome viruses (PRRSVs) currently circulating in the United States
SCINet Working Groups
SCINet working groups (WGs) support ARS researchers and their collaborators in using scientific computing methods and SCINet computational resources in their research. Common WG activities include hosting recurring virtual meetings and webinars, organizing training events, and participating in collaborative research or software development projects.
The March 26, 2026 SCINet Corner featured short presentations from four SCINet working groups. Check out the session recording to learn more about these working groups and how you can get involved!
Current Working Groups
- Ag100Pest Initiative (subgroup of AGR)
- Animal Behaviour AI Working Group
- Arthropod Genomics Research (AGR) Working Group
- Breeding AI and ML Working Group
- Geospatial Research Working Group
- Microbiome Working Group
- SCINet-Longterm Agroecosystem Research (LTAR) Phenology Working Group
- Protein Science Working Group
- Translational Omics Working Group
If you are interested in creating a working group, please compile the following:
- The working group’s name
- A description of the working group including its purpose and goals
- Contact information for people to reach out to if they want to learn more about or join the working group.
Send this information to the SCINet office at ARS-SCINet-Office@usda.gov.
Training
The Carpentries instructor training
SCINet is collaborating with The Carpentries to offer The Carpentries’ Instructor Training Course for ARS scientists. In this course, you will learn about evidence-based practices for effective and inclusive teaching, with a particular focus on teaching computational skills. There is no fee charged to course participants, but seats are limited.
If you are interested in becoming a Carpentries-certified instructor, please complete this form.
Coursera
The SCINet Office and the AI-COE are excited to provide training opportunities through Coursera. Coursera licenses are available to ARS scientists and support staff for training focused on scientific computing, data science, artificial intelligence, and related topics. Successful completion of courses and specializations result in widely recognized certificates and credentials.
Please visit the SCINet Coursera Training Page to request a license. Licenses will be assigned on a rolling basis and are active for three months. Users may be able to extend their licenses upon request.
Workshop Reports
Foundations in bioinformatics workshop series
Leads: Genome Informatics Facility at Iowa State University, ARS researchers (Sheina Sim, Craig Carlson, and Haley Arnold), and the SCINet Office
We again hosted our Foundations in bioinformatics workshop series, which provides a thorough introduction to modern bioinformatics concepts, best practices, and practical skills. For this offering, we expanded on previous material by introducing variant calling. This update to the content of the series provided participants with a more comprehensive overview of foundational bioinformatics concepts and analysis workflows.
The series consisted of five workshops:
- Introduction to modern bioinformatics: April 13, 15-16, 2026, 1-5 PM ET
- Genome assembly: April 20, 22-23, 2026, 1-5 PM ET
- Introduction to RNA-seq analysis: May 4, 6-7, 2026, 1-5 PM ET
- Genome annotation: May 11 & 13, 2026, 1-5 PM ET
- From reads to variants: GATK & Deepvariant: May 19 & 20, 2026, 2-5 PM ET
To sign up for the waitlist for future offerings, complete the registration form.
Automating bioinformatics workflows workshop series
Lead: Genome Informatics Facility at Iowa State University
We hosted a new workshop series, which expanded on the concepts introduced in our Foundations in bioinformatics workshop series by exploring tools and platforms used for automating and streamlining bioinformatics analyses. Participants were introduced to Galaxy, Nextflow, and Snakemake and how these workflow management tools can be used to automate the analysis of biological data and develop reproducible workflows on SCINet.
Series Outline:
- RNAseq and variant calling pipelines in Galaxy: June 22, 24-25, 2026, 1-5 PM ET
- Automating bioinformatics pipelines with Nextflow: July 7 & 9, 2026, 1-5 PM ET
- Introduction to Snakemake: July 21 & 23, 2026, 1-5 PM ET
To sign up for the waitlist for future offerings, please fill out this registration form.
Please help us improve our training offerings!
What scientific computing training do you need? The SCINet Office’s goal is to provide training opportunities and resources that meet the needs of ARS researchers, so we would be grateful if you could complete our short training request form and let us know how we can best help you learn the computing skills you need. Your feedback will help us decide where we should focus our efforts over the next year and beyond.
Training opportunities are continually being updated on the SCINet Upcoming Events webpage. For more information on any of the above trainings, registration questions, or suggestions, please email SCINet-training@usda.gov.
Support
Getting Started with SCINet is as easy as 1,2,3
If you do not already have a SCINet account, we hope you will consider joining the 2,300+ researchers who do. Follow the steps below to get started with SCINet.
- Request a SCINet account to gain access to computational and training resources.
- Read the SCINet FAQs covering helpful topics such as account management, accessing and installing software, obtaining storage space for your project(s), and how to get technical help.
- Visit the SCINet Forum to connect to other users, ask questions, and learn how SCINet can enable your research. P.S. Don’t forget to complete your annual USDA information security awareness training! This is required to maintain your account. For technical assistance with your SCINet account, please email scinet_vrsc@usda.gov.
Support email addresses
All requests for help with user accounts, login problems, resource requests, or support for the Ceres HPC cluster should be sent to the SCINet Virtual Research Support Core (VRSC) at scinet_vrsc@usda.gov. Help requests specific to the Atlas HPC cluster should be sent to help-usda@hpc.msstate.edu.
Many emails are currently being sent to other SCINet email inboxes. For the most expedient response to your support requests, be sure to send them to scinet_vrsc@usda.gov or to help-usda@hpc.msstate.edu for Atlas-specific requests.
SCINet User Tip
Checking your storage quotas with the updated my_quotas command
The commands available on Ceres and Atlas for checking your storage allocations, or quotas, have been updated to make the output more user-friendly and to make the command and its output more consistent between the two systems. Now, on both clusters, you can use the command my_quotas to see your current usage (GB), your allocated storage amount (GB), how much of your allocation remains (GB), how much of your quota you have consumed (%), and the overall status of your quota usage. On both clusters, this quota information will be reported for your home directory and the project directories of all SCINet projects of which you are a member. On Ceres, your quota information for Juno is also reported (i.e., there is no my_quotas command on Juno).
If you are near or over your quota for a SCINet project, the project’s PI or manager can submit a form to request a quota increase. The quota usage information needed to fill out the form is provided by the my_quotas command.
For further information about quotas and more about SCINet project management, see the SCINet Project Management guide. For more information on all available storage locations on SCINet infrastructure, see the Storage Locations guide.
Do you have tips to share? Email them to ARS-SCINet-Office@usda.gov to be included in future newsletters.
SCINet Corner
SCINet Corner is a VRSC-moderated virtual space for people to share knowledge, discuss best practices, learn about new opportunities, and explore resources to support progress on their projects.
The next SCINet Corner will be held on August 27, 2026, from 1 – 2 PM ET. At August’s event, we will discuss storage quotas on SCINet, including how to check your quota usage and tips for using disk space efficiently. And, as usual, we will include plenty of time for general questions and answers!
You can register for this and future SCINet Corners here.
Have a question that just can’t wait? Want to see what other users are doing? Reach out to the ever-expanding SCINet Forum community for ideas, support, or just someone to bounce ideas off of at https://forum.scinet.usda.gov/.
Connect
The SCINet Community
To see all the SCINet community updates and review past newsletters, visit the Newsletter Archive.
Contribute
Do you use SCINet for your research? We would love to share your story! Email ARS-SCINet-Office@usda.gov to contribute content, ask questions, or provide feedback on the SCINet newsletter or website.
SCINet Office
Haitao Huang, Computational Biologist
Moe Richert, Web Developer
Lavida Rogers, Training Coordinator
Heather Savoy, Computational Biologist
Brian Stucky, Computational Biologist, Acting Chief Scientific Information Officer
SCINet Leadership Team
Brian Stucky, Acting Chief Scientific Information Officer
Rob Butler, SCINet Program Manager
Hye-Seon Kim, Scientific Advisory Committee (SAC) Chair
Jeff Silverstein, Associate Administrator