Today’s guest post was contributed by Nele Haelterman, assistant professor at Baylor College of Medicine’s molecular and human genetics department. Nele combines team science, genetics, and neuroscience to study the mechanisms that drive joint pain. Nele is passionate about science communication and advocacy: she runs a blog for early career scientists (ecrLife) and promotes open, reproducible science (reproducibility 4 everyone). You can follow Nele on LinkedIn.
When you talk to scientists about standardizing an aspect of their data-related or experimental workflows, you can almost hear their collective internal groans. Most researchers greatly prefer being at the bench to test their next hypothesis compared to sitting behind a computer, completing spreadsheets to make sure their dataset complies with the most current standard. At the same time, researchers also aspire to generate findings and knowledge that can provide new scientific insights.
Advancing scientific knowledge requires researchers to be able to reproduce, compare, and build on each other’s findings. Standardization provides the foundation for this process, allowing us to compare and integrate data, and hence findings, across studies. It does this through the development and use of harmonized methods, workflows, and reporting criteria that make sure data are generated, processed, and described in consistent ways, regardless of where it was generated. Standardization ensures we all use the same names for the same tissue types, cells, genes, etc (terminology, ontology). In addition, it makes sure each shared dataset is accompanied by a complete description of how it was generated (metadata). Annotating datasets so they contain all of this information can be quite tedious, forming the basis of researchers’ mixed feelings about standards. However, as data generation speeds up across labs, standardization is the engine that will drive integration and innovation into the future.
Understanding the importance of consistent genome assembly and annotation for enabling direct comparisons across genomes and species, geneticists and genomicists have long been leaders in developing and adopting community standards. Despite these efforts, several significant gaps remain. For example, while de novo genome assembly based on long-read sequencing has become increasingly accessible to individual research groups, no consensus exists on the (meta)data that should be included in arthropod genome reports. Consequently, answering seemingly straightforward questions about your favorite gene’s evolutionary history can become surprisingly challenging.
To overcome this problem and catalyze the development of a community standard for arthropod genome reporting, Tvedte et al. reviewed 100 recently published arthropod genome papers for approaches used throughout the study. Reporting on this comparative analysis in the September issue of GENETICS, the team found broad consensus in genome assembly approaches and reported quality metrics across studies. In contrast, the projects displayed substantial variation in other key stages, including sample processing and sequencing, genome annotation, and others. In addition, the authors discovered that genome annotations were often not deposited in publicly accessible repositories, greatly reducing the reusability of the newly published genome. Together, the authors identified several opportunities for developing community standards and policies that would make arthropod genome resources more comparable and would greatly increase their long-term value for comparative and applied research.
References
Tvedte ES, Bucher G, Luecke DM et al. Toward standardization in arthropod and biodiversity genome projects. Genetics, Volume 234, Issue 1, September 2026, iyag172. https://doi.org/10.1093/genetics/iyag172