Using prototyping to choose a bioinformatics workflow management system
Author summary Data analysis involves many steps, as data are wrangled, processed, and analysed using a succession of unrelated software packages. Running the right steps, in the right order, and putting the right outputs in the right places, is a major source of frustration. Workflow management systems require that each data analysis step be “wrapped” in a structured way, describing its inputs, parameters, and outputs. By writing these wrappers, the scientist can focus on the meaning of each step, and how they fit together, which is the interesting part. The system uses these wrappers to decide what steps to run and how to run these and takes charge of running the steps, including reporting on errors. This makes it much easier to repeatedly run the analysis and to run it transparently upon different computers. To select a workflow management system, we surveyed available tools and chose 4 in which we developed prototype implementations to evaluate their suitability for our project. We conclude that many similar multistep data analysis workflows can be rewritten in a workflow management system, and we advocate prototyping as a low-cost (both time and effort) way of making an informed selection of software for use within a research project.
Using prototyping to choose a bioinformatics workflow management system | PLOS Computational Biology Article Authors Metrics Comments Media Coverage Reader Comments Figures Figures Abstract Workflow management systems represent, manage, and execute multistep computational analyses and offer many benefits to bioinformaticians. They provide a common language for describing analysis workflows, contributing to reproducibility and to building libraries of reusable components. They can support both incremental build and re-entrancy—the ability to selectively re-execute parts of a workflow in the pre
Explore this link on the map →saved by
related reading
- Challenges and recommendations to improve the installability and archival stability of omics computational tools | PLOS Biologyjournals.plos.org
- rnaseq: Parametersnf-co.re
- So where are we with deep learning for biochem?ladanuzhna.xyz
- How Software in the Life Sciences Actually Works (And Doesn’t Work) - New Sciencenewscience.org
- Gap Mapgap-map.org
- Paving the way for agents in biology \ Anthropicanthropic.com
- Tammy Taabassum — Product Designertaamannae.dev
- Dynamic Workflows — Vlad's Playbookdive.vladyslavpodoliako.com
- KnowledgeHut Blog | Resources, Career Guides, & Moreknowledgehut.com
- A Brief History of Bioinformatics Softwaresubstack.com
- Practical Cheminformatics Index - Practical Cheminformaticspatwalters.github.io
- Best Practices for Scientific Computing | PLOS Biologyjournals.plos.org