A simple command-line utility to calculate biological sequence (DNA or protein) sizes in a (multi) FASTA file. It gives averages, GC (or methionine) content, N50, N90, N95, number of N's, and total bases, and can also report by codon if requested.
Features
- sequence sizes (DNA or protein)
- GC content, in percentage (for each sequence and overall weighted average)
- methionine content, in absolute number and percentage (protein only)
- codon GC content (DNA only)
- multi FASTA input files
- reports average sequence size, total nucleotides, N50, N90, and N95
- by default, report shows sequence names sorted in descending size order
- report is tab-delimited text with results from one FASTA entry per line
- gzip-compressed input supported
Categories
Bio-InformaticsLicense
GNU General Public License version 3.0 (GPLv3)Follow mfsizes
Other Useful Business Software
Iris Powered By Generali - Iris puts your customer in control of their identity.
Iris Identity Protection API sends identity monitoring and alerts data into your existing digital environment – an ideal solution for businesses that are looking to offer their customers identity protection services without having to build a new product or app from scratch.
Rate This Project
Login To Rate This Project
User Reviews
Be the first to post a review of mfsizes!