Quick Start Guide
- Navigate to the Dashboard and go to the "Pan-Genome Comparison" tab.
- Click "Choose File" and select your FASTA file.
- Optionally fill in organism name, strain, and description.
- Click "Upload Genome" and wait for the confirmation message.
Supported File Formats
- .fasta — Standard FASTA format
- .fa — FASTA short form
- .fna — Nucleotide FASTA
- .fasta — Gene sequences (Core, Accessory, Unique)
- .xlsx — Excel annotations with COG categories
- .png — COG classification charts
- .txt — Summary statistics and mapping logs
Tutorial Guide (PDF)
Download the comprehensive tutorial with screenshots and step-by-step walkthroughs for every feature.
Troubleshooting
Tips & Best Practices
- Use properly formatted FASTA files.
- Include organism name for better results tracking.
- Ensure sequences are complete and valid.
- Select appropriate species for comparison.
- Save Comparison IDs for checking results later.
- Start with smaller genomes for testing.
- Use descriptive strain names for easy tracking.
- Download results immediately after completion.
- Check COG categories for functional insights.
Software References & Citations
PanWSGA integrates established bioinformatics utilities. When publishing research analyzed with PanWSGA, please cite the underlying tools appropriately:
- Roary: Page AJ, et al. (2015) Roary: rapid large-scale prokaryote pan-genome analysis. Bioinformatics, 31(22): 3691–3693.
- Panaroo: Tonkin-Hill G, et al. (2020) Producing accurate pangenomes with Panaroo. Genome Biology, 21(1): 180.
- PPanGGOLiN: Gautreau G, et al. (2020) PPanGGOLiN: Depicting microbial diversity via a partitioned pangenome graph. PLoS Comput Biol, 16(3): e1007773.
- MMseqs2: Steinegger M, Söding J (2017) MMseqs2 enables sensitive protein sequence searches for the analysis of massive data sets. Nat Biotechnol, 35(11): 1026–1028.
- BLAST / BLASTp: Altschul SF, et al. (1990) Basic local alignment search tool. J Mol Biol, 215(3): 403–410.
- eggNOG-mapper: Cantalapiedra CP, et al. (2021) eggNOG-mapper v2: functional annotation, orthology assignments and domain prediction. Mol Biol Evol, 38(12): 5825–5829.
Methodological Thresholds & Biological Interpretation
Unique Gene Finder Parameters & Benchmarks
The default parameters (Stage 1 k-mer filtering, Stage 2 nucleotide alignment identity ≥ 80% and coverage ≥ 80%, Stage 3 protein validation E-value ≤ 1e-5) were tuned empirically to balance sensitivity and specificity across bacterial pan-genomes. In comparative E. coli benchmark evaluations, PanWSGA identified core genome fractions (~80–85% core genes, ~3,800–4,200 core gene families) consistent with ranges reported by standard tools like Panaroo and PPanGGOLiN.
Biological Interpretation Guidelines
Genes unmatched to the selected reference database are designated as putative unique genes. Sequence absence relative to a reference database represents absence within that specific reference dataset (a direct genomic observation) rather than definitive proof of recent horizontal gene transfer. Similarly, eggNOG functional annotations represent predicted sequence homology rather than experimentally verified mechanistic function.