archAIons was evaluated against established taxonomic profiling approaches using nine held-out full-length Nanopore 16S mock-community samples.
Highest species-level F1 and lowest false-positive species count among the five evaluated profiling approaches at the predefined 0.01% reporting threshold.
The primary comparison is species level, at the predefined 0.01% reporting threshold, averaged over all nine mock-community samples. archAIons was the only approach that combined perfect recall (1.000) with the best precision (0.387), giving the highest species-level F1 (0.555) and the fewest false-positive species calls.
archAIons achieved the highest species-level F1 in all 9/9 individual mock-community samples. As reported in the manuscript, the mean paired advantage over the next-best approach, EMU, was +0.086 (bootstrap 95% CI [0.067, 0.105]; one-sided Wilcoxon signed-rank p = 0.002), and larger against the other three approaches (+0.30 to +0.48, all confidence intervals excluding zero).
| Method | Precision ↑ | Recall ↑ | F1 ↑ | FP count ↓ | FP mass % ↓ | L1 ↓ | Bray-Curtis ↓ | Unclass. % |
|---|---|---|---|---|---|---|---|---|
| archAIons | 0.387 | 1.000 | 0.555 | 16.2 | 6.56 | 0.426 | 0.239 | 21.8 |
| EMU | 0.309 | 1.000 | 0.469 | 23.1 | 5.29 | 0.438 | 0.219 | 0.1 |
| EPI2ME-Minimap | 0.155 | 0.807 | 0.255 | 47.6 | 37.19 | 1.098 | 0.549 | 0.0 |
| BugSeq | 0.140 | 0.939 | 0.243 | 63.4 | 14.29 | 0.476 | 0.250 | 9.0 |
| EPI2ME-Kraken | 0.039 | 0.524 | 0.072 | 127.4 | 55.65 | 1.317 | 0.658 | 0.0 |
Table 1. Headline species-level performance (mean over 9 mock samples, 0.01% threshold). L1 and Bray-Curtis are the protocol full-profile values (unclassified mass retained). “Unclass. %” is the mean unclassified/missing mass each approach reports. Ordered by F1.
Evaluated approaches: archAIons, EMU v3.6.2, BugSeq v7.0.0 and the two EPI2ME wf-metagenomics workflows (Kraken2 and minimap2). Every approach was run with its standard workflow at default settings on identical input reads, and scored under a single protocol frozen in advance of any tool ranking.
| Comparison | archAIons F1 | Competitor F1 | Mean paired Δ | Bootstrap 95% CI | archAIons higher | Wilcoxon p |
|---|---|---|---|---|---|---|
| archAIons vs EMU | 0.555 | 0.469 | +0.086 | [0.067, 0.105] | 9/9 | 0.002 |
| archAIons vs BugSeq | 0.555 | 0.243 | +0.312 | [0.256, 0.362] | 9/9 | 0.002 |
| archAIons vs EPI2ME-Minimap | 0.555 | 0.255 | +0.299 | [0.279, 0.317] | 9/9 | 0.002 |
| archAIons vs EPI2ME-Kraken | 0.555 | 0.072 | +0.483 | [0.424, 0.541] | 9/9 | 0.002 |
Table 10. Statistical support for the species-level F1 advantage (0.01% threshold, n = 9 mock samples). Paired per-sample differences (archAIons − competitor), bootstrap 95% CIs (10,000 resamples), and one-sided Wilcoxon signed-rank tests. The Wilcoxon p-value is bounded at 0.002 for n = 9 with all-concordant signs.
The manuscript stresses that the nine samples derive from three community designs with three replicates each: these paired statistics establish consistency of the ranking across the sampled conditions, not generalization to arbitrary community compositions.
At the predefined reporting threshold, archAIons detected all defined target species across the nine held-out mock-community samples.
As reported in the manuscript, archAIons and EMU were the only two approaches to recover every truth species in every sample in which it was present, including the two separately scored Streptomyces targets. BugSeq missed S. albidoflavus entirely, and both EPI2ME workflows missed both Streptomyces species; EPI2ME-Kraken additionally failed to detect several further species at species level.
On the mock communities, where the true composition is defined, archAIons reported a mean of 16.2 false-positive species per sample — the fewest of any approach evaluated.
The manuscript reports that most of archAIons’ residual false positives are sister species of true targets that 16S genuinely cannot separate — for example members of the Enterobacter cloacae complex reported alongside the true E. ludwigii — rather than taxa from unrelated genera. On false-positive abundance mass, archAIons (6.56%) was second-lowest, marginally above EMU (5.29%) and well below the other three approaches (14–56%).
False-positive claims here refer only to the mock communities with defined ground truth. Panel (c) of the figure below shows how many genera each approach reports on environmental samples, which have no ground truth: that is a difference in reporting breadth, not evidence that the additional taxa reported by other approaches are false.
The manuscript evaluated every approach at three reporting thresholds — 0.01%, 0.5% and 1% — with the predefined 0.01% threshold used for the primary benchmark. The threshold controls which detections are counted; it is applied identically to every approach and every sample.
archAIons had the highest species-level F1 at all three thresholds: 0.555 at 0.01%, 0.705 at 0.5% and 0.692 at 1%, ahead of EMU (0.469 / 0.699 / 0.676) at every step. Raising the threshold removes trace false positives and improves precision for every approach, while lowering recall — both because truth members whose true abundance falls below the cutoff can no longer be reported by any tool, and because of tool-specific under-calling. The manuscript therefore treats the higher thresholds as most informative for comparing precision and false-positive suppression, and keeps 0.01% as the primary comparison for recall.
The values below are genus-level results at the predefined 0.01% threshold, averaged over the nine mock-community samples.
For comparison, the manuscript reports genus-level F1 of 0.800 for EMU, 0.619 for EPI2ME-Minimap, 0.562 for BugSeq and 0.251 for EPI2ME-Kraken. On genus-level abundance the ranking is different and is reported as such: BugSeq achieved the lowest genus-level L1 (0.262) and Bray-Curtis (0.138), ahead of archAIons (0.324 and 0.182).
Beyond detecting the right organisms, a profiler has to place abundance mass correctly. Plotting predicted against true abundance for every truth taxon shows archAIons and EMU tracking the identity line across four orders of magnitude (log-scale Pearson r = 0.84 and 0.87 on detected taxa), with no false negatives for either approach, while the EPI2ME workflows leave many truth taxa undetected.
On the protocol full-profile denominator, archAIons produced the lowest species-level L1 error (0.426), narrowly ahead of EMU (0.438) and BugSeq (0.476), while retaining the largest honest unclassified mass of the three (21.8%, vs. 0.1% and 9.0%). On Bray-Curtis dissimilarity, EMU (0.219) edged archAIons (0.239) at species level because archAIons’ larger missing mass registers as under-prediction; the manuscript describes the two as effectively tied and both far ahead of the EPI2ME workflows (≥0.55). Under the classified-normalized convention, archAIons has the lowest species-level L1 (0.421) and Bray-Curtis (0.210).
The benchmark used publicly available full-length 16S Nanopore data (Oxford Nanopore R10.4.1) generated by Zhang et al. (NCBI BioProject PRJNA925180): two synthetic 12-organism communities at contrasting proportions and the ZymoBIOMICS microbial community standard, three replicates each, spanning 523,488 classified-eligible reads.
All nine mock samples were strictly held out from archAIons development, as described in the manuscript: they were not used to train the model, tune any threshold, select model or database versions, design reference neighborhoods, or otherwise optimize the method. The four reference approaches were run with their default configurations and are likewise not trained on these samples.
Scoring rules — the synonym table, the species- versus genus-level truth targets, the detection threshold, the denominator convention (unclassified reads retained as missing mass) and every metric definition — were frozen before any tool ranking was inspected, and applied identically to all five approaches.
The five manuscript figures, unmodified. Tap or click any figure to open it at full resolution.
This page presents a summary of the benchmark study. It is not the study itself. Readers should consult the manuscript for the complete methodology, statistical analysis, limitations and interpretation, including the parts that do not favour archAIons.
Raw sequencing data: NCBI Sequence Read Archive, BioProject PRJNA925180, originally generated by Zhang et al. (2023), Applied and Environmental Microbiology 89(10):e00605-23.
Questions about the benchmark: [email protected].