archAIons
Species benchmark Species detection False positives Thresholds Genus level Abundance Study design About the benchmark
Benchmark study · Preprint

Scientifically benchmarked taxonomic profiling

archAIons was evaluated against established taxonomic profiling approaches using nine held-out full-length Nanopore 16S mock-community samples.

100%
Species-level recall
Every defined target species recovered in every sample
0.555
Species-level F1
Mean over nine mock-community samples
16.2
False-positive species/sample
Mean above-threshold false-positive taxa

Highest species-level F1 and lowest false-positive species count among the five evaluated profiling approaches at the predefined 0.01% reporting threshold.

Primary comparison

Species-level accuracy across five profiling approaches

The primary comparison is species level, at the predefined 0.01% reporting threshold, averaged over all nine mock-community samples. archAIons was the only approach that combined perfect recall (1.000) with the best precision (0.387), giving the highest species-level F1 (0.555) and the fewest false-positive species calls.

archAIons achieved the highest species-level F1 in all 9/9 individual mock-community samples. As reported in the manuscript, the mean paired advantage over the next-best approach, EMU, was +0.086 (bootstrap 95% CI [0.067, 0.105]; one-sided Wilcoxon signed-rank p = 0.002), and larger against the other three approaches (+0.30 to +0.48, all confidence intervals excluding zero).

Figure 1. Species-level accuracy across nine mock samples (0.01% threshold). Bars are the mean over nine samples; points are individual samples. archAIons (red) attains the highest F1 (a) and precision (b) at perfect recall (c) and the lowest full-profile L1 abundance error (d), while EMU achieves marginally lower Bray-Curtis dissimilarity (e). E- = EPI2ME. Tap or click the figure to open it at full resolution.
Table 1. Headline species-level performance, mean over nine mock samples at the 0.01% threshold.
Method Precision ↑ Recall ↑ F1 ↑ FP count ↓ FP mass % ↓ L1 ↓ Bray-Curtis ↓ Unclass. %
archAIons0.3871.0000.55516.26.560.4260.23921.8
EMU0.3091.0000.46923.15.290.4380.2190.1
EPI2ME-Minimap0.1550.8070.25547.637.191.0980.5490.0
BugSeq0.1400.9390.24363.414.290.4760.2509.0
EPI2ME-Kraken0.0390.5240.072127.455.651.3170.6580.0
Swipe the table sideways to see all columns.

Table 1. Headline species-level performance (mean over 9 mock samples, 0.01% threshold). L1 and Bray-Curtis are the protocol full-profile values (unclassified mass retained). “Unclass. %” is the mean unclassified/missing mass each approach reports. Ordered by F1.

Evaluated approaches: archAIons, EMU v3.6.2, BugSeq v7.0.0 and the two EPI2ME wf-metagenomics workflows (Kraken2 and minimap2). Every approach was run with its standard workflow at default settings on identical input reads, and scored under a single protocol frozen in advance of any tool ranking.

Table 10. Statistical support for the species-level F1 advantage at the 0.01% threshold, n = 9 mock samples.
Comparison archAIons F1 Competitor F1 Mean paired Δ Bootstrap 95% CI archAIons higher Wilcoxon p
archAIons vs EMU0.5550.469+0.086[0.067, 0.105]9/90.002
archAIons vs BugSeq0.5550.243+0.312[0.256, 0.362]9/90.002
archAIons vs EPI2ME-Minimap0.5550.255+0.299[0.279, 0.317]9/90.002
archAIons vs EPI2ME-Kraken0.5550.072+0.483[0.424, 0.541]9/90.002
Swipe the table sideways to see all columns.

Table 10. Statistical support for the species-level F1 advantage (0.01% threshold, n = 9 mock samples). Paired per-sample differences (archAIons − competitor), bootstrap 95% CIs (10,000 resamples), and one-sided Wilcoxon signed-rank tests. The Wilcoxon p-value is bounded at 0.002 for n = 9 with all-concordant signs.

The manuscript stresses that the nine samples derive from three community designs with three replicates each: these paired statistics establish consistency of the ranking across the sampled conditions, not generalization to arbitrary community compositions.

Sensitivity

Detection across individual species

At the predefined reporting threshold, archAIons detected all defined target species across the nine held-out mock-community samples.

As reported in the manuscript, archAIons and EMU were the only two approaches to recover every truth species in every sample in which it was present, including the two separately scored Streptomyces targets. BugSeq missed S. albidoflavus entirely, and both EPI2ME workflows missed both Streptomyces species; EPI2ME-Kraken additionally failed to detect several further species at species level.

Figure 4. Per-species detection sensitivity. Cell value = percentage of mock samples in which each tool detected each truth species at ≥0.01% (species level). archAIons and EMU show complete high-sensitivity columns. Group targets (Achromobacter, Klebsiella) shown upright; species targets italicized. Tap or click the figure to open the full heatmap at full resolution.
Specificity

Reducing false-positive classifications

On the mock communities, where the true composition is defined, archAIons reported a mean of 16.2 false-positive species per sample — the fewest of any approach evaluated.

16.2
archAIons
false-positive species/sample
23.1
EMU
47.6
EPI2ME-Minimap
63.4
BugSeq
127.4
EPI2ME-Kraken

The manuscript reports that most of archAIons’ residual false positives are sister species of true targets that 16S genuinely cannot separate — for example members of the Enterobacter cloacae complex reported alongside the true E. ludwigii — rather than taxa from unrelated genera. On false-positive abundance mass, archAIons (6.56%) was second-lowest, marginally above EMU (5.29%) and well below the other three approaches (14–56%).

False-positive claims here refer only to the mock communities with defined ground truth. Panel (c) of the figure below shows how many genera each approach reports on environmental samples, which have no ground truth: that is a difference in reporting breadth, not evidence that the additional taxa reported by other approaches are false.

Figure 5. False-positive burden and environmental reporting breadth. (a) False-positive taxa per sample and (b) false-positive abundance mass on the mocks (species, 0.01%); lower is better. (c) Number of genera reported per environmental sample (≥0.01%); archAIons reports a far shorter tail. The environmental samples have no defined composition; the manuscript states explicitly that a shorter taxon list on environmental samples is not evidence of better performance, and that these data cannot distinguish false-positive suppression from limited generalization to environmental taxa. Tap or click the figure to open it at full resolution.
Robustness

Performance across reporting thresholds

The manuscript evaluated every approach at three reporting thresholds — 0.01%, 0.5% and 1% — with the predefined 0.01% threshold used for the primary benchmark. The threshold controls which detections are counted; it is applied identically to every approach and every sample.

archAIons had the highest species-level F1 at all three thresholds: 0.555 at 0.01%, 0.705 at 0.5% and 0.692 at 1%, ahead of EMU (0.469 / 0.699 / 0.676) at every step. Raising the threshold removes trace false positives and improves precision for every approach, while lowering recall — both because truth members whose true abundance falls below the cutoff can no longer be reported by any tool, and because of tool-specific under-calling. The manuscript therefore treats the higher thresholds as most informative for comparing precision and false-positive suppression, and keeps 0.01% as the primary comparison for recall.

Figure 3. Threshold robustness and the species-to-genus gap. (a) Species F1 and (b) species precision as functions of the detection threshold; archAIons (bold red) leads at every threshold. (c) Species vs. genus F1 at 0.01%; the height of the hatched (genus) bar above the solid (species) bar is the recoverable, wrong-species component of each tool’s error. Tap or click the figure to open it at full resolution.
Genus level

Genus-level performance

The values below are genus-level results at the predefined 0.01% threshold, averaged over the nine mock-community samples.

0.993
Genus-level F1
0.988
Genus-level precision
1.000
Genus-level recall
0.11
False-positive genera/sample

For comparison, the manuscript reports genus-level F1 of 0.800 for EMU, 0.619 for EPI2ME-Minimap, 0.562 for BugSeq and 0.251 for EPI2ME-Kraken. On genus-level abundance the ranking is different and is reported as such: BugSeq achieved the lowest genus-level L1 (0.262) and Bray-Curtis (0.138), ahead of archAIons (0.324 and 0.182).

Abundance

Taxonomic abundance estimation

Beyond detecting the right organisms, a profiler has to place abundance mass correctly. Plotting predicted against true abundance for every truth taxon shows archAIons and EMU tracking the identity line across four orders of magnitude (log-scale Pearson r = 0.84 and 0.87 on detected taxa), with no false negatives for either approach, while the EPI2ME workflows leave many truth taxa undetected.

On the protocol full-profile denominator, archAIons produced the lowest species-level L1 error (0.426), narrowly ahead of EMU (0.438) and BugSeq (0.476), while retaining the largest honest unclassified mass of the three (21.8%, vs. 0.1% and 9.0%). On Bray-Curtis dissimilarity, EMU (0.219) edged archAIons (0.239) at species level because archAIons’ larger missing mass registers as under-prediction; the manuscript describes the two as effectively tied and both far ahead of the EPI2ME workflows (≥0.55). Under the classified-normalized convention, archAIons has the lowest species-level L1 (0.421) and Bray-Curtis (0.210).

Figure 2. Predicted vs. true species abundance. Each point is one truth taxon in one mock sample (log-log). Filled = detected; open symbols on the floor = false negatives. Dashed line = perfect agreement. rlog is the Pearson correlation of detected taxa in log space. Tap or click the figure to open it at full resolution.
Study design

How the benchmark was run

The benchmark used publicly available full-length 16S Nanopore data (Oxford Nanopore R10.4.1) generated by Zhang et al. (NCBI BioProject PRJNA925180): two synthetic 12-organism communities at contrasting proportions and the ZymoBIOMICS microbial community standard, three replicates each, spanning 523,488 classified-eligible reads.

9
Held-out mock-community samples
Defined composition; three replicates of each community design.
3
Community designs
Two synthetic 12-organism communities (S1, S2) and the ZymoBIOMICS standard.
16S
Full-length Oxford Nanopore
R10.4.1 chemistry, full-length ~1,500 bp 16S rRNA amplicons.
5
Evaluated profiling approaches
archAIons, EMU, BugSeq, EPI2ME wf-metagenomics Kraken2 and minimap2.
2
Taxonomic levels evaluated
Species- and genus-level scoring, each with the full metric set.
0.01%
Primary reporting threshold
Predefined and identical for every approach and sample; 0.5% and 1% reported for robustness.

All nine mock samples were strictly held out from archAIons development, as described in the manuscript: they were not used to train the model, tune any threshold, select model or database versions, design reference neighborhoods, or otherwise optimize the method. The four reference approaches were run with their default configurations and are likewise not trained on these samples.

Scoring rules — the synonym table, the species- versus genus-level truth targets, the detection threshold, the denominator convention (unclassified reads retained as missing mass) and every metric definition — were frozen before any tool ranking was inspected, and applied identically to all five approaches.

Figures

Benchmark figures

The five manuscript figures, unmodified. Tap or click any figure to open it at full resolution.

Scientific transparency

About the benchmark

This page presents a summary of the benchmark study. It is not the study itself. Readers should consult the manuscript for the complete methodology, statistical analysis, limitations and interpretation, including the parts that do not favour archAIons.

  • The manuscript is a preprint and has not been peer reviewed.
  • archAIons was not the single best approach on every individual metric: EMU achieved marginally lower species-level false-positive mass and Bray-Curtis dissimilarity, and BugSeq placed genus-level abundance mass more accurately.
  • 100% species-level recall means every defined target species was detected in the nine mock-community samples. It is not a statement of overall accuracy: at the 0.01% threshold archAIons also reported a mean of 16.2 false-positive species per sample, and precision was 0.387.
  • The nine samples represent three community designs with three replicates each. The statistics demonstrate consistency across the sampled conditions rather than generalization to arbitrary communities, and the manuscript lists this among its limitations.
  • On environmental samples, which have no defined composition, the results describe differences in reporting behaviour, not accuracy.
  • The author of the study is the developer of archAIons and declares this competing interest in the manuscript. The scoring protocol was frozen in advance of any tool ranking and the complete per-sample outputs and scoring audit are released with the study so the comparison can be independently checked.

Raw sequencing data: NCBI Sequence Read Archive, BioProject PRJNA925180, originally generated by Zhang et al. (2023), Applied and Environmental Microbiology 89(10):e00605-23.

Questions about the benchmark: [email protected].