You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
No proprietary source code, model weights, feature logic, or training data is disclosed in this document. Only aggregate benchmark results and visualizations are presented.
Benchmark Summary
Metric
Full Dataset (5,549 rec)
IoT Dataset (1,504 rec)
Accuracy
0.9921
1.0000
Precision
0.9968
1.0000
Recall
0.9875
1.0000
F1-Score
0.9921
1.0000
FPR
0.0033
0.0000
AUC
0.9996
1.0000
ECE
0.0065
0.0033
Full Dataset: 2,797 malware + 2,752 benign (mixed seen/unseen — training was subsampled from this set)
IoT Dataset: 1,266 malware + 238 benign (syscalls from real IoT devices — 100% unseen by the model)
Classification Performance
Confusion Matrices
Full Dataset
IoT Dataset
TP / FP / FN / TN
2,762 / 9 / 35 / 2,743
1,266 / 0 / 0 / 238
Full Dataset
IoT
PROBABLE
5,072 (91.4%) acc=0.9998
1,489 (99.0%) acc=1.0000
UNCERTAIN
417 (7.5%) acc=0.9400
15 (1.0%) acc=1.0000
WEAK
57 (1.0%) acc=0.7193
—
REJECT
3 (0.1%) acc=0.3333
—
ROC & PR Curves
Full Dataset
IoT Dataset
Confusion Matrices
Full Dataset
IoT Dataset
Score Distributions
Full Dataset
IoT Dataset
Multi-Gate Analysis
Gate Confidence Profiles
Full Dataset
IoT Dataset
Gate Correlation
Full Dataset
IoT Dataset
Quantum Collapse Metrics
Full Dataset
IoT Dataset
Metric
Full Dataset
IoT Dataset
Collapse Magnitude
0.276 ± 0.040
0.293 ± 0.020
Collapse Coherence
0.691 ± 0.043
0.683 ± 0.020
Collapse Entropy
0.715 ± 0.058
0.735 ± 0.025
Calibration
Full Dataset
IoT Dataset
Temperature: 0.5928
ECE: 0.0065 (both datasets)
Decision Threshold Analysis
Full Dataset
IoT Dataset
Latency & Verdict Distribution
Full Dataset
IoT Dataset
Resource Profile
Metric
Value
Binary Size (compiled C)
23.4 KB
RAM Estimate
~47 KB
Flash Estimate (embedded)
~35 KB
C Engine Latency
~5 µs/rec @ 187K rec/s
Python Latency (full)
~783 µs/rec
Python Latency (IoT)
~9,027 µs/rec
Energy/Inference (x86-64)
~0.001 J
Energy/Inference (ARM M4)
~0.00001 J
Power Under Load (x86-64)
0.5–2 W
Power Under Load (ARM M4)
0.01–0.1 W
Data Integrity
All benchmarks were run with 100% unseen separation between training and evaluation:
IoT Dataset: 0 raw syscall sequences overlap with the training set — fully out-of-distribution
Full Dataset: Training set (_dataset_balanced.csv, 2,974 rec) is a proper subset of the full CSV (5,549 rec); the remaining 2,575 records are unseen by the model
No synthetic or augmented data was used in evaluation.
Generated by AXIOM-QUANTUM benchmarking suite. Full audit data available in the audit.json and whitebox_report.txt files within each dataset directory.