Which features actually carry the signal in encrypted traffic classification, and does performance degrade over time?
A harness for evaluating encrypted traffic classifiers under conditions the literature usually avoids: group-aware splits, temporal splits, feature-family ablation, and open-world evaluation. Built to test whether reported accuracy survives an honest setup.
Findings so far are on CESNET-TLS22 (30 classes, 189,190 flows, ten consecutive days in October 2021).
Gradient boosting, macro-F1, mean ± sd across folds. Chance floor is 0.0335.
| Setup | Features | macro-F1 |
|---|---|---|
| Grouped split (by capture day) | all three | 0.9331 ± 0.0163 |
| Temporal split (3-day window, test next day) | all three | 0.9425 ± 0.0077 |
| Temporal | packet sizes only | 0.9284 ± 0.0282 |
| Temporal | flow stats + timing | 0.8570 ± 0.0114 |
| Temporal | flow stats only | 0.7552 ± 0.0115 |
| Temporal | timing only | 0.7197 ± 0.0180 |
| Chance floor | — | 0.0335 ± 0.0010 |
Testing on a future day (0.9425) performs no worse than testing on held-out days (0.9331). Whatever drift exists over ten days in a single month is below the noise floor of this setup.
This is a null result and it is reported as one. It does not generalise beyond the timescale tested, ten consecutive days in one October is a narrow window, and a year-long evaluation is the obvious next test. But at this scale, the concern that in-domain results overstate temporal robustness is not supported.
Packet sizes alone recover 98.5% of full performance (0.9284 vs 0.9425). Removing them costs nine points (0.8570); removing everything else costs one.
Flow statistics and inter-packet timing each score around 0.72–0.76 in isolation, so they are not uninformative, they are redundant given packet sizes. This has a practical consequence: the cheapest feature to collect is also the one that matters.
The most useful methodological finding here came from a bug.
CESNET's AppSelection.ALL_KNOWN label-encodes applications against a list rebuilt per time period. An application that misses the minimum-support threshold on a given day drops out, shifting every index after it. Observed list sizes across the ten days: 183, 183, 183, 180, 181, 182, 182, 173, 181, 179.
Using the raw integer as a label therefore trains on one label scheme and evaluates against another. It degraded macro-F1 from 0.94 to 0.30 — and nothing crashed, no warning appeared, and the degraded numbers looked like a plausible drift finding. Resolving labels through the per-day name list fixes it.
Anyone using per-period label encoding across time periods should check this. scripts/verify_labels.py is the check.
pip install -e ".[dev]"
pip install cesnet-datazoo
pytest tests/ -v
python -m etc.run configs/experiments --all # synthetic, no download neededFor the CESNET results, the dataset path comes from data_root in
configs/base.yaml (default: data/). The dataset downloads itself on first
use — no manual download step.
python scripts/check_labels.py # demonstrates the label trap
python scripts/verify_labels.py # must PASS before anything else
python -m etc.run configs/experiments/11_top30_grouped.yaml
python scripts/summarise_results.pyFirst load downloads ~1.2 GB and takes a few minutes; subsequent loads use a
cache in data/cesnet_tls22/.
Group-aware splits by default, with the leaky option kept deliberately. random_flow splitting is available because reproducing the inflated numbers is necessary to quantify the gap — but every result carries a leaky flag and the runner prints [LEAKY] beside it. A leaky number cannot quietly become a headline number.
Fixed-width temporal windows. An expanding window would make early folds train on less data than later ones, so fold-to-fold variation would measure training-set size rather than drift. A regression test enforces the fixed width.
Macro AUC ignores classes absent from a fold. A class with no positives in the test set has an undefined one-vs-rest AUC; averaging it in as 0.5 drags the macro figure toward chance. With many classes and per-day folds this is the common case, not an edge case.
SNI never reaches a feature extractor. It is retained in the schema only so a test can assert its absence. Including it turns the task into string matching.
Macro-F1 and balanced accuracy lead. Accuracy is reported for comparability with prior work that reports nothing else.
configs/experiments/ one YAML per experiment; config + commit identifies a run
src/etc/data/ schema, loaders, split strategies, synthetic generator
src/etc/features/ feature families, independently ablatable
src/etc/models/ sklearn baselines; torch optional
src/etc/eval/ protocols and metrics
scripts/ verification and diagnostic tools
tests/ leakage tests
results/ JSONL, one record per fold
- Harness: schema, splits, leakage tests, config system, runner
- CESNET-TLS22 loader with per-day label resolution
- Feature ablation and temporal evaluation at ten-day scale
- CESNET-TLS-Year22 — twelve months, the real test of temporal drift
- ISCX VPN-nonVPN 2016 — the only dataset here with per-user structure, so the only place the leakage-gap experiment can run
- CESNET-QUIC22 — QUIC/TCP contrast
- Full 183-class results
- Sequence models (1D CNN, transformer)
Single dataset, one two-week capture period from 2021, TCP/TLS only, gradient boosting only, and a single random seed for the headline numbers. The feature-ablation result is well separated relative to fold variance; the temporal null result is bounded by the ten-day window and should not be read as a general claim about drift.
The group split here groups by capture day, not by user — CESNET-TLS22 is anonymised backbone traffic with no client identifier. A true user-level leakage test requires ISCX VPN-nonVPN 2016, which is not yet implemented. No claim about user-level leakage is made.
CESNET-TLS22: Luxemburk & Čejka, Fine-grained TLS services classification with reject option, Computer Networks, 2023. DOI 10.1016/j.comnet.2022.109467
MIT for code. Datasets carry their own terms.