Skip to content

Folders and files

NameName
Last commit message
Last commit date

Latest commit

 

History

8 Commits
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

etc-bench

Which features actually carry the signal in encrypted traffic classification, and does performance degrade over time?

A harness for evaluating encrypted traffic classifiers under conditions the literature usually avoids: group-aware splits, temporal splits, feature-family ablation, and open-world evaluation. Built to test whether reported accuracy survives an honest setup.

Findings so far are on CESNET-TLS22 (30 classes, 189,190 flows, ten consecutive days in October 2021).

Results

Gradient boosting, macro-F1, mean ± sd across folds. Chance floor is 0.0335.

Setup Features macro-F1
Grouped split (by capture day) all three 0.9331 ± 0.0163
Temporal split (3-day window, test next day) all three 0.9425 ± 0.0077
Temporal packet sizes only 0.9284 ± 0.0282
Temporal flow stats + timing 0.8570 ± 0.0114
Temporal flow stats only 0.7552 ± 0.0115
Temporal timing only 0.7197 ± 0.0180
Chance floor 0.0335 ± 0.0010

1. No temporal degradation at ten-day scale

Testing on a future day (0.9425) performs no worse than testing on held-out days (0.9331). Whatever drift exists over ten days in a single month is below the noise floor of this setup.

This is a null result and it is reported as one. It does not generalise beyond the timescale tested, ten consecutive days in one October is a narrow window, and a year-long evaluation is the obvious next test. But at this scale, the concern that in-domain results overstate temporal robustness is not supported.

2. Packet size sequences dominate; the other families are largely redundant

Packet sizes alone recover 98.5% of full performance (0.9284 vs 0.9425). Removing them costs nine points (0.8570); removing everything else costs one.

Flow statistics and inter-packet timing each score around 0.72–0.76 in isolation, so they are not uninformative, they are redundant given packet sizes. This has a practical consequence: the cheapest feature to collect is also the one that matters.

3. A label-encoding trap that silently destroys results

The most useful methodological finding here came from a bug.

CESNET's AppSelection.ALL_KNOWN label-encodes applications against a list rebuilt per time period. An application that misses the minimum-support threshold on a given day drops out, shifting every index after it. Observed list sizes across the ten days: 183, 183, 183, 180, 181, 182, 182, 173, 181, 179.

Using the raw integer as a label therefore trains on one label scheme and evaluates against another. It degraded macro-F1 from 0.94 to 0.30 — and nothing crashed, no warning appeared, and the degraded numbers looked like a plausible drift finding. Resolving labels through the per-day name list fixes it.

Anyone using per-period label encoding across time periods should check this. scripts/verify_labels.py is the check.

Reproducing

pip install -e ".[dev]"
pip install cesnet-datazoo
pytest tests/ -v

python -m etc.run configs/experiments --all     # synthetic, no download needed

For the CESNET results, the dataset path comes from data_root in configs/base.yaml (default: data/). The dataset downloads itself on first use — no manual download step.

python scripts/check_labels.py                  # demonstrates the label trap
python scripts/verify_labels.py                 # must PASS before anything else
python -m etc.run configs/experiments/11_top30_grouped.yaml
python scripts/summarise_results.py

First load downloads ~1.2 GB and takes a few minutes; subsequent loads use a cache in data/cesnet_tls22/.

Design decisions

Group-aware splits by default, with the leaky option kept deliberately. random_flow splitting is available because reproducing the inflated numbers is necessary to quantify the gap — but every result carries a leaky flag and the runner prints [LEAKY] beside it. A leaky number cannot quietly become a headline number.

Fixed-width temporal windows. An expanding window would make early folds train on less data than later ones, so fold-to-fold variation would measure training-set size rather than drift. A regression test enforces the fixed width.

Macro AUC ignores classes absent from a fold. A class with no positives in the test set has an undefined one-vs-rest AUC; averaging it in as 0.5 drags the macro figure toward chance. With many classes and per-day folds this is the common case, not an edge case.

SNI never reaches a feature extractor. It is retained in the schema only so a test can assert its absence. Including it turns the task into string matching.

Macro-F1 and balanced accuracy lead. Accuracy is reported for comparability with prior work that reports nothing else.

Layout

configs/experiments/   one YAML per experiment; config + commit identifies a run
src/etc/data/          schema, loaders, split strategies, synthetic generator
src/etc/features/      feature families, independently ablatable
src/etc/models/        sklearn baselines; torch optional
src/etc/eval/          protocols and metrics
scripts/               verification and diagnostic tools
tests/                 leakage tests
results/               JSONL, one record per fold

Status

  • Harness: schema, splits, leakage tests, config system, runner
  • CESNET-TLS22 loader with per-day label resolution
  • Feature ablation and temporal evaluation at ten-day scale
  • CESNET-TLS-Year22 — twelve months, the real test of temporal drift
  • ISCX VPN-nonVPN 2016 — the only dataset here with per-user structure, so the only place the leakage-gap experiment can run
  • CESNET-QUIC22 — QUIC/TCP contrast
  • Full 183-class results
  • Sequence models (1D CNN, transformer)

Limitations

Single dataset, one two-week capture period from 2021, TCP/TLS only, gradient boosting only, and a single random seed for the headline numbers. The feature-ablation result is well separated relative to fold variance; the temporal null result is bounded by the ten-day window and should not be read as a general claim about drift.

The group split here groups by capture day, not by user — CESNET-TLS22 is anonymised backbone traffic with no client identifier. A true user-level leakage test requires ISCX VPN-nonVPN 2016, which is not yet implemented. No claim about user-level leakage is made.

Citation

CESNET-TLS22: Luxemburk & Čejka, Fine-grained TLS services classification with reject option, Computer Networks, 2023. DOI 10.1016/j.comnet.2022.109467

Licence

MIT for code. Datasets carry their own terms.

About

Harness for evaluating encrypted traffic classification under group-aware and temporal splits. Finds that packet size sequences carry 98.5% of the signal, and documents a label-encoding bug that silently degrades results by two thirds.

Topics

Resources

Stars

2 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages