A distributed lookup table for caching Tiramisu program execution measurements on HPC clusters. Designed for Slurm-based clusters with Lustre shared filesystems where standard file locking (fcntl/flock) is unreliable across nodes.
Multiple users and up to 64+ concurrent worker nodes can safely read and write to the same database without coordination beyond the library itself.
pip install -e .The only runtime dependency is py-cpuinfo. For running unit tests, install with pip install -e ".[dev]".
from tirastore import TiraStore
# Connect (creates the DB if it doesn't exist)
store = TiraStore(
"/lustre/shared/tiramisu_cache.db",
source_project="autoscheduler-v2",
)
# Record a measurement
store.record(
program_name="blur",
program_source_code="void blur() { ... }",
tiralib_schedule_string="S(L0,L1,4,8,comps=['c1'])",
is_legal=True,
execution_times=[0.042, 0.039, 0.041],
)
# Look up a previous measurement
result = store.lookup("blur", "void blur() { ... }", "S(L0,L1,4,8,comps=['c1'])")
if result is not None:
print(result.is_legal) # True
print(result.execution_times) # [0.042, 0.039, 0.041]
print(result.hostname) # node042
print(result.source_project) # autoscheduler-v2from tirastore import TiraStore
store = TiraStore("/lustre/shared/cache.db", source_project="my_experiment")
def process(program_name, source_code, schedule):
# Check cache first
result = store.lookup(program_name, source_code, schedule)
if result is not None:
return result # Cache hit — skip computation
# Cache miss — run the actual measurement
is_legal, times = run_tiramisu_measurement(source_code, schedule)
# Store for future lookups
store.record(program_name, source_code, schedule, is_legal, times)
return store.lookup(program_name, source_code, schedule)Constructor parameters:
| Parameter | Type | Default | Description |
|---|---|---|---|
db_path |
str or Path |
(required) | Path to the SQLite database file. Created if it doesn't exist. |
source_project |
str |
"" |
Project name stored as metadata on every new record. |
cpu_model |
str or None |
None |
CPU model for the DB. Auto-detected if not provided. |
slurm_cpus |
str or None |
None |
SLURM_CPUS_PER_TASK. Read from environment if not provided. |
allow_cpu_mismatch |
bool |
False |
Allow writes even if the local CPU doesn't match the DB's CPU metadata. |
stale_lock_timeout |
float |
600.0 |
Seconds before a held lock is considered stale (holder crashed). |
Returns a LookupResult if the record exists, otherwise None.
LookupResult fields: is_legal, execution_times, hostname, username, creation_date, update_date, source_project. Call .to_dict() to get a plain dictionary.
store.record(program_name, program_source_code, tiralib_schedule_string, is_legal, execution_times, overwrite=False)
Stores a measurement result. Returns True if a write occurred, False if the record already existed (and overwrite was False).
- If
is_legal=True,execution_timesmust be a non-empty list of floats. - If
is_legal=False,execution_timescan beNone. - Set
overwrite=Trueto replace an existing record. - The schedule string is validated before recording. Invalid schedules raise
ValueError.
Batch-insert multiple schedules for the same program in one operation. Returns the number of records actually written.
Each entry in schedules is a dict with keys tiralib_schedule_string (str), is_legal (bool), and execution_times (list of float or None).
All entries are validated upfront — if any schedule is invalid, nothing is written.
store.record_many(
program_name="blur",
program_source_code="void blur() { ... }",
schedules=[
{"tiralib_schedule_string": "S(L0,L1,4,8,comps=['c1'])", "is_legal": True, "execution_times": [0.04]},
{"tiralib_schedule_string": "I(L0,L1,comps=['c1'])", "is_legal": False, "execution_times": None},
{"tiralib_schedule_string": "R(L0,comps=['c1'])", "is_legal": True, "execution_times": [0.03, 0.031]},
],
)Returns True if a record exists for the given input.
Retrieve the source code of a program by name. Returns a list of dicts, each with program_hash and source_code. Multiple entries are returned if different source versions share the same name.
sources = store.get_program_source("blur")
for s in sources:
print(s["program_hash"], s["source_code"][:80])Retrieve all records for a specific program (all schedules recorded for that program). Returns a list of LookupResult objects.
records = store.get_program_records("blur", "void blur() { ... }")
for r in records:
print(r.schedule, r.is_legal, r.execution_times)Create a snapshot copy of the database file. If no path is given, a timestamped file is created next to the database (e.g. store_20260221T153012Z.db). Returns the Path to the backup.
backup = store.backup() # auto-timestamped
backup = store.backup("/shared/backups/v1.db") # custom pathExport the entire database to JSON or JSONL. Returns the Path to the exported file.
fmt="json"— a single JSON object keyed by program name.fmt="jsonl"— one JSON object per line, one program per line.
If the same program name has multiple source-code versions, keys are suffixed with _v1, _v2, etc.
store.export("dataset.json", fmt="json")
store.export("dataset.jsonl", fmt="jsonl")Exported structure:
{
"blur": {
"Tiramisu_cpp": "void blur() { ... }",
"schedules_list": [
{"schedule_str": "S(L0,L1,4,8,comps=['c1'])", "is_legal": true, "execution_times": [0.04]},
{"schedule_str": "I(L0,L1,comps=['c1'])", "is_legal": false, "execution_times": null}
],
"program_name": "blur"
}
}| Method | Description |
|---|---|
store.count() |
Total number of records in the database. |
store.program_count() |
Total number of distinct programs. |
store.stats() |
Summary dict: total/legal/illegal counts, program count, distinct users, projects, CPU info. |
store.keys(limit=0, offset=0) |
List of record SHA-256 keys (with optional pagination). |
store.get(key) |
Retrieve a raw record dict (joined with program data) by its SHA-256 key. |
store.put(key, program_hash, schedule, result_json, overwrite=False) |
Low-level insert/update by raw key (admin use). |
store.delete(key) |
Delete a record by its SHA-256 key. |
| Property | Description |
|---|---|
store.writes_allowed |
bool — whether writes are permitted on this connection. |
store.cpu_model |
CPU model string stored in the database metadata. |
store.slurm_cpus |
SLURM_CPUS_PER_TASK value stored in the database metadata. |
These are also importable from tirastore for direct use:
| Function | Description |
|---|---|
normalize_schedule(schedule) |
Normalize a schedule string (strip whitespace, unify quote style). |
validate_schedule(schedule) |
Validate a schedule string. Returns (bool, reason_string). |
normalize_program(source) |
Normalize a program source string for hashing (strip comments, includes, whitespace). |
make_program_hash(source) |
Compute the SHA-256 hash of the normalized program source. |
All records are stored in a single SQLite file on the shared Lustre filesystem. This keeps the file count minimal (important for HPC storage quotas).
Program deduplication: Program source code is stored in a separate programs table, keyed by the SHA-256 hash of its normalized form. The records table references programs by hash. This means if you record 100 different optimization sequences for the same program, the source code is stored only once.
Program normalization: Before hashing, program source code is normalized by stripping block comments (/* ... */), single-line comments (// ...), #include directives, and all whitespace. This ensures that cosmetically different versions of the same program (different formatting, comments, includes) produce the same hash and share a single programs-table entry. The original readable source code is always what gets stored; normalization is only used for hash computation.
Schedule normalization: Schedule strings are normalized (whitespace removed, comp names single-quoted) before hashing and storing. This means R( L0 , comps=["c1"] ) and R(L0,comps=['c1']) are treated as identical.
Schedule validation: Schedule strings are validated against the grammar of known Tiramisu transformations (S, I, R, P, T2, T3, U, F) before recording. Invalid schedules are rejected with a descriptive error.
Record keying: The record key is the SHA-256 hex digest of the canonical JSON of {program_hash, normalized_schedule}. Identical logical inputs always produce the same key regardless of cosmetic differences in the source code or schedule formatting.
SQLite settings for Lustre:
journal_mode=DELETE— WAL mode uses shared memory/mmap which breaks on Lustre.synchronous=FULL— ensure durability.busy_timeout=0— we rely on our own lock, not SQLite's internal locking.
Since fcntl/flock don't work reliably across Lustre nodes, TiraStore uses an atomic hard-link based mutex:
- Acquire: Create a temp file with a unique name (hostname, PID, timestamp), then
link()it to the lock path.link()is atomic on POSIX/Lustre — it fails if the target already exists. On failure, retry with exponential backoff + random jitter. - Release:
unlink()the lock file. - Stale detection: If a lock is older than the stale timeout (default 10 minutes), assume the holder crashed and break the lock.
Every database operation follows this protocol:
acquire lock -> open SQLite connection -> read/write -> close connection -> release lock
The database stores the CPU model and SLURM_CPUS_PER_TASK at creation time because execution times are only comparable across identical hardware configurations.
When connecting to an existing database:
- If the local machine's CPU matches the DB metadata: full read/write access.
- If there's a mismatch: writes are blocked (lookups still work) and a warning is printed. Override with
allow_cpu_mismatch=Truefor admin tasks like importing data.
The database file is created with mode 666 (world-readable/writable) and the parent directory is set to mode 1777 (sticky bit). No common Unix group is required.
pip install -e ".[dev]"
pytest tests/ -vA test suite is included to validate locking and concurrency across actual Slurm cluster nodes. Each worker runs as a separate srun job to ensure execution on different nodes.
# Set your Lustre path and cluster partition
export TIRASTORE_TEST_DIR=/lustre/scratch/$USER/tirastore_test
export SLURM_PARTITION=batch
# Launch: 8 workers, 50 unique records each, 20 contested records
bash tests/multinode/launch_test.sh 8 50 20The test runs four phases per worker:
- Unique writes — each worker writes records only it owns.
- Contested writes — all workers attempt to write the same records. Only one writer should win per record.
- Cross-reads — each worker reads its own records, contested records, and samples from other workers.
- Overwrite test — one worker overwrites a contested record.
A verification script runs automatically at the end, checking record counts, data integrity, error-free execution, and that workers actually ran on distinct nodes.
Stores database-level configuration:
| Key | Description |
|---|---|
schema_version |
Schema version (currently 2). |
cpu_model |
CPU model string of the machine that created the DB. |
slurm_cpus |
SLURM_CPUS_PER_TASK at DB creation time. |
created_at |
ISO 8601 timestamp of DB creation. |
Stores program source code, deduplicated by content hash:
| Column | Type | Description |
|---|---|---|
program_hash |
TEXT PRIMARY KEY |
SHA-256 hex digest of the normalized source code. |
program_name |
TEXT |
Human-readable program name. |
source_code |
TEXT |
Original (readable) program source code. |
| Column | Type | Description |
|---|---|---|
key |
TEXT PRIMARY KEY |
SHA-256 hex digest of {program_hash, normalized_schedule}. |
program_hash |
TEXT |
Foreign key referencing programs(program_hash). |
schedule |
TEXT |
Normalized schedule string. |
result_json |
TEXT |
JSON of {is_legal, execution_times}. |
hostname |
TEXT |
Node that recorded this entry. |
username |
TEXT |
User that recorded this entry. |
creation_date |
TEXT |
ISO 8601 timestamp of initial creation. |
update_date |
TEXT |
ISO 8601 timestamp of last update. |
source_project |
TEXT |
Project name provided at connection time. |
tirastore/
__init__.py # Public API: TiraStore, LookupResult, normalize/validate helpers
tirastore.py # Main TiraStore class
_lock.py # HardLinkLock — distributed mutex via atomic link()
_store.py # SQLite storage backend (programs + records tables)
_keys.py # SHA-256 key/hash generation
_schedule.py # Schedule + program normalization and validation
tests/
test_keys.py # Key and hash generation tests
test_lock.py # Locking tests
test_schedule.py # Schedule/program normalization and validation tests
test_store.py # SQLite storage tests
test_tirastore.py # Integration tests
multinode/
launch_test.sh # Slurm launcher for concurrency test
test_worker.py # Per-node worker script
verify_results.py # Post-test verification