Module Responsibilities¶
Experiment configs (configs/experiments/*.yml) define one launchable
unit: dataset, stage, representation, resources, and the stage-specific
payload under config.
Typed config (graphids/exp/config.py) validates YAML and builds a
RunConfig. It owns the contract for ExperimentConfig, ResourceConfig,
stage payloads, output paths, MLflow tags/params, and the representation drift
check between top-level metadata and data.source.
Primitives (graphids/primitives*.py) are the public object factory
surface for YAML specs. Data, model, loss, scaler, representation, ID
encoding, and discovery primitives live here or are re-exported here.
Ray backend (graphids/exp/ray_backend.py) owns driver-side launch. It
translates GraphIDS run metadata to Ray Train configs, starts or connects to
Ray, and reports Ray result metrics.
Ray backend (graphids/exp/ray_backend.py) owns driver launch, Ray worker lifecycle,
object construction from YAML specs, journal/offline-ingest writes, and Lightning
fit/test execution with Ray Train.
SLURM submit (graphids/exp/slurm.py, graphids/cli/exp.py) validates an
experiment YAML, renders a Ray allocation sbatch script, and submits it. The
allocation starts Ray head/workers and runs
python -m graphids exp launch <yaml> --address "${RAY_ADDRESS}" after
sourcing scripts/slurm/_preamble.sh.
Data sources and datamodules (graphids/core/data/) own raw CAN loading,
temporal representation selection, cache paths, metadata, and Lightning
dataloaders. CANBusTemporalSource builds train/validation/test
TemporalData caches, and TemporalDataModule serves them through PyG's
TemporalDataLoader.
Preprocessing (graphids/core/data/preprocessing/) turns normalized CAN
rows into temporal event tables and PyG TemporalData. The live
representation is kind: temporal; windowed graph materialization and graph
budgeting are no longer part of the primary training path.
Models (graphids/core/models/) own Lightning modules and metrics. The
live model families consume temporal event batches: stateless event
classification, stateful recurrent classification, causal event attention, and
temporal VGAE-style event reconstruction.
Callbacks (graphids/core/callbacks.py) hold graphids-specific Lightning
policy such as Sha256ModelCheckpoint, tau-norm, and VRAM drift warnings.
MLflow (graphids/_mlflow.py) resolves the shared tracking URI, builds the
Lightning MLflow logger, and starts/stops MLflow system-metrics monitoring for
Lightning-created runs.
The live flow is:
configs/experiments/<run>.yml
-> ExperimentConfig.from_yaml
-> ExperimentConfig.build_run
-> gx exp launch OR gx exp submit
-> graphids.exp.ray_backend.launch_run
-> graphids.exp.ray_backend worker loop
-> fit | test
-> MLflow + .graphids/manifest.json + .graphids/events.jsonl
The old graphids/plan row renderer, gx run, gx plans submit, and
graphids/orchestrate.py row dispatcher are historical and should not be used
for new work.