Development#
Source code#
The source repository is hosted on GitHub:
git clone git@github.com:nismod/transport_flow_model.git
cd transport_flow_model
If SSH access is not configured, use the HTTPS URL instead:
git clone https://github.com/nismod/transport_flow_model.git
cd transport_flow_model
Development environment#
The project uses pixi for the development environment. From the repository
root, run project tasks through pixi run so commands use the pinned
dependencies from pyproject.toml and pixi.lock.
pixi install
For a direct editable install without Pixi, use:
pip install -e .
Pixi is the preferred workflow for contributors and automation agents because it also installs the configured development and documentation dependencies.
Tests#
Run the full test suite with:
pixi run test
This currently runs:
python -m pytest tests
Linting and formatting#
The project currently uses Ruff for linting and formatting.
Run lint checks with:
pixi run lint
Run formatting with:
pixi run format
These tasks are defined in pyproject.toml. If a new lint or formatting tool
is added later, add it there and prefer a Pixi task over documenting an ad hoc
command.
Type checking#
No static type checker is currently configured for this project. There is no
mypy, pyright, or equivalent Pixi task in pyproject.toml at the
moment.
For agents and CI automation, do not assume a type-check command exists. Add a type-checking dependency and a dedicated Pixi task before making type checks a required validation step.
Documentation#
Build the HTML documentation with:
pixi run docs
Run doctests with:
pixi run doctest
Benchmark data preparation#
The generated West Yorkshire benchmark dataset is created by:
pixi run prepare-benchmark-data
This task runs scripts/benchmark_preparation.py. It uses OSMnx to download
the West Yorkshire driving road network and residential/commercial land-use
polygons. The script writes generated data under
benchmark_data/west_yorkshire/:
osmnx_road_network.gpkgwithnodesandedgeslayers.osmnx_landuse_zones.gpkgwith all downloaded land-use polygons and the subset used for OD generation.processed_data/network/network.csvfor flow allocation.processed_data/od/od.csvfrom the radiation-model OD estimate.processed_data/damages/failure_set.csvfor disruption benchmarking.
The default OD generation uses the 10 largest residential and 10 largest
commercial polygons. This keeps preparation and later allocation runs practical
on the current pure-Python shortest-path implementation. Use
--max-zones-per-landuse to scale the OD matrix up or 0 to use all
downloaded zones:
pixi run prepare-benchmark-data -- --max-zones-per-landuse 25 --overwrite
Use --overwrite when regenerating the benchmark dataset. benchmark_data/
and OSMnx cache/ output are ignored by Git.
Benchmarking scripts#
The repository includes a stdlib-only benchmark harness for integration timing of the two command-line flow scripts:
scripts/flow_model/flow_allocation.pyscripts/flow_model/flow_disruptions.py
Run the smoke benchmark with the example config:
pixi run benchmark-scripts-smoke
Run the default benchmark task with:
pixi run benchmark-scripts
The benchmark script defaults to config.example.json and accepts additional
configs with repeated --config arguments:
python scripts/benchmark_flow_scripts.py \
--config config.example.json \
--config path/to/larger-config.json \
--repeats 3
Each benchmark run copies the configured input data into a temporary working directory and writes script outputs to temporary results paths. This keeps the tracked example data and results directories unchanged while still measuring the script workflow, including CSV input/output.
Benchmark outputs are written under benchmark_results/:
flow_script_benchmark.csvappends timing history.flow_script_benchmark_latest.jsonstores the latest run.
benchmark_results/ is ignored by Git. Use --keep-workdirs when debugging
script outputs from an individual benchmark run.
Useful benchmark options:
python scripts/benchmark_flow_scripts.py --help
python scripts/benchmark_flow_scripts.py --repeats 5
python scripts/benchmark_flow_scripts.py --max-seconds 30
python scripts/benchmark_flow_scripts.py --keep-workdirs
To benchmark the generated West Yorkshire data, prepare the data first and pass the West Yorkshire config:
pixi run prepare-benchmark-data -- --overwrite
python scripts/benchmark_flow_scripts.py --config config.west_yorkshire.json
Profiling scripts#
CPU/time flamegraphs are generated with py-spy through
scripts/profile_flow_scripts.py. The default profile config is
config.west_yorkshire.json and output SVGs are written to
profile_results/:
pixi run profile-flow-scripts
This writes:
profile_results/flow_allocation.svgprofile_results/flow_disruptions.svg
The default profile records a bounded 30-second sampling window per script at 10 samples per second. This avoids turning large benchmark profiles into much longer full-script runs. Increase the duration or sampling rate when more detail is needed:
pixi run profile-flow-scripts -- --duration 60 --rate 25
Profile one script at a time with:
pixi run profile-flow-allocation
pixi run profile-flow-disruptions
When profiling only disruptions, the wrapper runs allocation first so
flow_disruptions.py has the required flow_od_paths inputs. If those
outputs already exist and should be reused, pass --skip-setup-allocation:
pixi run profile-flow-disruptions -- --skip-setup-allocation
profile_results/ is ignored by Git.
Rust extension#
The package uses a Rust core. Build the native extension when developing or benchmarking the Rust code:
pixi run extension-build
This task runs maturin develop --manifest-path Cargo.toml and installs the
PyO3 module as transport_flow_model._core in the Pixi environment. The
public helper module is transport_flow_model.core.
The extension is structured in two layers:
A Python-independent Rust core for graph, OD, allocation, disruption, unit tests, and Criterion benchmarks.
A PyO3 wrapper that exchanges in-memory Arrow IPC streams with Python.
The Arrow IPC boundary keeps file I/O in the wrapper language. Python callers
can pass pyarrow.Table, pyarrow.RecordBatch, or pandas.DataFrame to
transport_flow_model.core.allocate_arrow and
transport_flow_model.core.disrupt_arrow. Other language wrappers can target
the same Arrow stream schemas without depending on Python data-frame internals.
Run unit tests with:
pixi run extension-test
Run microbenchmarks with:
pixi run extension-bench
The scaffold intentionally avoids extra graph/routing crates for now. Add new Rust dependencies only when they replace substantial local complexity or are needed for a specific algorithmic feature.
Pixi command reference#
The current Pixi tasks are:
Command |
Purpose |
|---|---|
|
Run |
|
Run Ruff lint checks. |
|
Format Python code with Ruff. |
|
Build Sphinx HTML documentation. |
|
Run Sphinx doctests. |
|
Generate the ignored West Yorkshire benchmark input dataset. |
|
Run a quick script-level benchmark using |
|
Run script-level timing benchmarks and write CSV/JSON outputs. |
|
Write flamegraphs for both allocation and disruption scripts. |
|
Write a flamegraph for allocation only. |
|
Write a flamegraph for disruption only. |
|
Build and install the PyO3 Rust extension. |
|
Run unit tests for the extension. |
|
Run Criterion benchmarks for the extension. |