Resource-aware object-centric process simulation.
RAOCPS takes an OCEL 2.0 event log, mines a simulation model from it, and replays that model to produce a new synthetic OCEL 2.0 log that behaves like the original one.
Because resources carry calendars and capabilities, waiting time in the simulated log emerges from contention and availability rather than being sampled from the original log.
Every run works inside a project directory: mining creates it next to the input log and writes the model into it, and simulation reads that directory, writes the simulated log back into it, and renders HTML reports beside it.
Requires Python ≥ 3.12.
uv syncimport raocps
# 1. Mine a simulation model. `employees` are resources, `products` and
# `customers` are dropped from the log entirely.
project_dir = raocps.mine_simulation_model(
ocel_path="data/order-management.sqlite",
resource_object_types=["employees"],
objects_to_delete=["products", "customers"],
)
# 2. Adjust how much to generate. Everything not named here keeps the value
# mining derived from the log.
raocps.define_generation_parameters(
project_dir=project_dir,
seed="test",
object_type="orders",
object_count=10,
)
# 3. Generate the objects and simulate them.
raocps.simulate(project_dir=project_dir)This writes <project_dir>/results/ocel-simulated.sqlite — a full OCEL 2.0 log
you can feed straight back into mine_simulation_model — plus reports under
<project_dir>/reports/.
Steps 2 and 3 can be repeated on the same project directory without re-mining: change a parameter, simulate again.
main.py and second_pass.py in the repository root are runnable versions of
this, over the logs in data/.
A run writes a small static site into <project_dir>/reports/: the discovered
Petri net, the original and generated object graphs, a side-by-side comparison
of the original and generated schema distributions, the resource calendars, and
a step-by-step replay of the simulation. Serve it with:
python -m http.server -d <project_dir>/reports
Then open http://localhost:8000 — index.html links to every report.
The replay is by far the largest thing a run writes, since it records the
marking after every simulation step. Pass reports= to skip it, or to skip
reports altogether:
raocps.simulate(project_dir, reports=raocps.Reports.WITHOUT_SIMULATION_STEPS)raocps exports three functions, two enums and two column names. Everything
else is internal.
Reads an event log and writes a simulation model. Creates a fresh project
directory next to the log — <log-stem>-simulation-<n>, counting up so an
existing one is never overwritten — and returns its path.
| Parameter | Type | Default | Meaning |
|---|---|---|---|
ocel_path |
Path | str |
required | The OCEL 2.0 log. Any format pm4py can read: .sqlite, .jsonocel, .xml. SQLite logs are validated against the OCEL 2.0 relational schema before anything else runs. |
resource_object_types |
list[str] |
required | Object types to treat as resources rather than as things flowing through the process — machines, workers, vehicles. They are lifted out of the log into resource profiles and calendars, and the activities that used them become resource requirements. Pass [] for a log without resources. |
objects_to_delete |
list[str] | None |
None |
Object types to drop from the log before mining. Use for types that add nothing to the control flow (reference data such as products or customers). Naming a type the log does not contain is an error. |
events_to_delete |
list[str] | None |
None |
Activities to drop from the log before mining, for trimming a process down to the part you care about. Naming an activity the log does not contain is an error. |
reports |
Reports |
Reports.ALL |
Which reports to render for the mining half of the run. |
calendar_granularity |
CalendarGranularity |
QUARTER_HOUR |
Slot width of the weekly resource calendars. A property of the model: it is written into the calendars as they are mined, and simulate infers it back from them. |
object_relationship_depth |
int | None |
None |
How far the schema distributions measure relationships out from each object type, which is exactly as far as the generator then follows them. Defaults to the diameter of the object type graph, i.e. far enough to reach every type from any other. A path longer than that has to revisit a type, and describes no group of objects a shorter one does not. A property of the model, fixed at mining time. |
allow_silent_objects |
bool |
False |
What to do about an activity no object type leads (see Silent objects): refuse the log, or invent a type for it. A property of the model, fixed at mining time. |
event_begin_time_column |
str | None |
None |
The event column holding the instant each activity began, which is what service times are measured from. None guesses it (see Activity lifecycle). Pass raocps.BEGIN_TIME_COLUMN for a log this package simulated with save_activity_times=True. A column the log does not have, or one with gaps, is an error rather than a silent fall back to guessing. |
event_enable_time_column |
str | None |
None |
The event column holding the instant each activity became enabled. Mining does not use it — it is read so that reports and analysis can separate waiting from service — so leaving it None costs nothing. |
Overrides the generation parameters of an already-mined project, then saves them
back into the project directory. Reads what mining left on disk first, so naming
one field does not reset the others; passing None (or omitting a field) keeps
the current value. Keyword-only.
| Parameter | Type | Meaning |
|---|---|---|
project_dir |
Path | str |
The project directory returned by mine_simulation_model. |
object_type |
str |
The type to generate instances of; every other object is generated by following relationships out from these. Defaults to the alphabetically first type in the log. |
seed |
str |
Seeds every random draw in generation and simulation. The same seed and the same model give the same simulated log. |
object_count |
int |
How many objects of object_type to generate. Defaults to 10 % of the count observed in the log, at least 1. |
max_objects |
int |
Ceiling on the total number of generated objects, across all types (default 2000). A guard against a schema distribution that grows without bound. |
raise_on_max_objects |
bool |
Whether hitting max_objects is an error (default True) or just stops generation. |
simulation_start |
str |
ISO timestamp the simulated log starts at (default "2026-01-19T00:00:00", a Monday). |
refinement_effort |
int |
Ceiling on the link swaps proposed while fitting the generated object graph to the mined schema, per generated object (default 40). Reaching it is logged as a warning: the fit was still improving when it ran out. |
refinement_patience |
int |
How many proposals in a row may change nothing before the fit counts as settled (default 2000). Raising it fits closer for steeply diminishing returns — twenty times the default roughly halves the remaining distance for ten times the time. |
Both are counted in proposals rather than seconds, so the same seed and model always give the same population back whatever the machine does. A run logs how many swaps it took, out of how many it tried and how many it was allowed, what the fit gained and how long it spent, so these two can be set from evidence:
Fitted the object graph to the mined schema in 7.6s: 188 swaps taken out of
11907 proposals tried (of 431480 allowed), schema distance 0.6995 -> 0.5306 (-24.2%)
Returns the resulting GenerationParameters, which is also what
<project_dir>/simulation-model/generation-parameters.json now holds — editing
that file by hand is equivalent.
Runs the second half of the pipeline against a mined project: generates the
object population from the schema distributions, plays it through the
object-centric Petri net respecting resource capabilities and calendars, and
writes the resulting log to
<project_dir>/results/ocel-simulated.sqlite.
| Parameter | Type | Default | Meaning |
|---|---|---|---|
project_dir |
Path | str |
required | The project directory returned by mine_simulation_model. |
reports |
Reports |
Reports.ALL |
Which reports to render. |
save_activity_times |
bool |
False |
Whether to record the rest of each activity's lifecycle in the simulated log — ocel:timestamp:begin and ocel:timestamp:enable beside the ocel:timestamp that holds when it terminated. Pass True whenever the log will be mined again (see Activity lifecycle). |
Safe to call repeatedly on the same project directory; each call overwrites the previous results.
How much of the report site a run should write.
| Member | Effect |
|---|---|
Reports.ALL |
Everything, including the step-by-step simulation replay. |
Reports.WITHOUT_SIMULATION_STEPS |
Everything except the replay — which is by far the largest thing a run writes, since it snapshots the marking after every step. |
Reports.NONE |
No reports at all; only the model and the simulated log. |
Slot width of the weekly resource calendars: QUARTER_HOUR, HALF_HOUR or
HOUR. Finer granularity models availability more precisely at the cost of a
larger calendar. An IntEnum valued in minutes, so it survives a JSON round
trip as a plain number. Default: QUARTER_HOUR.
The event columns simulate(save_activity_times=True) writes —
"ocel:timestamp:begin" and "ocel:timestamp:enable" — exported so that a
second pass can name them without repeating the strings.
<log-stem>-simulation-<n>/
simulation-model/ the mined model, as JSON
results/ ocel-simulated.sqlite, the generated objects,
the final marking
reports/ the static report site
This repository was developed with the help of AI coding agents.
-
uvrun pre-commit install -
Use
uvto run every command. -
Tests:
uv run pytest -
Linting:
uv run pre-commit run --all-files
Golden tests. Every scenario in tests/scenarios.py runs through the pipeline
twice — on its own log, then on the log the first run simulated — and each
pinned artifact is compared against tests/golden/.
- Only the
fasttier runs by default. Others:-m slow. --regen-goldenrewrites the golden files instead of comparing them. Review the diff before committing.--covadds branch coverage overraocps/;--cov-report=htmlwriteshtmlcov/. Off by default, and it only measures the tier you select.