openswmm_gymnasium.rewards#

openswmm_gymnasium.rewards#

Reward term registry + built-in terms. Plan §5.

All built-in terms are framed as B{cost-to-minimize} (positive contribution = bad outcome), except those flagged direction="maximize" (e.g. ReliabilityMargin). The env negates the aggregated cost so higher reward is better, per Gymnasium convention (plan §0 #5).

Built-in terms self-register with RewardRegistry at import time so RewardRegistry.get("flooding_volume") works without any explicit imports beyond this module.

author:

Caleb Buahin

copyright:

Copyright (c) 2026 Caleb Buahin

license:

MIT

class openswmm_gymnasium.rewards.CSOVolume(node_ids, name='cso_volume')[source]#

Bases: FloodingVolume

Overflow volume restricted to a set of CSO / relief nodes.

Mechanically identical to FloodingVolume but defaults the term name to "cso_volume" and B{requires} an explicit node_ids list — running over all nodes would conflate CSO with general flooding.

Parameters:
  • node_ids (Sequence[str])

  • name (str)

class openswmm_gymnasium.rewards.FloodingVolume(node_ids=None, name='flooding_volume')[source]#

Bases: object

Sum of overflow volume across a set of nodes, per env step.

Computed as sum_i overflow_rate_i * dt_seconds, where overflow_rate is read from openswmm.engine.Nodes.get_overflow (project flow units). If node_ids is None, all nodes in the model contribute.

@ivar name: "flooding_volume" by default. @ivar direction: Always "minimize".

Parameters:
  • node_ids (Sequence[str] | None)

  • name (str)

bind(adapter)[source]#
Parameters:

adapter (SolverAdapter)

Return type:

None

direction = 'minimize'#
reset()[source]#

No cross-step state; no-op.

Return type:

None

step(adapter, dt_seconds)[source]#
Parameters:
Return type:

float

class openswmm_gymnasium.rewards.PeakOutflow(link_ids, name='peak_outflow')[source]#

Bases: object

Running-max link flow; per-step contribution is the new-peak increment.

Each step, for each tracked link, computes max(0, |flow_now| - peak_so_far). When a new peak is set the increment equals the delta; otherwise the contribution is zero. Cumulative reward over the episode therefore equals the maximum absolute flow observed at each link (summed across links).

Plan §5.1 framing: minimise peak outflow.

@ivar name: "peak_outflow" by default. @ivar direction: Always "minimize".

Parameters:
  • link_ids (Sequence[str])

  • name (str)

bind(adapter)[source]#
Parameters:

adapter (SolverAdapter)

Return type:

None

direction = 'minimize'#
reset()[source]#
Return type:

None

step(adapter, dt_seconds)[source]#
Parameters:
Return type:

float

class openswmm_gymnasium.rewards.ReliabilityMargin(node_ids=None, name='reliability_margin')[source]#

Bases: object

Per-step minimum freeboard across a set of nodes.

Freeboard at a node = max_depth - depth. The term returns the minimum across the tracked set, clamped at zero (surcharged nodes yield 0 rather than a negative penalty — the FloodingVolume term already accounts for overflow). Direction: B{maximize}.

@ivar name: "reliability_margin" by default. @ivar direction: Always "maximize".

Parameters:
  • node_ids (Sequence[str] | None)

  • name (str)

bind(adapter)[source]#
Parameters:

adapter (SolverAdapter)

Return type:

None

direction = 'maximize'#
reset()[source]#

No cross-step state; no-op.

Return type:

None

step(adapter, dt_seconds)[source]#
Parameters:
Return type:

float

class openswmm_gymnasium.rewards.RewardRegistry[source]#

Bases: object

Singleton-style registry of reward-term classes.

@cvar _registry: Backing dict mapping term name to class.

classmethod clear()[source]#

Drop all registrations. Intended for tests only.

Return type:

None

classmethod get(name)[source]#

Look up a registered term class by name.

Parameters:

name (str) – Registered identifier.

Returns:

The class previously passed to register.

Return type:

type

Raises:

KeyError – If name is not registered.

classmethod names()[source]#

List all registered names in sorted order.

Return type:

list[str]

classmethod register(name, term_cls)[source]#

Register a reward-term class under name.

Parameters:
  • name (str) – Unique short identifier (typically matches the instance’s default name attribute).

  • term_cls (type) – The reward-term class to register.

Raises:

ValueError – If name is already registered to a different class.

Return type:

None

class openswmm_gymnasium.rewards.RewardTerm(*args, **kwargs)[source]#

Bases: Protocol

Interface every reward term must satisfy.

@cvar name: Unique short identifier for the term. @cvar direction: "minimize" or "maximize".

bind(adapter)[source]#

Resolve symbolic IDs against the freshly-opened engine.

Parameters:

adapter (SolverAdapter) – Adapter wrapping the open solver.

Return type:

None

direction: str#
name: str#
reset()[source]#

Clear any cross-step accumulators inside the term.

Return type:

None

step(adapter, dt_seconds)[source]#

Compute this term’s contribution for one env step().

Parameters:
  • adapter (SolverAdapter) – Adapter wrapping the running solver.

  • dt_seconds (float) – Elapsed real-time of the step, in seconds.

Returns:

Non-negative contribution. Sign flipping for the agent happens in the env, not in the term.

Return type:

float

class openswmm_gymnasium.rewards.SetpointSmoothness(link_ids, name='setpoint_smoothness')[source]#

Bases: object

L2 norm-squared of Δsetting between successive env steps.

Each step, for each tracked link, reads the current control setting and adds (setting_now - setting_prev) ** 2 to the contribution. The first step after reset returns zero (no previous setting yet). Direction: B{minimize}.

@ivar name: "setpoint_smoothness" by default. @ivar direction: Always "minimize".

Parameters:
  • link_ids (Sequence[str])

  • name (str)

bind(adapter)[source]#
Parameters:

adapter (SolverAdapter)

Return type:

None

direction = 'minimize'#
reset()[source]#
Return type:

None

step(adapter, dt_seconds)[source]#
Parameters:
Return type:

float