Rewards#

All built-in reward terms are framed as cost-to-minimise internally (positive contribution = bad outcome). The env negates the aggregated cost so higher reward is better, per Gymnasium convention.

Maximise-direction terms (e.g. ReliabilityMargin) declare direction = "maximize"; the env’s composer handles sign-flipping automatically.

Built-in terms#

Term

Direction

Description

FloodingVolume

minimise

Sum of overflow volume across nodes per step

CSOVolume

minimise

Same math, restricted to tagged CSO nodes

PeakOutflow

minimise

Running-max increment of link flow

ReliabilityMargin

maximise

Minimum freeboard across nodes

SetpointSmoothness

minimise

L2-norm-squared of Δsetting between steps

Custom terms#

Subclass RewardTerm (it is a runtime-checkable Protocol). Implement bind, reset, step:

class MyCustomTerm:
    name = "my_term"
    direction = "minimize"

    def bind(self, adapter):
        self._idx = adapter.links.get_index("OUT")

    def reset(self):
        pass

    def step(self, adapter, dt_seconds):
        return float(adapter.links.get_flow(self._idx)) * dt_seconds

Register with RewardRegistry if you want name-based lookup:

from openswmm_gymnasium.rewards import RewardRegistry

RewardRegistry.register("my_term", MyCustomTerm)
cls = RewardRegistry.get("my_term")