openswmm_gymnasium.wrappers#

openswmm_gymnasium.wrappers#

Composable gymnasium.Wrapper subclasses. Plan §6 layout / §10 P5.

  • RescaleBoxActions — recursively rescales every gymnasium.spaces.Box leaf of a (possibly nested) action space into a unit interval, so an agent that emits actions in [0, 1] (or any other window) is automatically translated into the env’s physical bounds.

  • MaskDesignAction — flatten a CIP+RTC env so the agent only sees the runtime portion; the design is sampled once at reset() and frozen for the episode.

  • MaskRuntimeAction — flatten a CIP+RTC env so the agent only sees the design portion; the runtime is filled with a fixed (midpoint) setting each step.

  • LinearScalarize — convert vector reward to scalar via dot product with a weight vector.

  • TchebycheffScalarize — convert vector reward to scalar via weighted-Tchebycheff distance to a utopia point.

  • ForecastObservation — append features from a user-supplied forecast callable to the observation vector.

  • RecordTrajectory — write one JSONL file per episode capturing actions / observations / rewards / info, consumed by the openswmm_gymnasium.viz module (§5.5) and the regression-test goldens (§8.3).

author:

Caleb Buahin

copyright:

Copyright (c) 2026 Caleb Buahin

license:

MIT

class openswmm_gymnasium.wrappers.ForecastObservation(env, forecast_fn, horizon)[source]#

Bases: Wrapper

Append forecast features to the observation each step.

The user supplies a callable forecast_fn(env, info) returning a 1-D float array of length horizon. The wrapped env’s observation space is extended with that many additional features (bounds [-inf, +inf]).

@ivar horizon: Number of forecast features appended each step.

Parameters:
  • env (gym.Env)

  • forecast_fn (Callable[[gym.Env, dict[str, Any]], np.ndarray])

  • horizon (int)

reset(*, seed=None, options=None)[source]#

Uses the reset() of the env that can be overwritten to change the returned data.

Parameters:
  • seed (int | None)

  • options (dict[str, Any] | None)

step(action)[source]#

Uses the step() of the env that can be overwritten to change the returned data.

Parameters:

action (Any)

class openswmm_gymnasium.wrappers.LinearScalarize(env, weights)[source]#

Bases: RewardWrapper

Scalar reward = dot(weights, vector_reward).

Weights need not sum to one; the wrapper does not normalise. Each component is multiplied by its corresponding weight before summation.

@ivar weights: 1-D weight vector, same length as the env’s reward

vector.

Parameters:
  • env (gym.Env)

  • weights (Sequence[float])

reward(reward)[source]#

Returns a modified environment reward.

Args:

reward: The env step() reward

Returns:

The modified reward

Parameters:

reward (ndarray)

Return type:

float

class openswmm_gymnasium.wrappers.MaskDesignAction(env, frozen_design=None)[source]#

Bases: _DictHalfMask

Hide design — agent acts only on the runtime portion.

The design is sampled once at reset() (from the env’s design subspace) and frozen for the episode. The chosen design appears in info["frozen_design_action"].

Parameters:
  • env (gym.Env)

  • frozen_design (Any | None)

reset(*, seed=None, options=None)[source]#

Uses the reset() of the env that can be overwritten to change the returned data.

Parameters:
  • seed (int | None)

  • options (dict[str, Any] | None)

step(action)[source]#

Uses the step() of the env that can be overwritten to change the returned data.

Parameters:

action (Any)

class openswmm_gymnasium.wrappers.MaskRuntimeAction(env)[source]#

Bases: _DictHalfMask

Hide runtime — agent acts only on the design portion.

Each step the runtime portion is filled with the per-Box midpoint of the env’s runtime subspace (deterministic, identity-like for typical settings in [0, 1]).

Parameters:

env (gym.Env)

step(action)[source]#

Uses the step() of the env that can be overwritten to change the returned data.

Parameters:

action (Any)

class openswmm_gymnasium.wrappers.RecordTrajectory(env, output_dir)[source]#

Bases: Wrapper

Persist one JSONL file per episode under output_dir.

File naming: episode_<00000>.jsonl, counter incrementing across resets within the wrapper’s lifetime. Each file contains:

  • One "reset" record at the top with the initial observation and reset info.

  • One "step" record per env step with action, observation, reward, terminated, truncated, and info.

@ivar output_dir: Directory where episode files are written. :type output_dir: pathlib.Path

Parameters:
  • env (gym.Env)

  • output_dir (str | os.PathLike)

close()[source]#

Closes the wrapper and env.

Return type:

None

reset(*, seed=None, options=None)[source]#

Uses the reset() of the env that can be overwritten to change the returned data.

Parameters:
  • seed (int | None)

  • options (dict[str, Any] | None)

step(action)[source]#

Uses the step() of the env that can be overwritten to change the returned data.

Parameters:

action (Any)

class openswmm_gymnasium.wrappers.RescaleBoxActions(env, src_low=0.0, src_high=1.0)[source]#

Bases: ActionWrapper

Rescale every gymnasium.spaces.Box leaf of the action space.

The wrapper presents an action space whose Box leaves all live in [src_low, src_high]; on step() the action is rescaled back into the env’s original [env_low, env_high] per-leaf bounds. Non-Box subspaces are left untouched.

Useful for plugging in agents (e.g. SAC, PPO with tanh-squashed output) whose policy naturally emits values in a fixed unit interval.

@ivar src_low: Lower bound for every Box leaf after wrapping. @ivar src_high: Upper bound for every Box leaf after wrapping.

Parameters:
  • env (gym.Env)

  • src_low (float)

  • src_high (float)

action(action)[source]#

Translate from [src_low, src_high] back into env bounds.

Parameters:

action (Any)

Return type:

Any

class openswmm_gymnasium.wrappers.TchebycheffScalarize(env, weights, utopia)[source]#

Bases: RewardWrapper

Weighted-Tchebycheff scalarisation.

For B{higher-is-better} vector reward (which is what the MO envs emit), this returns -max_d w_d * (utopia_d - r_d). The negation keeps the higher-is-better convention at the scalar level: a reward vector closer to the utopia point produces a smaller weighted gap and hence a larger (less negative) scalar.

@ivar weights: Per-objective weights. @ivar utopia: Per-objective best-case reward (high values for

higher-is-better envs).

Parameters:
  • env (gym.Env)

  • weights (Sequence[float])

  • utopia (Sequence[float])

reward(reward)[source]#

Returns a modified environment reward.

Args:

reward: The env step() reward

Returns:

The modified reward

Parameters:

reward (ndarray)

Return type:

float