Gym.NET is a high-performance, native C# (.NET 8/10) port of OpenAI Gym / Farama Gymnasium designed for reinforcement learning algorithmic development, benchmarking, and quantitative financial microstructure simulations. Operating as part of the SciSharp STACK ecosystem, it provides a standardized, strongly-typed environment suite with NumSharp as the underlying multidimensional array/tensor engine.
The goal of this project is to achieve complete architectural, behavioral, and mathematical parity with Farama Gymnasium (refs/Gymnasium), replacing legacy OpenAI Gym patterns with modern lifecycle contracts, comprehensive observation/action spaces, extensible wrapper hierarchies, and vectorized execution engines.
- Contract:
(NDArray Observation, Dict Information) Reset(int? seed = null, Dict options = null). - Behavior:
- Reseeds environment PRNG (
np.random.RandomState(seed)) whenseedis provided. - Accepts domain-specific options (
Dict options) for custom initial states. - Returns initial observation tensor accompanied by diagnostic information metadata.
- Reseeds environment PRNG (
-
Contract:
StepResult Step(object action)supporting 5-tuple deconstruction:var (observation, reward, terminated, truncated, info) = env.Step(action);
-
Behavior:
-
terminated(bool): Signals MDP terminal condition (e.g. agent succeeded or failed). -
truncated(bool): Signals out-of-bounds termination or time limit expiration (e.g.TimeLimitwrapper reached max steps). - Deprecates legacy
Doneflag to ensure proper Generalized Advantage Estimation ($\text{GAE}$ ) bootstrapping in downstream RL engines likePPO.Core.
-
-
Box: Continuous multi-dimensional bounded intervals with vectorized Gaussian, exponential, and uniform sampling. -
Discrete: Categorical integer space{start, ..., start + n - 1}with action masking support. -
MultiDiscrete: Vector of discrete categorical dimensions, each with independent bounds. -
MultiBinary: Multi-dimensional binary arrays with values in${0, 1}$ . -
TupleSpace: Cartesian product of heterogeneous subspaces. -
DictSpace: Key-value mapped composite subspaces. -
TextSpace: Bounded/variable-length character string space. -
SequenceSpace: Variable-length sequences of subspace elements. -
GraphSpace: Structured graph spaces containing node, edge, and link features. -
OneOfSpace: Exclusive union of alternative subspaces.
-
Base Contracts:
Wrapper,ObservationWrapper,ActionWrapper,RewardWrapper. -
Standard Suite:
-
TimeLimit: Enforces maximum episode step bounds and markstruncated = true. -
TransformObservation: Functional transformation applied to observation tensors. -
TransformReward: Functional scaling/clipping applied to rewards. -
ClipAction: Clips continuous actions to action space bounds. -
RescaleAction: Maps continuous actions affine-transformed to custom ranges (e.g.$[-1, 1]$ ). -
RecordEpisodeStatistics: Collects episode return, length, and execution time ininfo["episode"]. -
Autoreset: Automatically invokesReset()whenterminated or truncatedis encountered inStep(). -
OrderEnforcing: Enforces thatReset()is invoked prior toStep().
-
-
VectorEnv: Base vectorized environment abstraction. -
SyncVectorEnv: Sequential batch execution across$N$ environment instances. -
AsyncVectorEnv: Parallel batch execution across worker threads/tasks. -
Autoreset Semantics: Automatically captures
final_observationininfowhen sub-environments terminate while continuing seamless rollout batching.
CartPole-v1: Discrete classic control balancing cart and pole via Euler kinematics.Pendulum-v1: Continuous torque control pendulum swing-up.MountainCar-v0&MountainCarContinuous-v0: Underpowered car mountain ascent.Acrobot-v1: Two-link double pendulum.LunarLander-v3: Discrete and Continuous 2D lunar lander physics simulation (viaAether.Physics2D).BipedalWalker-v3: 4-joint bipedal locomotion robot.
- Environment constructor parameter:
string render_mode = null("human","rgb_array",null). - Rendering decoupled from core physics/math via
IEnvViewerandIEnvironmentViewerFactoryDelegate. - Implementations:
NullEnvViewer: High-speed headless simulation.AvaloniaEnvViewer: Cross-platform hardware-accelerated GUI (Gym.Rendering.Avalonia).WinFormEnvViewer: Windows Forms GUI (Gym.Rendering.WinForm).
- Single Responsibility (SRP): Segregate observation calculation, physics integration, reward calculation, and viewer rendering into focused classes.
- Open/Closed (OCP): Extend environment capabilities via
Wrappercomposition rather than modifying concrete environment classes. - Liskov Substitution (LSP): All concrete environments and wrappers must satisfy
IEnvand genericIEnv<TObs, TAct>contracts without surprising side effects. - Interface Segregation (ISP): Focused interfaces (
IEnv,ISpace,IWrapper,IVectorEnv,IEnvViewer). - Dependency Inversion (DIP): Environments depend upon abstract viewers via
IEnvironmentViewerFactoryDelegate, enabling headless testing.
- One level of indentation per method: Extract nested loops/conditionals into private descriptive methods.
-
Never use the
elsekeyword: Guard clauses, early returns, and polymorphic strategy dispatch. -
Wrap domain primitives: Wrap scalars and raw arrays in strongly typed Value Objects (
Observation,Action,Reward,EpisodeStats). - First-class collections: Classes containing collections must encapsulate collection behavior without extraneous properties.
- One dot per line: Demeter compliance across all submodules.
-
No abbreviations: Use explicit identifiers (
observation,terminated,truncated,actionSpace). -
Keep entities small: Target classes
$\le 100$ lines, methods$\le 15$ lines. - No bare getters/setters: Expose domain behavior instead of mutable data structures (Tell, Don't Ask).
- Comprehensive test coverage with MSTest / FluentAssertions.
- Golden comparison tests validating against deterministic Python Gymnasium trajectories and bounds.
- Tests serve as immutable behavioral contracts.