Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
23 changes: 22 additions & 1 deletion .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -28,8 +28,10 @@ jobs:
- uses: astral-sh/setup-uv@v5
with:
enable-cache: true
# With the checker's extra: tests/test_conformance_probe.py runs it, and
# skips without websockets.
- name: Sync
run: uv sync --frozen
run: uv sync --frozen --extra conformance
- name: Test
run: uv run pytest -q

Expand Down Expand Up @@ -70,6 +72,15 @@ jobs:
--port 8765 --steps 10 --horizon 4 --action-dim 7 --timeout 120 \
--client "/tmp/plugrl_client 127.0.0.1 8765 10"

# SPEC section 8.1 for the C++ client too: what it does with each chunk,
# and how it handles a resync, a stop and a text frame.
- name: ... and in probe mode, in every scenario
run: |
uv run --extra conformance plugrl-conformance \
--probe --scenario all --port 8769 --steps 12 --horizon 4 \
--action-dim 3 --action-dtype float64 --timeout 120 \
--client "/tmp/plugrl_client --probe 127.0.0.1 8769 3"

# The action dtype is the environment's and is never renegotiated, so a
# client has to read the typestr. A client that hard-codes float32 still
# passes every clause the server can see - it sends valid messages - and
Expand All @@ -94,3 +105,13 @@ jobs:
uv run --extra conformance plugrl-conformance \
--port 8767 --steps 10 --horizon 3 --action-dim 5 --timeout 120 \
--client "python examples/raw_client.py --host 127.0.0.1 --port 8767 --steps 10"

# Section 8.1: the probe env lets the checker see what the client did
# with each chunk, and the scenarios drive a resync, a stop and a text
# frame. The client is started once per scenario.
- name: ... and in probe mode, in every scenario
run: |
uv run --extra conformance plugrl-conformance \
--probe --scenario all --port 8768 --steps 12 --horizon 4 \
--action-dim 3 --action-dtype float64 --timeout 120 \
--client "python examples/raw_client.py --probe --batch 3 --host 127.0.0.1 --port 8768"
47 changes: 47 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -21,6 +21,8 @@ own known defects, and everything in it that can be checked by a test is.
| [`SPEC.md`](SPEC.md) | the protocol, in full |
| [`src/plugrl_protocol/`](src/plugrl_protocol/) | the message types and the msgpack codec — 101 lines |
| [`src/plugrl_protocol/conformance.py`](src/plugrl_protocol/conformance.py) | `plugrl-conformance`, which grades a client clause by clause |
| [`src/plugrl_protocol/server_conformance.py`](src/plugrl_protocol/server_conformance.py) | `plugrl-conformance-server`, which grades a server |
| [`src/plugrl_protocol/reuse.py`](src/plugrl_protocol/reuse.py) | the server's half of the `reuse-feedback-obs` feature |
| [`examples/`](examples/) | two env clients written against the spec, sharing no code with PlugRL |
| [`tests/`](tests/) | the spec's checkable clauses, as tests |

Expand All @@ -40,6 +42,37 @@ uv run --extra conformance plugrl-conformance \
It exits non-zero on a violation, so it can sit in a CI job. What it does not
require, because SPEC.md does not, is listed at the top of the module.

That form watches one connection, so it sees the messages and not what the
client did with them. To check that too, have the client run the probe
environment of [SPEC.md section 8.1](SPEC.md#81-the-probe-environment), a
few lines in any language, and add `--probe`:

```bash
uv run --extra conformance plugrl-conformance --probe --scenario all \
--client "./my_client --probe 127.0.0.1 8000"
```

The checker then works out from each feedback whether the client summed the
chunk's reward, sent the terminal observation, and applied the actions
time-major and in order. It also drives the connection: it speaks late, sends
a metadata frame larger than 1 MiB, closes for a resync, stops the run, and
answers with a text frame, starting the client once per scenario.

## Checking a server

The other direction: a client that drives a training server the way SPEC.md
lets a client behave, and checks what comes back. That covers the action
layout and `env_ids`, ragged batches, large frames, a resync close on each
kind of malformed message with the server staying up afterwards, and the
stop at the end of the run ([SPEC.md section 8.2](SPEC.md#82-checking-a-server)):

```bash
uv run --extra conformance plugrl-conformance-server --port 8000 --state-dim 3 --until-stop
```

[`examples/reference_server.py`](examples/reference_server.py) is a server
written against the specification that trains nothing, and passes.

## The protocol in one screen

Four message types, four lowercase strings. The server speaks first.
Expand Down Expand Up @@ -74,6 +107,20 @@ Three things a first implementation usually gets wrong, all specified:
- the reward in a `feedback` is the **sum over the action chunk**, not the
last step's ([section 5.4](SPEC.md#54-feedback--client-to-server)).

## Optional features

A server lists the additions to version 1 it implements in its metadata's
`features`, and a client uses one only when it is listed, so an old client
works with a new server and the other way round
([section 10](SPEC.md#10-versioning)).

There is one so far, `reuse-feedback-obs`
([section 10.1](SPEC.md#101-reuse-feedback-obs)). Without it every
observation crosses the link twice: once in the feedback that ends a chunk,
and again in the next infer. With it the infer can say "the one you already
have". `plugrl-conformance --features reuse-feedback-obs` offers it to a
client, and `plugrl-conformance-server` exercises it when the server lists it.

## Relationship to openpi

The serialization is [openpi](https://github.com/Physical-Intelligence/openpi)'s.
Expand Down
228 changes: 187 additions & 41 deletions SPEC.md
Original file line number Diff line number Diff line change
Expand Up @@ -287,6 +287,7 @@ there today:
| `policy` | the policy's class name |
| `action_horizon` | `H`, the number of steps in an action chunk |
| `action_dim` | the width of one action |
| `features` | the optional features of section 10.1 the server implements, as a list of strings |

`action_horizon` and `action_dim` are read off the policy, which is free
not to declare them. **A key that is absent means the server does not know,
Expand Down Expand Up @@ -327,6 +328,11 @@ both keys **MUST** share the same leading dimension `n`.
This is the normal case: the batch is whichever environments happen to need
a chunk, and it is ragged by construction.

An `infer` may also carry a `reuse` key, when the server offers the
`reuse-feedback-obs` feature of section 10.1. A row it marks carries no
observation: the server takes it from that environment's last `feedback`.
Without the feature the key is not sent.

> **Gap — `step_ids` is accepted and ignored.** The client sends, per
> environment, the number of completed action chunks in the current episode,
> reset to `0` when the episode ends. Neither server reads it. It is
Expand Down Expand Up @@ -655,40 +661,134 @@ A client conforms to version 1 if it:
differently;
- [ ] drops any held `feedback` when a connection closes, and never sends
`feedback` for an `action` that arrived on an earlier connection;
- [ ] treats a text frame as a fatal error.

`examples/conformance_server.py` checks the clauses above that one passive
connection can observe - framing, alternation, env indices, observation
shape, and the feedback payload's keys, dtypes and lengths - and reports
what it accepts but cannot require as a note rather than a failure. Both
reference clients pass it with one note: they send `text` as a msgpack
string array rather than a `<U` array, which is the section 3.4 Gap.

It does not check the rest, and an unexercised clause leaves no trace in its
report: every clause string in the harness names section 2, 3.4, 4.2, 4.4,
5.2, 5.4 or 7.5 - or is the catch-all `connection` - and the report iterates
only the clauses it touched. A client that breaks all of the following still
prints "no violations":

* the connection options of section 1.1 - the harness sets `compression` and
`max_size` on its own side and never inspects what the client offered;
* reading `metadata` before sending anything;
* chunk-summed reward, and the terminal observation on a done step;
* handling the two close reasons differently, dropping held `feedback`
across a reconnect, and treating a text frame as fatal - these three need
the harness to drive a close or send a text frame, which is more than
watching one well-behaved connection.

What no server-side harness can see at all is what a client does with the
`action` it receives. Honouring the time-major layout and consuming the
horizon in order are checked on the Python side, by `plugrl-env-client`'s
`tests/test_protocol_alternation.py::TestChunkSemantics`. Reading `env_ids`
is not checked anywhere: `plugrl-server` sends the key, and the Python
client discards it - `infer()` returns the whole `data` map and its only
caller takes `["action"]` out of it, and the string `env_ids` does not
occur anywhere in `plugrl-env-client`'s source or tests. The clause above
is a rule for clients in other languages, with no reference implementation
behind it.
- [ ] treats a text frame as a fatal error;
- [ ] uses a feature of section 10 only when the server lists it, and, for
`reuse-feedback-obs`, reuses only an observation the server holds.

`plugrl-conformance` (`examples/conformance_server.py` is a shim for it) has
two modes.

**By default it watches one connection.** It checks what that connection
shows: framing, alternation, env indices, observation shape, and the
feedback payload's keys, dtypes and lengths. It reports what it accepts but
cannot require as a note rather than a failure. Both reference clients pass
it with one note: they send `text` as a msgpack string array rather than a
`<U` array, which is the section 3.4 Gap. An unexercised clause leaves no
trace in its report, and from one connection to an unknown environment it
cannot tell what the client did with an action.

**With `--probe` the client runs the probe environment of section 8.1.**
From each feedback the checker then works out how many steps of the chunk
the client ran and which action it applied last, and it drives the
connection instead of watching it. In that mode it checks every clause of
the list above except these:

* reading `env_ids`: section 4.3 requires them to equal the `infer`'s
`env_indices`, so a client that ignores them behaves identically;
* tolerating a `feedback` env set that differs from the `infer` set: the
client chooses its sets, and the server cannot make them differ;
* requiring no particular `metadata` key: the checker sends all of them;
* dropping held `feedback` after a close other than a resync.

`--scenario` selects how the connection is driven:

* **`basic`**: the checker waits 0.3 s before speaking, so a client that
sends first is caught. Its `metadata` carries an unknown 2 MiB key, so a
client that kept its library's 1 MiB frame cap cannot receive it, and one
that rejects unknown keys fails. It checks that the handshake did not offer
`permessage-deflate`. It answers every `infer` with float64, time-major
actions whose values encode their position. Once the exchanges are done it
closes with `plugrl-server-stop`, after which the client must exit 0 and
not reconnect.
* **`resync`**: as `basic`, but halfway through it closes with
`plugrl-server-resync` right after sending an `action`, so the client is
holding a `feedback`. The client must reconnect, and its first message on
the new connection must be an `infer`.
* **`text`**: it answers the first `infer` with a text frame, which the
client must treat as fatal: it closes the connection and sends nothing
more.
* **`all`**: each in turn, starting the client once per scenario.

Both reference clients pass all three with the same note:
`examples/raw_client.py --probe` and `plugrl_client --probe`, the C++ one,
which CI checks on every change. The Python client's `--bug` option breaks
one clause at a time, and `tests/test_conformance_probe.py` checks that the
checker names each one.

### 8.1 The probe environment

A client that wants its handling of actions checked, and not only its
messages, runs this environment for `plugrl-conformance --probe`. It needs
no simulator, and is a few lines in any language.

* The client runs `n` probe envs, with indices `0` to `n - 1` as its
`env_indices`. Indices stay below 1000.
* Env `i`'s episode lasts `L_i = 3 + 2 * (i mod 3)` steps: 3, 5, 7, 3, 5,
... So within one chunk, different envs end their episodes at different
steps.
* Its observation has no images and two states:
* `t`: float64 `[n, 1]`, the number of steps taken in the current
episode, `0` after a reset;
* `a`: float64 `[n, d]`, the action the env applied on its last step,
where `d` is the action width. Its value after a reset is not checked,
and before the client knows the width it may have `d = 1`.

`text` may be anything.
* A step applies the action, adds 1 to `t`, stores the action in `a`, and
pays a reward of 1. It reports `terminated` when `t` reaches `L_i`. The
probe env never truncates.
* A reset sets `t` back to `0`.

The checker sends action values that encode their position: element
`[k, row, d]` for env `j` is `10000 k + 10 j + d`. A feedback therefore says
exactly what the client did with the chunk. `t` minus the `t` of the
`infer` it answered is the number of steps it ran, which must be between 1
and `H`. With a reward of 1 per step, the reward must equal that number,
which is the chunk-sum rule. `a` must be the chunk's action at the last step
run, for that env, which checks time-major order, the step order, and the
dtype. A `terminated` env must report `t = L_i`, its own terminal
observation, and its next `infer` must report `t = 0`.

### 8.2 Checking a server

The checklist above is for clients. A training server conforms if it:

- [ ] sends one `metadata` message, before anything else, on every
connection;
- [ ] answers each `infer` with an `action` that is time-major, `[H, n, *da]`,
with `env_ids` equal to the `infer`'s `env_indices`, in order, and that
agrees with the `action_horizon` and `action_dim` it declared;
- [ ] accepts what a client may send: a `feedback` env set unlike the
`infer`'s, an `n` that varies, and frames larger than 1 MiB;
- [ ] closes a connection with 1001 and `plugrl-server-resync` on a malformed
message, two `infer`s in a row, or an `info` it cannot split per
environment, and stays up for its other clients;
- [ ] keeps env indices connection-scoped;
- [ ] ends a run with 1001 and `plugrl-server-stop`;
- [ ] if it lists `reuse-feedback-obs`, answers an `infer` that reuses
observations, and closes with 1001 and `plugrl-server-resync` on one
that reuses an observation it does not hold.

`plugrl-conformance-server` checks each of these from the client's side of
the wire:

```bash
plugrl-conformance-server --port 8000 --state-dim 3 --until-stop
```

It sends a states-only observation, `states["obs"]` of `--state-dim` values
(`--state-key` and `--image-key` change that), so it can drive any server
whose policy reads one. `examples/reference_server.py` is a server written
against this document that trains nothing and passes every check. Its
`--bug` option breaks one clause at a time, and
`tests/test_server_conformance.py` checks that the grader names each one.
`plugrl-server` passes too. Its version from before plugrl-server#108 fails
two checks: an `info` it could not split closed that connection with 1011,
and the server then refused new connections.

What the server does with a frame is not visible from the wire, so section
7.6's SHOULD - saying something about feedback for an environment it holds
no step state for - is not checked.

---

Expand Down Expand Up @@ -723,9 +823,17 @@ or Rust — carries across unchanged.
## 10. Versioning

This is version 1. The server publishes it as `protocol_version` in the
metadata message (section 5.1), but there is no negotiation: a client
**MUST NOT** require the key, and the server does not adapt to a client's
version.
metadata message (section 5.1). A client **MUST NOT** require the key, and
the server does not adapt to a client's version.

What is negotiated is **features**. A feature is an addition to version 1
that a client may use and a server may offer, and nothing else changes for
either side when it is absent. The server lists the ones it implements in
`metadata`'s `features`. A client uses a feature only if it is listed there,
and a server that lists one still accepts every client that does not use it.
So a new client works with an old server, and an old client with a new one.
A change that cannot be made this way is breaking, and moves
`protocol_version` to 2.

> **Correction, 2026-09-11.** This paragraph used to say version 1 "has no
> version field on the wire" and attributed to section 5.1 a remark that the
Expand All @@ -747,7 +855,45 @@ test rather than a deployment: `tests/test_wire_format.py` holds the four
are not, and cannot be - this is the codec repository, and the axis order
and the chunk-sum rule are properties of how the two sides behave, not of
what the packer emits. Transposing the action chunk or redefining `rewards`
as the last step's reward passes every test here. They are covered instead
by `plugrl-env-client`'s
`tests/test_protocol_alternation.py::TestChunkSemantics`, which is the same
file section 8 points at.
as the last step's reward passes every codec test here. They are checked
instead by `plugrl-conformance --probe` (section 8.1), against any client.

### 10.1 `reuse-feedback-obs`

Without it, every observation crosses the link twice. The client sends it in
the `feedback` that ends a chunk, as the step's next observation, and then
sends the same observation again in the next `infer`, as the state the new
chunk starts from. Only after a reset do the two differ. That second copy is
the factor of two in plugrl-server's E43 cost model: an exchange costs a
fixed latency plus twice the observation's bytes over the link.

With the feature, the `infer` may say "the observation you already have":

```
{ "message_type": "infer",
"data": <observation, batched to the rows that are not reused>,
"env_indices": <ndarray "<i8" shape [n]>,
"step_ids": <ndarray "<i8" shape [n]>,
"reuse": <ndarray "|b1" shape [n]> }
```

* `reuse[i]` true means env `env_indices[i]` sends no observation. Its
observation is the `obs` row of the last `feedback` that named it on this
connection.
* `data` is batched to the rows where `reuse` is false, in their
`env_indices` order. With every row reused, its arrays have a leading
dimension of 0.
* A client **MUST** send `reuse` false for an env whose last feedback on
this connection reported `terminated` or `truncated`, since that feedback
carried the terminal observation and not the reset one. It **MUST** also
send it false for an env with no feedback on this connection yet, which
includes every env just after a reconnect.
* A server that lists the feature **MUST** close a connection whose `infer`
reuses an observation it does not hold - an env never fed back on this
connection, or one whose last feedback ended its episode - with 1001 and
`plugrl-server-resync`.
* An `infer` without `reuse` sends every observation, as in version 1.

Both checkers know the feature. `plugrl-conformance --features
reuse-feedback-obs` offers it to a client and checks how the client uses it.
`plugrl-conformance-server` exercises it when the server lists it.
Loading
Loading