diff --git a/README.md b/README.md index d6c5f6a..f8bd9e5 100644 --- a/README.md +++ b/README.md @@ -313,8 +313,9 @@ All operations are **eventually consistent**: sent alone. Receivers stage provisional chunks in shard-owned private ETS and replace visible state only after the complete slice and its terminal manifest are present. -- **`replicated_anti_entropy_interval`** — interval in milliseconds for stream - head advertisements and nonblocking control heartbeats. Defaults to 1,000. +- **`replicated_anti_entropy_interval`** — interval in milliseconds for retrying + unacknowledged stream heads and sending nonblocking control heartbeats. + Acknowledged streams do not advertise unchanged heads. Defaults to 1,000. - **`replicated_peer_lease_timeout`** — time without a dist-Erlang control heartbeat before state owned by that Group peer is purged. Defaults to 15,000 and must exceed the anti-entropy interval. Probes continue after expiry so a @@ -409,9 +410,10 @@ When Group starts (or a new Erlang node connects), shards exchange nodes. This handshake: 1. Validates that shard counts match (raises on mismatch). -2. Exchanges cluster membership lists. -3. Shard 0 exchanges protocol version, origin generation, and one complete - active named-cluster epoch snapshot per node. Matching data shards exchange +2. Uses constant-size discovery probes, regardless of the cluster count. +3. Shard 0 exchanges protocol version, origin generation, and a complete + active named-cluster epoch snapshot on discovery or authority repair. + Matching data shards exchange only constant-size lane/transport descriptors tied to that authority revision. @@ -419,6 +421,14 @@ Constant-size heartbeats renew the peer lease. If an origin generation or cluster-epoch revision changes, the receiver requests a fresh authoritative hello; if heartbeats stop, lease expiry purges that origin's complete local view and discovery probes allow it to rejoin later. +Each lease expiry advances a discovery probe epoch, so the sender can replay +previously acknowledged heads when the receiver has lost its view. Retries of +that same probe epoch leave acknowledged streams quiet. An authority repair +request carries one token per recovery attempt: its first reply is immediate, +while repeated replies for the same token and unchanged authority are limited +to one every five seconds. A new authority revision is sent immediately. A +heartbeat that first restores a lane route also advertises the heads of writes +made while that route was absent. Incremental cluster open/close controls are generation fenced, processed in revision order, and installed by shard 0 into one node-wide authority table. @@ -481,10 +491,21 @@ ETS view and batched into one delta message per target. Process-death registry and PG removals can share one record and retain their one-event-batch behavior. Receivers advance a cursor only across a contiguous sequence prefix. A gap -requests the missing suffix. Repeated head advertisements recover a dropped -tail even when no later write occurs. If the requested sequence is older than -the bounded oplog floor, the origin sends an exact snapshot of only its own -registry claims and PG memberships; absence from that snapshot is a delete. +requests the missing suffix. The sender retries only unacknowledged stream +heads, so a dropped tail repairs even when no later write occurs; an applied +cursor acknowledgement stops those retries. Discovery after a receiver lease +expiry re-advertises the sender's current heads, including previously +acknowledged streams whose receiver state may have been lost. If the requested +sequence is older than the bounded oplog floor, the origin sends an exact +snapshot of only its own registry claims and PG memberships; absence from that +snapshot is a delete. +Gap requests identify the advertised head. The sender ignores requests for an +already acknowledged stream or an older head, then advertises its current head +on the next retry. After a complete snapshot send, repeated gap requests for +the same head wait up to five seconds before retrying the full snapshot. The +wait is at most half the peer lease so provisional chunks can survive between +attempts; a fresh peer recovery probe clears the hold. An interrupted chunk or +commit resumes from its recorded offset. There are no leaders, quorum acknowledgements, per-entry replicated tombstones, or known-membership retention barriers. Oplog memory is bounded locally and diff --git a/lib/group/replica.ex b/lib/group/replica.ex index 7a3d794..e3f6510 100644 --- a/lib/group/replica.ex +++ b/lib/group/replica.ex @@ -8,6 +8,13 @@ defmodule Group.Replica do @replicated_registry_receiver_flush_timer :flush_replicated_registry_receiver_buffer @replica_broadcast_flush_timer :flush_replica_broadcast_buffer @anti_entropy_timer :group_replica_anti_entropy + @replica_ack_flush_timer :group_replica_ack_flush + @replica_ack_flush_interval 25 + @replica_ack_busy_retry_interval 1_000 + @discovery_hello_retry_interval 5_000 + @snapshot_retry_interval 5_000 + @replica_head_batch_target_bytes 64_000 + @replica_head_batches_per_peer_turn 16 @local_request_tag :group_local_request @local_reply_tag :group_local_reply @protocol_version Group.Replica.WireProtocol.version() @@ -57,9 +64,11 @@ defmodule Group.Replica do Dist Erlang remains the control plane: - - peer_connect / peer_connect_ack discover matching shards and clusters. + - peer_connect / peer_connect_ack discover matching shards with constant-size + messages, independent of the number of named clusters. - shard 0 exchanges replica_hello authority containing the origin generation - and complete active named-cluster epoch set exactly once per node. + and complete active named-cluster epoch set during discovery or authority + repair, rather than once per shard. - matching nonzero shards exchange constant-size replica_lane_hello messages containing their transport descriptor and the authority revision they use. - replica_cluster_open / replica_cluster_close fence named-cluster lifetimes. @@ -91,7 +100,10 @@ defmodule Group.Replica do Replica state uses the configured Group.Transport: - - heads advertises {stream, retained_floor, head}. + - heads advertises {stream, retained_floor, head} for an unacknowledged stream + with a sender-issued token. applied echoes that token with the receiver's + committed cursor; a rediscovery rotates it so delayed ACKs cannot suppress + repair after the receiver loses its state. - delta_batch carries one or more contiguous stream runs. - need requests the receiver's next missing sequence. - snapshot_chunk carries a byte-bounded part of one exact origin slice when @@ -102,7 +114,7 @@ defmodule Group.Replica do Every stream field is validated against the source node and current generation/epoch. An old generation, a closed epoch, a wrong shard, or a transitive claim for another node's pid is rejected. Control/data - reordering is safe: early messages are ignored and repeated heads repair them; + reordering is safe: early messages are ignored and unacknowledged heads repair them; late messages fail their generation or epoch fence. Snapshot chunks may be lost, duplicated, reordered, or mixed across retransmissions at the same stream head; exact row counts, set insertion, and conflicting-retransmission @@ -116,12 +128,14 @@ defmodule Group.Replica do ## Bounded recovery The oplog is bounded per shard, not by peer acknowledgements. A dropped tail - is found by periodic heads. A gap inside the retained range is repaired with + is found by retried, unacknowledged heads. A gap inside the retained range is repaired with bounded delta batches. A gap below the retained floor receives the existing full-sync primitive, narrowed to an exact origin/shard/cluster snapshot. Absence from that snapshot is deletion, so no tombstones are required. There is no leader, quorum, retention ACK, or requirement to know all members. + Cursor acknowledgements only suppress unchanged head advertisements; they + never prevent oplog pruning. A slow or disconnected peer cannot pin memory. When it returns it repairs from deltas when possible and a snapshot otherwise. @@ -226,6 +240,7 @@ defmodule Group.Replica do :replica_transport, :replica_transport_opts, :anti_entropy_ref, + :replica_ack_flush_ref, :pending_replicated_pg_started_at, :pending_replicated_pg_flush_ref, :pending_replicated_registry_started_at, @@ -240,6 +255,18 @@ defmodule Group.Replica do pending_replica_broadcast_ops: [], remote_shards: %{}, peer_last_seen: %{}, + peer_probe_epochs: %{}, + pending_peer_probes: %{}, + remote_probe_epochs: %{}, + peer_connect_ack_seen: %{}, + authority_request_tokens: %{}, + remote_authority_request_tokens: %{}, + discovery_hello_last_sent: %{}, + pending_replica_heads: %{}, + pending_replica_acks: %{}, + replica_ack_backoff: false, + replica_send_tokens: %{}, + replica_receive_tokens: %{}, cluster_control_dirty: %{}, authority_dirty_notified: MapSet.new(), pending_registry_reprojections: %{}, @@ -407,7 +434,7 @@ defmodule Group.Replica do send_remote_shard_message( state, remote_node, - {:peer_connect, self(), shard_index, num_shards, Data.my_clusters(name)} + peer_connect_message(state, remote_node) ) end @@ -538,13 +565,19 @@ defmodule Group.Replica do replica_view_current?(state, remote_node) do state = notify_replica_transport_peer_up(state, remote_node, transport_descriptor) + lane_needs_heads? = + Map.get(state.remote_shards, remote_node) != remote_pid or + not Map.has_key?(state.replica_send_tokens, remote_node) + + state = %{ + state + | remote_shards: Map.put(state.remote_shards, remote_node, remote_pid), + peer_last_seen: Map.put(state.peer_last_seen, remote_node, monotonic_millis()), + cluster_control_dirty: Map.delete(state.cluster_control_dirty, remote_node) + } + {:noreply, - %{ - state - | remote_shards: Map.put(state.remote_shards, remote_node, remote_pid), - peer_last_seen: Map.put(state.peer_last_seen, remote_node, monotonic_millis()), - cluster_control_dirty: Map.delete(state.cluster_control_dirty, remote_node) - }} + if(lane_needs_heads?, do: send_replica_heads(state, remote_node), else: state)} else {:noreply, state @@ -591,14 +624,25 @@ defmodule Group.Replica do cond do replica_authority_current?(state, remote_node, generation, epoch_revision) and replica_view_current?(state, remote_node) -> + lane_needs_heads? = + Map.get(state.remote_shards, remote_node) != remote_pid or + not Map.has_key?(state.replica_send_tokens, remote_node) + state = state |> notify_replica_transport_peer_up(remote_node, transport_descriptor) |> put_remote_shard(remote_node, remote_pid) - |> purge_remote_streams_outside_authority(remote_node) |> touch_replica_peer(remote_node) |> Map.update!(:cluster_control_dirty, &Map.delete(&1, remote_node)) - |> send_replica_heads(remote_node) + + state = + if lane_needs_heads? do + state + |> purge_remote_streams_outside_authority(remote_node) + |> send_replica_heads(remote_node) + else + state + end {:noreply, state} @@ -665,11 +709,12 @@ defmodule Group.Replica do # retirement must not recreate an unleased peer. Once exact authority # reaches this lane, repeat shard-local discovery immediately instead # of waiting for the next anti-entropy probe. + state = advance_peer_probe_epoch(state, remote_node) + send_remote_shard_message( state, remote_node, - {:peer_connect, self(), state.shard_index, state.num_shards, - Data.my_clusters(state.name)} + peer_connect_message(state, remote_node) ) state @@ -772,7 +817,14 @@ defmodule Group.Replica do end def handle_info({:replica_authority_dirty_local, remote_node}, %{shard_index: 0} = state) do - {:noreply, mark_cluster_control_dirty(state, remote_node)} + if replica_view_current?(state, remote_node) do + {:noreply, state} + else + {:noreply, + state + |> mark_cluster_control_dirty(remote_node) + |> request_replica_authority(remote_node)} + end end def handle_info({:replica_authority_dirty_local, remote_node}, state) do @@ -927,10 +979,17 @@ defmodule Group.Replica do compatible? and replica_authority_current?(state, remote_node, generation, epoch_revision) and replica_view_current?(state, remote_node) -> - state - |> notify_replica_transport_peer_up(remote_node, transport_descriptor) - |> put_remote_shard(remote_node, remote_pid) - |> touch_replica_peer(remote_node) + route_changed? = + Map.get(state.remote_shards, remote_node) != remote_pid or + not Map.has_key?(state.peer_last_seen, remote_node) + + state = + state + |> notify_replica_transport_peer_up(remote_node, transport_descriptor) + |> put_remote_shard(remote_node, remote_pid) + |> touch_replica_peer(remote_node) + + if route_changed?, do: send_replica_heads(state, remote_node), else: state compatible? and replica_authority_current?(state, remote_node, generation, epoch_revision) -> @@ -943,11 +1002,28 @@ defmodule Group.Replica do {:noreply, state} end - def handle_info({:replica_hello_request, remote_pid}, state) do + def handle_info({:replica_hello_request, remote_pid, request_token}, state) + when is_pid(remote_pid) and is_reference(request_token) do if state.shard_index == 0 do - {:noreply, send_replica_hello(state, node(remote_pid))} + remote_node = node(remote_pid) + + new_request? = + Map.get(state.remote_authority_request_tokens, remote_node) != + {remote_pid, request_token} + + state = %{ + state + | remote_authority_request_tokens: + Map.put( + state.remote_authority_request_tokens, + remote_node, + {remote_pid, request_token} + ) + } + + {:noreply, maybe_send_discovery_hello(state, remote_node, new_request?)} else - _ = send_local_control_message(state, {:replica_hello_request, remote_pid}) + _ = send_local_control_message(state, {:replica_hello_request, remote_pid, request_token}) {:noreply, state} end end @@ -983,24 +1059,31 @@ defmodule Group.Replica do offsets = case result do :complete -> - Map.delete(state.snapshot_send_offsets, snapshot_key) + if snapshot_head_pending?(state, snapshot_key) do + Map.put(state.snapshot_send_offsets, snapshot_key, {:sent, monotonic_millis()}) + else + Map.delete(state.snapshot_send_offsets, snapshot_key) + end {:resume, chunk_index} -> - if current_snapshot_send?(state, snapshot_key) do + if current_snapshot_send?(state, snapshot_key) and + snapshot_head_pending?(state, snapshot_key) do Map.put(state.snapshot_send_offsets, snapshot_key, {:chunk, chunk_index}) else Map.delete(state.snapshot_send_offsets, snapshot_key) end {:resume_commit, manifest} -> - if current_snapshot_send?(state, snapshot_key) do + if current_snapshot_send?(state, snapshot_key) and + snapshot_head_pending?(state, snapshot_key) do Map.put(state.snapshot_send_offsets, snapshot_key, {:commit, manifest}) else Map.delete(state.snapshot_send_offsets, snapshot_key) end :retry -> - if current_snapshot_send?(state, snapshot_key) do + if current_snapshot_send?(state, snapshot_key) and + snapshot_head_pending?(state, snapshot_key) do state.snapshot_send_offsets else Map.delete(state.snapshot_send_offsets, snapshot_key) @@ -1026,7 +1109,7 @@ defmodule Group.Replica do |> probe_replica_peers() |> request_quiet_cluster_hellos() |> broadcast_replica_heartbeats() - |> broadcast_replica_heads() + |> retry_pending_replica_heads() |> schedule_anti_entropy() else state @@ -1035,6 +1118,14 @@ defmodule Group.Replica do {:noreply, state} end + def handle_info({@replica_ack_flush_timer, ref}, state) do + if state.replica_ack_flush_ref == ref do + {:noreply, flush_replica_acks(%{state | replica_ack_flush_ref: nil})} + else + {:noreply, state} + end + end + def handle_info({@local_request_tag, caller_pid, ref, request}, state) when is_pid(caller_pid) and is_reference(ref) do {:noreply, process_local_request_turn(state, [{{:send, caller_pid, ref}, request}])} @@ -1053,82 +1144,150 @@ defmodule Group.Replica do # ===================================================================== def handle_info( - {:peer_connect, remote_pid, remote_shard_index, remote_num_shards, remote_clusters}, + {:peer_connect, remote_pid, remote_shard_index, remote_num_shards, probe_epoch, + remote_generation, remote_revision}, state ) - when remote_shard_index == state.shard_index do + when remote_shard_index == state.shard_index and is_pid(remote_pid) and + is_integer(probe_epoch) and probe_epoch >= 0 and + is_integer(remote_revision) and remote_revision >= 0 do state = flush_pending_replicated_message_barrier(state) if remote_num_shards != state.num_shards do raise "Group shard count mismatch: local=#{state.num_shards} remote=#{remote_num_shards} from #{node(remote_pid)}" end - %{name: name, shard_index: shard} = state remote_node = node(remote_pid) - # Compute shared clusters for diagnostics only. The generation-fenced hello - # is the sole authority that mutates peer and cluster membership. Keeping - # discovery hints side-effect free prevents a delayed pre-restart - # peer_connect from permanently re-adding stale cluster rows. - my_clusters = Data.my_clusters(name) - shared = compute_shared_clusters(my_clusters, remote_clusters) + probe_status = + case Map.get(state.remote_probe_epochs, remote_node) do + {^remote_pid, old_epoch} when probe_epoch < old_epoch -> :stale + {^remote_pid, ^probe_epoch} -> :duplicate + _ -> :new + end + + # The generation-fenced hello is the sole authority that mutates peer and + # cluster membership. Discovery remains constant-size even with many + # subclusters, and a stale probe cannot re-add retired cluster routes. # Replica peers are addressed by registered `{name, node}` and established # only by replica_hello. Do not remotely monitor the shard PID: creating a # remote monitor itself emits a distribution signal and may suspend on a # busy dist connection. - # Send ack with our cluster list - send_to_peer( - state, - remote_node, - {:peer_connect_ack, self(), shard, state.num_shards, my_clusters} - ) + if probe_status == :stale do + {:noreply, state} + else + # A repeated probe is a retry of one receiver recovery, not a new loss + # of its state. Only a newer epoch can invalidate an already-applied ACK. + state = + if probe_status == :new do + %{ + state + | remote_probe_epochs: + Map.put(state.remote_probe_epochs, remote_node, {remote_pid, probe_epoch}) + } + else + state + end - send_replica_hello(state, remote_node) + route_ready? = + Map.get(state.remote_shards, remote_node) == remote_pid and + replica_authority_current?(state, remote_node, remote_generation, remote_revision) and + replica_view_current?(state, remote_node) - log_once(state, fn -> - "#{log_prefix(state)} peer_connect from #{remote_node} (#{length(shared)} shared clusters)" - end) + state = + if probe_status == :new or not route_ready? do + maybe_send_discovery_hello(state, remote_node, probe_status == :new) + else + state + end - {:noreply, state} + state = + if probe_status == :new and route_ready? do + state + |> reset_replica_send_token(remote_node) + |> send_replica_heads(remote_node) + else + state + end + + # A route for an older receiver PID cannot certify that this receiver's + # already-ACKed streams were requeued. Keep its probe obligation alive + # until the reverse hello installs the new PID and a later ACK is ready. + send_to_peer( + state, + remote_node, + {:peer_connect_ack, self(), state.shard_index, state.num_shards, probe_epoch, + route_ready?} + ) + + log_once(state, fn -> + "#{log_prefix(state)} peer_connect from #{remote_node}" + end) + + {:noreply, state} + end end - def handle_info({:peer_connect, _remote_pid, _other_shard, _num_shards, _clusters}, state) do + def handle_info( + {:peer_connect, _remote_pid, _other_shard, _num_shards, _probe_epoch, _remote_generation, + _remote_revision}, + state + ) do state = flush_pending_replicated_message_barrier(state) # Wrong shard index, ignore {:noreply, state} end def handle_info( - {:peer_connect_ack, remote_pid, remote_shard_index, remote_num_shards, remote_clusters}, + {:peer_connect_ack, remote_pid, remote_shard_index, remote_num_shards, probe_epoch, + route_ready?}, state ) - when remote_shard_index == state.shard_index do + when remote_shard_index == state.shard_index and is_integer(probe_epoch) and + probe_epoch >= 0 and is_boolean(route_ready?) do state = flush_pending_replicated_message_barrier(state) if remote_num_shards != state.num_shards do raise "Group shard count mismatch: local=#{state.num_shards} remote=#{remote_num_shards} from #{node(remote_pid)}" end - %{name: name} = state remote_node = node(remote_pid) - # Discovery acknowledgements are hints only; replica_hello is the sole - # generation-fenced authority for peer and cluster membership. - my_clusters = Data.my_clusters(name) - shared = compute_shared_clusters(my_clusters, remote_clusters) + if probe_epoch != Map.get(state.peer_probe_epochs, remote_node, 0) do + {:noreply, state} + else + # The ACK proves the sender processed this recovery probe. Until then, + # keep retrying even if an older in-flight hello restores our route. + pending_peer_probes = + if route_ready? and Map.get(state.pending_peer_probes, remote_node) == probe_epoch and + Map.get(state.remote_shards, remote_node) == remote_pid do + Map.delete(state.pending_peer_probes, remote_node) + else + state.pending_peer_probes + end - log_once(state, fn -> - "#{log_prefix(state)} peer_connect_ack from #{remote_node} (#{length(shared)} shared clusters)" - end) + # Discovery acknowledgements are hints only; replica_hello is the sole + # generation-fenced authority for peer and cluster membership. + log_once(state, fn -> "#{log_prefix(state)} peer_connect_ack from #{remote_node}" end) - send_replica_hello(state, remote_node) + first_ack? = Map.get(state.peer_connect_ack_seen, remote_node) != remote_pid - {:noreply, state} + state = %{ + state + | peer_connect_ack_seen: Map.put(state.peer_connect_ack_seen, remote_node, remote_pid), + pending_peer_probes: pending_peer_probes + } + + {:noreply, maybe_send_discovery_hello(state, remote_node, first_ack?)} + end end - def handle_info({:peer_connect_ack, _remote_pid, _other_shard, _num_shards, _clusters}, state) do + def handle_info( + {:peer_connect_ack, _remote_pid, _other_shard, _num_shards, _probe_epoch, _route_ready?}, + state + ) do state = flush_pending_replicated_message_barrier(state) {:noreply, state} end @@ -1139,12 +1298,11 @@ defmodule Group.Replica do def handle_info({:nodeup, remote_node}, state) do state = flush_pending_replicated_message_barrier(state) - %{shard_index: shard, name: name} = state send_remote_shard_message( state, remote_node, - {:peer_connect, self(), shard, state.num_shards, Data.my_clusters(name)} + peer_connect_message(state, remote_node) ) {:noreply, state} @@ -1190,6 +1348,18 @@ defmodule Group.Replica do state | remote_shards: Map.delete(state.remote_shards, dead_node), peer_last_seen: Map.delete(state.peer_last_seen, dead_node), + peer_probe_epochs: Map.delete(state.peer_probe_epochs, dead_node), + pending_peer_probes: Map.delete(state.pending_peer_probes, dead_node), + remote_probe_epochs: Map.delete(state.remote_probe_epochs, dead_node), + peer_connect_ack_seen: Map.delete(state.peer_connect_ack_seen, dead_node), + authority_request_tokens: Map.delete(state.authority_request_tokens, dead_node), + remote_authority_request_tokens: + Map.delete(state.remote_authority_request_tokens, dead_node), + discovery_hello_last_sent: Map.delete(state.discovery_hello_last_sent, dead_node), + pending_replica_heads: Map.delete(state.pending_replica_heads, dead_node), + pending_replica_acks: Map.delete(state.pending_replica_acks, dead_node), + replica_send_tokens: Map.delete(state.replica_send_tokens, dead_node), + replica_receive_tokens: Map.delete(state.replica_receive_tokens, dead_node), cluster_control_dirty: Map.delete(state.cluster_control_dirty, dead_node), authority_dirty_notified: MapSet.delete(state.authority_dirty_notified, dead_node) } @@ -2408,12 +2578,14 @@ defmodule Group.Replica do defp take_priority_control_turn(state, remaining) do receive do - {:peer_connect, _remote_pid, _remote_shard_index, _remote_num_shards, _remote_clusters} = + {:peer_connect, _remote_pid, _remote_shard_index, _remote_num_shards, _probe_epoch, + _remote_generation, _remote_revision} = msg -> state = process_inline_priority_message(state, msg) take_priority_control_turn(state, remaining - 1) - {:peer_connect_ack, _remote_pid, _remote_shard_index, _remote_num_shards, _remote_clusters} = + {:peer_connect_ack, _remote_pid, _remote_shard_index, _remote_num_shards, _probe_epoch, + _route_ready?} = msg -> state = process_inline_priority_message(state, msg) take_priority_control_turn(state, remaining - 1) @@ -2452,7 +2624,7 @@ defmodule Group.Replica do state = process_inline_priority_message(state, msg) take_priority_control_turn(state, remaining - 1) - {:replica_hello_request, _remote_pid} = msg -> + {:replica_hello_request, _remote_pid, _request_token} = msg -> state = process_inline_priority_message(state, msg) take_priority_control_turn(state, remaining - 1) @@ -2636,7 +2808,7 @@ defmodule Group.Replica do "#{log_prefix_shard(state)} flush_replica_broadcast_buffer ops=#{length(ops)}" end) - send_replicated_batches(state, ops) + state = send_replicated_batches(state, ops) %{ state @@ -2867,8 +3039,8 @@ defmodule Group.Replica do defp send_replicated_batches(state, ops) do ops |> group_broadcast_ops_by_target(state, &sequenced_op_cluster/1) - |> Enum.each(fn {target_node, target_ops} -> - send_replica_delta_batch(state, target_node, target_ops) + |> Enum.reduce(state, fn {target_node, target_ops}, acc -> + send_replica_delta_batch(acc, target_node, target_ops) end) end @@ -2890,6 +3062,14 @@ defmodule Group.Replica do {stream_id, first_seq, records, head} end) + state = + Enum.reduce(runs, state, fn {stream_id, _first_seq, _records, head}, acc -> + {floor, _head, _applied} = + Data.replica_stream_head(acc.name, acc.shard_index, stream_id) + + retain_pending_replica_head(acc, target_node, {stream_id, floor, head}) + end) + outgoing_replica_message(state, target_node, {:delta_batch, WireProtocol.version(), runs}) end @@ -3131,8 +3311,83 @@ defmodule Group.Replica do state end + defp peer_connect_message(state, target_node) do + {:peer_connect, self(), state.shard_index, state.num_shards, + Map.get(state.peer_probe_epochs, target_node, 0), Data.generation(state.name), + Data.local_cluster_epoch_revision(state.name)} + end + + defp advance_peer_probe_epoch(state, target_node) do + probe_epoch = Map.get(state.peer_probe_epochs, target_node, 0) + 1 + + %{ + state + | peer_probe_epochs: Map.put(state.peer_probe_epochs, target_node, probe_epoch), + pending_peer_probes: Map.put(state.pending_peer_probes, target_node, probe_epoch) + } + end + + defp maybe_send_discovery_hello(state, target_node, force?) do + now = monotonic_millis() + last_sent = Map.get(state.discovery_hello_last_sent, target_node) + retry_interval = @discovery_hello_retry_interval + authority = {Data.generation(state.name), Data.local_cluster_epoch_revision(state.name)} + + if force? or is_nil(last_sent) or elem(last_sent, 1) != authority or + now - elem(last_sent, 0) >= retry_interval do + state = send_replica_hello(state, target_node) + + %{ + state + | discovery_hello_last_sent: + Map.put(state.discovery_hello_last_sent, target_node, {now, authority}) + } + else + state + end + end + + defp request_replica_authority(%{shard_index: 0} = state, remote_node) do + {request_token, tokens} = + case Map.fetch(state.authority_request_tokens, remote_node) do + {:ok, token} -> + {token, state.authority_request_tokens} + + :error -> + token = make_ref() + {token, Map.put(state.authority_request_tokens, remote_node, token)} + end + + send_remote_control_message( + state, + remote_node, + {:replica_hello_request, self(), request_token} + ) + + %{state | authority_request_tokens: tokens} + end + defp request_replica_authority(state, remote_node) do - send_remote_control_message(state, remote_node, {:replica_hello_request, self()}) + # A missed local fanout can leave this lane behind even though shard zero + # already installed the exact authority. Repair that view locally. + case Data.remote_replica_authority_hint(state.name, remote_node) do + {generation, revision} -> + if replica_exact_authority_current?(state, remote_node, generation, revision) and + Data.remote_cluster_epoch_revision(state.name, remote_node) == revision do + install_current_replica_lane(state, remote_node, generation) + else + request_authority_from_local_control(state, remote_node) + end + + nil -> + request_authority_from_local_control(state, remote_node) + end + end + + defp request_authority_from_local_control(state, remote_node) do + # All lanes share one exact authority. Route repair through its local owner + # so the remote node sees one stable request token, not one per lane. + _ = send_local_control_message(state, {:replica_authority_dirty_local, remote_node}) state end @@ -3194,7 +3449,16 @@ defmodule Group.Replica do Data.remote_cluster_epoch_observed_revision(state.name, remote_node) ) do :ok -> - reproject_pending_registry_keys(state, remote_node) + state = reproject_pending_registry_keys(state, remote_node) + + if replica_view_current?(state, remote_node) do + %{ + state + | authority_request_tokens: Map.delete(state.authority_request_tokens, remote_node) + } + else + state + end :stale -> state @@ -3240,7 +3504,16 @@ defmodule Group.Replica do end defp touch_replica_peer(state, remote_node) do - %{state | peer_last_seen: Map.put(state.peer_last_seen, remote_node, monotonic_millis())} + state = %{ + state + | peer_last_seen: Map.put(state.peer_last_seen, remote_node, monotonic_millis()) + } + + if replica_view_current?(state, remote_node) do + %{state | authority_request_tokens: Map.delete(state.authority_request_tokens, remote_node)} + else + state + end end defp ensure_replica_peer_retirement_deadline(state, remote_node) do @@ -3344,20 +3617,21 @@ defmodule Group.Replica do defp request_quiet_cluster_hellos(state) do now = monotonic_millis() - dirty = - Enum.reduce(state.cluster_control_dirty, %{}, fn {remote_node, last_activity}, acc -> - cond do - is_nil(Data.remote_generation(state.name, remote_node)) and - is_nil(Data.remote_replica_authority_hint(state.name, remote_node)) -> - acc + {state, dirty} = + Enum.reduce(state.cluster_control_dirty, {state, %{}}, fn + {remote_node, last_activity}, {acc_state, dirty} -> + cond do + is_nil(Data.remote_generation(state.name, remote_node)) and + is_nil(Data.remote_replica_authority_hint(state.name, remote_node)) -> + {acc_state, dirty} - now - last_activity >= state.replicated_anti_entropy_interval -> - request_replica_authority(state, remote_node) - Map.put(acc, remote_node, now) + now - last_activity >= state.replicated_anti_entropy_interval -> + {request_replica_authority(acc_state, remote_node), + Map.put(dirty, remote_node, now)} - true -> - Map.put(acc, remote_node, last_activity) - end + true -> + {acc_state, Map.put(dirty, remote_node, last_activity)} + end end) %{state | cluster_control_dirty: dirty} @@ -3381,12 +3655,12 @@ defmodule Group.Replica do defp probe_replica_peers(state) do Enum.each(Node.list(), fn remote_node -> - unless Map.has_key?(state.remote_shards, remote_node) do + if not Map.has_key?(state.remote_shards, remote_node) or + Map.has_key?(state.pending_peer_probes, remote_node) do send_remote_shard_message( state, remote_node, - {:peer_connect, self(), state.shard_index, state.num_shards, - Data.my_clusters(state.name)} + peer_connect_message(state, remote_node) ) end end) @@ -3394,40 +3668,110 @@ defmodule Group.Replica do state end - defp broadcast_replica_heads(state) do - peers = Map.keys(state.peer_last_seen) + defp retry_pending_replica_heads(state) do + Enum.reduce(Map.keys(state.pending_replica_heads), state, fn target_node, acc -> + send_pending_replica_heads(acc, target_node) + end) + end - heads_by_target = - state.name - |> Data.replica_stream_heads(state.shard_index) - |> Enum.reduce(%{}, fn {stream_id, _floor, _head} = head, acc -> - if current_local_replica_stream?(state, stream_id) do - targets = - case WireProtocol.stream_cluster(stream_id) do - nil -> - peers + defp retain_pending_replica_head(state, target_node, {stream_id, floor, head}) do + pending = + Map.update(state.pending_replica_heads, target_node, %{stream_id => {floor, head, nil}}, fn + streams -> + Map.update(streams, stream_id, {floor, head, nil}, fn + {_old_floor, old_head, _last_sent} when head > old_head -> {floor, head, nil} + existing -> existing + end) + end) - cluster -> - state.name - |> Data.cluster_nodes(cluster) - |> Enum.filter(&Map.has_key?(state.peer_last_seen, &1)) - end + %{ + state + | pending_replica_heads: pending, + replica_send_tokens: Map.put_new(state.replica_send_tokens, target_node, make_ref()) + } + end - Enum.reduce(targets, acc, fn target_node, inner -> - Map.update(inner, target_node, [head], &[head | &1]) - end) - else - acc - end + defp send_pending_replica_heads(state, target_node, only_streams \\ :all) do + streams = + state.pending_replica_heads + |> Map.get(target_node, %{}) + |> Map.filter(fn {stream_id, _pending} -> + replica_stream_target?(state, stream_id, target_node) end) - Enum.reduce(heads_by_target, state, fn {target_node, heads}, acc -> - outgoing_replica_message( - acc, - target_node, - {:heads, WireProtocol.version(), Enum.reverse(heads)} - ) - end) + state = put_pending_replica_streams(state, target_node, streams) + + due = + streams + |> Enum.filter(fn {stream_id, _pending} -> + only_streams == :all or MapSet.member?(only_streams, stream_id) + end) + |> Enum.sort_by(fn {stream_id, {_floor, _head, last_sent}} -> + {not is_nil(last_sent), last_sent || 0, stream_id} + end) + + send_pending_replica_head_batches(state, target_node, due, 0) + end + + defp send_pending_replica_head_batches(state, _target_node, [], _sent), do: state + + defp send_pending_replica_head_batches(state, _target_node, _due, sent) + when sent >= @replica_head_batches_per_peer_turn, + do: state + + defp send_pending_replica_head_batches(state, target_node, due, sent) do + {batch, rest} = take_replica_head_batch(due, [], 0) + + heads = + Enum.map(batch, fn {stream_id, {floor, head, _last_sent}} -> {stream_id, floor, head} end) + + case state.replica_transport.outgoing( + state.name, + target_node, + state.shard_index, + {:heads, WireProtocol.version(), Map.fetch!(state.replica_send_tokens, target_node), + heads}, + state.replica_transport_opts + ) do + :ok -> + now = monotonic_millis() + + streams = + Enum.reduce(batch, Map.fetch!(state.pending_replica_heads, target_node), fn + {stream_id, {_floor, sent_head, _last_sent}}, acc -> + Map.update!(acc, stream_id, fn {floor, head, _last_sent} -> + if head == sent_head, do: {floor, head, now}, else: {floor, head, nil} + end) + end) + + state + |> put_pending_replica_streams(target_node, streams) + |> send_pending_replica_head_batches(target_node, rest, sent + 1) + + result when result in [:busy, :disconnected] -> + state + end + end + + defp take_replica_head_batch([], batch, _bytes), do: {Enum.reverse(batch), []} + + defp take_replica_head_batch([entry | rest] = entries, batch, bytes) do + {stream_id, {floor, head, _last_sent}} = entry + entry_bytes = :erlang.external_size({stream_id, floor, head}) + + if batch != [] and bytes + entry_bytes > @replica_head_batch_target_bytes do + {Enum.reverse(batch), entries} + else + take_replica_head_batch(rest, [entry | batch], bytes + entry_bytes) + end + end + + defp put_pending_replica_streams(state, target_node, streams) when map_size(streams) == 0 do + %{state | pending_replica_heads: Map.delete(state.pending_replica_heads, target_node)} + end + + defp put_pending_replica_streams(state, target_node, streams) do + %{state | pending_replica_heads: Map.put(state.pending_replica_heads, target_node, streams)} end defp expire_stale_replica_peers(state) do @@ -3510,6 +3854,16 @@ defmodule Group.Replica do %{state | snapshot_send_offsets: offsets} end + defp discard_snapshot_send_offsets_for_target_stream(state, target_node, stream_id) do + offsets = + Map.reject(state.snapshot_send_offsets, fn + {{^target_node, ^stream_id, _head}, _offset} -> true + {_key, _offset} -> false + end) + + %{state | snapshot_send_offsets: offsets} + end + defp discard_snapshot_send_offsets_for_streams(state, stream_ids) do stream_ids = MapSet.new(stream_ids) @@ -3554,10 +3908,24 @@ defmodule Group.Replica do ) end + probe_epoch = Map.get(state.peer_probe_epochs, remote_node, 0) + 1 + %{ state | remote_shards: Map.delete(state.remote_shards, remote_node), peer_last_seen: Map.delete(state.peer_last_seen, remote_node), + peer_probe_epochs: Map.put(state.peer_probe_epochs, remote_node, probe_epoch), + pending_peer_probes: Map.put(state.pending_peer_probes, remote_node, probe_epoch), + remote_probe_epochs: Map.delete(state.remote_probe_epochs, remote_node), + peer_connect_ack_seen: Map.delete(state.peer_connect_ack_seen, remote_node), + authority_request_tokens: Map.delete(state.authority_request_tokens, remote_node), + remote_authority_request_tokens: + Map.delete(state.remote_authority_request_tokens, remote_node), + discovery_hello_last_sent: Map.delete(state.discovery_hello_last_sent, remote_node), + pending_replica_heads: Map.delete(state.pending_replica_heads, remote_node), + pending_replica_acks: Map.delete(state.pending_replica_acks, remote_node), + replica_send_tokens: Map.delete(state.replica_send_tokens, remote_node), + replica_receive_tokens: Map.delete(state.replica_receive_tokens, remote_node), cluster_control_dirty: Map.delete(state.cluster_control_dirty, remote_node), authority_dirty_notified: MapSet.delete(state.authority_dirty_notified, remote_node) } @@ -3570,18 +3938,46 @@ defmodule Group.Replica do defp send_replica_heads(state, target_node, clusters) do heads = replica_heads_for_clusters(state, target_node, clusters) - if heads == [] do + state = %{ state - else - outgoing_replica_message(state, target_node, {:heads, WireProtocol.version(), heads}) + | replica_send_tokens: Map.put_new(state.replica_send_tokens, target_node, make_ref()) + } + + state = + Enum.reduce(heads, state, fn head, acc -> + retain_pending_replica_head(acc, target_node, head) + end) + + case {clusters, heads} do + {:all, []} -> + outgoing_replica_message( + state, + target_node, + {:heads, WireProtocol.version(), Map.fetch!(state.replica_send_tokens, target_node), []} + ) + + {:all, _heads} -> + send_pending_replica_heads(state, target_node) + + {_clusters, []} -> + state + + {_clusters, heads} -> + stream_ids = MapSet.new(heads, &elem(&1, 0)) + send_pending_replica_heads(state, target_node, stream_ids) end end + defp reset_replica_send_token(state, target_node) do + state = discard_snapshot_send_offsets_for_target(state, target_node) + %{state | replica_send_tokens: Map.put(state.replica_send_tokens, target_node, make_ref())} + end + defp replica_heads_for_clusters(state, target_node, :all) do state.name |> Data.replica_stream_heads(state.shard_index) - |> Enum.filter(fn {stream_id, _floor, _head} -> - replica_stream_target?(state, stream_id, target_node) + |> Enum.filter(fn {stream_id, _floor, head} -> + head > 0 and replica_stream_target?(state, stream_id, target_node) end) end @@ -3640,15 +4036,42 @@ defmodule Group.Replica do end end - defp handle_replica_message(state, source_node, {:heads, version, heads}) - when version == @protocol_version and is_list(heads) do - if Enum.all?(heads, &valid_replica_head?/1) do + defp handle_replica_message(state, source_node, {:heads, version, token, heads}) + when version == @protocol_version and is_reference(token) and is_list(heads) do + if replica_view_current?(state, source_node) and + Map.has_key?(state.remote_shards, source_node) and + Enum.all?(heads, &valid_replica_head?/1) do + state = %{ + state + | replica_receive_tokens: Map.put(state.replica_receive_tokens, source_node, token) + } + handle_replica_heads(state, source_node, heads) else state end end + defp handle_replica_message( + state, + source_node, + {:applied, version, receiver_pid, receiver_generation, token, cursors} + ) + when version == @protocol_version and is_pid(receiver_pid) and is_reference(token) and + is_list(cursors) do + if node(receiver_pid) == source_node and + Map.get(state.remote_shards, source_node) == receiver_pid and + Data.remote_generation(state.name, source_node) == receiver_generation and + Map.get(state.replica_send_tokens, source_node) == token and + replica_view_current?(state, source_node) do + Enum.reduce(cursors, state, fn cursor, acc -> + acknowledge_replica_cursor(acc, source_node, cursor) + end) + else + state + end + end + defp handle_replica_message(state, source_node, {:delta_batch, version, runs}) when version == @protocol_version and is_list(runs) do if Enum.all?(runs, &valid_replica_delta_run?/1) do @@ -3662,18 +4085,6 @@ defmodule Group.Replica do end end - defp handle_replica_message(state, source_node, {:need, version, stream_id, next_seq}) - when version == @protocol_version and is_integer(next_seq) and next_seq > 0 do - if WireProtocol.valid_stream_id?(stream_id) and - WireProtocol.stream_origin(stream_id) == node() and - WireProtocol.stream_shard(stream_id) == state.shard_index and - replica_stream_target?(state, stream_id, source_node) do - send_replica_repair(state, source_node, stream_id, next_seq) - else - state - end - end - defp handle_replica_message(state, source_node, {:needs, version, needs}) when version == @protocol_version and is_list(needs) do if Enum.all?(needs, &valid_replica_need?/1) do @@ -3734,23 +4145,140 @@ defmodule Group.Replica do defp handle_replica_message(state, _source_node, _message), do: state defp handle_replica_heads(state, source_node, heads) do - needs = - Enum.flat_map(heads, fn {stream_id, _floor, head} -> + {state, needs} = + Enum.reduce(heads, {state, []}, fn {stream_id, _floor, head}, {acc, needs} -> if valid_remote_stream?(state, source_node, stream_id) do cursor = Data.replica_cursor(state.name, state.shard_index, stream_id) - if head > cursor, do: [{stream_id, cursor + 1}], else: [] + + if head > cursor do + {acc, [{stream_id, cursor + 1, head} | needs]} + else + {queue_replica_ack(acc, source_node, stream_id, cursor), needs} + end else - [] + {acc, needs} end end) needs + |> Enum.reverse() |> Enum.chunk_every(state.replicated_sender_buffer_size) |> Enum.reduce(state, fn chunk, acc -> outgoing_replica_message(acc, source_node, {:needs, WireProtocol.version(), chunk}) end) end + defp acknowledge_replica_cursor(state, target_node, {stream_id, cursor, receiver_epoch}) + when is_integer(cursor) and cursor >= 0 do + if WireProtocol.valid_stream_id?(stream_id) and + current_local_replica_stream?(state, stream_id) and + replica_stream_target?(state, stream_id, target_node) and + Data.remote_cluster_epoch( + state.name, + target_node, + WireProtocol.stream_cluster(stream_id) + ) == receiver_epoch do + case get_in(state.pending_replica_heads, [target_node, stream_id]) do + {_floor, head, _last_sent} when cursor >= head -> + streams = + state.pending_replica_heads |> Map.fetch!(target_node) |> Map.delete(stream_id) + + state + |> put_pending_replica_streams(target_node, streams) + |> discard_snapshot_send_offsets_for_target_stream(target_node, stream_id) + + _ -> + state + end + else + state + end + end + + defp acknowledge_replica_cursor(state, _target_node, _cursor), do: state + + defp queue_replica_ack(state, source_node, stream_id, cursor) do + epoch = Data.local_cluster_epoch(state.name, WireProtocol.stream_cluster(stream_id)) + + pending = + Map.update(state.pending_replica_acks, source_node, %{stream_id => {cursor, epoch}}, fn + streams -> + Map.update(streams, stream_id, {cursor, epoch}, fn + {old_cursor, ^epoch} -> {max(cursor, old_cursor), epoch} + {_old_cursor, _old_epoch} -> {cursor, epoch} + end) + end) + + state = %{state | pending_replica_acks: pending} + + cond do + state.replica_ack_backoff -> + schedule_replica_ack_flush(state, @replica_ack_busy_retry_interval) + + map_size(Map.fetch!(pending, source_node)) >= state.replicated_sender_buffer_size -> + flush_replica_acks(state) + + true -> + schedule_replica_ack_flush(state) + end + end + + defp schedule_replica_ack_flush(state, interval \\ @replica_ack_flush_interval) + + defp schedule_replica_ack_flush(%{replica_ack_flush_ref: ref} = state, _interval) + when is_reference(ref), + do: state + + defp schedule_replica_ack_flush(state, interval) do + ref = make_ref() + Process.send_after(self(), {@replica_ack_flush_timer, ref}, interval) + %{state | replica_ack_flush_ref: ref} + end + + defp flush_replica_acks(state) do + if is_reference(state.replica_ack_flush_ref) do + Process.cancel_timer(state.replica_ack_flush_ref) + end + + state = %{state | replica_ack_flush_ref: nil} + + pending = + Enum.reduce(state.pending_replica_acks, %{}, fn {source_node, streams}, acc -> + retained = flush_replica_ack_chunks(state, source_node, Map.to_list(streams)) + + if map_size(retained) == 0, do: acc, else: Map.put(acc, source_node, retained) + end) + + state = %{ + state + | pending_replica_acks: pending, + replica_ack_backoff: map_size(pending) > 0 + } + + if map_size(pending) == 0, + do: state, + else: schedule_replica_ack_flush(state, @replica_ack_busy_retry_interval) + end + + defp flush_replica_ack_chunks(_state, _source_node, []), do: %{} + + defp flush_replica_ack_chunks(state, source_node, streams) do + {chunk, rest} = Enum.split(streams, state.replicated_sender_buffer_size) + cursors = Enum.map(chunk, fn {stream_id, {cursor, epoch}} -> {stream_id, cursor, epoch} end) + + case state.replica_transport.outgoing( + state.name, + source_node, + state.shard_index, + {:applied, WireProtocol.version(), self(), Data.generation(state.name), + Map.get(state.replica_receive_tokens, source_node), cursors}, + state.replica_transport_opts + ) do + :ok -> flush_replica_ack_chunks(state, source_node, rest) + result when result in [:busy, :disconnected] -> Map.new(streams) + end + end + defp valid_snapshot_stream?(state, source_node, stream_id, snapshot_seq) do valid_remote_stream?(state, source_node, stream_id) and snapshot_seq > Data.replica_cursor(state.name, state.shard_index, stream_id) @@ -3808,8 +4336,9 @@ defmodule Group.Replica do defp valid_replica_head?(_head), do: false - defp valid_replica_need?({stream_id, next_seq}) do - WireProtocol.valid_stream_id?(stream_id) and is_integer(next_seq) and next_seq > 0 + defp valid_replica_need?({stream_id, next_seq, advertised_head}) do + WireProtocol.valid_stream_id?(stream_id) and is_integer(next_seq) and next_seq > 0 and + is_integer(advertised_head) and advertised_head >= next_seq end defp valid_replica_need?(_need), do: false @@ -4057,7 +4586,7 @@ defmodule Group.Replica do ) notify_snapshot_events(state.name, transfer.events) - state + queue_replica_ack(state, source_node, stream_id, transfer.snapshot_seq) else state end @@ -4072,10 +4601,10 @@ defmodule Group.Replica do case records do [] -> - state + queue_replica_ack(state, source_node, stream_id, cursor) [{first_seq, _mutations} | _] when first_seq > cursor + 1 -> - request_replica_need(state, source_node, stream_id, cursor + 1) + request_replica_need(state, source_node, stream_id, cursor + 1, advertised_head) _ -> {contiguous, _next_seq} = take_contiguous_replica_records(records, cursor + 1, []) @@ -4105,14 +4634,15 @@ defmodule Group.Replica do if rejected == [] do state else - request_replica_need(state, source_node, stream_id, cursor + 1) + request_replica_need(state, source_node, stream_id, cursor + 1, advertised_head) end {last_seq, _mutations} -> :ok = Data.put_replica_cursor(state.name, state.shard_index, stream_id, last_seq) + state = queue_replica_ack(state, source_node, stream_id, last_seq) if last_seq < advertised_head or length(accepted) < length(records) do - request_replica_need(state, source_node, stream_id, last_seq + 1) + request_replica_need(state, source_node, stream_id, last_seq + 1, advertised_head) else state end @@ -4344,30 +4874,32 @@ defmodule Group.Replica do end) end - defp send_replica_repair(state, target_node, stream_id, next_seq) do - send_replica_repairs(state, target_node, [{stream_id, next_seq}]) - end - - defp request_replica_need(state, target_node, stream_id, next_seq) do + defp request_replica_need(state, target_node, stream_id, next_seq, advertised_head) do outgoing_replica_message( state, target_node, - {:needs, WireProtocol.version(), [{stream_id, next_seq}]} + {:needs, WireProtocol.version(), [{stream_id, next_seq, advertised_head}]} ) end defp send_replica_repairs(state, target_node, needs) do {state, runs} = - Enum.reduce(needs, {state, []}, fn {stream_id, next_seq}, {acc, runs} -> - if WireProtocol.stream_origin(stream_id) == node() and - WireProtocol.stream_shard(stream_id) == acc.shard_index and - replica_stream_target?(acc, stream_id, target_node) do - case replica_repair(acc, target_node, stream_id, next_seq) do - {:run, run} -> {acc, [run | runs]} - {:state, acc} -> {acc, runs} - end - else - {acc, runs} + Enum.reduce(needs, {state, []}, fn {stream_id, next_seq, advertised_head}, {acc, runs} -> + case get_in(acc.pending_replica_heads, [target_node, stream_id]) do + {_floor, ^advertised_head, _last_sent} -> + if WireProtocol.stream_origin(stream_id) == node() and + WireProtocol.stream_shard(stream_id) == acc.shard_index and + replica_stream_target?(acc, stream_id, target_node) do + case replica_repair(acc, target_node, stream_id, next_seq) do + {:run, run} -> {acc, [run | runs]} + {:state, acc} -> {acc, runs} + end + else + {acc, runs} + end + + _ -> + {acc, runs} end end) @@ -4457,14 +4989,33 @@ defmodule Group.Replica do end defp send_replica_snapshot(state, target_node, stream_id, head) do - case state.snapshot_send do - {worker, _token, _snapshot_key} when is_pid(worker) -> - if Process.alive?(worker), - do: state, - else: start_replica_snapshot_send(state, target_node, stream_id, head) + snapshot_key = {target_node, stream_id, head} + now = monotonic_millis() - nil -> - start_replica_snapshot_send(state, target_node, stream_id, head) + retry_interval = + min(@snapshot_retry_interval, max(1, div(state.replicated_peer_lease_timeout, 2))) + + case Map.get(state.snapshot_send_offsets, snapshot_key) do + {:sent, sent_at} when now - sent_at < retry_interval -> + state + + _ -> + case state.snapshot_send do + {worker, _token, _snapshot_key} when is_pid(worker) -> + if Process.alive?(worker), + do: state, + else: start_replica_snapshot_send(state, target_node, stream_id, head) + + nil -> + start_replica_snapshot_send(state, target_node, stream_id, head) + end + end + end + + defp snapshot_head_pending?(state, {target_node, stream_id, head}) do + case get_in(state.pending_replica_heads, [target_node, stream_id]) do + {_floor, ^head, _last_sent} -> true + _ -> false end end @@ -4481,7 +5032,12 @@ defmodule Group.Replica do end) |> Map.new() - resume = Map.get(offsets, snapshot_key, {:chunk, 1}) + resume = + case Map.get(offsets, snapshot_key) do + {:sent, _sent_at} -> {:chunk, 1} + nil -> {:chunk, 1} + offset -> offset + end snapshot_context = %{ name: state.name, @@ -5829,12 +6385,6 @@ defmodule Group.Replica do defp exit_local_conflict_loser(_pid, _key, _winner_meta), do: :ok - defp compute_shared_clusters(my_clusters, remote_clusters) do - my_set = MapSet.new(my_clusters) - remote_set = MapSet.new(remote_clusters) - MapSet.intersection(my_set, remote_set) |> MapSet.to_list() - end - defp purge_cluster_entries(name, shard, cluster, target) do # Local disconnect uses :all to discard the complete replicated view. # Remote disconnects remove only data owned by the departing node. diff --git a/lib/group/replica/wire_protocol.ex b/lib/group/replica/wire_protocol.ex index 8219350..dc07624 100644 --- a/lib/group/replica/wire_protocol.ex +++ b/lib/group/replica/wire_protocol.ex @@ -1,7 +1,7 @@ defmodule Group.Replica.WireProtocol do @moduledoc false - @version 3 + @version 7 def version, do: @version diff --git a/test/README.md b/test/README.md index 82d7cbb..dc3a0db 100644 --- a/test/README.md +++ b/test/README.md @@ -83,6 +83,7 @@ distributed scenarios always get fresh peers; there is no shared peer pool. | `diagnostics_test.exs` | Failure output, shard snapshots and unreachable-peer diagnostics | | `distributed_test.exs` | Multi-node: replication, peer discovery, node disconnect cleanup, partition healing, conflict resolution, event ordering, rolling restarts, and adversarial replica-transport loss/busy/snapshot recovery | | `anti_entropy_fault_regression_test.exs` | Three-node regressions for hidden-winner projection, receiver restart eviction, nodedown/lease lane retirement, authority gaps and cross-lane races, in-flight conflict fencing, crash-journal replay, cursorless/interrupted snapshot repair, malformed ingress, and sideband rediscovery | +| `replica_ack_test.exs` | ACK-driven stream quiescence, lost ACK and need recovery, one-sided lease expiry, stale hello/probe races, PID and authority fencing, delta and snapshot repair, and absence of repeated full sends | | `replica_adversarial_test.exs` | Reproducible three-node mixed-operation state machines: drops, busy returns, duplication, reordering, bounded delay, oplog pruning, conflicts, owner death, and named-cluster epoch churn, followed by exact convergence/dead-owner/internal-index checks | | `replica_model_property_test.exs` | StreamData-generated and shrunk owner histories against an independent lifecycle oracle and scheduler-controlled replica transport | | `replica_snapshot_test.exs` | Pure single-pass byte-bounded streaming, suffix resume, receive staging, and event batching | diff --git a/test/anti_entropy_fault_regression_test.exs b/test/anti_entropy_fault_regression_test.exs index 5af9bd8..5830df6 100644 --- a/test/anti_entropy_fault_regression_test.exs +++ b/test/anti_entropy_fault_regression_test.exs @@ -478,7 +478,7 @@ defmodule Group.AntiEntropyFaultRegressionTest do version = Group.Replica.WireProtocol.version() frames = [ - {:heads, version, [:not_a_head]}, + {:heads, version, make_ref(), [:not_a_head]}, {:delta_batch, version, [:not_a_delta_run]}, {:need, version, :not_a_stream, 1}, {:needs, version, [:not_a_need]}, @@ -604,6 +604,16 @@ defmodule Group.AntiEntropyFaultRegressionTest do start_group_on_peers(context.peers, opts) + TestCluster.assert_eventually(fn -> + context.node_b in TestCluster.rpc!(context.node_a, Group, :nodes, [name]) + end) + + :ok = + TestCluster.rpc!(context.node_a, Group.TestReplicaTransport, :set_mode, [ + name, + {:drop_types, [:heads, :delta_batch]} + ]) + for index <- 1..4 do owner = TestCluster.spawn_register(context.node_a, name, "snapshot/isolation/#{index}", %{}) true = TestCluster.rpc!(context.node_a, Process, :exit, [owner, :kill]) @@ -614,7 +624,7 @@ defmodule Group.AntiEntropyFaultRegressionTest do stream_id = TestCluster.rpc!(context.node_a, Group.Replica.Data, :local_stream_id, [name, 0, nil]) - {floor, _head, _applied} = + {floor, head, _applied} = TestCluster.rpc!(context.node_a, Group.Replica.Data, :replica_stream_head, [ name, 0, @@ -623,6 +633,14 @@ defmodule Group.AntiEntropyFaultRegressionTest do assert floor > 1 + source_state = + TestCluster.rpc!(context.node_a, :sys, :get_state, [Group.Replica.shard_name(name, 0)]) + + assert Map.has_key?( + Map.get(source_state.pending_replica_heads, context.node_b, %{}), + stream_id + ) + shard = TestCluster.rpc!(context.node_a, Process, :whereis, [Group.Replica.shard_name(name, 0)]) @@ -657,7 +675,7 @@ defmodule Group.AntiEntropyFaultRegressionTest do name, context.node_b, 0, - {:needs, Group.Replica.WireProtocol.version(), [{stream_id, 1}]} + {:needs, Group.Replica.WireProtocol.version(), [{stream_id, 1, head}]} ]) assert_receive {:forwarded_trace, @@ -3263,7 +3281,7 @@ defmodule Group.AntiEntropyFaultRegressionTest do name, context.node_a, 0, - {:heads, Group.Replica.WireProtocol.version(), [{stream_id, floor, head}]} + {:heads, Group.Replica.WireProtocol.version(), make_ref(), [{stream_id, floor, head}]} ]) TestCluster.assert_eventually( @@ -4005,9 +4023,6 @@ defmodule Group.AntiEntropyFaultRegressionTest do {:ok, _pid} = TestCluster.start_group(context.node_a, opts) - source_control = - TestCluster.rpc!(context.node_a, Process, :whereis, [Group.Replica.shard_name(name, 0)]) - source_lane = TestCluster.rpc!(context.node_a, Process, :whereis, [Group.Replica.shard_name(name, 1)]) @@ -4044,8 +4059,7 @@ defmodule Group.AntiEntropyFaultRegressionTest do TestCluster.rpc!(context.node_b, Process, :info, [target_control, :messages]) Enum.any?(messages, fn - {:replica_hello, ^source_control, _version, ^generation, ^revision, _epochs, _transport, - _descriptor} -> + {:replica_authority_dirty_local, peer} when peer == context.node_a -> true _message -> @@ -4101,8 +4115,12 @@ defmodule Group.AntiEntropyFaultRegressionTest do TestCluster.rpc!(context.node_a, Process, :info, [source_lane, :messages]) Enum.any?(messages, fn - {:peer_connect, ^target_lane, 1, 2, _clusters} -> true - _message -> false + {:peer_connect, ^target_lane, 1, 2, probe_epoch, _generation, _revision} + when probe_epoch > 0 -> + true + + _message -> + false end) end) diff --git a/test/distributed_test.exs b/test/distributed_test.exs index f17a1c3..c85f0fe 100644 --- a/test/distributed_test.exs +++ b/test/distributed_test.exs @@ -4410,12 +4410,10 @@ defmodule Group.DistributedTest do ["authority-new"] ]) - remote_clusters = TestCluster.rpc!(node_b, Group.Replica.Data, :my_clusters, [name]) - Enum.each(b_lanes, fn {shard, b_pid} -> TestCluster.rpc!(node_a, :erlang, :send, [ shard_name(name, shard), - {:peer_connect_ack, b_pid, shard, shards, remote_clusters} + {:peer_connect_ack, b_pid, shard, shards, 0, true} ]) end) @@ -4614,6 +4612,92 @@ defmodule Group.DistributedTest do ) == latest_revision end + @tag timeout: 60_000 + test "authority requests from multiple lanes do not repeat the full catalog" do + peers = TestCluster.start_peers(2) + on_exit(fn -> TestCluster.stop_peers(peers) end) + + [{_, node_a}, {_, node_b}] = peers + name = :"anti_entropy_lane_requests_#{System.unique_integer([:positive])}" + shards = 3 + + opts = [ + name: name, + shards: shards, + replica_transport: Group.TestReplicaTransport, + replicated_anti_entropy_interval: 600_000, + replicated_peer_lease_timeout: 1_200_000 + ] + + start_group_on_peers(peers, opts) + + TestCluster.assert_eventually(fn -> + node_a in TestCluster.rpc!(node_b, Group, :nodes, [name]) + end) + + b_control = TestCluster.rpc!(node_b, Process, :whereis, [shard_name(name, 0)]) + a_control = TestCluster.rpc!(node_a, Process, :whereis, [shard_name(name, 0)]) + :ok = TestCluster.rpc!(node_b, :sys, :suspend, [b_control]) + + on_exit(fn -> + TestCluster.rpc!(node_b, TestCluster, :resume_if_alive, [b_control]) + end) + + :ok = + TestCluster.rpc!(node_b, Group.Replica.Data, :delete_remote_replica_info, [ + name, + 0, + node_a + ]) + + {generation, revision, epochs} = + TestCluster.rpc!(node_a, Group.Replica.Data, :local_replica_authority, [name]) + + for shard <- 1..(shards - 1) do + b_lane = TestCluster.rpc!(node_b, Process, :whereis, [shard_name(name, shard)]) + marker = make_ref() + + send( + b_lane, + {:replica_hello, a_control, Group.Replica.WireProtocol.version(), generation, revision, + epochs, Group.TestReplicaTransport.id(), + Group.TestReplicaTransport.descriptor(name, [])} + ) + + send( + b_lane, + {:group_dispatch, [a_control], {:group_dispatch, [self()], {:lane_processed, marker}}} + ) + + assert_receive {:lane_processed, ^marker}, 5_000 + end + + {:messages, messages} = TestCluster.rpc!(node_b, Process, :info, [b_control, :messages]) + + full_hellos = + Enum.count(messages, fn + {:replica_hello, ^a_control, _, _, _, _, _, _} -> true + _ -> false + end) + + assert full_hellos == 0 + + :ok = TestCluster.rpc!(node_b, :sys, :resume, [b_control]) + + TestCluster.assert_eventually(fn -> + source_revision = + TestCluster.rpc!(node_a, Group.Replica.Data, :local_cluster_epoch_revision, [name]) + + TestCluster.rpc!(node_b, Group.Replica.Data, :remote_cluster_epoch_exact_revision, [ + name, + node_a + ]) == source_revision + end) + + source_state = TestCluster.rpc!(node_a, :sys, :get_state, [a_control]) + assert {^b_control, _token} = Map.get(source_state.remote_authority_request_tokens, node_b) + end + @tag timeout: 60_000 test "a backlogged authority shard cannot block independent replica lanes" do peers = TestCluster.start_peers(2) @@ -4642,17 +4726,47 @@ defmodule Group.DistributedTest do Enum.each([1, 2], fn shard_index -> b_lane = TestCluster.rpc!(node_b, Process, :whereis, [shard_name(name, shard_index)]) - TestCluster.assert_eventually(fn -> - b_lane_state = TestCluster.rpc!(node_b, :sys, :get_state, [b_lane]) - - Map.has_key?(b_lane_state.peer_last_seen, node_a) and - TestCluster.rpc!( - node_b, - Group.Replica.Data, - :remote_view_generation, - [name, shard_index, node_a] - ) == a_generation - end) + TestCluster.assert_eventually( + fn -> + b_lane_state = TestCluster.rpc!(node_b, :sys, :get_state, [b_lane]) + + Map.has_key?(b_lane_state.peer_last_seen, node_a) and + TestCluster.rpc!( + node_b, + Group.Replica.Data, + :remote_view_generation, + [name, shard_index, node_a] + ) == a_generation + end, + diagnostic: fn -> + lane_state = TestCluster.rpc!(node_b, :sys, :get_state, [b_lane]) + + %{ + shard: shard_index, + seen: Map.keys(lane_state.peer_last_seen), + remote_shards: Map.keys(lane_state.remote_shards), + requests: lane_state.authority_request_tokens, + view: + TestCluster.rpc!(node_b, Group.Replica.Data, :remote_view_generation, [ + name, + shard_index, + node_a + ]), + generation: + TestCluster.rpc!(node_b, Group.Replica.Data, :remote_generation, [name, node_a]), + exact: + TestCluster.rpc!( + node_b, + Group.Replica.Data, + :remote_cluster_epoch_exact_revision, + [ + name, + node_a + ] + ) + } + end + ) end) b_control = TestCluster.rpc!(node_b, Process, :whereis, [shard_name(name, 0)]) @@ -4661,7 +4775,7 @@ defmodule Group.DistributedTest do TestCluster.rpc!(node_a, :erlang, :send, [ shard_name(name, 0), - {:peer_connect_ack, b_control, 0, shards, [nil]} + {:replica_hello_request, b_control, make_ref()} ]) TestCluster.assert_eventually(fn -> @@ -5499,7 +5613,7 @@ defmodule Group.DistributedTest do {:messages, sender_messages} = TestCluster.rpc!(node_a, Process, :info, [sender, :messages]) refute Enum.any?(sender_messages, fn - {:replica_hello_request, remote_pid} -> node(remote_pid) == node_b + {:replica_hello_request, remote_pid, _token} -> node(remote_pid) == node_b _ -> false end) end @@ -5579,7 +5693,7 @@ defmodule Group.DistributedTest do {:messages, sender_messages} = TestCluster.rpc!(node_a, Process, :info, [sender, :messages]) refute Enum.any?(sender_messages, fn - {:replica_hello_request, remote_pid} -> node(remote_pid) == node_b + {:replica_hello_request, remote_pid, _token} -> node(remote_pid) == node_b _ -> false end) @@ -5625,7 +5739,7 @@ defmodule Group.DistributedTest do {:messages, sender_messages} = TestCluster.rpc!(node_a, Process, :info, [sender, :messages]) refute Enum.any?(sender_messages, fn - {:replica_hello_request, remote_pid} -> node(remote_pid) == node_b + {:replica_hello_request, remote_pid, _token} -> node(remote_pid) == node_b _ -> false end) @@ -5705,6 +5819,19 @@ defmodule Group.DistributedTest do interval: 100 ) + TestCluster.assert_eventually( + fn -> + Enum.all?(0..3, fn shard -> + node_a + |> TestCluster.rpc!(:sys, :get_state, [Group.Replica.shard_name(name, shard)]) + |> Map.get(:pending_replica_heads) + |> map_size() == 0 + end) + end, + timeout: 30_000, + interval: 100 + ) + for node <- [node_a, node_b] do assert :ok = TestCluster.rpc!(node, Group.TestCluster, :assert_replica_consistent, [name]) diff --git a/test/formal/README.md b/test/formal/README.md index dc4e4e0..7be4626 100644 --- a/test/formal/README.md +++ b/test/formal/README.md @@ -10,6 +10,57 @@ contract. It covers: - exact per-origin snapshot fallback; and - fair convergence after healing. +`ReplicaAck.tla` checks the ACK-driven stream protocol against the guards in +`Group.Replica`. Its state maps to `pending_replica_heads`, +`replica_send_tokens`, `replica_receive_tokens`, queued `:applied` cursors, +`remote_probe_epochs`, the receiver's cursor and materialized row, the peer +PID/generation, and the named-cluster epoch. It explores receiver lease expiry, +process restart, and rejoin after an epoch change with old heads, needs, ACKs, +and repairs still in flight. The exact source guards are represented: + +- a new `peer_connect` probe for the same PID rotates the sender token and + requeues heads; a duplicate probe preserves both and elicits no full hello + when the sender already has the exact receiver view; +- `peer_connect_ack` echoes the probe epoch and certifies that the sender's + route and exact authority matched the receiver PID, generation, and revision + after any required reseed; an ACK for an uninstalled route or stale authority + keeps the recovery probe pending; +- `:applied` clears a pending head only for the current PID, generation, token, + epoch, stream target, and a cursor at least as high as the pending head; +- `:needs` can cause repair only while its advertised head is still pending; +- a retained prefix uses delta, while a pruned prefix uses a snapshot; a + separate prune step reflects the shard-wide oplog cap, which can advance + this stream's floor even when this stream has no new writes; a sent + snapshot is held until its retry timer or a newer head; and +- no head is sent for a quiet stream, while a full hello needs discovery or + authority work. + +The model deliberately does **not** add a PID/token/epoch guard to `:needs`: +the Elixir wire message does not carry those fields. `:delta_batch` is modeled +as one contiguous run and `:snapshot_chunk` plus `:snapshot_commit` as one +committed exact state. `SnapshotAssembly.tla` checks the latter's partial wire +delivery and staging; `GroupAntiEntropy.tla` checks broader authority and +multi-receiver data convergence. The finite ACK model checks that a quiet head +has current ACK evidence, visible state is an exact committed prefix, stale +needs cannot justify a full send, and every healed run eventually converges. +The explicit wire actions check safety under loss and reordering. For liveness, +a weakly fair `Repair` action represents a successful probe/hello/head/repair/ +ACK round after transport healing, using the same pending obligations. This +assumes a healthy transport eventually completes such rounds; it does not prove +independent fairness of every packet queue. +The default configuration checks liveness through a one-sided lease expiry +with one in-flight frame. `ReplicaAckReincarnation.cfg` and `ReplicaAckEpoch.cfg` +also check liveness when the receiver PID or named-cluster authority changes. +`ReplicaAckExtended.cfg` keeps two frames and crosses both changes for safety. +The extended instance omits the +temporal property to keep the much larger interleaving space practical. +The matrix also removes the ACK token fence, the stale-need pending guard, +the full-hello discovery guard, the route-independent recovery probe, and the +ready bit on a probe ACK, and the receiver authority revision in +isolated copies. TLC must reject each mutant using only its corresponding +invariant or convergence property; a parser failure or unrelated check cannot +count as detection. + `SnapshotAssembly.tla` separately models the non-atomic wire delivery of an exact snapshot. It explores independent provisional-chunk and terminal-commit loss, duplication, and reordering, source invalidation before commit emission, @@ -67,6 +118,12 @@ Run it with Java 17 or later and a current `tla2tools.jar`: ```bash TLA_JAR=/path/to/tla2tools.jar test/formal/check.sh +# Run the ACK stream model alone (a converged terminal state is allowed) +TLA_JAR=/path/to/tla2tools.jar \ + TLA_SPEC="$PWD/test/formal/ReplicaAck.tla" \ + TLA_CONFIG="$PWD/test/formal/ReplicaAck.cfg" \ + test/formal/check.sh + TLA_JAR=/path/to/tla2tools.jar \ TLA_SPEC="$PWD/test/formal/SnapshotAssembly.tla" \ TLA_CONFIG="$PWD/test/formal/SnapshotAssembly.cfg" \ @@ -83,13 +140,15 @@ TLA_JAR=/path/to/tla2tools.jar TLA_EXTENDED=1 test/formal/check_matrix.sh ``` `TLC_WORKERS` controls worker concurrency and defaults to 4. `TLA_CONFIG` can -point at an alternate finite configuration. +point at an alternate finite configuration. `TLA_METADIR` can move TLC's +working files outside the checkout. `ReplicaAck.cfg` allows the healthy +terminal state, where there is no work left to send. TLC proves the listed invariants and liveness property for the configured finite instance, not for arbitrary unbounded node and key sets. Larger models should be run periodically by increasing `Nodes`, `Origins`, `Keys`, `MaxSeq`, -`OplogBound`, and `MaxMessages`. `check_matrix.sh` runs the protocol, snapshot -assembly, peer-eviction, authority-projection, and authority-hint models; set +`OplogBound`, and `MaxMessages`. `check_matrix.sh` runs the protocol, ACK stream, +snapshot assembly, peer-eviction, authority-projection, and authority-hint models; set `TLA_EXTENDED=1` for the larger anti-entropy configuration. The checked three-node default explores 1,835,826 states, finds 490,236 diff --git a/test/formal/ReplicaAck.cfg b/test/formal/ReplicaAck.cfg new file mode 100644 index 0000000..0d1e685 --- /dev/null +++ b/test/formal/ReplicaAck.cfg @@ -0,0 +1,18 @@ +SPECIFICATION Spec +CHECK_DEADLOCK FALSE + +CONSTANTS + MaxSeq = 2 + MaxMessages = 1 + MaxProbe = 1 + MaxPid = 1 + MaxEpoch = 1 + +INVARIANTS + TypeOK + ExactCommittedPrefix + QuietHasCurrentAck + NoUnjustifiedFullSend + +PROPERTY + HealedConvergence diff --git a/test/formal/ReplicaAck.tla b/test/formal/ReplicaAck.tla new file mode 100644 index 0000000..1e291d8 --- /dev/null +++ b/test/formal/ReplicaAck.tla @@ -0,0 +1,487 @@ +--------------------------- MODULE ReplicaAck --------------------------- +EXTENDS Integers, FiniteSets, TLC + +(* +One origin stream and one receiver, following Group.Replica's actual ACK and +repair handlers. A record in net is one accepted transport frame; Drop models +:busy, :disconnected, or loss. A set retains frames for arbitrary duplicate +delivery, and independent frames can be delivered out of order. + +The receiver's exact state is a single row toggled by each origin write. This +gives both creation and deletion (and detects zombie rows). Snapshot chunks and +their commit are abstracted as one *committed* snapshot here; the separate +SnapshotAssembly model checks the non-atomic transfer and install boundary. + +The hello action models a sender-to-receiver replica_hello. One valid old hello +may already be in flight initially. When its probe has reached the sender, the +same action abstracts completion of the two-way peer_connect/hello exchange. +An old hello can restore the receiver's route without proving the sender saw a +new probe; this is an important recovery race. +*) + +CONSTANTS MaxSeq, MaxMessages, MaxProbe, MaxPid, MaxEpoch + +ASSUME /\ MaxSeq >= 2 + /\ MaxMessages >= 1 + /\ MaxProbe >= 1 + /\ MaxPid >= 1 + /\ MaxEpoch >= 1 + +NoToken == <<0, 0>> +Token(pid, probe) == <> +ValueAt(n) == n % 2 = 1 + +Frame(kind, n, aux, tok, probe, pid, gen, epoch) == + [kind |-> kind, n |-> n, aux |-> aux, tok |-> tok, + probe |-> probe, pid |-> pid, gen |-> gen, epoch |-> epoch] + +VARIABLES s, net +vars == <> + +Init == + /\ s = [phase |-> "faulting", + head |-> 0, floor |-> 1, cursor |-> 0, view |-> FALSE, + connected |-> TRUE, + probe |-> 0, probePending |-> FALSE, + seenProbe |-> 0, seenPid |-> 1, + receiverPid |-> 1, receiverGen |-> 1, receiverEpoch |-> 1, + knownPid |-> 1, knownGen |-> 1, knownEpoch |-> 1, + sendToken |-> Token(1, 0), receiveToken |-> NoToken, + pending |-> FALSE, + ackQueued |-> FALSE, ackCursor |-> 0, ackEpoch |-> 1, + helloDue |-> FALSE, helloRetryReady |-> TRUE, + snapshotBlocked |-> FALSE, + lastSnapshotJustified |-> TRUE, + lastHelloJustified |-> TRUE, + ackEvidence |-> [head |-> 0, tok |-> NoToken, + pid |-> 0, gen |-> 0, epoch |-> 0]] + /\ net \in {{}, + {Frame("hello", 0, 0, NoToken, 0, 1, 1, 1)}} + +(* retain_pending_replica_head sets last_sent to nil on a newer head. *) +Write == + /\ s.phase = "faulting" + /\ s.head < MaxSeq + /\ s' = [s EXCEPT !.head = @ + 1, + !.pending = TRUE, + !.snapshotBlocked = FALSE] + /\ UNCHANGED net + +(* Data.prune_replica_oplog has a shard-wide cap. Writes to other streams may + prune this stream's already-applied records without changing its head. *) +Prune == + /\ s.phase = "faulting" + /\ s.floor < s.head + 1 + /\ s' = [s EXCEPT !.floor = @ + 1] + /\ UNCHANGED net + +(* expire_replica_peer deletes the cursor, rows, receive token, and queued ACKs, + but the sender can still think its old ACK retired the head. *) +ReceiverExpire == + /\ s.phase = "faulting" + /\ s.connected + /\ s.probe < MaxProbe + /\ s' = [s EXCEPT !.connected = FALSE, + !.cursor = 0, !.view = FALSE, + !.probe = @ + 1, + !.probePending = TRUE, + !.receiveToken = NoToken, + !.ackQueued = FALSE] + /\ UNCHANGED net + +ReceiverRestart == + /\ s.phase = "faulting" + /\ s.receiverPid < MaxPid + /\ s.probe < MaxProbe + /\ s' = [s EXCEPT !.connected = FALSE, + !.cursor = 0, !.view = FALSE, + !.probe = @ + 1, + !.probePending = TRUE, + !.receiverPid = @ + 1, + !.receiverGen = @ + 1, + !.receiveToken = NoToken, + !.ackQueued = FALSE] + /\ UNCHANGED net + +(* A named-cluster leave/rejoin gives the receiver a new local epoch. *) +ReceiverRejoin == + /\ s.phase = "faulting" + /\ s.receiverEpoch < MaxEpoch + /\ s.probe < MaxProbe + /\ s' = [s EXCEPT !.connected = FALSE, + !.cursor = 0, !.view = FALSE, + !.probe = @ + 1, + !.probePending = TRUE, + !.receiverEpoch = @ + 1, + !.receiveToken = NoToken, + !.ackQueued = FALSE] + /\ UNCHANGED net + +ProbeFrame == + Frame("probe", 0, 0, NoToken, s.probe, + s.receiverPid, s.receiverGen, s.receiverEpoch) + +ProbeAllowed == ~s.connected \/ s.probePending + +SendProbe == + /\ ProbeAllowed + /\ Cardinality(net) < MaxMessages + /\ net' = net \union {ProbeFrame} + /\ UNCHANGED s + +FreshProbe(m) == + m.pid # s.seenPid \/ m.probe > s.seenProbe + +(* peer_connect: a new probe from the *same* shard PID invalidates an ACKed + stream by rotating replica_send_tokens and requeuing every current head. + A duplicate never does. A new PID is requeued when its hello installs. *) +DeliverProbe(m) == + /\ m \in net + /\ m.kind = "probe" + /\ net' = (net \ {m}) \union + (IF m.pid = s.seenPid /\ m.probe < s.seenProbe + THEN {} + ELSE {Frame("probe_ack", 0, + IF m.pid = s.knownPid /\ m.gen = s.knownGen /\ + m.epoch = s.knownEpoch + THEN 1 ELSE 0, NoToken, m.probe, + 0, 0, 0)}) + /\ IF m.pid = s.seenPid /\ m.probe < s.seenProbe + THEN UNCHANGED s + ELSE IF FreshProbe(m) + THEN s' = [s EXCEPT !.seenPid = m.pid, + !.seenProbe = m.probe, + !.helloDue = TRUE, + !.sendToken = + IF m.pid = s.knownPid + THEN Token(m.pid, m.probe) + ELSE @, + !.pending = + IF m.pid = s.knownPid + THEN s.head > 0 ELSE @] + ELSE s' = [s EXCEPT !.helloDue = + @ \/ (s.helloRetryReady /\ + ~(m.pid = s.knownPid /\ + m.gen = s.knownGen /\ + m.epoch = s.knownEpoch))] + +(* peer_connect_ack echoes the probe epoch and says whether the sender route + matched the receiver PID after any required reseed. A mismatched PID must + complete its reverse hello before a later ready ACK can retire the probe. *) +DeliverProbeAck(m) == + /\ m \in net + /\ m.kind = "probe_ack" + /\ net' = net \ {m} + /\ IF m.probe = s.probe /\ m.aux = 1 /\ s.connected + THEN s' = [s EXCEPT !.probePending = FALSE] + ELSE UNCHANGED s + +(* The retry timer represents the implementation's 5-second discovery hello + gate. A duplicate probe with an exact sender view needs only its ACK. *) +HelloRetryTimer == + /\ (~s.connected \/ s.probePending) + /\ ~s.helloRetryReady + /\ s' = [s EXCEPT !.helloRetryReady = TRUE] + /\ UNCHANGED net + +SendHello == + /\ s.helloDue + /\ Cardinality(net) < MaxMessages + /\ net' = net \union + {Frame("hello", 0, 0, NoToken, s.seenProbe, + s.seenPid, s.receiverGen, s.receiverEpoch)} + /\ s' = [s EXCEPT !.helloDue = FALSE, + !.helloRetryReady = FALSE, + !.lastHelloJustified = s.helloDue] + +(* A valid old hello can restore the receiver route even after its PID changes: + the wire hello has no receiver PID fence. Only the current two-way handshake + updates the sender's receiver PID/generation/epoch and requeues heads. *) +CurrentHandshake(m) == + m.pid = s.receiverPid /\ m.probe = s.probe /\ + s.seenPid = s.receiverPid /\ s.seenProbe = s.probe + +DeliverHello(m) == + /\ m \in net + /\ m.kind = "hello" + /\ net' = net \ {m} + /\ s' = [s EXCEPT !.connected = TRUE, + !.knownPid = + IF CurrentHandshake(m) + THEN s.receiverPid ELSE @, + !.knownGen = + IF CurrentHandshake(m) + THEN s.receiverGen ELSE @, + !.knownEpoch = + IF CurrentHandshake(m) + THEN s.receiverEpoch ELSE @, + !.pending = + IF CurrentHandshake(m) /\ s.head > 0 /\ + (s.knownPid # s.receiverPid \/ + s.knownGen # s.receiverGen \/ + s.knownEpoch # s.receiverEpoch) + THEN TRUE ELSE @] + +HeadFrame == + Frame("head", s.head, s.floor, s.sendToken, 0, 0, 0, 0) + +(* retry_pending_replica_heads enumerates only unacknowledged streams. *) +SendHead == + /\ s.pending + /\ s.head > 0 + /\ Cardinality(net) < MaxMessages + /\ net' = net \union {HeadFrame} + /\ UNCHANGED s + +NeedFrame(next, advertised) == + Frame("need", next, advertised, NoToken, 0, 0, 0, 0) + +DeliverHead(m) == + /\ m \in net + /\ m.kind = "head" + /\ net' = (net \ {m}) \union + (IF s.connected /\ m.n > s.cursor + THEN {NeedFrame(s.cursor + 1, m.n)} ELSE {}) + /\ IF s.connected + THEN s' = [s EXCEPT !.receiveToken = m.tok, + !.ackQueued = + IF m.n <= s.cursor THEN TRUE ELSE @, + !.ackCursor = + IF m.n <= s.cursor THEN s.cursor ELSE @, + !.ackEpoch = + IF m.n <= s.cursor THEN s.receiverEpoch ELSE @] + ELSE UNCHANGED s + +(* send_replica_repairs requires a pending entry with the exact advertised + head. A stale need cannot trigger delta or snapshot. The wire need has no + receiver PID, generation, token, or epoch; do not invent those guards. *) +ValidNeed(m) == + s.pending /\ m.aux = s.head /\ m.n <= s.head + +DeliverNeed(m) == + /\ m \in net + /\ m.kind = "need" + /\ net' = (net \ {m}) \union + (IF ValidNeed(m) /\ m.n >= s.floor + THEN {Frame("delta", m.n, s.head, NoToken, 0, 0, 0, 0)} + ELSE IF ValidNeed(m) /\ ~s.snapshotBlocked + THEN {Frame("snapshot", s.head, 0, NoToken, 0, 0, 0, 0)} + ELSE {}) + /\ IF ValidNeed(m) /\ m.n < s.floor /\ ~s.snapshotBlocked + THEN s' = [s EXCEPT !.snapshotBlocked = TRUE, + !.lastSnapshotJustified = + s.pending /\ m.aux = s.head /\ + m.n <= s.head /\ m.n < s.floor] + ELSE UNCHANGED s + +(* The 5-second retry bound is a timer action, not an unbounded resend per + repeated need. The pending-head condition is checked again at send time. *) +SnapshotRetryTimer == + /\ s.snapshotBlocked + /\ s.pending + /\ s' = [s EXCEPT !.snapshotBlocked = FALSE] + /\ UNCHANGED net + +DeliverDelta(m) == + /\ m \in net + /\ m.kind = "delta" + /\ net' = (net \ {m}) \union + (IF s.connected /\ m.n > s.cursor + 1 + THEN {NeedFrame(s.cursor + 1, m.aux)} ELSE {}) + /\ IF s.connected /\ m.n = s.cursor + 1 + THEN s' = [s EXCEPT !.cursor = m.aux, + !.view = ValueAt(m.aux), + !.ackQueued = TRUE, + !.ackCursor = m.aux, + !.ackEpoch = s.receiverEpoch] + ELSE IF s.connected /\ m.n <= s.cursor + THEN s' = [s EXCEPT !.ackQueued = TRUE, + !.ackCursor = s.cursor, + !.ackEpoch = s.receiverEpoch] + ELSE UNCHANGED s + +(* Only a complete exact snapshot is represented here. *) +DeliverSnapshot(m) == + /\ m \in net + /\ m.kind = "snapshot" + /\ net' = net \ {m} + /\ IF s.connected /\ m.n > s.cursor + THEN s' = [s EXCEPT !.cursor = m.n, + !.view = ValueAt(m.n), + !.ackQueued = TRUE, + !.ackCursor = m.n, + !.ackEpoch = s.receiverEpoch] + ELSE UNCHANGED s + +(* flush_replica_acks reads the latest receive token when it emits a batch, + rather than freezing the token when the ACK was first queued. *) +SendAck == + /\ s.ackQueued + /\ Cardinality(net) < MaxMessages + /\ net' = net \union + {Frame("ack", s.ackCursor, 0, s.receiveToken, 0, + s.receiverPid, s.receiverGen, s.ackEpoch)} + /\ s' = [s EXCEPT !.ackQueued = FALSE] + +(* handle_replica_message(:applied) and acknowledge_replica_cursor. *) +ValidAck(m) == + /\ s.pending + /\ m.pid = s.knownPid + /\ m.gen = s.knownGen + /\ m.tok = s.sendToken + /\ m.epoch = s.knownEpoch + /\ m.n >= s.head + +DeliverAck(m) == + /\ m \in net + /\ m.kind = "ack" + /\ net' = net \ {m} + /\ IF ValidAck(m) + THEN s' = [s EXCEPT !.pending = FALSE, + !.snapshotBlocked = FALSE, + !.ackEvidence = + [head |-> s.head, tok |-> m.tok, + pid |-> m.pid, gen |-> m.gen, + epoch |-> m.epoch]] + ELSE UNCHANGED s + +Drop(m) == + /\ s.phase = "faulting" + /\ m \in net + /\ net' = net \ {m} + /\ UNCHANGED s + +Heal == + /\ s.phase = "faulting" + /\ s' = [s EXCEPT !.phase = "healed"] + /\ UNCHANGED net + +(* One successful anti-entropy round after the transport heals. The wire + actions above still explore arbitrary loss/reordering for safety. This + action abstracts a fair completion of the same probe, hello, head, repair, + and ACK handlers; it is enabled only while their obligations exist. *) +Repair == + /\ s.phase = "healed" + /\ UNCHANGED net + /\ \/ /\ ProbeAllowed + /\ s' = [s EXCEPT !.connected = TRUE, + !.probePending = FALSE, + !.seenPid = s.receiverPid, + !.seenProbe = s.probe, + !.knownPid = s.receiverPid, + !.knownGen = s.receiverGen, + !.knownEpoch = s.receiverEpoch, + !.sendToken = + IF s.receiverPid = s.knownPid + THEN Token(s.receiverPid, s.probe) + ELSE @, + !.pending = s.head > 0] + \/ /\ ~ProbeAllowed + /\ s.connected + /\ s.pending + /\ s.cursor < s.head + /\ s' = [s EXCEPT !.cursor = s.head, + !.view = ValueAt(s.head), + !.receiveToken = s.sendToken, + !.ackQueued = TRUE, + !.ackCursor = s.head, + !.ackEpoch = s.receiverEpoch] + \/ /\ ~ProbeAllowed + /\ s.connected + /\ s.pending + /\ s.cursor >= s.head + /\ s.knownPid = s.receiverPid + /\ s.knownGen = s.receiverGen + /\ s.knownEpoch = s.receiverEpoch + /\ s' = [s EXCEPT !.pending = FALSE, + !.ackQueued = FALSE, + !.snapshotBlocked = FALSE, + !.ackEvidence = + [head |-> s.head, tok |-> s.sendToken, + pid |-> s.receiverPid, gen |-> s.receiverGen, + epoch |-> s.receiverEpoch]] + +Next == + \/ Write + \/ Prune + \/ ReceiverExpire + \/ ReceiverRestart + \/ ReceiverRejoin + \/ SendProbe + \/ \E m \in net : DeliverProbe(m) + \/ \E m \in net : DeliverProbeAck(m) + \/ HelloRetryTimer + \/ SendHello + \/ \E m \in net : DeliverHello(m) + \/ SendHead + \/ \E m \in net : DeliverHead(m) + \/ \E m \in net : DeliverNeed(m) + \/ SnapshotRetryTimer + \/ \E m \in net : DeliverDelta(m) + \/ \E m \in net : DeliverSnapshot(m) + \/ SendAck + \/ \E m \in net : DeliverAck(m) + \/ \E m \in net : Drop(m) + \/ Heal + \/ Repair + +TypeOK == + /\ s.phase \in {"faulting", "healed"} + /\ s.head \in 0..MaxSeq + /\ s.floor \in 1..(s.head + 1) + /\ s.cursor \in 0..MaxSeq + /\ s.view \in BOOLEAN + /\ s.connected \in BOOLEAN + /\ s.probe \in 0..MaxProbe + /\ s.probePending \in BOOLEAN + /\ s.seenProbe \in 0..MaxProbe + /\ s.seenPid \in 1..MaxPid + /\ s.receiverPid \in 1..MaxPid + /\ s.receiverGen \in 1..MaxPid + /\ s.receiverEpoch \in 1..MaxEpoch + /\ s.knownPid \in 1..MaxPid + /\ s.knownGen \in 1..MaxPid + /\ s.knownEpoch \in 1..MaxEpoch + /\ s.pending \in BOOLEAN + /\ s.ackQueued \in BOOLEAN + /\ s.ackCursor \in 0..MaxSeq + /\ s.ackEpoch \in 1..MaxEpoch + /\ s.helloDue \in BOOLEAN + /\ s.helloRetryReady \in BOOLEAN + /\ s.snapshotBlocked \in BOOLEAN + /\ Cardinality(net) <= MaxMessages + +(* Every visible row is the exact prefix at the receiver's committed cursor. + In particular, expiry and rejoin cannot leave a zombie row. *) +ExactCommittedPrefix == + /\ s.cursor <= s.head + /\ s.view = ValueAt(s.cursor) + +(* A quiet nonempty stream has evidence from the current sender token and + installed receiver identity/epoch. An old ACK can never quiet a reseed. *) +QuietHasCurrentAck == + (~s.pending /\ s.head > 0) => + /\ s.ackEvidence.head = s.head + /\ s.ackEvidence.tok = s.sendToken + /\ s.ackEvidence.pid = s.knownPid + /\ s.ackEvidence.gen = s.knownGen + /\ s.ackEvidence.epoch = s.knownEpoch + +NoUnjustifiedFullSend == + s.lastSnapshotJustified /\ s.lastHelloJustified + +Converged == + /\ s.connected + /\ s.cursor = s.head + /\ s.view = ValueAt(s.head) + /\ ~s.pending + /\ ~s.probePending + +HealedConvergence == s.phase = "healed" ~> Converged + +Spec == + /\ Init + /\ [][Next]_vars + /\ WF_vars(Repair) + +============================================================================= diff --git a/test/formal/ReplicaAckEpoch.cfg b/test/formal/ReplicaAckEpoch.cfg new file mode 100644 index 0000000..9e9863f --- /dev/null +++ b/test/formal/ReplicaAckEpoch.cfg @@ -0,0 +1,18 @@ +SPECIFICATION Spec +CHECK_DEADLOCK FALSE + +CONSTANTS + MaxSeq = 2 + MaxMessages = 1 + MaxProbe = 1 + MaxPid = 1 + MaxEpoch = 2 + +INVARIANTS + TypeOK + ExactCommittedPrefix + QuietHasCurrentAck + NoUnjustifiedFullSend + +PROPERTY + HealedConvergence diff --git a/test/formal/ReplicaAckExtended.cfg b/test/formal/ReplicaAckExtended.cfg new file mode 100644 index 0000000..0640699 --- /dev/null +++ b/test/formal/ReplicaAckExtended.cfg @@ -0,0 +1,15 @@ +SPECIFICATION Spec +CHECK_DEADLOCK FALSE + +CONSTANTS + MaxSeq = 2 + MaxMessages = 2 + MaxProbe = 2 + MaxPid = 2 + MaxEpoch = 2 + +INVARIANTS + TypeOK + ExactCommittedPrefix + QuietHasCurrentAck + NoUnjustifiedFullSend diff --git a/test/formal/ReplicaAckReincarnation.cfg b/test/formal/ReplicaAckReincarnation.cfg new file mode 100644 index 0000000..3032711 --- /dev/null +++ b/test/formal/ReplicaAckReincarnation.cfg @@ -0,0 +1,18 @@ +SPECIFICATION Spec +CHECK_DEADLOCK FALSE + +CONSTANTS + MaxSeq = 2 + MaxMessages = 1 + MaxProbe = 1 + MaxPid = 2 + MaxEpoch = 1 + +INVARIANTS + TypeOK + ExactCommittedPrefix + QuietHasCurrentAck + NoUnjustifiedFullSend + +PROPERTY + HealedConvergence diff --git a/test/formal/check.sh b/test/formal/check.sh index 1f5b1c5..605aba4 100755 --- a/test/formal/check.sh +++ b/test/formal/check.sh @@ -7,7 +7,7 @@ if [[ -z "${TLA_JAR:-}" ]]; then fi repo_root="$(cd "$(dirname "${BASH_SOURCE[0]}")/../.." && pwd)" -metadir="${repo_root}/tmp/tlc" +metadir="${TLA_METADIR:-${repo_root}/tmp/tlc}" config="${TLA_CONFIG:-${repo_root}/test/formal/GroupAntiEntropy.cfg}" spec="${TLA_SPEC:-${repo_root}/test/formal/GroupAntiEntropy.tla}" mkdir -p "${metadir}" diff --git a/test/formal/check_matrix.sh b/test/formal/check_matrix.sh index 68419de..dd37e68 100755 --- a/test/formal/check_matrix.sh +++ b/test/formal/check_matrix.sh @@ -14,11 +14,16 @@ run_check() { } run_check GroupAntiEntropy GroupAntiEntropy +run_check ReplicaAck ReplicaAck +run_check ReplicaAck ReplicaAckReincarnation +run_check ReplicaAck ReplicaAckEpoch +elixir "${script_dir}/check_replica_ack_mutations.exs" elixir "${script_dir}/check_snapshot_commit.exs" run_check PeerEviction PeerEviction run_check AuthorityProjection AuthorityProjection run_check AuthorityHint AuthorityHint if [[ "${TLA_EXTENDED:-0}" == "1" ]]; then + run_check ReplicaAck ReplicaAckExtended run_check GroupAntiEntropy GroupAntiEntropyExtended fi diff --git a/test/formal/check_replica_ack_mutations.exs b/test/formal/check_replica_ack_mutations.exs new file mode 100644 index 0000000..a1b3a8d --- /dev/null +++ b/test/formal/check_replica_ack_mutations.exs @@ -0,0 +1,83 @@ +defmodule Group.Formal.ReplicaAckQualification do + @moduledoc false + + @mutations [ + {"stale ACK token", " /\\ m.tok = s.sendToken\n", "", "INVARIANT", + "QuietHasCurrentAck", 12, "Invariant QuietHasCurrentAck is violated"}, + {"stale need full send", " s.pending /\\ m.aux = s.head /\\ m.n <= s.head\n", + " m.aux = s.head /\\ m.n <= s.head\n", "INVARIANT", "NoUnjustifiedFullSend", 12, + "Invariant NoUnjustifiedFullSend is violated"}, + {"gratuitous full hello", " /\\ s.helloDue\n", "", "INVARIANT", + "NoUnjustifiedFullSend", 12, "Invariant NoUnjustifiedFullSend is violated"}, + {"lost recovery probe after a stale hello", + "ProbeAllowed == ~s.connected \\/ s.probePending\n", + "ProbeAllowed == ~s.connected\n", "PROPERTY", "HealedConvergence", 13, + "Temporal property HealedConvergence was violated"}, + {"premature ACK for an uninstalled receiver PID", + "m.probe = s.probe /\\ m.aux = 1 /\\ s.connected", + "m.probe = s.probe /\\ s.connected", "PROPERTY", "HealedConvergence", 13, + "Temporal property HealedConvergence was violated"}, + {"premature ACK for stale receiver authority", + "m.epoch = s.knownEpoch\n THEN 1 ELSE 0", + "TRUE\n THEN 1 ELSE 0", "PROPERTY", "HealedConvergence", 13, + "Temporal property HealedConvergence was violated"} + ] + + def run do + directory = __DIR__ + source = File.read!(Path.join(directory, "ReplicaAck.tla")) + check = Path.join(directory, "check.sh") + + Enum.each(@mutations, fn {label, old, replacement, check_kind, check_name, expected_status, + violation} -> + unless length(:binary.matches(source, old)) == 1 do + raise "#{label} mutation must match exactly one guard" + end + + artifacts = + Path.join( + System.tmp_dir!(), + "group-replica-ack-#{System.pid()}-#{System.unique_integer([:positive])}" + ) + + File.mkdir_p!(artifacts) + spec = Path.join(artifacts, "ReplicaAck.tla") + config = Path.join(artifacts, "ReplicaAck.cfg") + File.write!(spec, String.replace(source, old, replacement)) + + File.write!( + config, + """ + SPECIFICATION Spec + CHECK_DEADLOCK FALSE + CONSTANTS + MaxSeq = 2 + MaxMessages = 2 + MaxProbe = 2 + MaxPid = 2 + MaxEpoch = 2 + #{check_kind} #{check_name} + """ + ) + + {output, status} = + System.cmd("bash", [check], + env: [ + {"TLA_SPEC", spec}, + {"TLA_CONFIG", config}, + {"TLA_METADIR", Path.join(artifacts, "tlc")} + ], + stderr_to_stdout: true + ) + + unless status == expected_status and String.contains?(output, violation) do + raise "#{label} mutation was not rejected by #{check_name} (TLC status #{status}):\n#{output}" + end + + IO.puts("#{check_name} rejects #{label}") + File.rm_rf!(artifacts) + end) + end +end + +Group.Formal.ReplicaAckQualification.run() diff --git a/test/group_test.exs b/test/group_test.exs index e848125..3ac3d71 100644 --- a/test/group_test.exs +++ b/test/group_test.exs @@ -1574,7 +1574,8 @@ defmodule GroupTest.Fairness do send( shard, - {:group_replica_frame, remote_node, {:heads, Group.Replica.WireProtocol.version(), []}} + {:group_replica_frame, remote_node, + {:heads, Group.Replica.WireProtocol.version(), make_ref(), []}} ) for _ <- 1..10_000 do diff --git a/test/jepsen/README.md b/test/jepsen/README.md index 53a1b49..1b105fd 100644 --- a/test/jepsen/README.md +++ b/test/jepsen/README.md @@ -15,7 +15,11 @@ prefix Jepsen: - creates a three-node registry conflict and requires every loser to die; - partitions the full Erlang mesh, or only selected directed replica lanes; - exercises isolation, all-way partition, and asymmetric one-way loss; -- resets transport sessions and kills/restarts complete BEAM nodes; and +- resets transport sessions and kills/restarts complete BEAM nodes; +- expires one receiver's replica lease while its sender remains connected, + requiring rediscovery of already acknowledged streams; +- qualifies recovery when that first probe is lost and an older hello restores + the receiver route before the next anti-entropy tick; and - uses a 16-entry oplog and 1 KiB snapshot target so repair crosses pruning and multi-chunk exact-snapshot paths. @@ -43,9 +47,11 @@ terminal snapshots. The independent checker requires: indexes inside every shard; - every admitted receiver stream at its independently captured origin head, including streams with no writes; -- no staged partial snapshot and no retained data for a retired origin; +- no staged partial snapshot, unacknowledged head, or retained data for a + retired origin; - coverage of delta batches, snapshot fallback, multi-chunk assembly, and - registry conflict termination; + registry conflict termination, plus an injected one-sided replica lease + expiry; - two identical quiescent observations per node; and - every acknowledged Group operation below the configured latency ceiling. diff --git a/test/jepsen/cursor_capture.exs b/test/jepsen/cursor_capture.exs index 7715f77..68531f8 100644 --- a/test/jepsen/cursor_capture.exs +++ b/test/jepsen/cursor_capture.exs @@ -45,7 +45,86 @@ try do end) end) + :ok = + await.(fn -> + Enum.all?(tl(nodes), fn target -> + state = rpc.(origin, :peer_head_state, [shard, target]) + state.connected? and state.pending_heads == 0 and is_reference(state.send_token) + end) + end) + healthy = capture.() + old_send_token = rpc.(origin, :peer_head_state, [shard, receiver]).send_token + :ok = :erpc.call(receiver, TestCluster, :expire_replica_lane, [:jepsen_group, shard, origin]) + + :ok = + await.(fn -> + state = rpc.(origin, :peer_head_state, [shard, receiver]) + + state.connected? and state.pending_heads == 0 and + state.send_token != old_send_token and + :erpc.call(receiver, Group.Replica.Data, :replica_cursor, [ + :jepsen_group, + shard, + stream + ]) == 1 + end) + + lease_recovered = capture.() + + # Withhold the first recovery probe, then let an older exact hello restore + # the route. The receiver must keep probing until the sender ACKs that probe + # epoch; otherwise an already-ACKed stream can remain absent forever. + source_shard = + :erpc.call(origin, Process, :whereis, [Group.Replica.shard_name(:jepsen_group, shard)]) + + {receiver_shard, probe_epoch, tick_ref} = + :erpc.call(receiver, TestCluster, :expire_replica_lane_without_probe, [ + :jepsen_group, + shard, + origin + ]) + + {generation, revision, epochs} = + :erpc.call(origin, Group.Replica.Data, :local_replica_authority, [:jepsen_group]) + + source_state = :erpc.call(origin, :sys, :get_state, [source_shard]) + transport = source_state.replica_transport + + descriptor = + :erpc.call(origin, transport, :descriptor, [ + :jepsen_group, + source_state.replica_transport_opts + ]) + + send( + receiver_shard, + {:replica_hello, source_shard, Group.Replica.WireProtocol.version(), generation, revision, + epochs, :erpc.call(origin, transport, :id, []), descriptor} + ) + + receiver_state = :erpc.call(receiver, :sys, :get_state, [receiver_shard]) + true = Map.get(receiver_state.remote_shards, origin) == source_shard + ^probe_epoch = Map.fetch!(receiver_state.pending_peer_probes, origin) + 0 = :erpc.call(receiver, Group.Replica.Data, :replica_cursor, [:jepsen_group, shard, stream]) + + send(receiver_shard, {:group_replica_anti_entropy, tick_ref}) + + :ok = + await.(fn -> + state = rpc.(origin, :peer_head_state, [shard, receiver]) + receiver_state = :erpc.call(receiver, :sys, :get_state, [receiver_shard]) + + state.connected? and state.pending_heads == 0 and + not Map.has_key?(receiver_state.pending_peer_probes, origin) and + :erpc.call(receiver, Group.Replica.Data, :replica_cursor, [ + :jepsen_group, + shard, + stream + ]) == 1 + end) + + stale_hello_recovered = capture.() :ok = rpc.(receiver, :freeze, []) corruptions = @@ -116,6 +195,8 @@ try do Group.Jepsen.EDN.encode(%{ pristine: pristine, healthy: healthy, + lease_recovered: lease_recovered, + stale_hello_recovered: stale_hello_recovered, zero: zero, closed: closed, restarted: restarted, diff --git a/test/jepsen/node.exs b/test/jepsen/node.exs index 86989df..912480c 100644 --- a/test/jepsen/node.exs +++ b/test/jepsen/node.exs @@ -401,6 +401,38 @@ defmodule Group.Jepsen.Transport.Control do :ok end + def expire_peer(target_node) do + shards = + Group.get_config(:jepsen_group).num_shards + |> then(&Range.new(0, &1 - 1)) + |> Enum.filter(fn shard -> + replica = Group.Replica.shard_name(:jepsen_group, shard) + state = :sys.get_state(replica) + + if Map.has_key?(state.remote_shards, target_node) do + now = System.monotonic_time(:millisecond) + + :sys.replace_state(replica, fn current -> + expired_at = now - current.replicated_peer_lease_timeout - 1 + %{current | peer_last_seen: Map.put(current.peer_last_seen, target_node, expired_at)} + end) + + state = :sys.get_state(replica) + send(Process.whereis(replica), {:group_replica_anti_entropy, state.anti_entropy_ref}) + _state = :sys.get_state(replica) + true + else + false + end + end) + + if shards != [] do + Stats.increment_persistent(:replica_lease_expiry_injected, length(shards)) + end + + shards + end + defp maybe_disconnect(target_node) do if profile() == :tcp do Group.TestTCPTransport.disconnect_peer(:jepsen_group, target_node) @@ -477,7 +509,10 @@ defmodule Group.Jepsen.ConflictEvidence do def handle_call({:record, event}, _from, state) do event = Map.put(event, :sequence, state.sequence + 1) encoded = event |> :erlang.term_to_binary() |> Base.encode64() - :ok = File.write(state.path, encoded <> "\n", [:append, :sync]) + # Closing each append preserves evidence across BEAM/container restarts. + # A per-event fsync can stall the recorder long enough to kill an owner + # during a burst of cluster cleanup evidence. + :ok = File.write(state.path, encoded <> "\n", [:append]) {:reply, event.sequence, %{state | events: [event | state.events], sequence: event.sequence}} end @@ -517,18 +552,21 @@ defmodule Group.Jepsen.Owner do revision: revision }) - case safe_group_call(state.token, fn -> - state.api.register(:jepsen_group, registry_key(key), meta, cluster_opts(cluster)) - end) do + {result, latency_us} = + timed_group_call(state.token, fn -> + state.api.register(:jepsen_group, registry_key(key), meta, cluster_opts(cluster)) + end) + + case result do :ok -> registration_result(attempt, :ok) entry = %{cluster: cluster, key: key, revision: revision} state = put_in(state.registrations[{cluster, key}], entry) - {:reply, {:ok, snapshot(state)}, state} + {:reply, {:ok, snapshot(state), latency_us}, state} {:error, reason} -> registration_result(attempt, if(reason == :taken, do: :fail, else: :unknown)) - {:reply, {:error, reason, snapshot(state)}, state} + {:reply, {:error, reason, snapshot(state), latency_us}, state} end end @@ -536,9 +574,12 @@ defmodule Group.Jepsen.Owner do owner_key = {cluster, key} if Map.has_key?(state.registrations, owner_key) do - case safe_group_call(state.token, fn -> - state.api.unregister(:jepsen_group, registry_key(key), cluster_opts(cluster)) - end) do + {result, latency_us} = + timed_group_call(state.token, fn -> + state.api.unregister(:jepsen_group, registry_key(key), cluster_opts(cluster)) + end) + + case result do :ok -> Group.Jepsen.ConflictEvidence.record(%{ kind: :unregister, @@ -548,29 +589,32 @@ defmodule Group.Jepsen.Owner do }) state = %{state | registrations: Map.delete(state.registrations, owner_key)} - {:reply, {:ok, snapshot(state)}, state} + {:reply, {:ok, snapshot(state), latency_us}, state} {:error, reason} -> - {:reply, {:error, reason, snapshot(state)}, state} + {:reply, {:error, reason, snapshot(state), latency_us}, state} end else - {:reply, {:error, :not_owned, snapshot(state)}, state} + {:reply, {:error, :not_owned, snapshot(state), 0}, state} end end def handle_call({:mutate, :join, cluster, key, revision}, _from, state) do meta = %{token: state.token, revision: revision} - case safe_group_call(state.token, fn -> - state.api.join(:jepsen_group, pg_key(key), meta, cluster_opts(cluster)) - end) do + {result, latency_us} = + timed_group_call(state.token, fn -> + state.api.join(:jepsen_group, pg_key(key), meta, cluster_opts(cluster)) + end) + + case result do :ok -> entry = %{cluster: cluster, key: key, revision: revision} state = put_in(state.memberships[{cluster, key}], entry) - {:reply, {:ok, snapshot(state)}, state} + {:reply, {:ok, snapshot(state), latency_us}, state} {:error, reason} -> - {:reply, {:error, reason, snapshot(state)}, state} + {:reply, {:error, reason, snapshot(state), latency_us}, state} end end @@ -578,18 +622,21 @@ defmodule Group.Jepsen.Owner do owner_key = {cluster, key} if Map.has_key?(state.memberships, owner_key) do - case safe_group_call(state.token, fn -> - state.api.leave(:jepsen_group, pg_key(key), cluster_opts(cluster)) - end) do + {result, latency_us} = + timed_group_call(state.token, fn -> + state.api.leave(:jepsen_group, pg_key(key), cluster_opts(cluster)) + end) + + case result do :ok -> state = %{state | memberships: Map.delete(state.memberships, owner_key)} - {:reply, {:ok, snapshot(state)}, state} + {:reply, {:ok, snapshot(state), latency_us}, state} {:error, reason} -> - {:reply, {:error, reason, snapshot(state)}, state} + {:reply, {:error, reason, snapshot(state), latency_us}, state} end else - {:reply, {:error, :not_owned, snapshot(state)}, state} + {:reply, {:error, :not_owned, snapshot(state), 0}, state} end end @@ -628,6 +675,12 @@ defmodule Group.Jepsen.Owner do defp sort_entries(entries), do: Enum.sort_by(entries, &{&1.cluster || "", &1.key}) + defp timed_group_call(token, fun) do + started_at = System.monotonic_time(:microsecond) + result = safe_group_call(token, fun) + {result, System.monotonic_time(:microsecond) - started_at} + end + defp safe_group_call(token, fun) do case fun.() do :ok -> @@ -701,16 +754,11 @@ defmodule Group.Jepsen.Driver do end def mutate(operation, logical_owner, cluster, key, revision) do - started_at = System.monotonic_time(:microsecond) - - response = - GenServer.call( - driver(logical_owner), - {:mutate, operation, logical_owner, cluster, key, revision}, - 10_000 - ) - - Map.put(response, :latency_us, System.monotonic_time(:microsecond) - started_at) + GenServer.call( + driver(logical_owner), + {:mutate, operation, logical_owner, cluster, key, revision}, + 10_000 + ) end def kill(logical_owner), do: GenServer.call(driver(logical_owner), {:kill, logical_owner}) @@ -765,12 +813,14 @@ defmodule Group.Jepsen.Driver do try do case GenServer.call(pid, {:mutate, operation, cluster, key, revision}, 8_000) do - {:ok, owner_state} -> - {:reply, %{status: :ok, owner: owner_state}, + {:ok, owner_state, latency_us} -> + {:reply, %{status: :ok, owner: owner_state, latency_us: latency_us}, put_owner_state(state, logical_owner, pid, owner_state)} - {:error, reason, owner_state} -> - response = Map.merge(failure_response(reason), %{owner: owner_state}) + {:error, reason, owner_state, latency_us} -> + response = + Map.merge(failure_response(reason), %{owner: owner_state, latency_us: latency_us}) + {:reply, response, put_owner_state(state, logical_owner, pid, owner_state)} end catch @@ -1048,17 +1098,33 @@ defmodule Group.Jepsen.Invariant do total + map_size(state.snapshot_transfers) end) + pending_head_count = + Enum.reduce(shards, 0, fn shard, total -> + state = :sys.get_state(Group.Replica.shard_name(:jepsen_group, shard)) + + total + + Enum.reduce(state.pending_replica_heads, 0, fn {_peer, streams}, count -> + count + map_size(streams) + end) + end) + oplog_entries = Enum.reduce(shards, 0, fn shard, total -> total + :ets.info(Data.replica_oplog_order_table(:jepsen_group, shard), :size) end) %{ - healthy: failures == [] and staging_count == 0, - errors: Enum.map(failures, & &1.message), + healthy: failures == [] and staging_count == 0 and pending_head_count == 0, + errors: + Enum.map(failures, & &1.message) ++ + if(pending_head_count > 0, + do: ["#{pending_head_count} unacknowledged replica heads"], + else: [] + ), failed_invariants: Enum.flat_map(failures, &List.wrap(&1.invariant)), injected_corruptions: injected_corruptions, snapshot_staging_count: staging_count, + pending_head_count: pending_head_count, oplog_entries: oplog_entries, oplog_max_entries_per_shard: config.replicated_oplog_max_entries, shard_mailbox_max: @@ -1654,6 +1720,10 @@ defmodule Group.Jepsen.Wire do :ok = Group.Jepsen.Transport.Control.reset(String.to_atom(target)) %{status: :ok} + ["transport", "expire-peer", target] -> + shards = Group.Jepsen.Transport.Control.expire_peer(String.to_atom(target)) + %{status: :ok, expired_shards: shards} + ["snapshot", key_count, clusters, retired] -> Group.Jepsen.Snapshot.capture( context.node_id, diff --git a/test/jepsen/src/group/jepsen/client.clj b/test/jepsen/src/group/jepsen/client.clj index e6de357..27b77b3 100644 --- a/test/jepsen/src/group/jepsen/client.clj +++ b/test/jepsen/src/group/jepsen/client.clj @@ -42,7 +42,7 @@ {:node node, :last-response response})))))))) (defn wait-listening! - ([node] (wait-listening! node 15000)) + ([node] (wait-listening! node 45000)) ([node timeout-ms] (let [deadline (+ (System/currentTimeMillis) timeout-ms)] (loop [] diff --git a/test/jepsen/src/group/jepsen/core.clj b/test/jepsen/src/group/jepsen/core.clj index b6a3f48..f8bc2cb 100644 --- a/test/jepsen/src/group/jepsen/core.clj +++ b/test/jepsen/src/group/jepsen/core.clj @@ -35,6 +35,8 @@ (gen/sleep fault-interval) {:type :info, :f :replica-reset} (gen/sleep fault-interval) + {:type :info, :f :replica-lease-expire} + (gen/sleep fault-interval) {:type :info, :f :kill-node} (gen/sleep fault-interval) {:type :info, :f :restart-node} @@ -72,6 +74,9 @@ (gen/sleep 0.25) (gen/nemesis {:type :info, :f :replica-partition-stop}) (gen/sleep 1) + (gen/log "Expiring n3 on n1 while n3 still considers n1 connected") + (gen/nemesis {:type :info, :f :replica-lease-expire, + :value {:receiver :n1, :source :n3, :required true}}) (gen/log "Restarting n2 after conflict resolution to prove durable qualification evidence") (gen/nemesis {:type :info, :f :kill-node, :value {:node :n2}}) (gen/nemesis {:type :info, :f :restart-node}) @@ -158,7 +163,10 @@ opts (assoc opts :clusters clusters :terminal-nodes terminal-nodes - :retired-nodes retired-nodes)] + :retired-nodes retired-nodes + :required-transport-events + (conj model/default-required-transport-events + :replica-lease-expiry-injected))] (merge tests/noop-test opts {:name (str "group lifecycle convergence (" (:transport opts) "/" diff --git a/test/jepsen/src/group/jepsen/model.clj b/test/jepsen/src/group/jepsen/model.clj index 8c72f8e..2848467 100644 --- a/test/jepsen/src/group/jepsen/model.clj +++ b/test/jepsen/src/group/jepsen/model.clj @@ -85,7 +85,7 @@ (defn stable-internal [snapshot] (select-keys (:internal snapshot) - [:healthy :errors :snapshot-staging-count :oplog-entries])) + [:healthy :errors :snapshot-staging-count :pending-head-count :oplog-entries])) (defn snapshot-fingerprint [test snapshot] {:owners (set (:owners snapshot)) diff --git a/test/jepsen/src/group/jepsen/nemesis.clj b/test/jepsen/src/group/jepsen/nemesis.clj index c0c3a2a..af2658c 100644 --- a/test/jepsen/src/group/jepsen/nemesis.clj +++ b/test/jepsen/src/group/jepsen/nemesis.clj @@ -99,7 +99,25 @@ (group-client/request! source ["transport" "reset" (str "group@" (name target))])) - (assoc op :type :info, :value {:source source, :target target})))) + (assoc op :type :info, :value {:source source, :target target})) + + :expire-peer + (let [nodes (vec (filter docker/running? (:nodes test))) + receiver (or (get-in op [:value :receiver]) + (when (seq nodes) (rand-nth nodes))) + sources (when receiver (vec (remove #(= receiver %) nodes))) + source (or (get-in op [:value :source]) + (when (seq sources) (rand-nth sources))) + result (when (and receiver source) + (group-client/request! + receiver + ["transport" "expire-peer" (str "group@" (name source))]))] + (when (and (get-in op [:value :required]) + (empty? (:expired-shards result))) + (throw (ex-info "required one-sided replica lease expiry was not injected" + {:receiver receiver, :source source, :result result}))) + (assoc op :type :info, :value {:receiver receiver, :source source, + :expired-shards (:expired-shards result)})))) (teardown! [_this test] (docker/heal-replica! (:nodes test)) @@ -165,7 +183,8 @@ {:replica-partition-start :start, :replica-partition-stop :stop, - :replica-reset :reset} + :replica-reset :reset, + :replica-lease-expire :expire-peer} (ReplicaNemesis. (atom nil)) {:kill-node :start, :restart-node :stop} diff --git a/test/jepsen/test/group/jepsen/cursor_capture_test.clj b/test/jepsen/test/group/jepsen/cursor_capture_test.clj index 813db2e..6fbc66e 100644 --- a/test/jepsen/test/group/jepsen/cursor_capture_test.clj +++ b/test/jepsen/test/group/jepsen/cursor_capture_test.clj @@ -13,7 +13,8 @@ (is (= 0 (:exit run)) (str (:out run) (:err run))) (when (zero? (:exit run)) (let [captures (edn/read-string (slurp output))] - (doseq [kind [:pristine :healthy :zero :closed :restarted :retired]] + (doseq [kind [:pristine :healthy :lease-recovered :stale-hello-recovered + :zero :closed :restarted :retired]] (is (empty? (streams/errors (get captures kind))) (str kind)) (is (every? #(true? (get-in % [:internal :healthy])) (vals (get captures kind))))) (doseq [snapshots (:corruptions captures)] diff --git a/test/mutation/run.exs b/test/mutation/run.exs index 115fff1..163df79 100644 --- a/test/mutation/run.exs +++ b/test/mutation/run.exs @@ -19,7 +19,7 @@ defmodule Group.MutationCampaign do " WireProtocol.stream_generation(stream_id) ==\n" <> " Data.remote_generation(state.name, source_node) and", faulty_source: " true and", - test: ["test/distributed_test.exs:5487"] + test: ["test/distributed_test.exs:5842"] }, %{ name: "accept_old_epoch", @@ -28,14 +28,14 @@ defmodule Group.MutationCampaign do " WireProtocol.stream_epoch(stream_id) ==\n" <> " Data.remote_cluster_epoch(state.name, source_node, cluster) and", faulty_source: " true and", - test: ["test/distributed_test.exs:4890"] + test: ["test/distributed_test.exs:5062"] }, %{ name: "advance_cursor_across_gap", file: "lib/group/replica.ex", correct_source: """ [{first_seq, _mutations} | _] when first_seq > cursor + 1 -> - request_replica_need(state, source_node, stream_id, cursor + 1) + request_replica_need(state, source_node, stream_id, cursor + 1, advertised_head) """, faulty_source: """ [{first_seq, _mutations} | _] when first_seq > cursor + 1 -> @@ -55,7 +55,7 @@ defmodule Group.MutationCampaign do advertised_head ) """, - test: ["test/distributed_test.exs:5564"] + test: ["test/replica_ack_test.exs:483"] }, %{ name: "registry_snapshot_is_additive", @@ -68,20 +68,14 @@ defmodule Group.MutationCampaign do acc = Enum.reduce([], acc, fn """, - test: [ - "test/replica_snapshot_distributed_test.exs:16", - "test/distributed_test.exs:4094" - ] + test: ["test/group_test.exs:2742"] }, %{ name: "pg_snapshot_is_additive", file: "lib/group/replica.ex", correct_source: " if Snapshot.member_pg?(staging_table, key, pid) do", faulty_source: " if Process.alive?(self()) do", - test: [ - "test/replica_snapshot_distributed_test.exs:16", - "test/distributed_test.exs:4094" - ] + test: ["test/distributed_test.exs:4152"] }, %{ name: "commit_incomplete_snapshot", @@ -92,7 +86,7 @@ defmodule Group.MutationCampaign do faulty_source: " if MapSet.size(transfer.received) >= 1 and chunk_count >= 1 and\n" <> " registry_count >= 0 and pg_count >= 0 do", - test: ["test/replica_snapshot_distributed_test.exs:223"] + test: ["test/replica_snapshot_distributed_test.exs:225"] }, %{ name: "commit_snapshot_without_terminal_manifest", @@ -122,7 +116,7 @@ defmodule Group.MutationCampaign do faulty_source: " MapSet.member?(transfer.received, chunk_index) and\n" <> " Process.alive?(self()) ->", - test: ["test/replica_snapshot_distributed_test.exs:385"] + test: ["test/replica_snapshot_distributed_test.exs:387"] }, %{ name: "retain_conflicting_snapshot_manifest", @@ -135,7 +129,7 @@ defmodule Group.MutationCampaign do {:ok, state, _conflicting_transfer} -> state """, - test: ["test/replica_snapshot_distributed_test.exs:191"] + test: ["test/replica_snapshot_distributed_test.exs:193"] }, %{ name: "commit_snapshot_after_source_changes_during_scan", @@ -161,7 +155,7 @@ defmodule Group.MutationCampaign do file: "lib/group/replica.ex", correct_source: " _event_buffer = Snapshot.finish_event_buffer(event_buffer)", faulty_source: " _event_buffer = event_buffer", - test: ["test/replica_snapshot_distributed_test.exs:223"] + test: ["test/replica_snapshot_distributed_test.exs:225"] }, %{ name: "allow_duplicate_snapshot_rows", @@ -173,7 +167,7 @@ defmodule Group.MutationCampaign do faulty_source: """ if :ets.insert(table, objects) and size_before >= 0 do """, - test: ["test/replica_snapshot_distributed_test.exs:385"] + test: ["test/replica_snapshot_distributed_test.exs:387"] }, %{ name: "do_not_supersede_partial_snapshot", @@ -187,7 +181,7 @@ defmodule Group.MutationCampaign do " %{snapshot_seq: existing_seq} when existing_seq < snapshot_seq ->\n" <> " _ = existing_seq\n" <> " {:ignore, state}", - test: ["test/replica_snapshot_distributed_test.exs:328"] + test: ["test/replica_snapshot_distributed_test.exs:330"] }, %{ name: "accept_stale_snapshot_authority", @@ -203,7 +197,7 @@ defmodule Group.MutationCampaign do snapshot_seq > Data.replica_cursor(state.name, state.shard_index, stream_id) end """, - test: ["test/replica_snapshot_distributed_test.exs:556"] + test: ["test/replica_snapshot_distributed_test.exs:558"] }, %{ name: "disable_snapshot_staging_expiry", @@ -224,7 +218,7 @@ defmodule Group.MutationCampaign do acc end """, - test: ["test/replica_snapshot_distributed_test.exs:482"] + test: ["test/replica_snapshot_distributed_test.exs:484"] }, %{ name: "disable_below_floor_snapshot", @@ -247,7 +241,7 @@ defmodule Group.MutationCampaign do " append_process_down_records(state, reason_by_pid, pending_reg, pending_pg)\n", faulty_source: " sequenced_downs =\n if false,\n do: append_process_down_records(state, reason_by_pid, pending_reg, pending_pg),\n else: []\n", - test: ["test/distributed_test.exs:4004"] + test: ["test/distributed_test.exs:4005"] }, %{ name: "do_not_exit_conflict_loser", @@ -267,7 +261,7 @@ defmodule Group.MutationCampaign do file: "lib/group/replica/data.ex", correct_source: " hint_generation == generation and\n", faulty_source: " false and hint_generation == generation and\n", - test: ["test/anti_entropy_fault_regression_test.exs:2199"] + test: ["test/anti_entropy_fault_regression_test.exs:2226"] }, %{ name: "heartbeat_does_not_fence_newer_generation", @@ -279,7 +273,7 @@ defmodule Group.MutationCampaign do " not is_nil(hint_generation) and\n" <> " WireProtocol.generation_newer?(generation, hint_generation) and\n" <> " Process.get(:fence_newer_generation, false) ->\n", - test: ["test/anti_entropy_fault_regression_test.exs:2362"] + test: ["test/anti_entropy_fault_regression_test.exs:2389"] }, %{ name: "drop_new_generation_authority_hint", @@ -292,7 +286,7 @@ defmodule Group.MutationCampaign do " # below are being updated. Existing materialized projections may remain\n" <> " # visible until exact-authority repair or bounded lease retirement.\n" <> " _ = {state.name, remote_node, generation, revision}", - test: ["test/anti_entropy_fault_regression_test.exs:2362"] + test: ["test/anti_entropy_fault_regression_test.exs:2389"] }, %{ name: "accept_authority_older_than_generation_hint", @@ -301,7 +295,7 @@ defmodule Group.MutationCampaign do faulty_source: " _ = hinted_stale?\n" <> " known_stale? or revision_stale?", - test: ["test/anti_entropy_fault_regression_test.exs:2362"] + test: ["test/anti_entropy_fault_regression_test.exs:2389"] }, %{ name: "install_lane_view_behind_generation_hint", @@ -310,7 +304,7 @@ defmodule Group.MutationCampaign do " remote_replica_authority_hint(state.name, remote_node) == {generation, observed} do", faulty_source: " elem(remote_replica_authority_hint(state.name, remote_node), 1) == observed do", - test: ["test/group_test.exs:3101"] + test: ["test/group_test.exs:3105"] }, %{ name: "install_incremental_after_newer_hint", @@ -321,7 +315,7 @@ defmodule Group.MutationCampaign do faulty_source: " Process.get(:ignore_incremental_authority_race, true) and\n" <> " is_tuple(remote_replica_authority_hint(name, remote_node))\n", - test: ["test/group_test.exs:3035"] + test: ["test/group_test.exs:3039"] }, %{ name: "accept_hint_without_exact_authority", @@ -332,7 +326,7 @@ defmodule Group.MutationCampaign do faulty_source: " (is_nil(hint_generation) or\n" <> " WireProtocol.generation_newer?(generation, hint_generation)) ->\n", - test: ["test/anti_entropy_fault_regression_test.exs:3458"] + test: ["test/anti_entropy_fault_regression_test.exs:3489"] }, %{ name: "admit_retired_lane_route_without_authority", @@ -345,7 +339,7 @@ defmodule Group.MutationCampaign do faulty_source: " state = put_remote_shard(state, remote_node, remote_pid)\n" <> " {:noreply, request_replica_authority(state, remote_node)}", - test: ["test/anti_entropy_fault_regression_test.exs:3458"] + test: ["test/anti_entropy_fault_regression_test.exs:3489"] }, %{ name: "do_not_restore_hint_lease_after_lane_restart", @@ -357,20 +351,20 @@ defmodule Group.MutationCampaign do " # crash in that window cannot strand the peer forever.\n" <> " {{{:remote_authority_hint, :\"$1\"}, :_, :_}, [], [:\"$1\"]}\n", faulty_source: " {{{:remote_view_info, shard, :\"$1\"}, :_, :_, :_}, [], [:\"$1\"]}\n", - test: ["test/group_test.exs:2974"] + test: ["test/group_test.exs:2978"] }, %{ name: "retain_retired_authority_repair", file: "lib/group/replica.ex", correct_source: - " is_nil(Data.remote_generation(state.name, remote_node)) and\n" <> - " is_nil(Data.remote_replica_authority_hint(state.name, remote_node)) ->\n" <> - " acc\n", + " is_nil(Data.remote_generation(state.name, remote_node)) and\n" <> + " is_nil(Data.remote_replica_authority_hint(state.name, remote_node)) ->\n" <> + " {acc_state, dirty}\n", faulty_source: - " is_nil(Data.remote_generation(state.name, remote_node)) and\n" <> - " is_nil(Data.remote_replica_authority_hint(state.name, remote_node)) ->\n" <> - " Map.put(acc, remote_node, last_activity)\n", - test: ["test/anti_entropy_fault_regression_test.exs:3458"] + " is_nil(Data.remote_generation(state.name, remote_node)) and\n" <> + " is_nil(Data.remote_replica_authority_hint(state.name, remote_node)) ->\n" <> + " {acc_state, Map.put(dirty, remote_node, last_activity)}\n", + test: ["test/anti_entropy_fault_regression_test.exs:3489"] }, %{ name: "skip_authority_fanout", @@ -385,7 +379,7 @@ defmodule Group.MutationCampaign do faulty_source: """ :ok """, - test: ["test/distributed_test.exs:5709"] + test: ["test/distributed_test.exs:6064"] }, %{ name: "wait_for_periodic_lane_probe_after_authority_fanout", @@ -397,11 +391,12 @@ defmodule Group.MutationCampaign do # retirement must not recreate an unleased peer. Once exact authority # reaches this lane, repeat shard-local discovery immediately instead # of waiting for the next anti-entropy probe. + state = advance_peer_probe_epoch(state, remote_node) + send_remote_shard_message( state, remote_node, - {:peer_connect, self(), state.shard_index, state.num_shards, - Data.my_clusters(state.name)} + peer_connect_message(state, remote_node) ) state @@ -409,10 +404,11 @@ defmodule Group.MutationCampaign do """, faulty_source: """ else + _ = advance_peer_probe_epoch(state, remote_node) state end """, - test: ["test/anti_entropy_fault_regression_test.exs:3963"] + test: ["test/anti_entropy_fault_regression_test.exs:3994"] }, %{ name: "assume_authority_fanout_reaches_late_lane", @@ -442,7 +438,7 @@ defmodule Group.MutationCampaign do state end """, - test: ["test/anti_entropy_fault_regression_test.exs:3963"] + test: ["test/anti_entropy_fault_regression_test.exs:3994"] }, %{ name: "skip_generation_purge", @@ -474,35 +470,29 @@ defmodule Group.MutationCampaign do defp maybe_purge_remote_generation(state, _remote_node, _old_generation, _generation), do: state """, - test: ["test/distributed_test.exs:5709"] + test: ["test/distributed_test.exs:6064"] }, %{ - name: "disable_periodic_heads", + name: "disable_pending_head_retry", file: "lib/group/replica.ex", - correct_source: """ - defp broadcast_replica_heads(state) do - peers = Map.keys(state.peer_last_seen) - """, - faulty_source: """ - defp broadcast_replica_heads(state) do - _ = state.peer_last_seen - peers = [] - """, - test: ["test/distributed_test.exs:4059"] + correct_source: " |> retry_pending_replica_heads()", + faulty_source: + " |> then(fn current -> if current.replica_ack_backoff, do: retry_pending_replica_heads(current), else: current end)", + test: ["test/replica_ack_test.exs:87"] }, %{ name: "skip_journal_crash_repair", file: "lib/group/replica.ex", correct_source: ":ok = Data.repair_local_replica_journal(name, shard_index)", faulty_source: ":ok", - test: ["test/group_test.exs:2624"] + test: ["test/group_test.exs:2628"] }, %{ name: "skip_index_crash_repair", file: "lib/group/replica.ex", correct_source: ":ok = Data.repair_shard_indexes(name, shard_index)", faulty_source: ":ok", - test: ["test/group_test.exs:2670"] + test: ["test/group_test.exs:2674"] }, %{ name: "skip_pg_count_projection_update", @@ -518,14 +508,14 @@ defmodule Group.MutationCampaign do faulty_source: " _ = {table, count_key, total_delta, local_delta}\n" <> " [total_count, local_count] = [0, 0]\n", - test: ["test/group_test.exs:1966"] + test: ["test/group_test.exs:1970"] }, %{ name: "retain_stale_pg_counts_on_shard_repair", file: "lib/group/replica/data.ex", correct_source: " :ets.delete_all_objects(pg_counts)", faulty_source: " _ = pg_counts", - test: ["test/group_test.exs:2670"] + test: ["test/group_test.exs:2674"] }, %{ name: "consult_stale_pg_counts_during_snapshot_repair", @@ -550,7 +540,7 @@ defmodule Group.MutationCampaign do :ok end """, - test: ["test/anti_entropy_fault_regression_test.exs:3278"] + test: ["test/anti_entropy_fault_regression_test.exs:3305"] }, %{ name: "replay_local_journal_before_pg_count_repair", @@ -565,7 +555,7 @@ defmodule Group.MutationCampaign do state = replay_local_journal(state) :ok = Data.repair_shard_indexes(name, shard_index) """, - test: ["test/group_test.exs:3171"] + test: ["test/group_test.exs:3175"] }, %{ name: "skip_inactive_cluster_repair", @@ -575,7 +565,7 @@ defmodule Group.MutationCampaign do " if Process.get(:run_primary_replica_repair, false),\n" <> " do: repair_primary_replica_rows(name, shard),\n" <> " else: :ok", - test: ["test/group_test.exs:2778"] + test: ["test/group_test.exs:2782"] }, %{ name: "skip_closed_cluster_completion", @@ -591,7 +581,7 @@ defmodule Group.MutationCampaign do faulty_source: """ _completed_clusters = [] """, - test: ["test/group_test.exs:2778"] + test: ["test/group_test.exs:2782"] }, %{ name: "accept_unfenced_cluster_disconnect", @@ -606,7 +596,7 @@ defmodule Group.MutationCampaign do " Process.get(:accept_unfenced_cluster_disconnect, true)\n" <> " end)\n\n" <> " case epochs do\n", - test: ["test/group_test.exs:1528"] + test: ["test/group_test.exs:1531"] }, %{ name: "accept_completed_cluster_disconnect", @@ -621,14 +611,14 @@ defmodule Group.MutationCampaign do faulty_source: " _ = cluster\n" <> " Process.get(:accept_completed_cluster_disconnect, true)", - test: ["test/group_test.exs:1528"] + test: ["test/group_test.exs:1531"] }, %{ name: "acknowledge_wrong_cluster_close_epoch", file: "lib/group/replica/data.ex", correct_source: " [{^cluster, ^request_epoch, pending_shards}] ->\n", faulty_source: " [{^cluster, _stored_epoch, pending_shards}] ->\n", - test: ["test/group_test.exs:2819"] + test: ["test/group_test.exs:2823"] }, %{ name: "accept_shared_authority_before_lane_install", @@ -639,7 +629,7 @@ defmodule Group.MutationCampaign do faulty_source: " WireProtocol.stream_shard(stream_id) == state.shard_index and\n" <> " true and", - test: ["test/anti_entropy_fault_regression_test.exs:1313"] + test: ["test/anti_entropy_fault_regression_test.exs:1340"] }, %{ name: "apply_incremental_authority_across_revision_gap", @@ -648,21 +638,23 @@ defmodule Group.MutationCampaign do faulty_source: " if contiguous_cluster_controls?(accepted, next_revision) or\n" <> " Enum.any?(accepted, fn {revision, _epochs} -> revision >= next_revision end) do", - test: ["test/anti_entropy_fault_regression_test.exs:1505"] + test: ["test/anti_entropy_fault_regression_test.exs:1532"] }, %{ name: "allow_non_owner_lane_to_mutate_shared_authority", file: "lib/group/replica.ex", - correct_source: " if state.shard_index == 0 do\n remote_node = node(remote_pid)", - faulty_source: " if true do\n remote_node = node(remote_pid)", - test: ["test/anti_entropy_fault_regression_test.exs:1505"] + correct_source: + " if state.shard_index == 0 do\n remote_node = node(remote_pid)\n\n # Open and close controls", + faulty_source: + " if true do\n remote_node = node(remote_pid)\n\n # Open and close controls", + test: ["test/anti_entropy_fault_regression_test.exs:1532"] }, %{ name: "crash_lane_when_local_authority_owner_is_missing", file: "lib/group/replica.ex", correct_source: " _ = send_local_control_message(state, control)", faulty_source: " send(shard_name(state.name, 0), control)", - test: ["test/anti_entropy_fault_regression_test.exs:1452"] + test: ["test/anti_entropy_fault_regression_test.exs:1479"] }, %{ name: "retire_local_owner_after_remote_authority_changed", @@ -671,7 +663,7 @@ defmodule Group.MutationCampaign do faulty_source: " Process.get(:skip_remote_registry_authority, true) or\n" <> " registry_winner_authoritative?(state, cluster, winner) ->", - test: ["test/anti_entropy_fault_regression_test.exs:1664"] + test: ["test/anti_entropy_fault_regression_test.exs:1691"] }, %{ name: "skip_registry_reprojection_after_authority_restore", @@ -696,7 +688,7 @@ defmodule Group.MutationCampaign do state end """, - test: ["test/anti_entropy_fault_regression_test.exs:1847"] + test: ["test/anti_entropy_fault_regression_test.exs:1874"] }, %{ name: "retain_registry_reprojection_after_peer_expiry", @@ -712,7 +704,7 @@ defmodule Group.MutationCampaign do state = discard_snapshot_transfers_for_source(state, remote_node) state = discard_snapshot_send_offsets_for_target(state, remote_node) """, - test: ["test/anti_entropy_fault_regression_test.exs:2023"] + test: ["test/anti_entropy_fault_regression_test.exs:2050"] }, %{ name: "retain_registry_reprojection_after_nodedown", @@ -730,7 +722,7 @@ defmodule Group.MutationCampaign do state = discard_snapshot_transfers_for_source(state, dead_node) state = discard_snapshot_send_offsets_for_target(state, dead_node) """, - test: ["test/anti_entropy_fault_regression_test.exs:3741"] + test: ["test/anti_entropy_fault_regression_test.exs:3772"] }, %{ name: "separate_exact_authority_from_cluster_projection", @@ -739,7 +731,7 @@ defmodule Group.MutationCampaign do " replace_remote_cluster_projection(state.name, remote_node, current_epochs)\n", faulty_source: " _ = {&replace_remote_cluster_projection/3, state.name, remote_node, current_epochs}\n", - test: ["test/anti_entropy_fault_regression_test.exs:3777"] + test: ["test/anti_entropy_fault_regression_test.exs:3808"] }, %{ name: "separate_local_activation_from_cluster_projection", @@ -748,7 +740,7 @@ defmodule Group.MutationCampaign do " if durable?, do: project_activated_local_clusters(state.name, clusters)\n", faulty_source: " _ = {durable?, &project_activated_local_clusters/2, state.name, clusters}\n", - test: ["test/anti_entropy_fault_regression_test.exs:3832"] + test: ["test/anti_entropy_fault_regression_test.exs:3863"] }, %{ name: "drop_durable_cluster_deactivation_cleanup", @@ -757,11 +749,11 @@ defmodule Group.MutationCampaign do " cast_cluster_lifecycle(\n" <> " state.name,\n" <> " 0..(state.num_shards - 1),\n" <> - " {:cluster_disconnect, clusters, epochs}\n" <> + " {:cluster_disconnect, clusters, epochs, local_cluster_epoch_revision(state.name)}\n" <> " )\n", faulty_source: " _ = {&cast_cluster_lifecycle/3, state.name, state.num_shards, clusters, epochs}\n", - test: ["test/anti_entropy_fault_regression_test.exs:3881"] + test: ["test/anti_entropy_fault_regression_test.exs:3912"] }, %{ name: "delete_close_marker_before_terminal_route_cleanup", @@ -771,7 +763,7 @@ defmodule Group.MutationCampaign do " :ets.delete(closed_local_cluster_epochs_table(state.name), cluster)\n", faulty_source: " :ets.delete(closed_local_cluster_epochs_table(state.name), cluster)\n", - test: ["test/group_test.exs:2819"] + test: ["test/group_test.exs:2823"] }, %{ name: "retire_peer_authority_before_terminal_route_cleanup", @@ -781,7 +773,7 @@ defmodule Group.MutationCampaign do " :ok = delete_peer_routes(name, remote_node)\n", faulty_source: " :ets.delete(replication_meta_table(name), {:remote_generation, remote_node})\n", - test: ["test/group_test.exs:2873"] + test: ["test/group_test.exs:2877"] }, %{ name: "stale_peer_cleanup_removes_rediscovered_routes", @@ -793,7 +785,7 @@ defmodule Group.MutationCampaign do " if Process.get(:purge_rediscovered_peer_routes, true) or\n" <> " (is_nil(remote_generation(state.name, dead_node)) and\n" <> " is_nil(remote_replica_authority_hint(state.name, dead_node))) do\n", - test: ["test/group_test.exs:2926"] + test: ["test/group_test.exs:2930"] }, %{ name: "stale_restart_cleanup_removes_reactivated_routes", @@ -801,7 +793,7 @@ defmodule Group.MutationCampaign do correct_source: " Enum.filter(clusters, &is_nil(local_cluster_epoch(state.name, &1)))\n", faulty_source: " clusters\n", - test: ["test/group_test.exs:2853"] + test: ["test/group_test.exs:2857"] }, %{ name: "retain_authority_repair_after_nodedown", @@ -812,7 +804,7 @@ defmodule Group.MutationCampaign do faulty_source: " cluster_control_dirty: state.cluster_control_dirty,\n" <> " authority_dirty_notified: MapSet.delete(state.authority_dirty_notified, dead_node)\n", - test: ["test/group_test.exs:2912"] + test: ["test/group_test.exs:2916"] }, %{ name: "retain_receive_cursor_for_inactive_local_cluster", @@ -823,14 +815,14 @@ defmodule Group.MutationCampaign do faulty_source: " WireProtocol.stream_origin(stream_id) != node() and\n" <> " true and", - test: ["test/anti_entropy_fault_regression_test.exs:2876"] + test: ["test/anti_entropy_fault_regression_test.exs:2903"] }, %{ name: "retire_shared_authority_with_live_lanes", file: "lib/group/replica/data.ex", correct_source: " result =\n if remaining_lanes == 0 do", faulty_source: " _ = remaining_lanes\n\n result =\n if true do", - test: ["test/anti_entropy_fault_regression_test.exs:857"] + test: ["test/anti_entropy_fault_regression_test.exs:884"] }, %{ name: "shard_zero_deletes_sibling_restart_views", @@ -893,7 +885,7 @@ defmodule Group.MutationCampaign do " Enum.split_while(contiguous, fn {_seq, mutations} ->\n" <> " valid_replica_mutations?(%{state | num_shards: 1}, stream_id, mutations)\n" <> " end)", - test: ["test/anti_entropy_fault_regression_test.exs:512"] + test: ["test/anti_entropy_fault_regression_test.exs:521"] }, %{ name: "restore_unsequenced_cluster_disconnect", @@ -930,7 +922,7 @@ defmodule Group.MutationCampaign do {:noreply, state} end """, - test: ["test/group_test.exs:2466"] + test: ["test/group_test.exs:2470"] }, %{ name: "skip_cursorless_restart_authority_repair", @@ -940,7 +932,7 @@ defmodule Group.MutationCampaign do " if Process.get(:run_primary_replica_repair, false),\n" <> " do: repair_primary_replica_rows(name, shard),\n" <> " else: :ok", - test: ["test/anti_entropy_fault_regression_test.exs:3166"] + test: ["test/anti_entropy_fault_regression_test.exs:3193"] }, %{ name: "project_stale_claims_before_restart_repair", @@ -955,7 +947,7 @@ defmodule Group.MutationCampaign do {state, _events} = rebuild_registry_projections(state) :ok = Data.repair_shard_indexes(name, shard_index) """, - test: ["test/anti_entropy_fault_regression_test.exs:3598"] + test: ["test/anti_entropy_fault_regression_test.exs:3629"] }, %{ name: "skip_interrupted_snapshot_install_repair", @@ -965,7 +957,7 @@ defmodule Group.MutationCampaign do " if Process.get(:run_snapshot_install_repair, false),\n" <> " do: repair_interrupted_snapshot_installs(name, shard),\n" <> " else: :ok", - test: ["test/anti_entropy_fault_regression_test.exs:3278"] + test: ["test/anti_entropy_fault_regression_test.exs:3305"] }, %{ name: "retain_cursorless_remote_registry_claims", @@ -974,20 +966,28 @@ defmodule Group.MutationCampaign do " :ets.member(replica_cursor_table(name, shard), stream_id)\n else\n false\n end\n end\n\n defp valid_remote_pg_authority?", faulty_source: " is_tuple(stream_id)\n else\n false\n end\n end\n\n defp valid_remote_pg_authority?", - test: ["test/anti_entropy_fault_regression_test.exs:3166"] + test: ["test/anti_entropy_fault_regression_test.exs:3193"] }, %{ name: "restart_snapshot_from_first_chunk_after_busy", file: "lib/group/replica.ex", - correct_source: " resume = Map.get(offsets, snapshot_key, {:chunk, 1})", + correct_source: """ + resume = + case Map.get(offsets, snapshot_key) do + {:sent, _sent_at} -> {:chunk, 1} + nil -> {:chunk, 1} + offset -> offset + end + """, faulty_source: """ resume = case Map.get(offsets, snapshot_key) do + {:sent, _sent_at} -> {:chunk, 1} {:commit, _manifest} = commit -> commit _chunk_resume -> {:chunk, 1} end """, - test: ["test/replica_snapshot_distributed_test.exs:877"] + test: ["test/replica_snapshot_distributed_test.exs:879"] }, %{ name: "drain_oversized_ingress_batch_without_yield", diff --git a/test/replica_ack_scale_test.exs b/test/replica_ack_scale_test.exs new file mode 100644 index 0000000..fc8a47c --- /dev/null +++ b/test/replica_ack_scale_test.exs @@ -0,0 +1,104 @@ +defmodule Group.ReplicaAckScaleTest do + use ExUnit.Case, async: false + + @moduletag :capture_log + @moduletag timeout: 360_000 + + alias Group.TestCluster + + @tag skip: System.get_env("GROUP_SCALE_TEST") != "1" + test "20,000 connected subclusters advertise only unsettled heads" do + peers = TestCluster.start_peers(2) + on_exit(fn -> TestCluster.stop_peers(peers) end) + [{_, source}, {_, receiver}] = peers + + name = :"replica_ack_scale_#{System.unique_integer([:positive])}" + clusters = for index <- 1..20_000, do: "tenant/#{index}" + + opts = [ + name: name, + shards: 4, + replica_transport: Group.TestReplicaTransport, + replicated_anti_entropy_interval: 5_000, + replicated_peer_lease_timeout: 120_000 + ] + + for node <- [source, receiver] do + {:ok, _pid} = TestCluster.start_group(node, opts) + end + + for batch <- Enum.chunk_every(clusters, 1_000) do + tasks = + for node <- [source, receiver] do + Task.async(fn -> TestCluster.connect_many_concurrently(node, name, batch) end) + end + + assert [:ok, :ok] = Task.await_many(tasks, 120_000) + end + + TestCluster.assert_eventually( + fn -> + Enum.all?([source, receiver], fn node -> + length(TestCluster.rpc!(node, Group.Replica.Data, :my_clusters, [name])) == 20_001 and + length(TestCluster.rpc!(node, Group, :nodes, [name, "tenant/20000"])) == 2 + end) + end, + timeout: 120_000, + interval: 100 + ) + + entries = + for index <- 1..12 do + cluster = "tenant/#{index * 1_666}" + key = "sprite/#{index}" + pid = TestCluster.spawn_register_in_cluster(source, name, key, %{index: index}, cluster) + {cluster, key, pid} + end + + TestCluster.assert_eventually( + fn -> + Enum.all?(entries, fn {cluster, key, pid} -> + match?( + {^pid, _meta}, + TestCluster.rpc!(receiver, Group, :lookup, [name, key, [cluster: cluster]]) + ) + end) + end, + timeout: 120_000, + interval: 100 + ) + + TestCluster.assert_eventually( + fn -> + Enum.all?(0..3, fn shard -> + source + |> TestCluster.rpc!(:sys, :get_state, [Group.Replica.shard_name(name, shard)]) + |> Map.get(:pending_replica_heads) + |> map_size() == 0 + end) + end, + timeout: 120_000, + interval: 100 + ) + + :ok = + TestCluster.rpc!(source, Group.TestReplicaTransport, :set_mode, [ + name, + {:capture_pass, [:heads]} + ]) + + for shard <- 0..3 do + shard_name = Group.Replica.shard_name(name, shard) + state = TestCluster.rpc!(source, :sys, :get_state, [shard_name]) + shard_pid = TestCluster.rpc!(source, Process, :whereis, [shard_name]) + send(shard_pid, {:group_replica_anti_entropy, state.anti_entropy_ref}) + + TestCluster.assert_eventually(fn -> + TestCluster.rpc!(source, :sys, :get_state, [shard_name]).anti_entropy_ref != + state.anti_entropy_ref + end) + end + + assert [] == TestCluster.rpc!(source, Group.TestReplicaTransport, :captured, [name]) + end +end diff --git a/test/replica_ack_test.exs b/test/replica_ack_test.exs new file mode 100644 index 0000000..58fc91b --- /dev/null +++ b/test/replica_ack_test.exs @@ -0,0 +1,1531 @@ +defmodule Group.ReplicaAckTest do + use ExUnit.Case, async: false + + @moduletag :capture_log + @moduletag timeout: 60_000 + + alias Group.TestCluster + + setup do + peers = TestCluster.start_peers(2) + on_exit(fn -> TestCluster.stop_peers(peers) end) + [{_, source}, {_, receiver}] = peers + + name = :"replica_ack_#{System.unique_integer([:positive])}" + + opts = [ + name: name, + shards: 1, + replica_transport: Group.TestReplicaTransport, + replicated_anti_entropy_interval: 25, + replicated_peer_lease_timeout: 5_000 + ] + + for node <- [source, receiver] do + {:ok, _pid} = TestCluster.start_group(node, opts) + :ok = TestCluster.rpc!(node, Group, :connect, [name, "org"]) + end + + TestCluster.assert_eventually(fn -> + Enum.all?([source, receiver], fn node -> + length(TestCluster.rpc!(node, Group, :nodes, [name, "org"])) == 2 + end) + end) + + on_exit(fn -> + for node <- [source, receiver] do + TestCluster.rpc!(node, Group.TestReplicaTransport, :clear, [name]) + end + end) + + {:ok, name: name, source: source, receiver: receiver} + end + + test "settled streams do not keep advertising heads", context do + owner = + TestCluster.spawn_join_in_cluster( + context.source, + context.name, + "sprite", + %{up: true}, + "org" + ) + + TestCluster.assert_eventually(fn -> + match?( + [{^owner, %{up: true}}], + TestCluster.rpc!(context.receiver, Group, :members, [ + context.name, + "sprite", + [cluster: "org"] + ]) + ) + end) + + TestCluster.assert_eventually(fn -> + context.source + |> TestCluster.rpc!(:sys, :get_state, [Group.Replica.shard_name(context.name, 0)]) + |> Map.get(:pending_replica_heads) + |> Map.get(context.receiver, %{}) + |> map_size() == 0 + end) + + :ok = + TestCluster.rpc!(context.source, Group.TestReplicaTransport, :set_mode, [ + context.name, + {:capture_pass, [:heads]} + ]) + + drive_anti_entropy(context.source, context.name, 5) + + assert [] == + TestCluster.rpc!(context.source, Group.TestReplicaTransport, :captured, [ + context.name + ]) + end + + test "a lost applied ACK retries the head and then becomes quiet", context do + :ok = + TestCluster.rpc!(context.receiver, Group.TestReplicaTransport, :set_mode, [ + context.name, + {:drop_types, [:applied]} + ]) + + owner = + TestCluster.spawn_join_in_cluster( + context.source, + context.name, + "sprite", + %{up: true}, + "org" + ) + + TestCluster.assert_eventually(fn -> + match?( + [{^owner, _meta}], + TestCluster.rpc!(context.receiver, Group, :members, [ + context.name, + "sprite", + [cluster: "org"] + ]) + ) + end) + + :ok = + TestCluster.rpc!(context.source, Group.TestReplicaTransport, :set_mode, [ + context.name, + {:capture_pass, [:heads]} + ]) + + TestCluster.assert_eventually(fn -> + context.source + |> TestCluster.rpc!(Group.TestReplicaTransport, :captured, [context.name]) + |> length() >= 2 + end) + + :ok = + TestCluster.rpc!(context.receiver, Group.TestReplicaTransport, :set_mode, [ + context.name, + :pass + ]) + + TestCluster.assert_eventually(fn -> + context.source + |> TestCluster.rpc!(:sys, :get_state, [Group.Replica.shard_name(context.name, 0)]) + |> Map.get(:pending_replica_heads) + |> Map.get(context.receiver, %{}) + |> map_size() == 0 + end) + + :ok = + TestCluster.rpc!(context.source, Group.TestReplicaTransport, :clear_captured, [context.name]) + + drive_anti_entropy(context.source, context.name, 5) + + assert [] == + TestCluster.rpc!(context.source, Group.TestReplicaTransport, :captured, [ + context.name + ]) + end + + test "busy ACK transport backs off while pending heads continue repair", context do + :ok = + TestCluster.rpc!(context.receiver, Group.TestReplicaTransport, :set_mode, [ + context.name, + {:capture_busy_types, [:applied]} + ]) + + owner = + TestCluster.spawn_join_in_cluster( + context.source, + context.name, + "sprite", + %{up: true}, + "org" + ) + + TestCluster.assert_eventually(fn -> + match?( + [{^owner, _meta}], + TestCluster.rpc!(context.receiver, Group, :members, [ + context.name, + "sprite", + [cluster: "org"] + ]) + ) + end) + + TestCluster.assert_eventually(fn -> + context.receiver + |> TestCluster.rpc!(Group.TestReplicaTransport, :captured, [context.name]) + |> length() > 0 + end) + + drive_anti_entropy(context.source, context.name, 5) + TestCluster.flush_shards(context.receiver, context.name) + + attempts = + TestCluster.rpc!(context.receiver, Group.TestReplicaTransport, :captured, [context.name]) + + assert length(attempts) <= 2 + + :ok = + TestCluster.rpc!(context.receiver, Group.TestReplicaTransport, :set_mode, [ + context.name, + :pass + ]) + + TestCluster.assert_eventually(fn -> + context.source + |> TestCluster.rpc!(:sys, :get_state, [Group.Replica.shard_name(context.name, 0)]) + |> Map.get(:pending_replica_heads) + |> Map.get(context.receiver, %{}) + |> map_size() == 0 + end) + end + + test "repeated discovery probes do not resend an acknowledged catalog", context do + owner = + TestCluster.spawn_join_in_cluster(context.source, context.name, "member", %{}, "org") + + TestCluster.assert_eventually(fn -> + match?( + [{^owner, _}], + TestCluster.rpc!(context.receiver, Group, :members, [ + context.name, + "member", + [cluster: "org"] + ]) + ) + end) + + source_shard = + TestCluster.rpc!(context.source, Process, :whereis, [ + Group.Replica.shard_name(context.name, 0) + ]) + + receiver_shard = + TestCluster.rpc!(context.receiver, Process, :whereis, [ + Group.Replica.shard_name(context.name, 0) + ]) + + # Prime one discovery round so the sender has seen this receiver lane. + send(source_shard, peer_connect(context, receiver_shard, 0)) + _state = TestCluster.rpc!(context.source, :sys, :get_state, [source_shard]) + + TestCluster.assert_eventually(fn -> + context.source + |> TestCluster.rpc!(:sys, :get_state, [source_shard]) + |> Map.get(:pending_replica_heads) + |> Map.get(context.receiver, %{}) + |> map_size() == 0 + end) + + TestCluster.flush_shards(context.receiver, context.name) + :ok = TestCluster.rpc!(context.receiver, :sys, :suspend, [receiver_shard]) + + on_exit(fn -> + TestCluster.rpc!(context.receiver, TestCluster, :resume_if_alive, [receiver_shard]) + end) + + :ok = + TestCluster.rpc!(context.source, Group.TestReplicaTransport, :set_mode, [ + context.name, + {:capture_pass, [:heads]} + ]) + + # Even after the discovery retry window, duplicate probes only need a + # constant-size ACK when the sender already has this exact receiver view. + :ok = + TestCluster.rpc!(context.source, TestCluster, :backdate_discovery_hello, [ + context.name, + 0, + context.receiver, + 10_000 + ]) + + for _ <- 1..5 do + send(source_shard, peer_connect(context, receiver_shard, 0)) + _state = TestCluster.rpc!(context.source, :sys, :get_state, [source_shard]) + end + + assert [] == + TestCluster.rpc!(context.source, Group.TestReplicaTransport, :captured, [ + context.name + ]) + + {:messages, messages} = + TestCluster.rpc!(context.receiver, Process, :info, [receiver_shard, :messages]) + + refute Enum.any?(messages, fn + {:replica_hello, _, _, _, _, _, _, _} -> true + _ -> false + end) + end + + test "a probe from an uninstalled receiver PID cannot certify head reseeding", context do + owner = + TestCluster.spawn_join_in_cluster(context.source, context.name, "member", %{}, "org") + + TestCluster.assert_eventually(fn -> + match?( + [{^owner, _}], + TestCluster.rpc!(context.receiver, Group, :members, [ + context.name, + "member", + [cluster: "org"] + ]) + ) + end) + + source_shard = + TestCluster.rpc!(context.source, Process, :whereis, [ + Group.Replica.shard_name(context.name, 0) + ]) + + receiver_shard = + TestCluster.rpc!(context.receiver, Process, :whereis, [ + Group.Replica.shard_name(context.name, 0) + ]) + + TestCluster.assert_eventually(fn -> + state = TestCluster.rpc!(context.source, :sys, :get_state, [source_shard]) + map_size(Map.get(state.pending_replica_heads, context.receiver, %{})) == 0 + end) + + :ok = TestCluster.rpc!(context.receiver, :sys, :suspend, [receiver_shard]) + + on_exit(fn -> + TestCluster.rpc!(context.receiver, TestCluster, :resume_if_alive, [receiver_shard]) + end) + + new_receiver_pid = TestCluster.spawn_trace_forwarder(context.receiver, self()) + send(source_shard, peer_connect(context, new_receiver_pid, 1)) + source_state = TestCluster.rpc!(context.source, :sys, :get_state, [source_shard]) + + assert map_size(Map.get(source_state.pending_replica_heads, context.receiver, %{})) == 0 + + {:messages, messages} = + TestCluster.rpc!(context.receiver, Process, :info, [receiver_shard, :messages]) + + assert Enum.any?(messages, fn + {:peer_connect_ack, ^source_shard, 0, 1, 1, false} -> true + _ -> false + end) + + {:peer_connect, ^receiver_shard, 0, 1, 2, generation, revision} = + peer_connect(context, receiver_shard, 2) + + send( + source_shard, + {:peer_connect, receiver_shard, 0, 1, 2, generation, revision + 1} + ) + + source_state = TestCluster.rpc!(context.source, :sys, :get_state, [source_shard]) + assert map_size(Map.get(source_state.pending_replica_heads, context.receiver, %{})) == 0 + + {:messages, messages} = + TestCluster.rpc!(context.receiver, Process, :info, [receiver_shard, :messages]) + + assert Enum.any?(messages, fn + {:peer_connect_ack, ^source_shard, 0, 1, 2, false} -> true + _ -> false + end) + end + + test "a lost discovery hello is retried without a catalog on every probe", context do + source_shard = + TestCluster.rpc!(context.source, Process, :whereis, [ + Group.Replica.shard_name(context.name, 0) + ]) + + receiver_shard = + TestCluster.rpc!(context.receiver, Process, :whereis, [ + Group.Replica.shard_name(context.name, 0) + ]) + + TestCluster.flush_shards(context.receiver, context.name) + :ok = TestCluster.rpc!(context.receiver, :sys, :suspend, [receiver_shard]) + + on_exit(fn -> + TestCluster.rpc!(context.receiver, TestCluster, :resume_if_alive, [receiver_shard]) + end) + + send(source_shard, peer_connect(context, receiver_shard, 1)) + _state = TestCluster.rpc!(context.source, :sys, :get_state, [source_shard]) + + for _ <- 1..5 do + send(source_shard, peer_connect(context, receiver_shard, 1)) + _state = TestCluster.rpc!(context.source, :sys, :get_state, [source_shard]) + end + + request_token = make_ref() + + for _ <- 1..5 do + send(source_shard, {:replica_hello_request, receiver_shard, request_token}) + _state = TestCluster.rpc!(context.source, :sys, :get_state, [source_shard]) + end + + assert 2 == queued_replica_hellos(context.receiver, receiver_shard) + + :ok = + TestCluster.rpc!(context.source, TestCluster, :backdate_discovery_hello, [ + context.name, + 0, + context.receiver, + 10_000 + ]) + + send(source_shard, {:replica_hello_request, receiver_shard, request_token}) + _state = TestCluster.rpc!(context.source, :sys, :get_state, [source_shard]) + + assert 3 == queued_replica_hellos(context.receiver, receiver_shard) + + :ok = TestCluster.rpc!(context.source, Group, :connect, [context.name, "new-org"]) + send(source_shard, {:replica_hello_request, receiver_shard, request_token}) + _state = TestCluster.rpc!(context.source, :sys, :get_state, [source_shard]) + + assert 4 == queued_replica_hellos(context.receiver, receiver_shard) + + send(source_shard, {:replica_hello_request, receiver_shard, make_ref()}) + _state = TestCluster.rpc!(context.source, :sys, :get_state, [source_shard]) + + assert 5 == queued_replica_hellos(context.receiver, receiver_shard) + + :ok = + TestCluster.rpc!(context.source, TestCluster, :forget_peer_connect_ack, [ + context.name, + 0, + context.receiver + ]) + + for _ <- 1..5 do + send(source_shard, {:peer_connect_ack, receiver_shard, 0, 1, 0, true}) + _state = TestCluster.rpc!(context.source, :sys, :get_state, [source_shard]) + end + + assert 6 == queued_replica_hellos(context.receiver, receiver_shard) + end + + test "a heartbeat that restores a lane advertises writes missed without a route", context do + source_shard = Group.Replica.shard_name(context.name, 0) + + receiver_shard = + TestCluster.rpc!(context.receiver, Process, :whereis, [ + Group.Replica.shard_name(context.name, 0) + ]) + + :ok = TestCluster.rpc!(context.receiver, :sys, :suspend, [receiver_shard]) + + on_exit(fn -> + TestCluster.rpc!(context.receiver, TestCluster, :resume_if_alive, [receiver_shard]) + end) + + :ok = + TestCluster.rpc!(context.source, TestCluster, :forget_replica_peer_route, [ + context.name, + 0, + context.receiver + ]) + + owner = TestCluster.spawn_join(context.source, context.name, "missed", %{}) + + source_state = TestCluster.rpc!(context.source, :sys, :get_state, [source_shard]) + assert Map.get(source_state.pending_replica_heads, context.receiver, %{}) == %{} + + {generation, revision, _epochs} = + TestCluster.rpc!(context.receiver, Group.Replica.Data, :local_replica_authority, [ + context.name + ]) + + send( + TestCluster.rpc!(context.source, Process, :whereis, [source_shard]), + {:replica_heartbeat, receiver_shard, Group.Replica.WireProtocol.version(), generation, + revision, Group.TestReplicaTransport.id(), + Group.TestReplicaTransport.descriptor(context.name, [])} + ) + + TestCluster.assert_eventually(fn -> + state = TestCluster.rpc!(context.source, :sys, :get_state, [source_shard]) + map_size(Map.get(state.pending_replica_heads, context.receiver, %{})) > 0 + end) + + :ok = TestCluster.rpc!(context.receiver, :sys, :resume, [receiver_shard]) + + TestCluster.assert_eventually(fn -> + match?( + [{^owner, _}], + TestCluster.rpc!(context.receiver, Group, :members, [context.name, "missed"]) + ) + end) + end + + test "an out-of-order delta cannot advance past a missing mutation", context do + first = + TestCluster.spawn_join_in_cluster(context.source, context.name, "first", %{seq: 1}, "org") + + TestCluster.assert_eventually(fn -> + match?( + [{^first, %{seq: 1}}], + TestCluster.rpc!(context.receiver, Group, :members, [ + context.name, + "first", + [cluster: "org"] + ]) + ) + end) + + stream = + TestCluster.rpc!(context.source, Group.Replica.Data, :local_stream_id, [ + context.name, + 0, + "org" + ]) + + :ok = + TestCluster.rpc!(context.source, Group.TestReplicaTransport, :set_mode, [ + context.name, + {:capture_drop_types, [:delta_batch]} + ]) + + second = + TestCluster.spawn_join_in_cluster(context.source, context.name, "second", %{seq: 2}, "org") + + TestCluster.flush_shards(context.source, context.name) + + third = + TestCluster.spawn_join_in_cluster(context.source, context.name, "third", %{seq: 3}, "org") + + TestCluster.flush_shards(context.source, context.name) + + TestCluster.assert_eventually(fn -> captured_record(context, stream, 3) != nil end) + record = captured_record(context, stream, 3) + + send( + {Group.Replica.shard_name(context.name, 0), context.receiver}, + {:group_replica_frame, context.source, + {:delta_batch, Group.Replica.WireProtocol.version(), [{stream, 3, [record], 3}]}} + ) + + TestCluster.flush_shards(context.receiver, context.name) + + assert 1 == + TestCluster.rpc!(context.receiver, Group.Replica.Data, :replica_cursor, [ + context.name, + 0, + stream + ]) + + :ok = + TestCluster.rpc!(context.source, Group.TestReplicaTransport, :set_mode, [ + context.name, + :pass + ]) + + TestCluster.assert_eventually(fn -> + match?( + [{^second, %{seq: 2}}], + TestCluster.rpc!(context.receiver, Group, :members, [ + context.name, + "second", + [cluster: "org"] + ]) + ) and + match?( + [{^third, %{seq: 3}}], + TestCluster.rpc!(context.receiver, Group, :members, [ + context.name, + "third", + [cluster: "org"] + ]) + ) + end) + end + + test "a dropped final leave repairs without a zombie", context do + owner = + TestCluster.spawn_join_in_cluster( + context.source, + context.name, + "sprite", + %{up: true}, + "org" + ) + + TestCluster.assert_eventually(fn -> + match?( + [{^owner, _meta}], + TestCluster.rpc!(context.receiver, Group, :members, [ + context.name, + "sprite", + [cluster: "org"] + ]) + ) + end) + + :ok = + TestCluster.rpc!(context.source, Group.TestReplicaTransport, :set_mode, [ + context.name, + {:drop_types, [:delta_batch]} + ]) + + true = TestCluster.rpc!(context.source, Process, :exit, [owner, :kill]) + TestCluster.flush_shards(context.source, context.name) + + TestCluster.assert_eventually(fn -> + context.source + |> TestCluster.rpc!(:sys, :get_state, [Group.Replica.shard_name(context.name, 0)]) + |> Map.get(:pending_replica_heads) + |> Map.get(context.receiver, %{}) + |> map_size() > 0 + end) + + :ok = + TestCluster.rpc!(context.source, Group.TestReplicaTransport, :set_mode, [ + context.name, + :pass + ]) + + TestCluster.assert_eventually(fn -> + TestCluster.rpc!(context.receiver, Group, :members, [ + context.name, + "sprite", + [cluster: "org"] + ]) == [] + end) + end + + test "receiver cursor loss and shard restart reseed an acknowledged stream", context do + owner = + TestCluster.spawn_join_in_cluster( + context.source, + context.name, + "sprite", + %{up: true}, + "org" + ) + + TestCluster.assert_eventually(fn -> + match?( + [{^owner, _meta}], + TestCluster.rpc!(context.receiver, Group, :members, [ + context.name, + "sprite", + [cluster: "org"] + ]) + ) + end) + + TestCluster.assert_eventually(fn -> + context.source + |> TestCluster.rpc!(:sys, :get_state, [Group.Replica.shard_name(context.name, 0)]) + |> Map.get(:pending_replica_heads) + |> Map.get(context.receiver, %{}) + |> map_size() == 0 + end) + + stream_id = + TestCluster.rpc!(context.source, Group.Replica.Data, :local_stream_id, [ + context.name, + 0, + "org" + ]) + + :ok = + TestCluster.rpc!(context.receiver, Group.Replica.Data, :delete_replica_cursor, [ + context.name, + 0, + stream_id + ]) + + old_lane = + TestCluster.rpc!(context.receiver, Process, :whereis, [ + Group.Replica.shard_name(context.name, 0) + ]) + + true = TestCluster.rpc!(context.receiver, Process, :exit, [old_lane, :kill]) + + TestCluster.assert_eventually(fn -> + case TestCluster.rpc!(context.receiver, Process, :whereis, [ + Group.Replica.shard_name(context.name, 0) + ]) do + pid when is_pid(pid) -> pid != old_lane + _ -> false + end + end) + + TestCluster.assert_eventually(fn -> + match?( + [{^owner, _meta}], + TestCluster.rpc!(context.receiver, Group, :members, [ + context.name, + "sprite", + [cluster: "org"] + ]) + ) + end) + end + + test "one-sided lease expiry reseeds a stream and rejects a delayed ACK", context do + :ok = + TestCluster.rpc!(context.receiver, Group.TestReplicaTransport, :set_mode, [ + context.name, + {:capture_pass, [:applied]} + ]) + + owner = + TestCluster.spawn_join_in_cluster( + context.source, + context.name, + "sprite", + %{up: true}, + "org" + ) + + TestCluster.assert_eventually(fn -> + match?( + [{^owner, _meta}], + TestCluster.rpc!(context.receiver, Group, :members, [ + context.name, + "sprite", + [cluster: "org"] + ]) + ) + end) + + TestCluster.assert_eventually(fn -> + context.source + |> TestCluster.rpc!(:sys, :get_state, [Group.Replica.shard_name(context.name, 0)]) + |> Map.get(:pending_replica_heads) + |> Map.get(context.receiver, %{}) + |> map_size() == 0 + end) + + {_target, 0, old_ack} = + context.receiver + |> TestCluster.rpc!(Group.TestReplicaTransport, :captured, [context.name]) + |> Enum.find(fn {_target, shard, frame} -> shard == 0 and elem(frame, 0) == :applied end) + + :ok = + TestCluster.rpc!(context.source, Group.TestReplicaTransport, :set_mode, [ + context.name, + {:drop_types, [:heads]} + ]) + + :ok = + TestCluster.rpc!(context.receiver, Group.TestCluster, :expire_replica_lane, [ + context.name, + 0, + context.source + ]) + + TestCluster.assert_eventually(fn -> + TestCluster.rpc!(context.receiver, Group, :members, [ + context.name, + "sprite", + [cluster: "org"] + ]) == [] + end) + + TestCluster.assert_eventually(fn -> + context.source + |> TestCluster.rpc!(:sys, :get_state, [Group.Replica.shard_name(context.name, 0)]) + |> Map.get(:pending_replica_heads) + |> Map.get(context.receiver, %{}) + |> map_size() > 0 + end) + + :ok = + TestCluster.rpc!(context.source, Group.Transport, :incoming, [ + context.name, + context.receiver, + 0, + old_ack + ]) + + TestCluster.flush_shards(context.source, context.name) + + source_state = + TestCluster.rpc!(context.source, :sys, :get_state, [ + Group.Replica.shard_name(context.name, 0) + ]) + + assert map_size(Map.get(source_state.pending_replica_heads, context.receiver, %{})) > 0 + + :ok = + TestCluster.rpc!(context.source, Group.TestReplicaTransport, :set_mode, [ + context.name, + :pass + ]) + + TestCluster.assert_eventually(fn -> + match?( + [{^owner, _meta}], + TestCluster.rpc!(context.receiver, Group, :members, [ + context.name, + "sprite", + [cluster: "org"] + ]) + ) + end) + end + + test "a stale hello after expiry cannot suppress an unacknowledged recovery probe", context do + owner = + TestCluster.spawn_join_in_cluster(context.source, context.name, "sprite", %{}, "org") + + TestCluster.assert_eventually(fn -> + match?( + [{^owner, _}], + TestCluster.rpc!(context.receiver, Group, :members, [ + context.name, + "sprite", + [cluster: "org"] + ]) + ) + end) + + TestCluster.assert_eventually(fn -> + state = + TestCluster.rpc!(context.source, :sys, :get_state, [ + Group.Replica.shard_name(context.name, 0) + ]) + + map_size(Map.get(state.pending_replica_heads, context.receiver, %{})) == 0 + end) + + source_shard = + TestCluster.rpc!(context.source, Process, :whereis, [ + Group.Replica.shard_name(context.name, 0) + ]) + + :ok = TestCluster.rpc!(context.source, :sys, :suspend, [source_shard]) + + on_exit(fn -> + TestCluster.rpc!(context.source, TestCluster, :resume_if_alive, [source_shard]) + end) + + {receiver_shard, probe_epoch, tick_ref} = + TestCluster.rpc!( + context.receiver, + TestCluster, + :expire_replica_lane_without_probe, + [context.name, 0, context.source] + ) + + {generation, revision, epochs} = + TestCluster.rpc!(context.source, Group.Replica.Data, :local_replica_authority, [ + context.name + ]) + + send( + receiver_shard, + {:replica_hello, source_shard, Group.Replica.WireProtocol.version(), generation, revision, + epochs, Group.TestReplicaTransport.id(), + Group.TestReplicaTransport.descriptor(context.name, [])} + ) + + receiver_state = TestCluster.rpc!(context.receiver, :sys, :get_state, [receiver_shard]) + assert Map.get(receiver_state.remote_shards, context.source) == source_shard + + assert [] == + TestCluster.rpc!(context.receiver, Group, :members, [ + context.name, + "sprite", + [cluster: "org"] + ]) + + send(receiver_shard, {:peer_connect_ack, source_shard, 0, 1, probe_epoch - 1, true}) + receiver_state = TestCluster.rpc!(context.receiver, :sys, :get_state, [receiver_shard]) + assert Map.get(receiver_state.pending_peer_probes, context.source) == probe_epoch + + send(receiver_shard, {:peer_connect_ack, source_shard, 0, 1, probe_epoch, false}) + receiver_state = TestCluster.rpc!(context.receiver, :sys, :get_state, [receiver_shard]) + assert Map.get(receiver_state.pending_peer_probes, context.source) == probe_epoch + + send(receiver_shard, {:group_replica_anti_entropy, tick_ref}) + _receiver_state = TestCluster.rpc!(context.receiver, :sys, :get_state, [receiver_shard]) + + {:messages, messages} = + TestCluster.rpc!(context.source, Process, :info, [source_shard, :messages]) + + assert Enum.any?(messages, fn + {:peer_connect, ^receiver_shard, 0, 1, ^probe_epoch, _generation, _revision} -> true + _ -> false + end) + + :ok = TestCluster.rpc!(context.source, :sys, :resume, [source_shard]) + + TestCluster.assert_eventually(fn -> + match?( + [{^owner, _}], + TestCluster.rpc!(context.receiver, Group, :members, [ + context.name, + "sprite", + [cluster: "org"] + ]) + ) + end) + + TestCluster.assert_eventually(fn -> + state = TestCluster.rpc!(context.receiver, :sys, :get_state, [receiver_shard]) + not Map.has_key?(state.pending_peer_probes, context.source) + end) + end + + test "a second one-sided expiry advances the recovery epoch and rejects the first ACK", + context do + owner = + TestCluster.spawn_join_in_cluster(context.source, context.name, "sprite", %{}, "org") + + TestCluster.assert_eventually(fn -> + match?( + [{^owner, _}], + TestCluster.rpc!(context.receiver, Group, :members, [ + context.name, + "sprite", + [cluster: "org"] + ]) + ) + end) + + source_shard = Group.Replica.shard_name(context.name, 0) + receiver_shard = Group.Replica.shard_name(context.name, 0) + + TestCluster.assert_eventually(fn -> + state = TestCluster.rpc!(context.source, :sys, :get_state, [source_shard]) + map_size(Map.get(state.pending_replica_heads, context.receiver, %{})) == 0 + end) + + :ok = + TestCluster.rpc!(context.source, Group.TestReplicaTransport, :set_mode, [ + context.name, + {:drop_types, [:heads]} + ]) + + :ok = + TestCluster.rpc!(context.receiver, TestCluster, :expire_replica_lane, [ + context.name, + 0, + context.source + ]) + + TestCluster.assert_eventually(fn -> + receiver_state = TestCluster.rpc!(context.receiver, :sys, :get_state, [receiver_shard]) + source_state = TestCluster.rpc!(context.source, :sys, :get_state, [source_shard]) + + Map.get(receiver_state.peer_probe_epochs, context.source) == 1 and + map_size(Map.get(source_state.pending_replica_heads, context.receiver, %{})) > 0 + end) + + first_token = + context.source + |> TestCluster.rpc!(:sys, :get_state, [source_shard]) + |> Map.fetch!(:replica_send_tokens) + |> Map.fetch!(context.receiver) + + :ok = + TestCluster.rpc!(context.receiver, TestCluster, :expire_replica_lane, [ + context.name, + 0, + context.source + ]) + + TestCluster.assert_eventually(fn -> + receiver_state = TestCluster.rpc!(context.receiver, :sys, :get_state, [receiver_shard]) + source_state = TestCluster.rpc!(context.source, :sys, :get_state, [source_shard]) + + Map.get(receiver_state.peer_probe_epochs, context.source) == 2 and + Map.get(source_state.replica_send_tokens, context.receiver) != first_token + end) + + second_state = TestCluster.rpc!(context.source, :sys, :get_state, [source_shard]) + second_token = Map.fetch!(second_state.replica_send_tokens, context.receiver) + assert second_token != first_token + + receiver_pid = TestCluster.rpc!(context.receiver, Process, :whereis, [receiver_shard]) + + send( + TestCluster.rpc!(context.source, Process, :whereis, [source_shard]), + peer_connect(context, receiver_pid, 1) + ) + + later_state = TestCluster.rpc!(context.source, :sys, :get_state, [source_shard]) + assert Map.fetch!(later_state.replica_send_tokens, context.receiver) == second_token + assert map_size(Map.get(later_state.pending_replica_heads, context.receiver, %{})) > 0 + + :ok = + TestCluster.rpc!(context.source, Group.TestReplicaTransport, :set_mode, [ + context.name, + :pass + ]) + + TestCluster.assert_eventually(fn -> + match?( + [{^owner, _}], + TestCluster.rpc!(context.receiver, Group, :members, [ + context.name, + "sprite", + [cluster: "org"] + ]) + ) + end) + end + + defp queued_replica_hellos(node, receiver_shard) do + {:messages, messages} = TestCluster.rpc!(node, Process, :info, [receiver_shard, :messages]) + + Enum.count(messages, fn + {:replica_hello, _, _, _, _, _, _, _} -> true + _ -> false + end) + end + + test "an ACK from a closed subscription cannot clear a reopened stream", context do + :ok = + TestCluster.rpc!(context.receiver, Group.TestReplicaTransport, :set_mode, [ + context.name, + {:capture_drop_types, [:applied]} + ]) + + owner = + TestCluster.spawn_join_in_cluster( + context.source, + context.name, + "sprite", + %{up: true}, + "org" + ) + + TestCluster.assert_eventually(fn -> + match?( + [{^owner, _meta}], + TestCluster.rpc!(context.receiver, Group, :members, [ + context.name, + "sprite", + [cluster: "org"] + ]) + ) + end) + + TestCluster.assert_eventually(fn -> + context.receiver + |> TestCluster.rpc!(Group.TestReplicaTransport, :captured, [context.name]) + |> Enum.any?(fn {_target, _shard, frame} -> elem(frame, 0) == :applied end) + end) + + {_target, 0, old_ack} = + context.receiver + |> TestCluster.rpc!(Group.TestReplicaTransport, :captured, [context.name]) + |> Enum.find(fn {_target, shard, frame} -> shard == 0 and elem(frame, 0) == :applied end) + + old_epoch = + TestCluster.rpc!(context.receiver, Group.Replica.Data, :local_cluster_epoch, [ + context.name, + "org" + ]) + + :ok = TestCluster.rpc!(context.receiver, Group, :disconnect, [context.name, "org"]) + :ok = TestCluster.rpc!(context.receiver, Group, :connect, [context.name, "org"]) + + TestCluster.assert_eventually(fn -> + TestCluster.rpc!(context.source, Group.Replica.Data, :remote_cluster_epoch, [ + context.name, + context.receiver, + "org" + ]) != old_epoch + end) + + TestCluster.assert_eventually(fn -> + context.source + |> TestCluster.rpc!(:sys, :get_state, [Group.Replica.shard_name(context.name, 0)]) + |> Map.get(:pending_replica_heads) + |> Map.get(context.receiver, %{}) + |> map_size() > 0 + end) + + :ok = + TestCluster.rpc!(context.source, Group.Transport, :incoming, [ + context.name, + context.receiver, + 0, + old_ack + ]) + + TestCluster.flush_shards(context.source, context.name) + + source_state = + TestCluster.rpc!(context.source, :sys, :get_state, [ + Group.Replica.shard_name(context.name, 0) + ]) + + assert map_size(Map.get(source_state.pending_replica_heads, context.receiver, %{})) > 0 + + :ok = + TestCluster.rpc!(context.receiver, Group.TestReplicaTransport, :set_mode, [ + context.name, + :pass + ]) + + TestCluster.assert_eventually(fn -> + match?( + [{^owner, _meta}], + TestCluster.rpc!(context.receiver, Group, :members, [ + context.name, + "sprite", + [cluster: "org"] + ]) + ) + end) + + TestCluster.assert_eventually(fn -> + context.source + |> TestCluster.rpc!(:sys, :get_state, [Group.Replica.shard_name(context.name, 0)]) + |> Map.get(:pending_replica_heads) + |> Map.get(context.receiver, %{}) + |> map_size() == 0 + end) + end + + test "sender shard restart rebuilds a lost pending head", context do + :ok = + TestCluster.rpc!(context.source, Group.TestReplicaTransport, :set_mode, [ + context.name, + {:drop_types, [:delta_batch, :heads]} + ]) + + owner = + TestCluster.spawn_join_in_cluster( + context.source, + context.name, + "sprite", + %{up: true}, + "org" + ) + + TestCluster.assert_eventually(fn -> + context.source + |> TestCluster.rpc!(:sys, :get_state, [Group.Replica.shard_name(context.name, 0)]) + |> Map.get(:pending_replica_heads) + |> Map.get(context.receiver, %{}) + |> map_size() > 0 + end) + + assert [] == + TestCluster.rpc!(context.receiver, Group, :members, [ + context.name, + "sprite", + [cluster: "org"] + ]) + + old_lane = + TestCluster.rpc!(context.source, Process, :whereis, [ + Group.Replica.shard_name(context.name, 0) + ]) + + true = TestCluster.rpc!(context.source, Process, :exit, [old_lane, :kill]) + + TestCluster.assert_eventually(fn -> + case TestCluster.rpc!(context.source, Process, :whereis, [ + Group.Replica.shard_name(context.name, 0) + ]) do + pid when is_pid(pid) -> pid != old_lane + _ -> false + end + end) + + :ok = + TestCluster.rpc!(context.source, Group.TestReplicaTransport, :set_mode, [ + context.name, + :pass + ]) + + TestCluster.assert_eventually(fn -> + match?( + [{^owner, _meta}], + TestCluster.rpc!(context.receiver, Group, :members, [ + context.name, + "sprite", + [cluster: "org"] + ]) + ) + end) + end + + test "an unacknowledged pruned stream repairs by exact snapshot", context do + name = :"replica_ack_snapshot_#{System.unique_integer([:positive])}" + + opts = [ + name: name, + shards: 1, + replica_transport: Group.TestReplicaTransport, + replicated_oplog_max_entries: 2, + replicated_anti_entropy_interval: 25, + replicated_peer_lease_timeout: 5_000 + ] + + for node <- [context.source, context.receiver] do + {:ok, _pid} = TestCluster.start_group(node, opts) + :ok = TestCluster.rpc!(node, Group, :connect, [name, "snapshot-org"]) + end + + TestCluster.assert_eventually(fn -> + length(TestCluster.rpc!(context.source, Group, :nodes, [name, "snapshot-org"])) == 2 + end) + + :ok = + TestCluster.rpc!(context.source, Group.TestReplicaTransport, :set_mode, [ + name, + {:drop_types, [:delta_batch, :heads]} + ]) + + entries = + for index <- 1..8 do + key = "member/#{index}" + + {key, + TestCluster.spawn_join_in_cluster( + context.source, + name, + key, + %{index: index}, + "snapshot-org" + )} + end + + TestCluster.flush_shards(context.source, name) + + stream_id = + TestCluster.rpc!(context.source, Group.Replica.Data, :local_stream_id, [ + name, + 0, + "snapshot-org" + ]) + + {floor, head, _applied} = + TestCluster.rpc!(context.source, Group.Replica.Data, :replica_stream_head, [ + name, + 0, + stream_id + ]) + + assert floor > 1 + assert head >= length(entries) + + :ok = + TestCluster.rpc!(context.source, Group.TestReplicaTransport, :set_mode, [ + name, + {:capture_pass, [:snapshot_commit]} + ]) + + TestCluster.assert_eventually(fn -> + Enum.all?(entries, fn {key, pid} -> + match?( + [{^pid, _meta}], + TestCluster.rpc!(context.receiver, Group, :members, [ + name, + key, + [cluster: "snapshot-org"] + ]) + ) + end) + end) + + TestCluster.assert_eventually(fn -> + context.source + |> TestCluster.rpc!(:sys, :get_state, [Group.Replica.shard_name(name, 0)]) + |> Map.get(:pending_replica_heads) + |> Map.get(context.receiver, %{}) + |> map_size() == 0 + end) + + assert Enum.any?( + TestCluster.rpc!(context.source, Group.TestReplicaTransport, :captured, [name]), + fn + {_target, _shard, {:snapshot_commit, _, ^stream_id, _, _, _, _}} -> true + _ -> false + end + ) + + commits_before = + context.source + |> TestCluster.rpc!(Group.TestReplicaTransport, :captured, [name]) + |> Enum.count(fn + {_target, _shard, {:snapshot_commit, _, ^stream_id, _, _, _, _}} -> true + _ -> false + end) + + source_shard = + TestCluster.rpc!(context.source, Process, :whereis, [ + Group.Replica.shard_name(name, 0) + ]) + + receiver_shard = + TestCluster.rpc!(context.receiver, Process, :whereis, [ + Group.Replica.shard_name(name, 0) + ]) + + # A delayed need from before the applied ACK cannot start another full snapshot. + send( + source_shard, + {:group_replica_frame, receiver_shard, + {:needs, Group.Replica.WireProtocol.version(), [{stream_id, 1, head}]}} + ) + + source_state = TestCluster.rpc!(context.source, :sys, :get_state, [source_shard]) + assert source_state.snapshot_send == nil + + assert commits_before == + context.source + |> TestCluster.rpc!(Group.TestReplicaTransport, :captured, [name]) + |> Enum.count(fn + {_target, _shard, {:snapshot_commit, _, ^stream_id, _, _, _, _}} -> true + _ -> false + end) + + :ok = + TestCluster.rpc!(context.source, Group.TestReplicaTransport, :set_mode, [ + name, + {:capture_drop_types, [:heads, :delta_batch, :snapshot_commit]} + ]) + + new_member = + TestCluster.spawn_join_in_cluster( + context.source, + name, + "member/new", + %{}, + "snapshot-org" + ) + + TestCluster.assert_eventually(fn -> + source_state = TestCluster.rpc!(context.source, :sys, :get_state, [source_shard]) + Map.has_key?(Map.get(source_state.pending_replica_heads, context.receiver, %{}), stream_id) + end) + + # This need describes the previous head. A newer pending head must not + # turn it into a full snapshot request. + send( + source_shard, + {:group_replica_frame, receiver_shard, + {:needs, Group.Replica.WireProtocol.version(), [{stream_id, 1, head}]}} + ) + + source_state = TestCluster.rpc!(context.source, :sys, :get_state, [source_shard]) + assert source_state.snapshot_send == nil + assert snapshot_commits(context.source, name, stream_id) == commits_before + + :ok = TestCluster.rpc!(context.source, Group.TestReplicaTransport, :set_mode, [name, :pass]) + drive_anti_entropy(context.source, name, 1) + + TestCluster.assert_eventually(fn -> + match?( + [{^new_member, _meta}], + TestCluster.rpc!(context.receiver, Group, :members, [ + name, + "member/new", + [cluster: "snapshot-org"] + ]) + ) + end) + end + + test "a lost snapshot commit retries after a bound, not on every need", context do + name = :"replica_ack_snapshot_retry_#{System.unique_integer([:positive])}" + + opts = [ + name: name, + shards: 1, + replica_transport: Group.TestReplicaTransport, + replicated_oplog_max_entries: 2, + replicated_anti_entropy_interval: 2_000, + replicated_peer_lease_timeout: 5_000 + ] + + for node <- [context.source, context.receiver] do + {:ok, _pid} = TestCluster.start_group(node, opts) + :ok = TestCluster.rpc!(node, Group, :connect, [name, "snapshot-org"]) + end + + TestCluster.assert_eventually(fn -> + length(TestCluster.rpc!(context.source, Group, :nodes, [name, "snapshot-org"])) == 2 + end) + + :ok = + TestCluster.rpc!(context.source, Group.TestReplicaTransport, :set_mode, [ + name, + {:drop_types, [:heads, :delta_batch]} + ]) + + for index <- 1..8 do + TestCluster.spawn_join_in_cluster( + context.source, + name, + "member/#{index}", + %{index: index}, + "snapshot-org" + ) + end + + TestCluster.flush_shards(context.source, name) + + stream_id = + TestCluster.rpc!(context.source, Group.Replica.Data, :local_stream_id, [ + name, + 0, + "snapshot-org" + ]) + + {floor, head, _applied} = + TestCluster.rpc!( + context.source, + Group.Replica.Data, + :replica_stream_head, + [name, 0, stream_id] + ) + + assert floor > 1 + + :ok = + TestCluster.rpc!(context.source, Group.TestReplicaTransport, :set_mode, [ + name, + {:capture_drop_types, [:snapshot_commit]} + ]) + + source_shard = + TestCluster.rpc!(context.source, Process, :whereis, [ + Group.Replica.shard_name(name, 0) + ]) + + receiver_shard = + TestCluster.rpc!(context.receiver, Process, :whereis, [ + Group.Replica.shard_name(name, 0) + ]) + + drive_anti_entropy(context.source, name, 1) + + TestCluster.assert_eventually(fn -> + snapshot_commits(context.source, name, stream_id) == 1 and + TestCluster.rpc!(context.source, :sys, :get_state, [source_shard]).snapshot_send == nil + end) + + need = + {:group_replica_frame, receiver_shard, + {:needs, Group.Replica.WireProtocol.version(), [{stream_id, 1, head}]}} + + for _ <- 1..3 do + send(source_shard, need) + state = TestCluster.rpc!(context.source, :sys, :get_state, [source_shard]) + assert state.snapshot_send == nil + end + + assert snapshot_commits(context.source, name, stream_id) == 1 + + :ok = + TestCluster.rpc!(context.source, TestCluster, :backdate_completed_snapshot, [ + name, + 0, + context.receiver, + stream_id, + head, + 3_000 + ]) + + send(source_shard, need) + + TestCluster.assert_eventually(fn -> + snapshot_commits(context.source, name, stream_id) == 2 + end) + + TestCluster.assert_eventually(fn -> + TestCluster.rpc!(context.source, :sys, :get_state, [source_shard]).snapshot_send == nil + end) + + source_state = TestCluster.rpc!(context.source, :sys, :get_state, [source_shard]) + + {^receiver_shard, probe_epoch} = + Map.fetch!(source_state.remote_probe_epochs, context.receiver) + + # A receiver that has lost its view must not wait for the old retry hold. + send(source_shard, peer_connect(context, receiver_shard, probe_epoch + 1, name)) + + TestCluster.assert_eventually(fn -> + snapshot_commits(context.source, name, stream_id) == 3 + end) + end + + defp peer_connect(context, receiver_pid, probe_epoch, name \\ nil) do + name = name || context.name + + {generation, revision, _epochs} = + TestCluster.rpc!(context.receiver, Group.Replica.Data, :local_replica_authority, [ + name + ]) + + {:peer_connect, receiver_pid, 0, 1, probe_epoch, generation, revision} + end + + defp snapshot_commits(source, name, stream_id) do + source + |> TestCluster.rpc!(Group.TestReplicaTransport, :captured, [name]) + |> Enum.count(fn + {_target, _shard, {:snapshot_commit, _, ^stream_id, _, _, _, _}} -> true + _ -> false + end) + end + + defp captured_record(context, stream, sequence) do + context.source + |> TestCluster.rpc!(Group.TestReplicaTransport, :captured, [context.name]) + |> Enum.find_value(fn + {target, 0, {:delta_batch, _version, runs}} when target == context.receiver -> + Enum.find_value(runs, fn + {^stream, _first, records, _head} -> + Enum.find(records, fn {seq, _mutations} -> seq == sequence end) + + _other -> + nil + end) + + _other -> + nil + end) + end + + defp drive_anti_entropy(node, name, turns) do + shard_name = Group.Replica.shard_name(name, 0) + + for _ <- 1..turns do + state = TestCluster.rpc!(node, :sys, :get_state, [shard_name]) + shard = TestCluster.rpc!(node, Process, :whereis, [shard_name]) + send(shard, {:group_replica_anti_entropy, state.anti_entropy_ref}) + + TestCluster.assert_eventually(fn -> + TestCluster.rpc!(node, :sys, :get_state, [shard_name]).anti_entropy_ref != + state.anti_entropy_ref + end) + + TestCluster.flush_shards(node, name) + end + end +end diff --git a/test/replica_snapshot_distributed_test.exs b/test/replica_snapshot_distributed_test.exs index e94a736..7e87e15 100644 --- a/test/replica_snapshot_distributed_test.exs +++ b/test/replica_snapshot_distributed_test.exs @@ -127,7 +127,8 @@ defmodule Group.ReplicaSnapshotDistributedTest do name, node_b, 0, - {:need, Group.Replica.WireProtocol.version(), stream_id, 1} + {:needs, Group.Replica.WireProtocol.version(), + [{stream_id, 1, local_head(node_a, name, stream_id)}]} ]) assert_receive {:replica_transport_paused, worker, ^name, :snapshot_chunk}, 5_000 @@ -174,7 +175,8 @@ defmodule Group.ReplicaSnapshotDistributedTest do name, node_b, 0, - {:need, Group.Replica.WireProtocol.version(), stream_id, 1} + {:needs, Group.Replica.WireProtocol.version(), + [{stream_id, 1, local_head(node_a, name, stream_id)}]} ]) TestCluster.assert_eventually(fn -> @@ -986,7 +988,8 @@ defmodule Group.ReplicaSnapshotDistributedTest do name, node_b, 0, - {:needs, Group.Replica.WireProtocol.version(), [{stream_id, next_seq}]} + {:needs, Group.Replica.WireProtocol.version(), + [{stream_id, next_seq, local_head(node_a, name, stream_id)}]} ]) TestCluster.assert_eventually( @@ -1048,7 +1051,8 @@ defmodule Group.ReplicaSnapshotDistributedTest do name, node_b, 0, - {:needs, Group.Replica.WireProtocol.version(), [{stream_id, 1}]} + {:needs, Group.Replica.WireProtocol.version(), + [{stream_id, 1, local_head(node_a, name, stream_id)}]} ]) source = TestCluster.rpc!(node_a, Process, :whereis, [Group.Replica.shard_name(name, 0)]) @@ -1118,6 +1122,13 @@ defmodule Group.ReplicaSnapshotDistributedTest do TestCluster.rpc!(node, Group.Replica.Data, :local_stream_id, [name, 0, cluster]) end + defp local_head(node, name, stream_id) do + {_floor, head, _applied} = + TestCluster.rpc!(node, Group.Replica.Data, :replica_stream_head, [name, 0, stream_id]) + + head + end + defp replica_cursor(node, name, stream_id) do TestCluster.rpc!(node, Group.Replica.Data, :replica_cursor, [name, 0, stream_id]) end diff --git a/test/support/jepsen_cursor_capture.ex b/test/support/jepsen_cursor_capture.ex index df8aceb..d3937f3 100644 --- a/test/support/jepsen_cursor_capture.ex +++ b/test/support/jepsen_cursor_capture.ex @@ -72,6 +72,16 @@ defmodule Group.JepsenCursorCapture do :ok end + def peer_head_state(shard, remote_node) do + state = :sys.get_state(Group.Replica.shard_name(:jepsen_group, shard)) + + %{ + connected?: Map.has_key?(state.remote_shards, remote_node), + pending_heads: state.pending_replica_heads |> Map.get(remote_node, %{}) |> map_size(), + send_token: Map.get(state.replica_send_tokens, remote_node) + } + end + def freeze do for shard <- 0..1, do: :sys.suspend(Group.Replica.shard_name(:jepsen_group, shard)) :ok diff --git a/test/support/replica_model_scheduler.ex b/test/support/replica_model_scheduler.ex index fc766ff..88d61f1 100644 --- a/test/support/replica_model_scheduler.ex +++ b/test/support/replica_model_scheduler.ex @@ -214,7 +214,8 @@ defmodule Group.ReplicaModelScheduler do converged?(state) end, timeout: 15_000, - interval: 25 + interval: 25, + diagnostic: fn -> convergence_diagnostic(state) end ) Enum.each(state.nodes, fn {_id, node} -> @@ -436,6 +437,70 @@ defmodule Group.ReplicaModelScheduler do end) end + defp convergence_diagnostic(state) do + expected_regs = ReplicaLifecycleModel.expected_registrations(state.model) + expected_members = ReplicaLifecycleModel.expected_memberships(state.model) + pid_to_owner = Map.new(state.owners, fn {owner_id, %{pid: pid}} -> {pid, owner_id} end) + + for {node_id, node} <- state.nodes do + regs = + for {cluster, key} = scope <- state.model.seen_registration_keys, + actual = + (case TestCluster.rpc!(node, Group, :lookup, [ + state.name, + key, + cluster_opts(cluster) + ]) do + nil -> nil + {pid, meta} -> {Map.get(pid_to_owner, pid, {:unknown_pid, pid}), meta} + end), + actual != Map.get(expected_regs, scope) do + {scope, Map.get(expected_regs, scope), actual} + end + + members = + for {cluster, key} = scope <- state.model.seen_membership_keys, + actual = + node + |> TestCluster.rpc!(Group, :members, [state.name, key, cluster_opts(cluster)]) + |> Enum.map(fn {pid, meta} -> + {Map.get(pid_to_owner, pid, {:unknown_pid, pid}), meta} + end) + |> Enum.sort(), + actual != Map.get(expected_members, scope, []) do + {scope, Map.get(expected_members, scope, []), actual} + end + + shards = + for shard <- 0..(Keyword.fetch!(state.group_opts, :shards) - 1) do + shard_state = + TestCluster.rpc!(node, :sys, :get_state, [Group.Replica.shard_name(state.name, shard)]) + + {shard, + %{ + remote_shards: Map.keys(shard_state.remote_shards), + peer_last_seen: Map.keys(shard_state.peer_last_seen), + send_tokens: Map.keys(shard_state.replica_send_tokens), + pending: + Map.new(shard_state.pending_replica_heads, fn {peer, streams} -> + {peer, map_size(streams)} + end), + heads: + TestCluster.rpc!(node, Group.Replica.Data, :replica_stream_heads, [ + state.name, + shard + ]), + cursors: + TestCluster.rpc!(node, :ets, :tab2list, [ + Group.Replica.Data.replica_cursor_table(state.name, shard) + ]) + }} + end + + {node_id, %{registrations: regs, memberships: members, shards: shards}} + end + end + defp registrations_match?(node, name, keys, expected, pid_to_owner) do Enum.all?(keys, fn {cluster, key} = scope -> actual = diff --git a/test/support/test_cluster.ex b/test/support/test_cluster.ex index 697a16f..72e3085 100644 --- a/test/support/test_cluster.ex +++ b/test/support/test_cluster.ex @@ -112,6 +112,104 @@ defmodule Group.TestCluster do :ok end + @doc false + def backdate_discovery_hello(name, shard, remote_node, milliseconds) do + replica = Process.whereis(Group.Replica.shard_name(name, shard)) + + :sys.replace_state(replica, fn state -> + {last_sent, authority} = Map.fetch!(state.discovery_hello_last_sent, remote_node) + + %{ + state + | discovery_hello_last_sent: + Map.put( + state.discovery_hello_last_sent, + remote_node, + {last_sent - milliseconds, authority} + ) + } + end) + + :ok + end + + @doc false + def backdate_completed_snapshot(name, shard, remote_node, stream_id, head, milliseconds) do + replica = Process.whereis(Group.Replica.shard_name(name, shard)) + key = {remote_node, stream_id, head} + + :sys.replace_state(replica, fn state -> + {:sent, sent_at} = Map.fetch!(state.snapshot_send_offsets, key) + + %{ + state + | snapshot_send_offsets: + Map.put(state.snapshot_send_offsets, key, {:sent, sent_at - milliseconds}) + } + end) + + :ok + end + + @doc false + def forget_replica_peer_route(name, shard, remote_node) do + replica = Process.whereis(Group.Replica.shard_name(name, shard)) + + :sys.replace_state(replica, fn state -> + %{ + state + | remote_shards: Map.delete(state.remote_shards, remote_node), + peer_last_seen: Map.delete(state.peer_last_seen, remote_node) + } + end) + + :ok + end + + @doc false + def expire_replica_lane_without_probe(name, shard, remote_node) do + # Reproduce the state immediately after lease expiry while withholding the + # peer_connect that the timer would normally send in the same callback. + replica = Process.whereis(Group.Replica.shard_name(name, shard)) + + state = + :sys.replace_state(replica, fn state -> + probe_epoch = Map.get(state.peer_probe_epochs, remote_node, 0) + 1 + + %{ + state + | remote_shards: Map.delete(state.remote_shards, remote_node), + peer_last_seen: Map.delete(state.peer_last_seen, remote_node), + peer_probe_epochs: Map.put(state.peer_probe_epochs, remote_node, probe_epoch), + pending_peer_probes: Map.put(state.pending_peer_probes, remote_node, probe_epoch), + replica_receive_tokens: Map.delete(state.replica_receive_tokens, remote_node), + pending_replica_acks: Map.delete(state.pending_replica_acks, remote_node), + anti_entropy_ref: make_ref() + } + end) + + Group.Replica.Data.delete_replica_cursors_for_origin(name, shard, remote_node) + Group.Replica.Data.purge_node(name, shard, remote_node) + Group.Replica.Data.purge_registry_claims_for_origin(name, shard, remote_node) + + if Group.Replica.Data.expire_remote_replica_lane(name, shard, remote_node) == :node_retired do + Group.Replica.Data.purge_cluster_node(name, remote_node) + end + + {replica, Map.fetch!(state.peer_probe_epochs, remote_node), state.anti_entropy_ref} + end + + @doc false + def forget_peer_connect_ack(name, shard, remote_node) do + replica = Process.whereis(Group.Replica.shard_name(name, shard)) + + :sys.replace_state(replica, fn state -> + %{state | peer_connect_ack_seen: Map.delete(state.peer_connect_ack_seen, remote_node)} + end) + + :ok + end + @doc false def put_pending_registry_reprojection(replica, remote_node, cluster, key) do :sys.replace_state(replica, fn state -> diff --git a/test/support/test_replica_transport.ex b/test/support/test_replica_transport.ex index 23ac30d..a078d55 100644 --- a/test/support/test_replica_transport.ex +++ b/test/support/test_replica_transport.ex @@ -17,6 +17,8 @@ defmodule Group.TestReplicaTransport do :drop_types, :duplicate_types, :capture_drop, + :capture_drop_types, + :capture_busy_types, :capture_pass, :busy_delta_above ]) or @@ -106,6 +108,22 @@ defmodule Group.TestReplicaTransport do if message_type(message) in types, do: capture(group, target_node, shard, message) :ok + {:capture_drop_types, types} -> + if message_type(message) in types do + capture(group, target_node, shard, message) + :ok + else + forward(group, target_node, shard, message) + end + + {:capture_busy_types, types} -> + if message_type(message) in types do + capture(group, target_node, shard, message) + :busy + else + forward(group, target_node, shard, message) + end + {:capture_pass, types} -> if message_type(message) in types, do: capture(group, target_node, shard, message) forward(group, target_node, shard, message)