Spock Release Notes
Spock 6.0.0-beta.2
This section lists what changed since 6.0.0-beta.1, for those testing the beta. Each item is covered in full in the 6.0.0 notes below. This section will be removed at general availability.
Upgrading from beta 1
- Beta 1 and beta 2 both install extension version
6.0.0, but the catalog differs between them, so a plainALTER EXTENSION spock UPDATEdoes nothing on a beta 1 node. Update a beta 1 node in two steps instead:
- Pause DDL on the cluster.
- Install the beta 2 binaries and restart the node.
-
In every database where Spock is installed, run:
ALTER EXTENSION spock UPDATE TO '6.0.0-beta1-to-beta2'; ALTER EXTENSION spock UPDATE TO '6.0.0'; -
Repeat for each node, then resume DDL.
-
On each primary, check that every Spock slot has the
failoverflag:SELECT slot_name, failover FROM pg_replication_slots WHERE plugin = 'spock_output';For any slot still showing
false, pause the subscriber that holds it and runSELECT spock.slot_enable_failover();on the primary.
The first command adds the objects beta 2 needs, removes
spock.wait_for_apply_worker(), and sets the failover flag on existing
slots. The second only sets the version name back to 6.0.0.
The failover flag matters on PostgreSQL 17 and 18, where native slot sync
copies only flagged slots to a standby. Beta 1's upgrade from 5.x did not
flag the slots that already existed, so a beta 1 node upgraded from 5.x may
have slots that would be missing on the standby after a failover. Slots
created by 6.0.0 already have the flag. The first command flags the rest,
but skips any slot in use at that moment and prints a NOTICE naming it. On
PostgreSQL 16 and older, or with Spock's own slot sync worker on 17, the
flag is not used and step 5 can be skipped.
Until a database has been updated, AutoDDL, adding tables to a replication
set and structure sync fail there with catalog "spock.reserved_object"
does not exist. Running the commands on a node that already has beta 2's
catalog changes nothing. These two update steps will be removed at general
availability.
* The upgrade from 5.0.x now starts at 5.0.12. A 5.0.11 node still
upgrades in one ALTER EXTENSION spock UPDATE, through 5.0.12.
* On a server with PostgreSQL's fix for CVE-2026-6471, spock_output must be
listed in output_plugin_libraries. See Upgrading below.
GUC changes
- New
spock.sync_timeout(seconds, default0,USERSET): the time limit for a single wait in the node management routines.0keeps each routine's built-in limit. - New
spock.failover_slots_naptime(milliseconds, default1000,SIGHUP) andspock.failover_slots_feedback_naptime(milliseconds, default10000,SIGHUP): how often Spock's failover-slots worker syncs slots. It used to wait a fixed 60 seconds. spock.pause_timeout: the 300-second maximum is removed.spock.restart_delay_default,spock.restart_delay_on_exceptionandspock.output_delayare now declared in milliseconds, so they accept values with units such as'5s'.spock.restart_delay_defaultnow applies to every apply-worker restart, including after a lost connection or a temporary error.
SQL function and catalog changes
- New table
spock.reserved_object, listing the schemas and extensions Spock keeps out of the structure-sync dump, out of replication sets, or out of DDL replication. - New
spock.reserved_object_add(p_name name, p_kind text, p_exclude_from_dump boolean DEFAULT true, p_block_in_repset boolean DEFAULT true, p_replicate_ddl boolean DEFAULT NULL). - New
spock.reserved_object_remove(p_name name, p_kind text). - New
spock.slot_enable_failover() RETURNS integer, called by the 5.x upgrade. - Removed
spock.wait_for_apply_worker(p_subbid bigint, timeout int). spock.repset_add_all_tables()now skips a table it cannot add, with a warning, instead of failing the whole call.
Other changes
- Direct DDL against the
spockandsnowflakeschemas is refused while DDL replication is on. - Large objects stored by lolor can be replicated, and Zodan
add_nodecopies them to the new node. - Zodan
add_nodecan add a 6.0.0 node to a cluster of 5.0.9 or later nodes, and copies the source node's replication sets to the new node. forward_originscannot be active on a subscription while another subscription on the node is enabled.- The apply worker restarts and tries again after a temporary error, such as a deadlock or lock timeout, instead of treating it as a data exception.
pg_upgradeof a Spock node on PostgreSQL 17 and later no longer fails withresource manager with ID 144 not registered.- Bug fixes from 5.0.11 and 5.0.12.
- Builds against PostgreSQL 19 beta 3 and 4.
Spock 6.0.0
The on-disk catalog format and the GUC surface both change in this release;
see Upgrading below before running ALTER EXTENSION spock UPDATE.
Highlights
- Catastrophic node failure recovery — new ACE-based workflow with origin-preserving repair prevents silent data divergence after node loss.
- Progress tracking moved to WAL + shared memory — eliminates
heavyweight catalog writes in the apply hot path;
spock.progressis now a view backed byspock.apply_group_progress(). - More granular conflict classification — seven conflict types (up from four), with origin-aware suppression and DELETE conflicts now resolved through timestamp-based resolution.
- Automatic node metadata propagation — a node's
location,country, andinfo(including its conflict-resolutiontiebreaker) now propagate automatically to every direct subscriber on change, via the newspock.node_alter()function or a plainUPDATE; no more silently disagreeing tiebreakers from a forgottenspock.node_refresh_info(). - Per-subscription conflict statistics on PostgreSQL 18+ via a custom pgstat kind.
- Liveness and feedback refactor — TCP keepalive replaces the fragile
wal_sender_timeoutworkaround; newspock.apply_idle_timeoutGUC. - Logical slot failover uses native PostgreSQL slotsync on PG17+ / PG18+.
- Rolling upgrade support — protocol negotiation (v4 for 5.0.x, v5 for 6.0+) enables zero-downtime rolling upgrades.
- Exception handling refactor — stable behaviour under TRANSDISCARD /
SUB_DISABLE; original error messages preserved in
spock.exception_log. - Replay queue spills to disk instead of refetching from the publisher.
- Reserved schemas and extensions are configurable — a new
spock.reserved_objectcatalog controls which schemas and extensions are left out of structure sync, replication sets and DDL replication. - Mixed-version node addition — Zodan
add_nodecan add a 6.0.0 node to a cluster of 5.0.9 or later nodes. - Large object replication — tables of the lolor extension can join a
replication set, and Zodan
add_nodecopies large objects to the new node. - PostgreSQL 19 support — compatibility work across core patches and
version-specific API changes (replication-origin session state folded into
replorigin_xact_state,CLUSTERfolded intoREPACK, thepgstatand recovery-conflict signalling changes, and theRepOriginIdrename). Spock builds against PostgreSQL 19 beta 3 and 4. This support is preliminary: PostgreSQL 19 is still in beta, and Spock on 19 is not supported for production use.
Catastrophic node failure recovery
Spock 6.0 documents and supports a structured recovery workflow for the scenario where a node fails permanently and one or more surviving nodes are lagging behind. Using the Active Consistency Engine (ACE), operators can:
- Identify which rows are missing on lagging survivors using
table-diff --preserve-origin. - Repair the lagging node from a fully-synchronized survivor using
table-repair --recovery-mode --source-of-truth --preserve-origin.
The --preserve-origin flag is important: it ensures that repaired rows
retain their original origin ID and commit timestamp, so replication
metadata stays correct and the cluster remains conflict-free. Without it,
repaired rows would appear as local changes and could trigger false
conflicts.
Both single-node and multi-node failure scenarios are covered. See Catastrophic Node Failure Recovery for the full guide.
Replication progress tracking
Replication progress is no longer tracked in the spock.progress catalog
table. Spock now tracks it live in shared memory and persists a snapshot
to $PGDATA/spock/resource.dat. On apply-worker startup, Spock
reconciles the snapshot against the durable replication-origin LSN: if
resource.dat is stale, remote_commit_lsn is advanced from the origin
and stale timestamp fields are cleared (to be refreshed by subsequent
apply).
A custom resource manager (id 144) emits one WAL record per
SpockGroupHash entry at each resource.dat dump event, visible via
pg_waldump for debugging, not used for state recovery.
The catalog surface has been restructured accordingly:
spock.progressis now a view over the newspock.apply_group_progress()set-returning function, filtered to the current database. Code that read the previous table of the same name continues to work, but writes are no longer possible. Useapply_group_progress()directly (or the view) for live data, and use the newspock.read_peer_progress(slot, provider_node_id, subscriber_node_id)function during slot snapshot import.spock.lag_trackerview redefined. The oldlast_received_lsncolumn has been renamed tocommit_lsnto reflect what it actually is, and a newreceived_lsncolumn reports the last LSN sent by the publisher (matching thepgoutputprotocol convention).remote_insert_lsnis now updated on every incoming WAL record rather than only on COMMIT, so lag readings are more responsive.
Exception handling improvements
- Fixed TRANSDISCARD / SUB_DISABLE handling during the commit phase in retry mode, preventing apply-worker death loops.
- Eliminated "(unknown action)" in error-context strings and NULL error
messages in
exception_log. The originally-failing row now carries the real error message; bystander rows in the same transaction refer to the entry containing the root cause message. - Added
initial_operationfield to track which DML operation caused a transaction discard. - Subscriptions stuck in an unrecoverable synchronization stage are now disabled rather than restarted indefinitely.
Exception replay: spill to disk
spock.exception_replay_queue_size has changed semantics.
- Before: a single byte threshold (default 4 MiB ≈ 4194304). When the in-memory queue exceeded the threshold, Spock could not replay the transaction and would refetch it from the publisher.
- Now: a soft cap (default
4MB, unitMB). When the queue exceeds the cap, subsequent entries spill to a temporary file on disk. Set to0to disable spilling and allow unbounded memory use (improves throughput for very large transactions; allocation failure falls back tospock.exception_behaviour).
spock.exception_behaviour continues to control the action on
unrecoverable errors:
| Value | Action |
|---|---|
discard |
Skip the failed transaction and continue replication. |
transdiscard |
Roll back the failed transaction and continue. |
sub_disable |
Disable the subscription and exit cleanly. |
Logical Slot Failover
- On PostgreSQL 17+, Spock creates all logical replication slots with
the
FAILOVERflag, allowing PostgreSQL's built-in slotsync worker (sync_replication_slots = on) to automatically synchronize them to physical standbys. - On PostgreSQL 18+, Spock's own
spock_failover_slotsbackground worker is no longer registered. The native PostgreSQL slotsync worker fully replaces it. See the Logical Slot Failover section of the configuration guide for requiredpostgresql.confsettings. - On PostgreSQL 17, Spock's worker remains active but automatically
yields to the native slotsync worker if
sync_replication_slots = onis set, preventing conflicts.
More granular conflict classification
v5.0 classified conflicts into four types. v6.0 expands this to seven, adopting PostgreSQL-native conflict-type naming:
| v5.0 type | v6.0 type(s) | What changed |
|---|---|---|
insert_exists |
insert_exists |
No change. |
update_update |
update_exists / update_origin_differs |
Spock now distinguishes between an update whose local origin matches the incoming origin (suppressed silently) and one whose local origin differs (evaluated through resolution). update_exists is used for update constraint violations. |
update_delete |
update_missing |
Renamed for clarity. |
delete_delete |
delete_missing / delete_origin_differs |
Spock now logs whether the missing row was last modified by a different origin. |
| (new) | delete_exists |
A remote DELETE arrives, but the local row has been updated more recently than the delete's timestamp. The newer local version is preserved. |
Existing spock.resolutions data is migrated automatically during
ALTER EXTENSION spock UPDATE (see Catalog changes below).
DELETE operations now use conflict resolution
In v5.0, when a DELETE found a matching local row, it was applied
immediately without consulting conflict resolution. In v6.0, DELETE
operations go through timestamp-based resolution, just like updates. This
enables the new delete_exists classification — Spock can determine
whether a delete should be applied or whether a newer local version should
be preserved.
Automatic node metadata propagation
A node's location, country, and info (including a tiebreaker key
used to break same-timestamp conflicts -- see Tiebreaker in
conflict_types.md) previously had to be manually
re-fetched onto every other node with spock.node_refresh_info() after
any change; a node left unrefreshed would silently keep disagreeing with
the rest of the cluster about that peer's tiebreaker.
v6.0 propagates such a change automatically to every node that
subscribes to it directly, via a new AFTER UPDATE trigger on
spock.node and a custom logical-replication message -- deliberately
direct-subscribers-only, not forwarded through a cascade topology, so it
introduces no message-ordering dependency on the existing replication
stream. The new spock.node_alter() function is the recommended way to
make the change: it validates a tiebreaker value (must be a whole JSON
number that fits a 32-bit integer) and merges an info patch rather than
replacing it wholesale; a raw UPDATE spock.node on the node's own row
continues to work and propagates the same way. spock.node_refresh_info()
updates a node whose cached row missed a change, for example because the
transaction carrying it was skipped or discarded, and spock.sub_resync_table()
warns when the provider's row differs from the cached one. See
Automatic Node Metadata Propagation
and spock.node_alter for
details.
Cascade replication origin tracking
v6.0 adds support for tracking and forwarding replication origins in cascade topologies (Node A → Node B → Node C). The origin name is now included in ORIGIN messages when the protocol version is 5 or higher. This ensures that conflict evaluation on Node C has accurate origin information even when changes pass through intermediate Node B.
If Node C also has a disabled subscription directly to Node A, Spock now advances that subscription's replication origin as forwarded changes from Node A arrive through Node B. When the direct subscription is later enabled, it starts from where the forwarded changes left off, so rows already received through Node B are not sent again.
A subscription cannot activate forward_origins while another subscription
is already enabled on the same node, and a subscription cannot be enabled
while another has forwarding active; clear forward_origins
(spock.sub_alter_options) or disable the other subscription first. See
Origin Forwarding for
details.
Per-subscription conflict statistics
On PostgreSQL 18+, Spock registers a custom pgstat kind
(SPOCK_PGSTAT_KIND_LRCONFLICTS) to count replication conflicts per
subscription. New SQL functions
spock.get_subscription_stats(subid, peer_node_id) and
spock.reset_subscription_stats(subid) expose per-subscription apply
statistics. On PostgreSQL versions before 18 these functions raise a
feature-not-supported error.
Liveness and feedback refactoring
The wal_sender_timeout-based liveness mechanism has been replaced with a
cleaner two-layer model:
- TCP keepalive as the primary liveness detector (detects dead network or host in ~25 seconds).
wal_sender_timeout=0on replication connections so the walsender never disconnects due to missing feedback.spock.apply_idle_timeout(default 300s) as a subscriber-side safety net for hung-but-connected walsenders.
This replaces the maybe_send_feedback() workaround that coupled
subscriber behaviour to a server GUC it cannot control.
Read-only mode refactoring
Read-only mode has been refactored to use PostgreSQL's native
DefaultXactReadOnly mechanism instead of manual command-type filtering:
READONLY_USERrenamed toREADONLY_LOCAL(the old name is retained as a backward-compatible alias) to clarify that it restricts local (non-replication) connections.- In
READONLY_ALLmode, the apply worker sends keepalive feedback but skips applying new transactions. - Read-only mode changes are detected before applying the next
transaction (
CHECK_FOR_INTERRUPTSmoved to top of the replay loop).
Rolling upgrade via protocol negotiation
Spock now negotiates an output-plugin protocol version between publisher and subscriber (Protocol 4 for v5.0.x, Protocol 5 for v6.0+). This enables zero-downtime rolling upgrades: a v6.0 subscriber can apply from a v5 publisher and vice versa, with protocol-5-only features (such as cascade origin forwarding) activated only when both sides support them.
An incompatible upstream is now rejected early: create_subscription()
runs the version check at DDL time when synchronize_structure is
requested, and a structure sync from a newer-major upstream is refused
(pg_dump cannot read a newer server) instead of failing opaquely after a
slot and snapshot were created.
Zodan: data-loss fix when adding a node under write load
When adding a node while existing nodes were handling writes, the new node could permanently miss rows or accumulate stale updates. The root cause was that apply workers were not paused during slot creation (sub-second operation), so the captured resume LSN could fall before the COPY boundary. The fix pauses apply workers during slot creation and captures the resume LSN deterministically.
Additional Zodan fixes in 6.0:
- Slots and subscriptions are now named consistently as
sub_{provider}_{subscriber}via the newspock.gen_sub_name()helper. remove_nodedrops subscriptions from thespock.subscriptioncatalog instead of guessing names.- Phase 3 catch-up wait aligned and
spock.sync_event()parameterised to avoid the Phase 9 handshake deadlock. - Mixed versions —
add_nodecan add a 6.0.0 node to a cluster whose nodes run 5.0.9 or later, instead of requiring every node to be upgraded first. Existing nodes may already be mixed. The new node must run the same or a newer version than every existing node. - Replication sets are copied — the new node gets the source node's replication sets, including user-created sets, column lists, row filters, and tables removed from a set. Before, the subscriptions named only the three built-in sets, so tables in a user-created set were neither sent nor received.
- Large objects — the new node may have lolor installed as long as its tables are empty, and must have it when the source replicates lolor tables. The initial data sync copies the large objects.
- Longer waits — the new
spock.sync_timeoutGUC raises the time limit on each wait in the node management routines, which was fixed at about three minutes and too short for large databases.spock.pause_timeoutno longer has a 300-second maximum. - The old Python versions,
zodan.pyandzodremove.py, are removed. Use the SQL procedures.
AutoDDL improvements
AutoDDL has been refactored and hardened:
- Centralized in
spock_autoddl.c— all AutoDDL plumbing moved into a single module for maintainability. - Simplified control flow — dropped the recursive tag-matching heuristic in favour of clearer early-exit filters.
- Extension script guard — internal subcommands during
CREATE EXTENSION/ALTER EXTENSION UPDATEare no longer replicated. - Transaction guards — AutoDDL hook no longer fails when utility
commands (e.g.,
CLUSTERwithout a table name) execute outside a transaction context. - Filtered non-transactional DDL — commands that cannot execute inside
transaction blocks on subscribers (e.g.,
CLUSTERon partitioned tables,VACUUM) are filtered on the publisher side before being queued. - Temporary-relation DDL not replicated —
CREATE TABLE … (LIKE temp)andCREATE TABLE … AS SELECT … FROM tempare skipped, since the temporary relation does not exist on the subscriber. - Replication-set stickiness on
ALTER TABLE— withspock.include_ddl_repsetenabled,ALTER TABLEno longer evicts a table from a user-defined replication set unless it adds or drops a primary key / replica identity. InlineADD COLUMN … PRIMARY KEYandREPLICA IDENTITY FULLtables are now classified correctly. - Safe
search_pathinterpolation — replicated DDL ships a validSET search_patheven when the session path is empty. - Direct DDL against
spockandsnowflakeis refused (behaviour change) — whilespock.enable_ddl_replicationis on, a statement whose target schema is one of the built-in extension-owned schemas now fails outright:
ERROR: cannot run DDL against schema snowflake while DDL replication is enabled
DETAIL: The schema is managed by an extension and its objects are not replicated.
HINT: For a deliberate node-local change, run the statement with spock.enable_ddl_replication = off.
These schemas are populated by extension scripts, where AutoDDL is
already suppressed, so a runtime statement against them is almost always
a mistake. Previously such a statement half-applied: the command text
replicated to peers while the relation was silently kept out of every
replication set. See Upgrading for what this affects.
* An extension's own DDL during DROP EXTENSION stays node-local —
with spock.allow_ddl_from_functions on, DDL an extension ran while being
dropped was queued for replication, even though every peer runs the same
cleanup when it applies the DROP EXTENSION. For lolor this left peers
unable to drop the extension and stalled apply. The DROP EXTENSION
itself still replicates.
Reserved schemas and extensions
Spock keeps some schemas and extensions out of the structure-sync dump and
out of replication sets. That list used to be fixed in the code. It is now
the spock.reserved_object catalog, with three settings per row:
exclude_from_dump— left out of the structure-sync dump.block_in_repset— its tables may not be added to a replication set.replicate_ddl(schemas only) — whenfalse, DDL against the schema runs only on the node where it is issued.
The built-in rows cannot be changed or removed:
| Object | exclude_from_dump | block_in_repset | replicate_ddl |
|---|---|---|---|
spock |
yes | yes | yes |
snowflake |
yes | yes | yes |
lolor |
yes | no | yes |
coldfront |
yes | no | yes |
pgedge_ace |
yes | yes | no |
Add your own with spock.reserved_object_add() and remove them with
spock.reserved_object_remove(). Your rows are kept across dump and
restore. See
Reserved Schemas and Extensions
for details.
Exclusion from the dump now applies whether or not the object exists on the subscriber. Before, an object present only on the provider was never excluded.
repset_add_all_tables() skips tables it cannot add
spock.repset_add_all_tables() used to fail, and add nothing, when any one
table had no replica identity index. It now adds the tables it can and
prints a warning naming each table it skipped. Tables in a reserved schema
or owned by a reserved extension are skipped the same way. Naming a
reserved schema in the call is still an error, and
spock.repset_add_table() still fails for a table it cannot add.
The message now says "replica identity index" instead of blaming a missing
primary key, since a table with a primary key and REPLICA IDENTITY FULL
or NOTHING is also rejected.
Memory and stability
- Replay queue spills to disk (see Exception replay above) instead of triggering a worker restart, eliminating the unpredictable restarts and replication lag spikes seen with large transactions.
- Resource leak fixes:
- Row filter processing no longer leaks resource-owner entries when many rows are filtered out.
remotetupleis now freed after use in the conflict reporting path.- Shared memory initialization centralized in
spock_shmem.cwith proper handling of startup, checkpointer, and background-worker processes.
Bug fixes specific to 6.0
- WAL recovery crash after apply-worker failure — when an apply worker crashed and triggered a server restart, WAL recovery could fail due to improperly initialized shared memory structures. Shared memory init has been refactored to handle startup, checkpointer, and background-worker processes correctly.
remote_insert_lsnlost after crash recovery — the value was not included in WAL records and defaulted to 0 after crash recovery, breaking lag tracking inspock.lag_tracker.- Generated columns — tables containing
GENERATED ALWAYS AS … STOREDcolumns caused replication errors because generated columns were treated as regular columns during COPY, protocol messages, default-value filling, and conflict resolution. Now properly excluded. - Replica index lookup with
REPLICA IDENTITY FULL— cached index information was not invalidated before callingRelationGetReplicaIndex(), causing incorrect index usage when switching between replica identity modes. - Hash table iteration safety — fixed removal-during-iteration
corruption, missing NULL checks for
HASH_FIXED_SIZEtables, and torn reads of 64-bit counters without spinlock protection. - Skip LSN mechanism — a successfully skipped transaction in
SUB_DISABLEmode would incorrectly disable the subscription again due to LSN mismatch between the skip point (BEGIN LSN) and the clear point (COMMIT LSN). Exception-handling state is now cleared after a successful skip. - Table sync failure detection — COPY failures during table sync were
silently swallowed and the sync worker retried indefinitely.
PQgetResult()is now called afterPQputCopyEnd()to detect failures, and a newSYNC_STATUS_FAILEDstate stops retry loops and surfaces clear error messages. - Un-appliable TRUNCATE loop — in
discardmode a TRUNCATE that failed on the subscriber (e.g. the relation does not exist) crash-looped the apply worker. It is now wrapped in a subtransaction, logged, and discarded like other DML. pg_upgradefailed on PostgreSQL 17 and later —pg_upgradedecodes each logical slot's remaining WAL to check it is drained, and Spock's resource manager was not registered in the serverspg_upgradestarts. The check failed withresource manager with ID 144 not registered.- Apply idle timeout treated as a data exception — when
spock.apply_idle_timeoutfired in the middle of a transaction, the error went down the data-exception path, so the subscription could be disabled or a good transaction discarded. It is now handled like a lost connection. - Clock skew in progress tracking — an origin commit timestamp ahead of
the subscriber's clock, or forwarded transactions whose timestamps do not
rise with the provider's WAL position, tripped an assertion in
assert-enabled builds.
remote_commit_lsnnow advances independently of the timestamp fields. - Replay entries restored from the spill file were not terminated — a truncated or malformed record read back from disk could be scanned past the end of its buffer instead of being rejected.
- Peer sessions use a fixed time format — Spock reads commit timestamps
from other nodes as text, and their meaning depended on each side's
DateStyleandTimeZone. Where the two differed, the same text could mean a different moment, which matters when a timestamp picks the LSN to advance a slot to duringadd_node. Peer connections now requestdatestyle=ISO, and a peer timestamp without an explicit UTC offset is rejected. - Initial sync with a row filter on a table with a dropped column — rows were read at the wrong offsets after a dropped column, giving corrupt rows or a crash.
- Sync worker lost the original error — a failure while creating the
slot during table sync was reported as
errstart was not calledinstead of the real error.
Security
- DSN password obfuscation — passwords in DSN connection strings are no longer visible in error messages or log output. A message-filter hook wipes password characters from error messages before they are written.
Other notable changes
- Protocol cleanup — removed unused SPI API support from the Spock protocol; fixed multiple protocol parameter validation issues.
- Prevented adding tables from ignored schemas to replication sets.
- Snowflake removal — removed snowflake-related functions and documentation since it is now in its own repository.
- SpockCtrl removal — removed the SpockCtrl utility code.
- Windows support removal — dropped Win32-specific code paths; Spock targets POSIX-only environments.
--exclude-extension— extended use of the pg_dump option (PostgreSQL 17+).- PostgreSQL 18 support — compatibility work across core patches, initdb, row filters, and compiler warnings.
- Source tree restructured under
src/andinclude/directories. - Packaging — the repository now builds the
spock60RPM and DEB packages itself. Debian bullseye, which has reached end of life, is no longer built. - Documentation — a new guide on sizing PostgreSQL resources for Spock
by node and database count (see Sizing), a
SECURITY.mdexplaining how to report vulnerabilities, and theoutput_plugin_librariesrequirement in the README.
GUC changes
New
spock.apply_idle_timeout(int seconds, default300,SIGHUP) — safety net for detecting a hung walsender that keeps the TCP connection alive but stops sending data. The timer resets on any received message. Set to0to disable and rely solely on TCP keepalive for liveness detection.spock.output_delay(int milliseconds, default0, range 0–60000,SIGHUP) — artificial delay in the publisher-side output plugin. Used to reproduce conflict and lag scenarios in tests.spock.resolutions_retention_days(int days, default100,SIGHUP) — TTL for rows inspock.resolutions. Rows older than this are deleted periodically by the apply worker. Set to0to disable automatic cleanup. Complementsspock.cleanup_resolutions(days)and the new index onresolutions(log_time).spock.enable_quiet_mode(bool, defaultoff,SIGHUP) — when enabled, downgrades DDL replication INFO/WARNING messages to LOG level and suppresses dependent-object reporting inDROP CASCADEoperations. Intended for regression tests and production environments where less verbose output is desired.spock.log_verbosity(enumnormal/debug1/debug2, defaultnormal,SUSET) — promotes Spock's ownDEBUG1/DEBUG2messages toLOGso they survive a normallog_min_messages.spock.apply_change_logging(enumnone/key_only/verbose, defaultnone,SUSET) — emits per-change JSON from the apply worker (action, table, key, origin, commit timestamp; plus row data inverbose).spock.read_retry_countallows for configuring the number of retries when an expected row is not found. The default matches the previously hard-coded value of 5 (there is 1ms of sleep between each retry). Setting it to 0 disables retries. This helps when a node has a large lag and we do not want to slow down processing.spock.sync_timeout(int seconds, default0,USERSET) — the time limit for one wait in the node management routines: for a sync event to arrive, for a peer to catch up, or for a subscription to start replicating. Raise it on large databases, where catching up takes longer than the built-in limits allow.0keeps each routine's built-in limit. It can be set for one operation, for exampleSET spock.sync_timeout = '2h'.spock.failover_slots_naptime(int milliseconds, default1000, range 1–3600000,SIGHUP) — how long Spock's failover-slots worker sleeps between slot sync passes. It used to be a fixed 60 seconds, which let a standby's slots fall up to a minute behind.spock.failover_slots_feedback_naptime(int milliseconds, default10000, range 1–3600000,SIGHUP) — the shorter retry interval the failover-slots worker uses while waiting for the standby to receive the WAL a slot needs. Both failover-slots settings matter only where Spock's own worker runs: PostgreSQL 15 and 16, and 17 whensync_replication_slotsis off.
Removed
spock.use_spi— no longer used; the legacy SPI apply path is gone.spock.feedback_frequency— replaced by internal pacing.spock.batch_inserts- unused.
Changed
spock.exception_replay_queue_size: default4194304→4, unit changed from bytes to MB, semantics changed from a hard threshold (with publisher refetch) to a soft cap (with disk spill). See above.spock.pause_timeout: the 300-second maximum is removed, for very large or long-running transactions. The default is still 10 seconds.spock.restart_delay_default,spock.restart_delay_on_exceptionandspock.output_delayare now declared in milliseconds, so they accept values with units ('5s') andSHOWreports the unit. Their values do not change.spock.restart_delay_defaultnow applies to every apply-worker restart, including after a lost connection or a temporary error, so a problem that does not clear cannot become a tight restart loop.
Catalog changes
- New:
spock.sub_id_generatorsequence (replaces inline oid-from-counter generation for subscription rows). - New:
spock.reserved_objecttable, the list of reserved schemas and extensions (see Reserved schemas and extensions above). Built-in rows are protected by a trigger, and only user-added rows are dumped. - New index
spock.resolutions(log_time)to supportspock.cleanup_resolutions(). spock.resolutions.conflict_typevalues renamed during upgrade:update_update→update_exists,update_delete→update_missing,delete_delete→delete_missing.insert_existsis unchanged.spock.subscription.sub_skip_schemais now correctly typed astext[](the column was added astextin 5.0.2 but the C code always treated the bytes astext[]; the 5.0.7 step migration relabels the catalog so no rewrite or downtime is required).
New functions
spock.apply_group_progress()— set-returning health view (backsspock.progress).spock.read_peer_progress(slot, provider_node_id, subscriber_node_id)— used by zodan when importing a slot snapshot.spock.get_subscription_stats(subid, peer_node_id)/spock.reset_subscription_stats(subid)— per-subscription apply statistics.spock.cleanup_resolutions(days int DEFAULT NULL)— TTL-based cleanup of the resolutions log. Default permissions revoke EXECUTE from PUBLIC.spock.sub_alter_options(subscription_name name, options text[])— bulk subscription option changes, with input validation and no-op restart skipping.spock.reserved_object_add(p_name name, p_kind text, p_exclude_from_dump boolean DEFAULT true, p_block_in_repset boolean DEFAULT true, p_replicate_ddl boolean DEFAULT NULL)— reserve a schema or extension, or change an existing row.spock.reserved_object_remove(p_name name, p_kind text)— remove a reserved schema or extension you added.spock.slot_enable_failover() RETURNS integer— sets thefailoverflag on Spock's logical slots on PostgreSQL 17 and later, and returns how many it changed. The upgrade from 5.x calls it. EXECUTE is revoked from PUBLIC.spock.node_alter(p_location text, p_country text, p_info_patch jsonb)— validated way to change the local node's own location/country/info (merging info, validating atiebreakerpatch); see Automatic node metadata propagation above.spock.node_info_emit()/spock.node_info_apply()/spock.node_info_broadcast()(trigger function, backing the newnode_info_broadcast_triggeronspock.node) — internal plumbing behind automatic node metadata propagation; not intended to be called directly (EXECUTErevoked fromPUBLIC).
Removed functions
spock.convert_column_to_int8(regclass, smallint)— superseded.spock.convert_sequence_to_snowflake(regclass)— superseded.spock.wait_for_apply_worker(p_subbid bigint, timeout int)— had no callers since SpockCtrl was removed.
Bug fixes carried forward from the 5.0.x line
These shipped in 5.0.6 – 5.0.12 and are included in 6.0.0:
- Handle upstream connection loss cleanly without replication-origin
advance leak. Previously a stale libpq socket fd produced an
epoll_ctl(EINVAL)cascade and the recovery path could silently advance the replication origin past an in-flight remote transaction, causing it to be missed on reconnect (and, withspock.exception_behaviour = transdiscard, its rows would land inspock.exception_logwitherror_message = unknown). The apply worker nowPG_RE_THROWs on connection-class errors and the manager respawns it from the last durably-committed remote LSN. - Preserve the original error message in
spock.exception_logwhen applying undertransdiscard/sub_disable. The originally-failing row now carries the real error message; bystander rows in the same transaction are recorded asunavailableinstead ofunknown. - zodan: name slots and subscriptions consistently
(
sub_{provider}_{subscriber}) via the newspock.gen_sub_name()helper;remove_nodedrops subscriptions from thespock.subscriptioncatalog instead of guessing names; pause apply workers during slot creation to prevent data loss when adding a new node; align Phase 3 catch-up wait and parameterisespock.sync_event()to avoid the Phase 9 handshake deadlock. - Prevent data loss on
resync_tablewhen the subscriber is read-only (SPOC-440). - Fix stale
local_tuplepointer in the INSERT exception path for a missing relation;SUB_DISABLEnow exits cleanly instead of dereferencing freed memory. - Suppress repeated
hot_standby_feedback offERROR spam when a standby is used for reporting (not failover). - Fix stack-use-after-scope in
spock_connect_base()(ASan finding). - Fix resource owner bug
apply_replay_bytesint→uint64overflow fix (b4ba9bdf)- A lost provider connection under
SUB_DISABLEcould disable the subscription: the resent transaction was mistaken for one that had failed to apply. - The apply worker retries after temporary errors — deadlocks,
lock_timeout, running out of a resource (SQLSTATE class 53), or a provider that is restarting or in recovery (57P02, 57P03) — instead of treating them as data exceptions that disable the subscription or discard the transaction. The worker exits without advancing the replication origin and the provider resends the transaction. - Restarting an apply worker in the middle of a transaction could disable
the subscription under
SUB_DISABLE. The failure marker is now cleared when the worker exits onSIGTERM. - An error between transactions left the apply worker in replay mode for
the next transaction, which had never failed, so under
TRANSDISCARDorSUB_DISABLEit could be discarded or disable the subscription. - Two subscriptions whose names share a prefix (such as
subandsub_parallel) shared one exception-log slot. The lookup now compares the whole name, and no longer reads past the end of the array. spock.get_lsn_from_commit_ts()could hang on an idle node. Its WAL scan now stops at the end of WAL as it was when the scan began.- The failover-slots worker died whenever the walreceiver reconnected to
the primary, when
spock.primary_dsnwas not set. It now waits for the next cycle. - Two retry messages lowered from LOG to DEBUG1 still reached the server log with default settings, because the level check was backwards.
- A column that exists on the provider but not on the subscriber put the apply worker in a restart loop. It is now handled like a missing table, and the error names every missing column.
- Tuple data from the provider is checked against the local column definition before use, so a tuple that does not fit is rejected instead of applied. Names and attribute counts in RELATION messages are checked too, and a subscriber built without lz4 rejects an lz4-compressed value.
Upgrading
The upgrade from 5.0.12 to 6.0.0 is a single ALTER EXTENSION spock UPDATE
once the binaries are swapped. Earlier 5.0.x releases upgrade in the same
single step, through 5.0.12. The upgrade:
- drops the legacy
spock.progresstable and recreates it as a view, - replaces the
spock.lag_trackerview definition, - migrates
spock.resolutions.conflict_typevalues, - adds the new functions, the
sub_id_generatorsequence and thespock.reserved_objectcatalog, - drops
spock.wait_for_apply_worker(), - adds the
node_info_broadcast_triggertrigger onspock.node, - refreshes the parallel-safety attributes on
spock.md5_agg_sfuncandspock.spock_gen_slot_name, - turns on the
failoverflag for existing logical slots by callingspock.slot_enable_failover(), so PostgreSQL 17 and 18 slot synchronization picks them up. This is a no-op on PostgreSQL 16 and older and on a standby, and it skips slots that are in use at the time. See Logical Slot Failover for how to handle any skipped slots.
Check your runbooks and automation for direct DDL against the spock or
snowflake schemas before upgrading. Statements such as
DROP TABLE snowflake.x or CREATE INDEX on a spock table succeeded on
5.0.x and now raise an error (see AutoDDL improvements above). Any
procedure that needs to keep doing this must set
spock.enable_ddl_replication = off for the session first, which is also
what the error hint says. The restriction covers only the built-in
extension-owned schemas: schemas you reserve yourself with
spock.reserved_object_add() are unaffected, and pgedge_ace continues to
accept DDL and keep it node-local.
PostgreSQL's 2026 security fix for CVE-2026-6471 added the
output_plugin_libraries parameter, and its default leaves out
spock_output, so logical decoding fails on a server with the fix. Set
output_plugin_libraries = 'pgoutput, test_decoding, spock_output' on every
node, physical standbys included, and only on servers that have the
parameter. See Configuring Spock.
The apply worker now restarts and tries again after a temporary error such
as a deadlock or lock timeout. If the problem never clears, expect the
worker to restart every spock.restart_delay_default (5 seconds by default)
instead of a disabled subscription or a discarded transaction.
Spock 5.0.12
Upgrade Notes
-
There are no schema changes in 5.0.12.
-
PostgreSQL's 2026 security fix for CVE-2026-6471 (back-patched to every supported major) added the
output_plugin_librariesparameter, and a library may no longer be used as an output plugin unless it is listed there. Its default excludesspock_output, so on a server carrying the fix logical decoding fails withlibrary "spock_output" may not be used as an output plugin. The check runs whenever decoding starts, not only at slot creation, so this stops an established cluster after a minor-version upgrade - not just new subscriptions. Setoutput_plugin_libraries = 'pgoutput, test_decoding, spock_output'on every node, physical standbys included, and only on servers that have the parameter -
on older releases an unrecognised parameter stops the server from starting. See Configuring Spock. The regression and TAP test suites and the Docker test image now check for the parameter and set it themselves.
-
When the apply worker fails for a reason that has nothing to do with the data (a deadlock, a lock timeout, a provider that is restarting), it now restarts and tries again instead of treating the failure as an exception. See Bug Fixes. If the problem never goes away, you will see the apply worker restart every
spock.restart_delay_default(5 seconds by default) rather than a disabled subscription or a discarded transaction.
Bug Fixes
-
The apply worker now retries after temporary errors. Some errors have nothing to do with the replicated data: PostgreSQL picked the apply worker as a deadlock victim,
lock_timeoutfired, the server briefly ran out of a resource (SQLSTATE class 53), or the provider was restarting or still in recovery (57P02, 57P03). All of these succeed if tried again. Before, they went down the same path as a permanent data error, so depending onspock.exception_behaviourthe subscription was disabled or the transaction was written tospock.exception_logand thrown away, even though the provider would happily have sent it again. Now the worker exits without moving the replication origin, the manager starts it again, and the provider resends the transaction. Every restart path, including the existing one for a lost connection, now waitsspock.restart_delay_defaultbefore starting again, so a problem that does not clear cannot turn into a tight restart loop. -
Tables with
GENERATED ALWAYS AS ... STOREDcolumns caused replication errors. Generated columns were handled like ordinary columns during the initial table copy, in the replication protocol, when filling in defaults for missing columns, and during conflict resolution. They are now left out everywhere, and the subscriber computes them from the column definition. -
AutoDDL failed after utility commands that run outside a transaction block, such as
CLUSTERwith no table name. These commands commit after each table, so by the time the AutoDDL hook ran there was no open transaction and no snapshot, and the catalog lookups and the queue insert failed. AutoDDL now starts a transaction and takes a snapshot itself when it needs to. -
spock.get_lsn_from_commit_ts()could hang on an idle node. The function scans the node's WAL for the last local commit at or before a timestamp, but the scan had no stopping point, so once it caught up it sat waiting for WAL that nobody was going to write. Anadd_noderun was seen stuck behind this call for over eleven minutes. The scan now stops at the end of WAL as it stood when it began, takes the commit timestamp and origin from the WAL record instead ofpg_commit_ts, counts local two-phase commits written by other applications, and refuses to run on a server in recovery. A timestamp with no matching commit still returns the slot'srestart_lsn, now logged at DEBUG1. -
A failure between transactions could make the apply worker mishandle the next one. When a transaction fails to apply, the worker restarts, the provider resends it, and the worker runs it again in what the code calls replay mode: each row is tried in a subtransaction so the failing row can be written to
spock.exception_logandspock.exception_behaviourcan decide what to do with it. Replay mode is meant for that one transaction. But an error raised between transactions switched the worker into replay mode with no failed transaction to attach it to, and the flag stayed set for whatever the provider sent next. That transaction, which had never failed, ran as a replay and underTRANSDISCARDorSUB_DISABLEwas thrown away or used to disable the subscription. The worker now turns replay mode on only for the transaction that actually failed, and any other transaction clears the recorded failure. -
A replayed transaction that applied nothing was still committed. Under
TRANSDISCARDandSUB_DISABLEevery row of a replay runs in a subtransaction that is always rolled back, so a replay applies no rows. The check for "replayed but hit no error" ran only after the replication origin had moved forward and the transaction had committed, so the origin ended up past a transaction that was never applied. The transaction was lost while the log said it had been discarded on purpose. The check now runs first and the transaction is rolled back instead, so the provider resends it. Transactions skipped withspock.sub_alter_skiplsnare exempt. -
Two subscriptions whose names share a prefix (for example
subandsub_parallel) shared one exception-log slot, because the lookup compared only the first few characters. Sharing the slot means sharing the commit LSN that marks a transaction as having already failed, so a failure recorded by one subscription could push the other into replay mode for a transaction that never failed. The lookup now compares the whole name. The same code also read, and could have written, one entry past the end of the array. -
The failover-slots worker died whenever the walreceiver reconnected to the primary. Without
spock.primary_dsnthe worker takes the primary's connection string from the walreceiver, which blanks that string for the whole of every connection attempt, not only at standby startup. The connection then failed, the worker exited, and failover slots stopped syncing until it restarted. The worker now waits for the next cycle. -
Two retry messages for serialization failures are lowered from LOG to DEBUG1, but the check that decides whether to print them at the new level was backwards. With default settings they went to the server log no matter what
log_min_messagessaid. Also marked a deliberate switch fall-through in the apply worker so clang stops warning about it on every build before PostgreSQL 19. -
Restarting an apply worker in the middle of a transaction could disable the subscription. The worker records the commit LSN of every transaction it begins, not only ones that fail, and a clean shutdown left the marker behind, so the provider's resend of the interrupted transaction looked like a transaction that had already failed. Under
SUB_DISABLEthat disabled the subscription for no real error. The marker is now cleared when the worker exits onSIGTERM. -
A column that exists on the provider but not on the subscriber put the apply worker in a restart loop. The error was raised while opening the table, before the per-row subtransaction exception handling needs, so
spock.exception_behaviournever got a say and the error came back on every retry without end. A missing column is now handled like a missing table: the first attempt raises the error, the retry discards the change and logs it through the normal exception path. The message also names every column the local table is missing, so repairing a drifted schema no longer costs one column per apply attempt. -
Tuple data arriving on the wire is now checked against the local column definition before it is used. Each attribute is copied out of the message into its own buffer and its length checked, so a tuple that does not fit the local column is rejected rather than applied. That is what happens after a provider-only
ALTER COLUMN TYPEwith DDL replication off: the apply worker exits and retries until the two schemas agree. Names and attribute counts in the RELATION message are validated too, and a subscriber built without lz4 now rejects an lz4-compressed value instead of storing it.
Other Changes
-
Refreshed the
attoptionsserver patch for theheap_update()change in PostgreSQL 17.11, 18.5 and 19 beta 3. One hunk's context lines no longer matched and the patch was rejected. Only context lines changed. -
Spock builds against PostgreSQL 19 beta 3 and 4, with a new
compat/19layer for the API changes in that beta (CLUSTERfolded intoREPACK, the recovery-conflict signalling changes, tuple-descriptor finalisation, and the flattenedReorderBufferTXNcommit-time field).CREATE EXTENSIONaccepts major version 19, and CI and the release packaging cover it. This is preliminary: PostgreSQL 19 is still in beta, the compatibility layer will change before it is released, and Spock on 19 is not supported for production use. -
Dropped Debian bullseye from the platforms the release workflow builds packages for. It has reached end of life.
-
Removed the old Z0DAN Python files and their TAP test.
add_nodeandremove_nodeare provided by the SQL implementation. -
New documentation on running logical slot failover under Patroni on PostgreSQL 17 and later. It covers the parameters you need, the settings to put in the bootstrap DCS, and what to do about
synchronized_standby_slotson a switchover. A sample Patronion_role_changecallback,samples/set_synchronized_standby_slots.sh, resetssynchronized_standby_slotsafter a promotion. See Logical Slot Failover. -
Documented the
wait_if_disabledargument ofspock.wait_for_sync_event(). -
TAP suite: tests share one set of log and wait helpers instead of each carrying a copy, cluster startup waits for the servers to be ready instead of sleeping 17 seconds, and a single test can be run on its own with
prove. New tests cover retry after temporary errors, real deadlocks, replay mode carrying over between transactions,spock.get_lsn_from_commit_ts(), a worker restart in the middle of a transaction, and a column missing on the subscriber.
Spock 5.0.11
New Features
-
Add
spock.use_native_failover_slots(default off, PGC_POSTMASTER). When enabled, spock marks logical slots with the FAILOVER flag on PG17+, yields to PostgreSQL's native slotsync worker on PG17 whensync_replication_slots=on, and does not register spock's own failover-slot worker on PG18+. Off preserves the existing worker-based behavior. Read on the subscriber node that creates the logical replication slot; changing it requires a server restart. See Logical Slot Failover for setup steps and the required post-promotionsynchronized_standby_slotsrunbook. -
Sync failover slots to the standby every 1s by default instead of a hard-coded 60s, shrinking the window in which a promotion can find a stale slot. The interval is now tunable via
spock.failover_slots_naptime(and the feedback-wait retry viaspock.failover_slots_feedback_naptime), both SIGHUP-settable in milliseconds -
Never copy the
pgedge_aceschema to a node joining viaadd_node. ACE state is node-local, so its objects are now excluded from both the schema copy and the data sync, even when an ACE table belongs to a replication set.
Bug Fixes
- Fixed a bug where a transient provider connection loss could incorrectly disable a subscription under SUB_DISABLE exception handling. A transaction retransmitted after a reconnect is no longer misclassified as an apply failure, so replication resumes normally instead of stopping.
Spock 5.0.10
Bug Fixes
- Improve apply performance on tables with a unique index over a frequently NULL column. Rows with a NULL in such a key are now recognized as non-conflicting without an unnecessary index scan, avoiding a slowdown that grew with table size.
- Correctly handle conflicts on
NULLS NOT DISTINCTunique indexes. On these indexes (PostgreSQL 15+) two NULL values are treated as equal, so matching NULL-keyed rows are now resolved through normal conflict resolution (last-update-wins) instead of failing with a duplicate-key error. - Fix a
could not open relationerror in the apply worker when an index is dropped while replication is running. The worker now refreshes its cached view of a table's indexes before using it. - Fix gradual memory growth during a transaction that logs to
spock.exception_log. Memory used while recording exceptions is now released per row instead of accumulating until the transaction commits. - Record the real cause of a discarded transaction in
spock.exception_loginstead of the opaque placeholderunavailable. The failing command's message is now stored (prefixed with its SQLSTATE where informative), and the other rows of the transaction note that they were discarded as collateral, making it possible to provide the root cause of the exception. - Fix ZODAN (
add_node) version checking. Spock versions were compared as text, so5.0.10sorted below5.0.4and a valid node was wrongly rejected. Versions are now compared numerically, and nodes may differ in patch level as long as their major.minor matches, allowing rolling upgrades. - Fix a spurious "tiebreaker values are equal" WARNING on timestamp-tied conflicts in single-writer topologies (and on node-id hash collisions). The tiebreaker now compares the applying node against the remote sender, and node-id collisions are detected at node creation.
Spock 5.0.9
New Features
- Initial PostgreSQL 19 support. Adds the server-side patches required to build and run Spock against PostgreSQL 19.
Bug Fixes
- Fix silently dropped writes after a synchronous-standby failover. Spock could report an incorrect replication position to the publisher, so a promoted standby resumed from the wrong point and lost the changes in between. The position reported to the publisher is now always taken from the publisher's own WAL stream.
- Fix
spock_failover_slotsstalling on PG17+ standbys. Slot synchronization could hold replication-slot and proc-array locks exclusively long enough to block every backend trying to start up on the standby, freezing the worker and any new connections. The sync path now uses shorter, shared locks matching PostgreSQL's own slot code. - Fix manager-worker respawn loop when databases are dropped or have
connections disabled (
datconnlimit = -2) concurrently with Spock starting a per-database worker. The database OID is now revalidated immediately before the worker is registered to avoid race. - Reduce per-row work and memory growth in the apply path. Replicated changes no longer re-open a table's indexes for every row, and transient memory is freed per row instead of building up over the transaction.
Performance
- Skip building conflict log messages that would not be logged. When the log
level is suppressed, Spock no longer formats the conflicting rows for an
ereport()that would be discarded, speeding up apply on high-conflict workloads. Recording conflicts inspock.resolutionsis unaffected.
Spock 5.0.8
Bug Fixes
- Fix subscriber crash on transactions larger than 2 GB.
apply_replay_byteswas declared asint, causing signed-integer overflow and a crash when a single replicated transaction exceeded 2 GB of WAL data. Changed touint64. - The apply worker now exits cleanly when the upstream connection dies
(firewall reload, walsender SIGKILL/RST, walsender ping timeout) and the
manager respawns it from the last durably-committed remote LSN. Previously
a stale libpq socket fd produced an
epoll_ctl()cascade with a follow-onerror during exception handlingper disconnect, and a corner of the recovery path could silently advance the replication origin past the in-flight remote transaction, causing it to be skipped on reconnect. - Removed native PG17+/PG18 slot-sync integration. Spock no longer creates slots with (FAILOVER), does not defer to PostgreSQL's slotsync worker, and keeps spock_failover_slots active on PG18; the spock worker is once again the only path for failover-slot sync on 5.x.
- spock_failover_slots: handle primary disconnects and post-promotion edge cases. Reconnects to the primary on transient failure during sync; fixes a PG15/16 PANIC where a freshly promoted standby could enter a crash loop because the synchronized slot's restart_lsn was below the standby's WAL floor; tolerates invalid/zero LSNs and catalog_xmin values that previously tripped assertions.
Spock 5.0.7
New Features
Logical Slot Failover Improvements
- On PostgreSQL 17+, Spock now creates all logical replication slots with
the
FAILOVERflag, allowing PostgreSQL's built-in slotsync worker (sync_replication_slots = on) to automatically synchronize them to physical standbys. - On PostgreSQL 18+, Spock's own
spock_failover_slotsbackground worker is no longer registered. The native PostgreSQL slotsync worker fully replaces it. See the Logical Slot Failover section in the configuration guide for requiredpostgresql.confsettings. - On PostgreSQL 17, Spock's worker remains active but automatically yields
to the native slotsync worker if
sync_replication_slots = onis set, preventing conflicts.
Bug Fixes
- Preserve the original error message in
spock.exception_logwhen applying inTRANSDISCARDorSUB_DISABLEmode. Previously the original error was written only to the server log and lost before the retry pass; every row in the discarded transaction was logged witherror_message = NULL(stored as "unknown"). The originally-failing row now carries the real error message, while the bystander rows in the same transaction are recorded as "unavailable" instead of the misleading "unknown" fallback. - Fix missing data when adding a new node. An apply worker that committed
between
ensure_replication_slot_snapshotandadjust_progress_infocould advancespock.progresspast the COPY snapshot boundary, causing the new node to permanently skip those changes. Apply workers are now paused via an atomic flag and condition variable inSpockContextfor the duration of slot creation and resumed onceadjust_progress_infocompletes. - Fix apply worker crash with
epoll_ctl() failed: Invalid argumentafter a provider connection died. The socket fd captured beforestream_replay:was never refreshed, so a stale fd was passed toWaitLatchOrSocketon re-entry. The fd is now refreshed atstream_replay:and the worker exits cleanly so the postmaster can restart and reconnect it. - Fix PG15/16 failover slot loss on promotion: reconnect with retry and guard against a zero WAL flush LSN.
- ZODAN:
create_sub_on_new_node_to_src_node(Phase 9 ofadd_node) was generating subscription names assub_{subscriber}_{provider}instead of the expectedsub_{provider}_{subscriber}convention, causingremove_nodeto fail when looking up subscriptions by name. A newspock.gen_sub_name(provider_node, subscriber_node)helper is now used in place of inline name concatenations. - ZODAN: the
remove_nodecleanup loop tried to drop subscriptions by guessing names and nodes, which broke whenevercross_wireandadd_nodeused different naming conventions. It now performs a per-nodeDROPof every subscription read directly from thespock.subscriptioncatalog. - Since 5.0.2 there was a bug with spock.subscription.sub_skip_schema having the incorrect type, it should be text[]. Fixed, including in the spock-5.0.6--5.0.7.sql migration script to repair.
- Fix a rare bug where under high concurrency when using delta apply columns a ResourceOwner exception may occur.
- Have more meaningful messages appear in spock.exception_log.
Operational Improvements
- Documentation: filled out
snowflake.md; revisedupgrading_spock.md; fixed event-trigger syntax (EXECUTE FUNCTIONinstead of the deprecatedEXECUTE PROCEDURE); standardized function-reference filenames to match function names exactly; removed obsoletebatch_inserts.md - Test infrastructure aligned with
main:run_tests.shandSpockTest.pmnow use absoluteTESTLOGDIRpaths, derive per-test log filenames, and isolatepsqlfrom a developer'spsqlrc/history. New regression tests added for failover slots (018), exception-handling/TRANSDISCARD error quality (013), and the stale-fd-after-connection-death scenario (019). - Avoid unnecessarily waiting after sync_event()
Spock 5.0.6
New Features
- New
spock.feedback_frequencyGUC that controls how often feedback is sent to the WAL sender. Feedback is sent every n messages, where n is the configured value. Note that feedback is also sent everywal_sender_timeout / 2seconds. - New
spock.log_origin_changeGUC to control logging of row origin changes to the PostgreSQL log. Origin changes caused by replication are no longer written to thespock.resolutionstable, as they are informational and not true conflicts. Three modes are available:none— Do not log origin changes (default)remote_only_differs— Log only when a row from one remote publisher is updated by a different remote publishersince_sub_creation— Log origin changes for tuples modified after subscription creation (suppresses noise from pg_restored data)
- New
sub_created_atcolumn onspock.subscriptionto help distinguish pre-existing data (e.g. from pg_restore) from post-subscription data. - COPY TO is considered read-only and can now be run when a node is in read-only mode.
Performance Improvements
- Deferred
spock.progresscatalog writes. The progress table was previously updated on every committed transaction and every keepalive, causing significant table bloat and I/O overhead. Progress catalog writes are now batched and flushed at most once per second from the main apply loop. Shared memory is still updated immediately internally for correctness.
Bug Fixes
- Fix initdb assertion failure in attoptions patch for PG15/16/17.
- Fix two bugs in the table re-sync routine: WAL sending is now switched off during truncate and re-sync to prevent data loss, and the infinite wait was fixed when no more DML is committed by using the last committed LSN instead of the last received LSN.
- Use NULL for unknown
local_origininspock.resolutionsinstead of an invalid origin ID when origin cannot be determined (e.g. pg_dump, frozen transactions, truncated commit timestamps). Also fixed off-by-one errors inspock_conflict_row_to_json()that were overwriting thelocal_originNULL flag. - Fix Zodan initialization issue:
present_final_cluster_statenow executes a COMMIT to allow newly created subscriptions to update their state, and final cluster state now checks all subscriptions across the cluster. - Suppress hot_standby_feedback off error messages in log in case a read replica is used for reporting and not failover (it is required for failover).
- Fix bug when applying changes to a table that has been dropped.
Operational Improvements
add_nodeis now restricted to run only on a new (uninitialized) node, preventing accidental misuse.
Spock 5.0.5 on Feb 12, 2026
- Fix segfault that occurs when using new Postgres minor releases like 18.2.
- Zero Downtime Add Node (Zodan) minor bug fixes and improvements
- Updated documentation
Spock 5.0.4 on Oct 8, 2025
- Reduce memory usage for transactions with many inserts.
- When a subscriber’s apply worker updates a row that came from the same
origin, spock will no longer log it in the
spock.resolutionstable, reversing behavior that has been in place since 5.0.0. - Improved handling to block replicating DDL when adding an extension.
- Improved documentation
-
Zero Downtime Add Node Improvements:
- New health checks and verifications (ex: version compatibility) before and during the add node process.
- New remove node SQL procedure (
spock.remove_node()in samples/Zodan/zodremove.sql) and python script (samples/Zodan/zodremove.py). This also handles removing nodes that were partially added when the user decided to undo this work. - Handle DSN strings that contain quotes.
-
Bug fixes:
- Log messages containing credentials will now obfuscate password information.
- Fix bug when the subscriber receives DML for tables that do not exist.
This case will be handled according to the configured
spock.exception_behavioursetting (SUB_DISABLE,DISCARD,TRANSDISCARD). - Fix bug where spock incorrectly outputs a message that DDL was replicated when a transaction is executing in repair mode.
v5.0.3 on Sep 26, 2025
- Spock 5.0.3 adds support for Postgres 18.
- When using row filters with Postgres 18 such as in the functions spock.repset_add_table() or spock.repset_add_partition(), allowable filters are now stricter. Expressions may not use UDFs nor reference another table, similar to native logical replication in Postgres.
v5.0.2 on Sept 22, 2025
- Improved logging for all Zodan phases in both the stored procedure and Python examples.
- You can use the new Zodan skip_schema parameter to exclude schemas when
you're adding a node, preventing local extension metadata from being copied.
- Added skip_schema option to sub_create. When synchronize_structure = true, schemas listed in skip_schema are not synced.
- dblink-based stored procedure for add node leverages skip_schema to skip preexisting schemas on the new node, avoiding failures on already existing schemas.
- Python add_node now mirrors stored procedure semantics, including skip_schema handling, structure sync, and logging. It works as a direct alternative to the procedure and no longer requires the dblink extension.
- Spock has been updated to use the PostgreSQL License.
v5.0.1 on Aug 27, 2025
- Bug fix for an incorrect commit timestamp being used for the case of updating a row that was inserted in the same transaction. A consequence was possible incorrect resolution handling for updates.
- Use the default search_path in the replicate_ddl() function.
- Prevent false positives for conflict detection when using partial unique indexes.
- New sample files for adding nodes with zero downtime.
- Python example
- Added enhanced add node support in stored procedures in
samples/zodan.sql
- Add a second node when there is only one node.
- Extended support for adding 3rd, 4th and subsequent nodes.
- Chain adding new nodes off of the previously added new node (eg: add N2 off N1, then add N3 off N2).
v5.0 on July 15, 2025
- Spock functions and stored procedures now support node additions and major
PostgreSQL version. This means:
- existing nodes are able to maintain full read and write capability while a new node is populated and added to the cluster.
- You can perform PostgreSQL major version upgrades as a rolling upgrade by adding a new node with the new major PostgreSQL version, and then removing old nodes hosting the previous version.
- Exception handling performance improvements are now managed with the spock.exception_replay_queue_size GUC.
- Previously, replication lag was estimated on the source node; this meant that if there were no transactions being replicated, the reported lag could continue to increase. Lag tracking is now calculated at the target node, with improved accuracy.
- Spock 5.0 implements LSN Checkpointing with
spock.sync()andspock.wait_for_sync_event(). This feature allows you to identify a checkpoint in the source node WAL files, and watch for the LSN of the checkpoint on a replica node. This allows you to guarantee that a DDL change, has replicated from the source node to all other nodes before publishing an update. - The
spockctrlcommand line utility and sample workflows simplify the management of a Spock multi-master replication setup for PostgreSQL.spockctrlprovides a convenient interface for:- node management
- replication set management
- subscription management
- ad-hoc SQL execution
- workflow automation
- Previously, replicated
DELETEstatements that attempted to delete a missing row were logged as exceptions. Since the purpose of aDELETEstatement is to remove a row, we no longer log these as exceptions. Instead these are now logged in theResolutionstable. INSERTconflicts resulting from a duplicate primary key or identity replica are now transformed into anUPDATEthat updates all columns of the existing row, using Last-Write-Wins (LWW) logic. The transaction is then logged in the node’sResolutionstable, as either:keep localif the local node’sINSERThas a later timestamp than the arrivingINSERTapply remoteif the arrivingINSERTfrom the remote node had a later timestamp
- In a cluster composed of distributed and physical replica nodes, Spock 5.0
improves performance by tracking the Log Sequence Numbers (LSNs) of
transactions that have been applied locally but are still waiting for
confirmation from physical replicas. A final
COMMITconfirmation is provided only after those LSNs are confirmed on the physical replica. This provides a two-phase acknowledgment: - Once when the target node has received and applied the transaction.
- Once when the physical replica confirms the commit.
- The
spock.check_all_uc_indexesGUC is an experimental feature (disabledby default); use this feature at your own risk. If this GUC isenabled, Spock will continue to check unique constraint indexes, after checking the primary key / replica identity index. Only one conflict will be resolved, using Last-Write-Wins logic. If a second conflict occurs, an exception is recorded in thespock.exception_logtable.
Version 4.1
- Hardening Parallel Slots for OLTP production use.
- Commit Order
- Skip LSN
- Optionally stop replicating in an Error
- Enhancements to Automatic DDL replication
Version 4.0
- Full re-work of parallel slots implementation to support mixed OLTP workloads
- Improved support for delta_apply columns to support various data types
- Improved regression test coverage
- Support for Large Object Logical Replication
- Support for pg17
Our current production version is v3.3 and includes the following enhancements over v3.2:
- Automatic replication of DDL statements
Version 3.2
- Support for pg14
- Support for Snowflake Sequences
- Support for setting a database to ReadOnly
- A couple small bug fixes from pgLogical
- Native support for Failover Slots via integrating pg_failover_slots extension
- Parallel slots support for insert only workloads
Version 3.1
- Support for both pg15 and pg16
- Prelim testing for online upgrades between pg15 & pg16
- Regression testing improvements
- Improved support for in-region shadow nodes (in different AZ's)
- Improved and documented support for replication and maintaining partitioned tables.
Version 3.0 (Beta) includes the following important enhancements beyond the BDR/pg_logical base:
- Support for pg15 (support for pg10 thru pg14 dropped)
- Support for Asynchronous Multi-Master Replication with conflict resolution
- Conflict-free delta-apply columns
- Replication of partitioned tables (to help support geo-sharding)
- Making database clusters location aware (to help support geo-sharding)
- Better error handling for conflict resolution
- Better management & monitoring stats and integration
- A 'pii' table for making it easy for personally identifiable data to be kept in country
- Better support for minimizing system interruption during switch-over and failover