Redact PII while you migrate & sync
Seed staging, dev, analytics, and vendor handoffs from production without letting personal data leave with the rows.
You need a realistic copy of production in a place production data isn't allowed to go — a staging database, a developer laptop, an analytics warehouse, a vendor's environment. The schema, the row shapes, and the referential structure all have to survive; the emails, card numbers, and national IDs must not. sluice does this inline with --redact: PII is transformed between the source reader and the target writer, so the sensitive value never lands on the target. There is no separate scrubbing pass to forget to run.
Read Where redaction applies before relying on this for backups. Redaction reaches every path that copies rows in bulk, and the live CDC apply stream — but not backup incremental or backup stream run. A redacted backup chain is therefore a series of fulls, not a full plus incrementals. sluice v0.144.0 refuses every combination that would silently produce a mixed chain; earlier versions produced one at exit 0.
How --redact works #
Each rule names a column and a strategy:
--redact '[schema.]table.column=STRATEGY[:options]'
The flag is repeatable — pass it once per column. Rules are applied in the bulk-copy hot path and the CDC apply path alike; the strategy's output replaces the source value verbatim at the named column before it reaches the target. When no --redact is configured the pipeline short-circuits before any per-row work, so operators who don't use redaction pay nothing for the feature.
sluice migrate \
--source-driver postgres --source "$SRC" \
--target-driver postgres --target "$DST" \
--redact users.email=hash:sha256
Every rule also has a YAML form under a redactions: block in the config file (see Configuration). CLI and YAML mix, and the CLI rule wins: the --redact flags are registered first, the YAML block is merged in after them, and a YAML rule for a schema.table.column the CLI already declared is skipped, with a WARN naming the column (since v0.96.0; before that the YAML rule silently overwrote the CLI one). Keep the bulk in version-controlled YAML; reach for the flag for per-environment overrides (--redact users.email=null in staging) — that override is exactly what the precedence exists for.
The YAML spelling of each spec #
Every CLI spec has a YAML form: the strategy name goes in strategy: and the colon-separated options become sibling keys. The Required column is enforced at config load.
| CLI spec | YAML keys | Required |
|---|---|---|
null | strategy: "null" — MUST be quoted; bare null is YAML's null literal | — |
static:<value> | strategy: static, value: <value> | — (value may be omitted or empty for an explicit empty-out) |
hash:sha256 / hash:hmac-sha256[:<key>] | strategy: hash, algo: sha256 | hmac-sha256, key: <keyset key> | algo |
truncate:<n> | strategy: truncate, length: <n> | length |
mask:inner:<m1>,<m2>[,<char>] / mask:outer:… | strategy: mask, form: inner | outer, m1: <n>, m2: <n>, char: <one rune> | m1, m2 |
mask:<preset> | strategy: mask, form: ssn | pan | pan-relaxed | email | ca-sin | uk-nin | iban | uuid | form (no m1 / m2 / char) |
randomize:int:<min>,<max> | strategy: randomize, form: int, min: <n>, max: <n> | min, max |
randomize:pan[:<brand>] | strategy: randomize, form: pan, brand: visa | mastercard | amex | form |
randomize:iban[:<country-code>] | strategy: randomize, form: iban, country_code: DE | GB | FR | form |
randomize:email / us-phone / uuid / ssn / ca-sin / uk-nin | strategy: randomize, form: <name> | form |
randomize:dict:<name> | strategy: randomize, form: dict, dict: <name> | dict |
tokenize:dict:<name>[:<key>] | strategy: tokenize, dict: <name>, key: <keyset key> | dict |
strategy: truncate without length: fails exactly like --redact users.email=truncate. Through v0.153.1 an omitted length, m1, m2, min or max decoded to 0 and the rule ran at exit 0: truncate emptied every row, randomize / int wrote 0 to every row, mask / inner masked the whole value, and mask / outer masked nothing — the source value shipped unchanged under a rule that declared it masked — on migrate, sync, backup and preview, while the identical --redact spec was refused. If a config that loaded on v0.153.1 refuses on v0.153.2, the rule it names was never doing what it said; under mask / outer, treat every target or backup written under it as unredacted. Two edges: a key that is present with the value 0 is honoured (length: 0, m1: 0, min: 0), exactly as truncate:0 / mask:inner:0,4 / randomize:int:0,9 are on the CLI; and a numeric key present on a form that takes none — m1: 0 on a preset mask, min: 0 on form: email — is refused as spurious (before v0.153.2 it was silently ignored). An unknown key name is refused by the loader itself.Where redaction applies #
Redaction reaches every path that copies rows in bulk, and the live CDC apply stream. It does not reach the two commands that archive change events, and that asymmetry is load-bearing rather than incidental — their events come off the source's CDC pump verbatim, with no seam to apply a rule at:
| Command | Behaviour |
|---|---|
sluice migrate | One-shot bulk copy — every row passes through the redactor. |
sluice sync start | Both phases honour --redact: the cold-start snapshot copy and the live CDC apply stream. |
sluice backup full | Redaction is applied at chunk-write time, so the archive on disk is PII-clean and a later restore copies it through unchanged. |
sluice backup incremental | No --redact flag and no redaction. Change events are archived off the CDC pump verbatim. |
sluice backup stream run | No redaction, same as backup incremental; its rotation-born segment fulls are unredacted too. |
sluice restore | No --redact flag. Redaction happens on the way in, never on the way out — restoring an unredacted archive yields unredacted rows whatever the target is configured with. |
sluice schema preview | No data moves — it annotates the generated CREATE TABLE DDL with which columns are redacted (see below). |
A redacted backup chain is a series of fulls #
Because only backup full redacts, extending a redacted full with an incremental would archive plaintext for every row touched after the snapshot — and restore it. Through v0.143.0 that happened silently, at exit 0. Since v0.144.0 a redacted full records the fact in its manifest (a rule count plus a fingerprint of the policy — never the rules, and never the column names, so a manifest shipped off-site does not enumerate where your PII lives) and every operation that would mix or hide the chain's posture refuses with SLUICE-E-BACKUP-REDACTED-CHAIN:
backup incrementalandbackup stream runrefuse to extend it, before the change window opens.- A resumed
backup fullrefuses under different--redactrules, which would otherwise leave one manifest listing chunks written under two policies. restore, chain restore,backup verify,export-as-parquetand the from-backup broker refuse a chain whose links disagree.backup compactrefuses a compaction that would merge segments of a chain carrying the marker. Compaction re-attributes one segment's manifest to several segments' merged data, which would produce a chain claiming to be redacted over plaintext — and one the read doors above would then pass. Usebackup pruneto reclaim space instead. This door is defense-in-depth: no sluice binary can currently produce a marker-carrying chain of more than one segment (both extenders refuse, and an older binary can neither write the marker nor read past the format-version stamp), so what it refuses is a hand-assembled lineage. It also runs after compaction's "fewer than two eligible segments" check, so a chain with nothing to merge exits 0 without firing.sync start --position-from-manifestrefuses unless the sync carries the same--redactrules. Resuming CDC off a redacted chain without them overwrites each restored redacted value with the plaintext one, on a live target.
Chains taken before v0.144.0 carry no marker, so a pre-existing mix cannot be detected retroactively — a v0.143.0-or-earlier chain rooted in a redacted full still restores its incrementals' rows unredacted. If you have one, re-take it.
To keep a chain PII-clean, take periodic backup full --redact runs. To accept plaintext in the change window, start an unredacted chain deliberately so the choice is explicit.
The strategy families #
sluice ships 26 strategies across five families. Pick the one that matches the column's shape.
Constant & foundational #
| Strategy | Behaviour |
|---|---|
null | Replace with NULL. Refuses on NOT NULL columns — use static: there instead. |
static:<value> | Replace every value with one literal constant. |
truncate:<n> | Keep the first N runes (rune-counted; UTF-8 and emoji safe). |
hash:sha256 | SHA-256 hex digest — deterministic, no key required. |
hash:hmac-sha256 | Keyed HMAC-SHA256 hex digest — requires --keyset-source (see below). |
Format-preserving masks #
Generic masks keep some characters and blank the rest (default mask char X):
| Strategy | Behaviour |
|---|---|
mask:inner:<m1>,<m2>[,<char>] | Keep first M1 + last M2 runes; mask the middle. mask:inner:4,4 on 4111111111111111 → 4111XXXXXXXX1111. |
mask:outer:<m1>,<m2>[,<char>] | Mask the first M1 + last M2; keep the middle. |
Country- and format-specific presets validate the input shape and preserve just the non-identifying part:
| Preset | Behaviour |
|---|---|
mask:ssn | US SSN — preserve last 4 (XXX-XX-NNNN). |
mask:pan / mask:pan-relaxed | Card PAN — preserve first 6 + last 4. mask:pan requires a valid Luhn checksum; mask:pan-relaxed skips the check. |
mask:email | First char of the local part + masked middle + full @domain. |
mask:ca-sin | Canadian SIN — preserve last 3 (Luhn-validated). |
mask:uk-nin | UK National Insurance number — keep prefix letters + suffix, mask the digits. |
mask:iban | IBAN — preserve country code, check digits, 2 BBAN, and last 4. |
mask:uuid | UUID — preserve hyphens + first 4 + last 4 hex. See the caveat below. |
mask:uuid on a native uuid column. The masked output contains X characters that aren't valid hex, so a target column typed uuid (Postgres) refuses at preflight — before any data moves — unless you also map that column to text with --type-override=table.col=text.Realistic synthetic values (randomize) #
The randomize:* generators produce fresh, valid-shape fake values — ideal when staging needs data that looks real. Output is replay-stable per source row: the same source primary key always regenerates the same target value across CDC resume, cold-start re-apply, and backup → restore (ADR-0039).
| Strategy | Output |
|---|---|
randomize:int:<min>,<max> | Integer in [min, max] inclusive. |
randomize:email | rand-local@rand-domain.test (IETF-reserved TLD). |
randomize:us-phone | NANP-valid XXX-XXX-XXXX. |
randomize:uuid | RFC 4122 UUIDv4 (passes strict UUID column validation). |
randomize:ssn | US SSN avoiding reserved ranges. |
randomize:pan[:<brand>] | Luhn-valid card PAN; optional visa / mastercard / amex. |
randomize:ca-sin | Luhn-valid Canadian SIN. |
randomize:uk-nin | UK NIN matching the HMRC prefix alphabet. |
randomize:iban[:<country>] | IBAN with mod-97 check digits; optional DE / GB / FR. |
randomize:* rule needs a primary key on the source table — the replay seed is derived from the row's PK. The pipeline refuses loudly at startup if a randomize:* rule targets a heap (no-PK) table; add a PK on the source, or pick a non-random strategy.Dictionary strategies #
Dictionary strategies map source values into a named lookup table declared in YAML (ADR-0040):
| Strategy | Keyed by | Use case |
|---|---|---|
randomize:dict:<name> | Source PK (replay-stable) | Per-row random pick with controlled cardinality. |
tokenize:dict:<name> | Source value (HMAC) | Stable per-value surrogates — the same input value maps to the same dict entry in every table and column. |
The distinction is the point: randomize:dict can send two rows with the same value but different PKs to different entries, whereas tokenize:dict guarantees every occurrence of a value (anywhere) maps to the same surrogate — so analytics joins on the redacted column stay coherent. Dictionaries must be declared in YAML; a CLI reference to an undeclared dict name refuses at parse time.
Determinism #
Redaction output is deterministic, which is what makes it safe to re-run — CDC resume and backup → restore reproduce identical surrogates on the same data. There are four contracts:
| Semantics | Strategies | Guarantee |
|---|---|---|
| Stateless | null, static:, truncate:, hash:sha256, all mask:* | Same input → same output on any sluice run, anywhere. |
| Keyed | hash:hmac-sha256 | Same input + same keyset key → same output. |
| PK-keyed replay-stable | randomize:* (incl. randomize:dict) | Same source row (table + column + PK) → same output across re-runs. |
| Input-keyed cross-stream | tokenize:dict | Same input value + same key → same output across tables, columns, and streams. |
To correlate a redacted column across tables (a join key), use tokenize:dict or hash:hmac-sha256; the other strategies don't carry cross-table consistency on the same source value.
The operator keyset (--keyset-source) #
The two keyed strategies — hash:hmac-sha256 and tokenize:dict — resolve their HMAC secret from an operator-controlled keyset (ADR-0041). Any rule using either strategy requires --keyset-source; sluice refuses loudly at preflight otherwise. The keyset is a small YAML document holding one or more named keys, each with generations so old surrogates keep resolving after a rotation. It resolves from three sources:
# keyset YAML on disk
--keyset-source=file:/etc/sluice/keyset.yaml
# keyset YAML in an env var (container / secret-manager friendly)
--keyset-source=env:SLUICE_KEYSET
# sluice-managed sluice_keysets table on a DSN — shared across streams
--keyset-source=db:postgres://user:pw@host:5432/keysetdb
A rule names which key it uses via the trailing :<keyname> segment (or a YAML key: field); omit it to use the keyset's declared default or its sole entry. Two rules that name the same key produce cross-consistent surrogates:
--redact users.email=hash:hmac-sha256:customer_pii
--redact users.first_name=tokenize:dict:first_names:customer_pii
The db: form is the cross-stream stability primitive: two streams pointing at the same keyset DSN turn alice@example.com into the same surrogate on staging-1 and staging-2. For two independent installs to agree (cross-org exchange), install the same file: keyset at both ends.
Preflight refusals #
sluice runs these checks before any data movement, on every lane that redacts rows — migrate (single- and multi-database), sync start cold start (single- and multi-database), schema add-table, and backup full. When more than one fires in the same run they aggregate into a single combined error, so you see the full picture in one pass.
- Selector resolution. Every rule's
[schema.]table.columnmust resolve to a real column in the source schema, and a schema-qualified rule must resolve to a table in that namespace. A rule that resolves to nothing is refused as a typo class — it would apply to nothing and ship the column in clear. Two namespace shapes matter:- Single-database MySQL is flat scope. Tables carry no namespace there, so a qualified rule such as
source_db.users.emailcan never match a bulk-copied or backed-up row (only CDC rows carry the binlog database name). Through v0.137.4 that rule passed preflight and the column shipped unredacted through bulk copy and backup at exit 0 while CDC rows were redacted. Since v0.138.0 it is refused: use the bareusers.emailform, or select the database explicitly (--include-database=source_db) so tables are namespaced. - Multi-database / multi-schema runs validate a qualified rule in the pass for its own namespace; a qualified rule naming a namespace outside the selected set is refused up front.
- Single-database MySQL is flat scope. Tables carry no namespace there, so a qualified rule such as
mask:uuidon a UUID-typed column refuses unless--type-override=col=textshort-circuits the target column type.randomize:*on a table with no primary key refuses; add a PK to the source or pick a non-random strategy.hash:hmac-sha256/tokenize:dictwith no resolvable keyset refuses — supply--keyset-sourceand reference a key via the rule'skey:option (or rely on the keyset default / sole entry).
Preview redaction before you run #
sluice schema preview annotates the generated DDL so you can eyeball which columns are covered before moving a single row. The annotation is comment-only — the CREATE TABLE itself is unchanged, so the output stays drop-in usable:
sluice schema preview \
--source-driver postgres --source "$SRC" \
--target-driver postgres --target "$DST" \
--redact users.email=hash:sha256 \
--redact users.ssn=mask:ssn
CREATE TABLE users (
id SERIAL PRIMARY KEY,
email TEXT NOT NULL, -- REDACTED via hash:sha256
ssn TEXT, -- REDACTED via mask:ssn
...
);
The audit log #
Every command that moves rows emits exactly one INFO line at startup recording the configured redaction surface — the scope, the column count, and the distinct strategy names:
sluice: redaction configured scope=migrate columns=5 strategies=[hash:sha256 mask:pan randomize:email tokenize:dict:first_names truncate:4]
Per-column rules are deliberately not logged — the mapping itself is sensitive (--redact billing.credit_card=truncate:4 reveals which column holds card numbers), and per-row surrogates are never logged. When a keyset loads, a second line records its source scheme and per-key generations, with any DSN credentials redacted.
Worked examples #
Mask and hash for an analytics copy #
sluice migrate \
--source-driver postgres --source "$SRC" \
--target-driver postgres --target "$DST" \
--redact users.email=hash:sha256 \
--redact users.phone=mask:inner:3,4 \
--redact users.ssn=randomize:ssn
Realistic synthetic data for a live staging sync #
Redaction is honoured on the CDC stream too, so staging stays continuously fresh and continuously scrubbed. YAML config plus a stream id:
# sluice.yaml
redactions:
- table: users.email
strategy: randomize
form: email
- table: users.phone
strategy: randomize
form: us-phone
- table: customers.pan
strategy: randomize
form: pan
brand: visa
sluice sync start -c sluice.yaml \
--source-driver postgres --source "$SRC" \
--target-driver postgres --target "$DST" \
--stream-id staging-refresh
Cross-table stable surrogates for a vendor handoff #
Use tokenize:dict with one shared key so a customer's name is the same token in every table — the vendor can still join, but never sees the real value:
# sluice.yaml
dictionaries:
first_names:
entries: [Alpha, Bravo, Charlie, Delta, Echo, Foxtrot]
redactions:
- table: users.first_name
strategy: tokenize
dict: first_names
key: customer_pii
- table: orders.customer_first_name
strategy: tokenize
dict: first_names
key: customer_pii
sluice migrate -c sluice.yaml \
--source-driver postgres --source "$SRC" \
--target-driver postgres --target "$DST" \
--keyset-source=file:/etc/sluice/keyset.yaml
What redaction is not #
- Not a PII discovery scanner. sluice redacts the columns you name; it does not crawl the schema to find which columns hold personal data. Identifying them is your (or your compliance team's) job.
- Not encryption at rest. Redaction transforms values in flight so the sensitive original never reaches the target or the backup. Protecting the keyset secret and the target storage itself is your storage layer's responsibility — sluice does not encrypt the key bytes at rest.
Next steps #
- Configuration — the YAML
redactions:,dictionaries:, and keyset blocks in full. - Command reference — the flag set for
migrate,sync,backup, andschema preview. - Getting started — install sluice and run your first migration and sync.
