The reorder's renumbering and the enactment of a new rank 0 on the
organizations lived in the HTTP handler. They are now
entitlements.ReorderTiers, which takes a ReorderTiersInput: the ladder, the
order, and, when the rank 0 changed on a ladder some organization type
defaults to, an Enact value naming the types, the disposition, the operator
and the readers it needs.
ReorderTiers runs the same steps in the same order: it opens BeginRuleChange
(the exclusive rendezvous, no pool lock), renumbers the ladder and commits.
With Enact set it then reads the ladder's name for the reason (the id stands
in when the read fails), composes "plan-ladder tier reorder (<ladder>)", and
for each type lists its organizations and enacts each through EnactDefault
against the new rank 0. A failed listing returns the result so far and a
PlanActError that now also carries the organization type.
ReorderPlanLadderTiers keeps the session, the form parse, the full preview,
the stale check and the disposition gate. tierReorderAnswer returns the
stale sentence, the rank-0 sentence, the toast naming the types' display
names and the failed organizations' answer; tierReorderFailure returns and
logs the sentences for a failed begin, renumbering, commit and listing. The
failures of each type's organizations are logged after the enactment instead
of after each type; the lines and their order are the same.
One departure, as for the other moved tier acts: a failed renumbering ("This
page was out of date; showing the current tier order. Try again.") and the
other failures before the commit are now answered after the act's rollback,
which releases the rendezvous before the page is read, where the handler
rendered them before its deferred rollback. Only reads move: no write, no
sentence and no lock changes, and the lock is held for less time. Under
concurrency a request queued on the rendezvous can now commit before the
page is read. The stale-candidate case is unchanged: the enactment still
classifies against the rank 0 the reorder saved.
Docs correction that rides along: the reorder entry in docs/database-locks.md
names entitlements.ReorderTiers.
ReorderPlanLadderTiers had a cognitive complexity of 30, the limit, with no
exception to delete. Measured: 14 at this commit; ReorderTiers 4,
enactReorder 6 (the largest new helper), tierReorderAnswer 7,
tierReorderFailure 2.
No output changes. Names and results of ./internal/server,
./internal/entitlements and ./internal/testkit/... are equal at the parent
and at this commit after a database reset on each side: 1779 results on
each side, 1778 passing and the one existing skip. No test, golden or
testdata file is edited and none ran with -update, so every golden is
byte-equal. The seeded operator walkthroughs (TestPlanLaddersWalkthrough and
TestOrgTypesWalkthrough among them) give equal names and results on both
sides, and the screens of operator-plan-ladders, plan-ladder-detail (and its
open and invalid states) and org-types, desktop and mobile, are
byte-identical at the parent and at this commit. Every string literal of the
old operator_plan_ladders.go and operator_org_types.go is still found in a
non-test file of internal/server or internal/entitlements.
15 breaks (24 records), each caught by the old code and by the new code
where its text exists on both sides: a stale submission says so, a new rank
0 on no type's default says so, the reorder's reason names the ladder,
organizations are enacted against the new rank 0, a failed organization's
answer, a new rank 0 on a default ladder is enacted, the gate counts the
organizations, a failed begin is logged, a failed listing after the reorder,
the toast names the types by their display names, the reorder renumbers the
ladder, the reorder takes the exclusive rendezvous, a reorder in contention
fails generically, a failed read of the types reads as no type's default,
and a cosmetic reorder says the order was updated. None was dropped.
Not probed: faults inside the reorder's transaction (the renumbering and the
commit; no querier makes them fail), so the page-out-of-date sentence and
its log line are read but not seen, and the move of its answer after the
rollback is named, not shown; the label's fallback to the ladder id when the
ladder read fails, which no querier refuses; a listing failure on the second
of two types, since the listing fault refuses every listing and the first
type fails; and the browser end-to-end plan-management tests, which were not
run here.
23 KiB
title, audience, summary
| title | audience | summary | |
|---|---|---|---|
| Database Locks |
|
Every advisory lock and serializing row lock the console takes, how each key is built, and the order each path takes them in. |
Database Locks
This page describes the locks as the code takes them today. It records what each lock serializes, who takes it, and the order each path takes several. It does not describe a target design. Two pieces of planned work change the page: #186 gives every advisory key one family prefix and writes one lock order, and #165 adds a lock on the usage row under the FedWiki workspace lock. Whoever lands either edits this page in the same change.
The code names functions and sqlc queries; so does this page. Line numbers are left out because they move.
Every advisory lock is transaction-scoped (pg_advisory_xact_lock), so
Postgres releases it at commit or rollback. No code takes a session-level
advisory lock, LOCK TABLE, or an explicit FOR SHARE.
The locks
Advisory locks
| Lock | Key | Mode | Taken by | Serializes |
|---|---|---|---|---|
| Materialization rendezvous | hashtextextended('entitlement_materialization', 0) |
shared for a transaction that materializes a pool; exclusive for a rule change and for a change to a ladder's tiers (tier add, tier delete, tier removal, the reorder's renumbering) | the entitlements.Materializing and entitlements.RuleChange transaction kinds (BeginMaterializing, BeginRuleChange), as the transaction's first statement |
A rule change or a tier change against every materializing transaction: each one finishes before the change or starts after it, so a conferral reads one set of rules and one order of each ladder. Materializing transactions do not wait on each other here. |
| Subscription | hashtext(<Stripe subscription id>)::bigint |
exclusive | fulfillment.ReconcileSubscription |
Two reconciles of one subscription (the checkout return and a webhook), so they converge instead of racing the exclusion constraint on subscription periods. |
| Invoice | hashtext(<Stripe invoice id>)::bigint |
exclusive | invoiceReconcile.converge in internal/integrations/stripe/workflows |
Two reconciles of one invoice (two events, or an event and the sweep), so the second finds what the first wrote. |
| Product sync | hashtext(<product id>)::bigint |
exclusive | catalogsync.Sync in internal/integrations/stripe/catalogsync |
Two sync clicks for one product, so the mapping writes and the outbox enqueue happen once. |
| FedWiki workspace | hashtext(<workspace id>)::bigint |
exclusive | withWorkspaceTx in internal/integrations/fedwiki/web, for the status-change and swap handlers |
One workspace's rotations of its active site, so two tabs cannot interleave the cooldown check and stamp or the choice of the site to park. |
| Lifecycle instance | hashtextextended('lifecycle:' || <instance id>, 0) |
exclusive | the LockLifecycleInstance query, run by integration.RecordLifecycleRequest, RecordFirstLifecycleRequest and the FedWiki activity helper finishAllWith |
One instance's requests: the latest-generation check and the insert are one step, and an outcome is recorded against a request that is still accepted. |
| Desired-state address | hashtextextended('desired-state:' || <connection id> || ':' || <recipient kind> || ':' || <recipient id>, 0) |
exclusive | the LockDesiredStateAddress query, run by lockAddress in internal/workflows/desiredstate |
Every write to the content or generation of the record at one (connection, recipient) address, and to its projection, including an address that has no record yet. |
| Domains registry | the constant registryAllocationLock (the ASCII bytes of "domains") |
exclusive | domains.Registry.WithLock |
Every mutating allocation in the registry, one lock for the whole namespace, because the disjointness rules span rows no database constraint can compare. |
Row locks
| Lock | Row | Mode | Taken by | Serializes |
|---|---|---|---|---|
| Pool row | core.resource_pools row of one pool |
FOR UPDATE |
the functions core.confer, core.end_conferral, core.align_conferral_shape, core.update_conferral_bounds and core.settle_obligation; LockPools (through the LockResourcePool query), called by RevokeGrantTx, ManageGrantTx, ResumeReplacedGrantsTx, settlePool, EnactDefault and RemoveTier; and LockPoolOfGrant (through the LockGrantPool query), called by ExpireGrantTx and ExtendWaitingGrantTx |
Everything that changes a pool's positions or entitlements: conferral, endings, grant acts, tier removal, a default change, rule-change settlement and grant resumption. |
| Entitlement set row and rule row | core.entitlement_sets row and core.entitlement_set_rules row |
FOR UPDATE |
core.commit_rule_change (the set row also by the LockEntitlementSet query, once per batch) |
A rule change against a writer that changes the set or the rule without holding the rendezvous. |
| Change obligation row | core.entitlement_set_change_obligations row of one (change, pool) |
FOR UPDATE |
core.settle_obligation |
Two passes settling one owed pool, so each effect is written once. |
| Grant lineage head | the core.grants row no other grant extends |
FOR UPDATE OF g |
the GetGrantLineageHeadForUpdate query, run by RevokeGrantTx, ExtendWaitingGrantTx and ManageGrantTx |
Acts on one lineage, so an extension that committed while another act waited is visible to it. |
| Usage row | core.numeric_entitlement_usage row of one (pool, resource key) |
FOR UPDATE in LockNumericUsage; the row lock of the statement in AtomicIncrementUsage, AtomicDecrementUsage, RaiseNumericUsageTo and BoundNumericUsage |
usage.BoundWorkspace takes the explicit lock; the atomic statements lock the row for the length of their transaction |
The counter against a boot repair that counts rows and then writes the counter. |
| FedWiki site row | fedwiki.sites row of one domain |
FOR UPDATE |
the LockSiteStatusByDomain query, run by sweepLocally, RecordSiteStatusActivity and RecordSwapActivity |
Two paths that move one site out of the active status (a member's request and the downgrade sweep), so only the first to commit releases the site's slot. |
| Desired-state record | core.desired_state_records row |
FOR UPDATE |
the LockDesiredStateRecord query (after the address lock); StampDesiredStateRecordsDispatched, which locks the records in identifier order |
The record's content and generation within one recompute; the dispatch stamp against another stamp. |
| Desired-state notice | core.desired_state_outbox rows |
FOR UPDATE SKIP LOCKED |
the ClaimDesiredStateOutbox query |
Two relay passes, so each claims different rows; a pass skips rows another still holds. Triggers write the rows through core.insert_desired_state_notice inside the transaction that changed the source, and they take no lock beyond the insert. |
| Resource-key row | core.resource_keys row of one key |
the update lock of SetResourceKeyProvider |
integration.RegisterProviders at boot, stamping keys in byte order |
A stamp that changes a key's provider against an open entitlement insert of that key, which holds the row's FOR KEY SHARE through the foreign key. The materializer writes a pool's keys in byte order too, so the two wait in one direction. |
How advisory keys are built
Postgres takes one 64-bit key for an advisory lock, so every lock above maps its identity onto a bigint. The code builds those keys four ways.
hashtextextendedof a fixed name. The rendezvous. One key for the whole database.hashtextextendedof a prefixed string. The lifecycle instance (lifecycle:and the instance id) and the desired-state address (desired-state:, the connection, the recipient kind and the recipient id). The prefix keeps the two families from sharing keys with each other.hashtextof a bare id, widened to bigint. The subscription, the invoice, the product sync and the FedWiki workspace.hashtextreturns 32 bits, so these four families share one 32-bit key space with no prefix between them; a collision serializes two unrelated objects briefly and cannot lose a write. The ids differ in shape (Stripe ids against UUIDs), which makes a cross-family collision unlikely, not impossible.- A literal constant. The domains registry. Its value is larger than any 32-bit hash, so it cannot collide with family 3.
The order each path takes locks
Each path below lists its locks in the order it takes them. "Rendezvous"
means the shared lock a transaction of the entitlements.Materializing kind
(BeginMaterializing) takes as its first statement; the entitlements.RuleChange
kind takes the exclusive one. Those two kinds are what set
SET LOCAL lock_timeout = '5s', which bounds every later lock wait in the
transaction. A plain transaction sets no timeout, so its lock waits last until
its context ends. A wait past the bound surfaces as SQLSTATE 55P03. The rule
change, tier removal and default change render it as the refusal for a change
in progress; the other paths show their own failure copy. Materializing
transactions run at READ COMMITTED.
Rule change and drain
- Commit (
RuleChangeBatchCommit): the exclusive rendezvous (BeginRuleChange), then the set row (LockEntitlementSet, once for the batch), then, for each delta in resource-key order, the set row again (the same lock) and the rule row insidecore.commit_rule_change. At or below the sync cap the same transaction then takes each carrying pool in ascendingpool_idorder (ListPoolsCarryingSet):settlePoollocks the pool row throughLockPools, one pool per call, then the obligation row and the pool row again insidecore.settle_obligation, once per act. - Drain (
DrainChange, called by the commit's own drain above the cap and by the poller's activity): one short transaction per owed pool, opened throughBeginMaterializing. It takes the rendezvous, the pool row (settlePool, throughLockPools), then the obligation row of each owed act in the batch's order. A failed pool rolls back, andrecordObligationFailurerecords the attempt in a second transaction that takes the rendezvous alone.
Grant acts
All of these run in a transaction from BeginMaterializing: the operator
handlers, ExpireGrantActivity and the demo seed.
- Issue or confer a grant (
ConferGrantTx, whichConferGrant, provisioning andReapplyDefaultsForPoolcall): the rendezvous, the grant insert, then the pool row insidecore.confer, then materialization. - Revoke (
RevokeGrantTx): the rendezvous, the organization's default pool row (LockPools), the lineage head row, thencore.end_conferral(its pool locks are already held), resumption (the pool row again), the floor restoration (core.confer) and materialization. - Expire (
ExpireGrantTx): the rendezvous, then the pool row throughLockPoolOfGrant, then the grant update. It does not take the lineage head row. A grant that names no organization has no pool, so the lock, resumption and materialization are skipped. - Extend a waiting grant (
ExtendWaitingGrantTx): the rendezvous, the pool row (LockPoolOfGrant), the lineage head row. - Manage (
ManageGrantTx): the rendezvous, the pool row (LockPools), the lineage head row, then eitherExtendGrantTx(inside itcore.conferandcore.end_conferral, whose pool lock is already held) orExtendWaitingGrantTx(the same pool row again). - Resume replaced grants (
ResumeReplacedGrantsTx): called inside the transactions above and in the subscription reconcile. It takes the pool row first and reads the resumable grants only after the lock is held, so an expiry, revocation, mark change or extension is either wholly visible to it or waits for it. Each resumed grant callscore.confer.
Subscription reconcile
ReconcileSubscription reads the subscription from Stripe before it opens a
transaction. The transaction then takes the rendezvous, then the subscription
advisory lock, then the rest by branch.
- Active:
reconcileItemsconfers a missing product (core.confer), changes a quantity (core.update_conferral_bounds) and ends a surplus provision (core.end_conferral). Each of those functions takes the pool row of the provision it touches, so the pool row is first taken at the first of those calls. Materialization follows.convergeScheduledChangesthen writes the scheduled-change ledger. A spent Stripe subscription schedule it finds is released after the commit (releaseSpentSchedule), with no lock held. When the release succeeds and the schedule held a downgrade row still scheduled,supersedeReleasedDowngradesopens a second transaction, takes the subscription advisory lock and supersedes those rows; a released schedule that held none opens no second transaction. It takes no rendezvous and no pool row: it writes only scheduled-change rows. - Ended:
core.end_conferralby subscription (the pool row of each provision it ends), then, only when something ended, resumption (the pool row), then the floor restoration (ReapplyDefaultsIfVacant, which confers throughcore.confer), then materialization. - Suspended:
core.sync_source_status, which takes no pool lock, then materialization with no pool row lock held.
Tier add, tier removal, default change and reorder
The tier add, the tier removal, the tier delete and the reorder's
renumbering run in transactions from BeginRuleChange, which takes the
exclusive rendezvous: each changes a ladder's tiers, and core.confer reads a
ladder's ranks and a product's shape in separate statements, so no conferral
may run while one of them commits (#213). The default change and each
organization's enactment run in transactions from BeginMaterializing.
- Tier add (
entitlements.AddTier): the exclusive rendezvous, the tier insert, then for each live provision of the product in creation ordercore.align_conferral_shape(the pool row of that provision) and materialization. The pools are locked in provision order, not inpool_idorder. - Tier removal with holders (
entitlements.RemoveTier): the exclusive rendezvous, then the pool row of every holder throughLockPools, which sorts and deduplicates them, so they lock in ascendingpool_idorder, then the holders are read again under the locks, the tier is deleted and the ranks renumbered. Then, for each pool in the same order, the incumbents end (core.end_conferral), the remaining positions align (core.align_conferral_shape), the floor is restored when the disposition is migrate, and the pool materializes. - Tier delete without holders (
entitlements.DeleteTier): the exclusive rendezvous only, with no pool lock, around the tier delete and the renumbering. - Default change (
entitlements.ChangeOrgTypeDefaultandEnactDefault): the default is saved outside any lock beyond the row update. Then each organization is its own transaction: the rendezvous, the organization's default pool row (LockPools), then the bucket's work (ReapplyDefaultsIfVacant, orConferGrantTxfor a grandfathered incumbent, orcore.end_conferral,ReapplyDefaultsIfVacantand materialization for a migration). A failure on one organization leaves the others alone. - Reorder (
entitlements.ReorderTiers): the exclusive rendezvous around the renumbering, committed with no pool lock. When the rank-0 tier changed, each affected organization then runsEnactDefaultunder the shared rendezvous, as a default change does.
Invoice reconcile
invoiceReconcile.converge opens a plain transaction, with no rendezvous, and
takes the invoice advisory lock. It writes invoice, line and payment rows
under that lock and touches no pool.
Product sync
catalogsync.Sync opens a plain transaction, takes the product advisory
lock, re-reads the price and product mappings, writes the pending mappings
and enqueues the outbox rows. It takes no pool lock and no rendezvous. A sync
that commits nothing returns with its transaction open, and
SyncProductToStripe rolls it back after it has rendered its answer, so the
lock is held until then.
FedWiki at the handler
withWorkspaceTx takes the workspace advisory lock, then runs the handler's
statements. A swap (swapActiveSite) and a rotation with a cooldown
(changeStatus) then record a lifecycle request, which takes the lifecycle
instance lock of the activated site. A swap records a second request for the
parked site, which takes that instance's lock after the first. The swap policy
row is upserted last. The handler takes no usage-row lock and no site-row
lock. A status change without a cooldown skips the workspace lock and records
its request in a transaction of its own, taking only the instance lock.
When recording an outcome
The workflow activities that finish a lifecycle request run
finishAllWith: the lifecycle instance locks, in the order the handler
recorded the requests (the activated site first, then the parked site), then
the activity's own writes.
RecordSiteStatusActivity: the site row (LockSiteStatusByDomain), then the usage row throughAtomicDecrementUsagewhen the site leaves the active status.RecordSwapActivity: the site row of the parked site, then the site row of the activated site, then the usage row (AtomicIncrementUsageorAtomicDecrementUsage, by what each site's status calls for).RecordSiteDeletedActivity: the usage row throughAtomicDecrementUsagewhen the deleted row held a slot.
So a transaction that holds a site row and the usage row takes the site row first.
In the downgrade sweep and at boot
- Park (
parkSite): the farm call first, with no lock held, then one transaction (sweepLocally) that takes the site row and, in the same transaction, decrements the usage row. - Reactivate (
reactivateSite): the usage increment as a statement of its own, before the farm call, then a second transaction that takes the site row. The two never overlap, so they form no cycle. - Boot repair (
usage.BoundWorkspace): one transaction that takes the usage row withLockNumericUsage, counts the site slots without locking any site row, then writes the bounded counter. - Provider registration (
RegisterProviders): the resource-key rows in byte order. - Domains reconciliation at boot (
reconcile_domains.goin the FedWiki package): oneRegistry.WithLocktransaction per orphaned placement or adopted site.
Lifecycle intake
RecordLifecycleRequest takes the instance lock first, then reads the latest
generation, then inserts. RecordFirstLifecycleRequest takes the same lock
before calling it, which is harmless because the lock is re-entrant within a
transaction. Called with a *sql.DB it opens its own transaction; called with
a transaction it uses the caller's, which is how the handler records under the
workspace lock.
Desired-state relay
- Claim (
ProcessDesiredStateOutbox):ClaimDesiredStateOutboxis one statement that locks the pending rowsSKIP LOCKEDand leases them. The relay never takes the rendezvous and never locks a source row. - One recompute (
recomputeAddress, and the same sequence inreadBackItemandwritePushAnswer): one READ COMMITTED transaction that takes the address advisory lock, then the record row (LockDesiredStateRecord), then computes the content in a new statement, then updates or inserts the record. Every transaction that writes a record's content or generation, or a projection, takes the address lock first. - Dispatch stamp (
StampDesiredStateRecordsDispatched): one statement that locks the records in identifier order before it writes. It takes no address lock.
Registry
domains.Registry.WithLock opens a plain transaction and takes the registry
lock before anything else. Every composite operation, including the FedWiki
create, release and sync projections, runs on the *Tx it passes, so the lock
is taken once. The comment on registryAllocationLock states the rule for a
transaction that would touch both the registry and a usage row: registry lock
first. No transaction does both today: the usage increment and decrement are
workflow activities that run their own statements, apart from the transaction
that allocates or releases the name.
Rules every path keeps
- A transaction that materializes a pool takes the shared rendezvous as its
first statement, through
BeginMaterializing. A rule change or an operator's tier change (the tier add, the tier delete, the tier removal, the reorder's renumbering) takes the exclusive one throughBeginRuleChange, and may materialize under it; the demonstration seed adds its ladder's tiers on the plain handle, outside the rendezvous.core.assert_rendezvousraisesmaterialization_rendezvous_missingwhen a function that materializes or settles runs without it. - The pool row is locked before positions or entitlements change, except in the suspended branch of the subscription reconcile (see "Where the paths differ"). The five conferral and settlement functions lock the row themselves, so a caller that goes straight to one of them still holds it.
- A transaction that locks several pools takes them in ascending
pool_idorder, throughLockPools. The rule change and the tier removal do; tier add does not (see "Where the paths differ"). The deferred drain locks one pool per transaction. - The grant lineage head row is locked after the pool row, never before.
- A transaction that holds a site row and the usage row takes the site row first.
- Lifecycle instance locks of one transaction are taken in the order the requests were recorded.
- A desired-state write takes the address lock, then the record row, then computes content in a new statement.
- Every advisory lock belongs to its transaction, so a commit or a rollback releases it, and an error never leaves one behind.
Where the paths differ
- The plan-ending paths take the pool lock at different points (#180).
Revocation and expiry lock the pool before they write. The subscription
reconcile's ended branch ends positions first, so the pool row is taken
inside
core.end_conferral, and takes it again for resumption. Each path then runs resumption, the floor restoration and materialization in the same order. - Advisory keys are built four ways (#186). The keys follow the four conventions above.
- The suspended branch of the reconcile materializes without a pool row lock. It holds the shared rendezvous and the subscription lock only.
- Tier add locks pools in provision order, where tier removal sorts them.
- The FedWiki usage counter has no explicit lock in the request paths. The atomic statements lock the row for their own duration. #165 adds a usage-row lock under the workspace lock, and its position in the order above belongs on this page when it lands.