Adding, deleting or reordering tiers while the entitlement lock is held answers "see server logs" instead of the refusal that says to try again #216

Open
opened 2026-10-10 23:59:47 +00:00 by cgalo5758 · 0 comments
Owner

What happens

Every transaction that recomputes an organization's entitlements first takes the materialization lock, a database advisory lock. Most take it in shared mode. A rule change takes it in exclusive mode, so the two never overlap. Each waits at most 5 seconds (the lock timeout) and then gives up. Tier add, tier delete and the reorder open their transactions this way, and when that wait times out each one reports a plain failure:

  • Tier add: "Failed to add tier; see server logs."
  • Tier delete: "Failed to remove tier; see server logs."
  • Reorder: "Failed to reorder tiers."

The tier removal and the default change answer the same wait with "Another entitlement set change is in progress. Try again." (tierRemovalRefusal and dispositionFailure, both of which check db.IsLockContention). Tier add also answers "Failed to add tier; see server logs." for a lock wait or deadlock in its later steps, including the deadlock described in #194.

Today the wait times out only when a rule change holds the lock for longer than 5 seconds. #213 will make it common. Its fix opens the transactions that change a ladder's tiers (add, delete, removal and reorder) with the exclusive lock. Each of these acts will then wait for every conferral, grant, subscription reconcile and recomputation already running.

What should happen

A timed-out wait or a deadlock in tier add, tier delete or the reorder is answered with "Another entitlement set change is in progress. Try again.", the same answer as the removal and the default change. The entitlements spec ("A transaction that materializes a pool is ordered against a rule change") names tier add as one of the transactions it binds and requires that refusal for a timed-out wait and for a deadlock. Tier delete and the reorder take the same lock and should answer the same way. The full error is still logged.

Where

internal/server/operator_plan_ladders.go: CreatePlanLadderTier, DeletePlanLadderTier and ReorderPlanLadderTiers, in the failure branches for opening and committing the transaction, and in tier add's create, align and materialize steps. TestPlanActsRefuseWhileARuleChangeHoldsTheRendezvous in internal/server/plan_acts_concurrency_db_test.go holds the exclusive lock while each act runs, and its recorded answers in internal/server/testdata/plan_acts/contention.json will change with the fix.

Steps

The test above reproduces it: it holds the lock with entitlements.BeginRuleChange and sends each act.

Why it matters

An operator is told an act failed and is sent to the server logs for a wait that clears by itself, when trying again would succeed. Once #213 is fixed, this answer will appear whenever a tier change meets ordinary traffic.

### What happens Every transaction that recomputes an organization's entitlements first takes the materialization lock, a database advisory lock. Most take it in shared mode. A rule change takes it in exclusive mode, so the two never overlap. Each waits at most 5 seconds (the lock timeout) and then gives up. Tier add, tier delete and the reorder open their transactions this way, and when that wait times out each one reports a plain failure: - Tier add: "Failed to add tier; see server logs." - Tier delete: "Failed to remove tier; see server logs." - Reorder: "Failed to reorder tiers." The tier removal and the default change answer the same wait with "Another entitlement set change is in progress. Try again." (`tierRemovalRefusal` and `dispositionFailure`, both of which check `db.IsLockContention`). Tier add also answers "Failed to add tier; see server logs." for a lock wait or deadlock in its later steps, including the deadlock described in #194. Today the wait times out only when a rule change holds the lock for longer than 5 seconds. #213 will make it common. Its fix opens the transactions that change a ladder's tiers (add, delete, removal and reorder) with the exclusive lock. Each of these acts will then wait for every conferral, grant, subscription reconcile and recomputation already running. ### What should happen A timed-out wait or a deadlock in tier add, tier delete or the reorder is answered with "Another entitlement set change is in progress. Try again.", the same answer as the removal and the default change. The entitlements spec ("A transaction that materializes a pool is ordered against a rule change") names tier add as one of the transactions it binds and requires that refusal for a timed-out wait and for a deadlock. Tier delete and the reorder take the same lock and should answer the same way. The full error is still logged. ### Where `internal/server/operator_plan_ladders.go`: `CreatePlanLadderTier`, `DeletePlanLadderTier` and `ReorderPlanLadderTiers`, in the failure branches for opening and committing the transaction, and in tier add's create, align and materialize steps. `TestPlanActsRefuseWhileARuleChangeHoldsTheRendezvous` in `internal/server/plan_acts_concurrency_db_test.go` holds the exclusive lock while each act runs, and its recorded answers in `internal/server/testdata/plan_acts/contention.json` will change with the fix. ### Steps The test above reproduces it: it holds the lock with `entitlements.BeginRuleChange` and sends each act. ### Why it matters An operator is told an act failed and is sent to the server logs for a wait that clears by itself, when trying again would succeed. Once #213 is fixed, this answer will appear whenever a tier change meets ordinary traffic.
cgalo5758 added the
kind
bug
area/catalogarea/operator-ui
labels 2026-10-10 23:59:47 +00:00
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: wiki-cafe/member-console#216