What Liquibase is actually tracking
Liquibase manages schema the way git manages source: a sequence of small, ordered, idempotent changes applied to a target. Everything else in this piece (the pipeline, the lock, the logging context) exists to make that sequence safe to run more than once, and safe to run from more than one place at a time.
- changelog
- A file (XML / YAML / JSON / SQL) listing changes to apply, in order.
- changeset
- An atomic unit inside a changelog, keyed by the triple
(id, author, filepath). Holds one or more database-agnostic changes (createTable,addColumn) each translated into dialect-specific SQL by a per-databaseSqlGenerator. DATABASECHANGELOG- A table Liquibase creates in your target database. Every changeset that runs successfully gets a row here. On the next run, Liquibase reads this table and skips anything already recorded, which is what makes
updatesafe to re-run. DATABASECHANGELOGLOCK- A one-row mutex table. Before touching schema, Liquibase must flip that row's
LOCKEDcolumn toTRUE; when done, flip it back. Stops two concurrent runs (two app instances booting at once) from racing on the same schema.
That lock table is the thing at the center of Case B. Before we get there, two pieces of internal machinery need explaining: the pipeline that executes a command, and the context object that carries state through it.
The command pipeline
Every operation (update, updateSql, rollback, tag) runs as a pipeline of small, single-purpose CommandStep objects, wired together automatically by declared dependencies rather than hand-assembled.
Two-phase execution
A CommandScope holds one invocation's argument values plus a bag of typed dependencies (Database, DatabaseChangeLog, LockService) that steps hand to each other. Each CommandStep declares what it requiredDependencies() (needs already present) and what it providedDependencies() (produces for later steps). Given a command name, CommandFactory resolves those declarations transitively into an ordered pipeline, a small dependency-injection graph solved once per command.
Execution then runs in two passes over that same pipeline: run() on every step, in order; then cleanUp() on every step that implements it, in reverse order. That reversal is inert almost everywhere. It is the entire mechanism behind Case B.
LockServiceCommandStep gets inserted purely because UpdateCommandStep asked for LockService.One wrinkle matters later: UpdateCommandStep (plain update) deliberately opts out of this auto-wired lock. It drops LockService.class from its own requiredDependencies() and acquires the lock itself, lazily, only after confirming there's real work to do, a fast-check optimization that lets a no-op update return without ever touching the lock table at all. updateCount, updateSql, and updateToTag did not opt out. They still let the pipeline acquire the lock for them. That asymmetry is where Case B starts.
Scope & MDC
Scope is Liquibase's own context-propagation mechanism, conceptually a ThreadLocal-backed stack, the same job Go's context.Context does. Scope.child(values, runnable) pushes a nested scope for a block's duration; Scope.enter() / exit() is the manual open/close pair used across coarser boundaries, like CLI startup and shutdown.
MDC (Mapped Diagnostic Context, the same term SLF4J and Logback use) is structured-logging data attached to the current context: key/value pairs like changeSetId=… that get stamped onto every log line automatically, without threading a parameter through every method that logs. Scope.addMdcValue(key, value, removeWhenScopeExits) registers one of these, tagged for cleanup when its owning scope exits.
The detail that makes Case A possible: Liquibase stores the active scope-id in an InheritableThreadLocal. A thread spawned from inside a scope inherits that same scope-id: both threads are logically “in” the same scope, for MDC purposes, without either calling Scope.child(). That's intentional and useful, a worker pool started mid-update should share its parent's logging context. It's also exactly the crack the bug lived in.
Shared bookkeeping, separate threads
Concept: cleanup scoped to an ID instead of to the thread that did the work
The bookkeeping for “which MDC entries were added under scope X, so I know what to close when X exits” lived in one field, shared by every thread in the process:
private static final Map<String, List<MdcObject>> addedMdcEntries
= new ConcurrentHashMap<>();
Keyed by scopeId, not by which thread added the entry. So when a parent thread and a child thread share an inherited scope-id, they're also, invisibly, sharing one cleanup ledger.
job under the shared scope-id S; the parent's exit(S) has no way to know that entry belongs to a different thread, and closes it anyway.Partition the ledger itself by thread, so a parent's exit can only ever reach entries its own thread registered:
private static final Map<String, List<MdcObject>> addedMdcEntries
= new ConcurrentHashMap<>();
private static final ThreadLocal<Map<String, List<MdcObject>>> addedMdcEntries
= new ThreadLocal<>();
liquibase/Scope.java: field declaration
Every read and write of that map moves behind .get() on the calling thread's own slot. exit() now only ever sees, and only ever closes, entries the exiting thread itself added; a parent exiting can no longer see the child's map at all, let alone close entries in it. Standard remedy for shared mutable state hiding behind an implicit context: stop sharing it, key it per-thread instead.
A first attempt to reproduce it gave the child thread a fresh scope through Scope.child([:]). Parent and child never shared an id, so the race could not trigger and the check passed on the broken code too. The bug only appears when the child inherits the parent's scope id, records its value under that id, and the parent exits while the child is still running.
Who owns the lock?
Concept: a local record of ownership versus the live state of the lock
AbstractUpdateCommandStep tracked lock ownership with one ThreadLocal<Boolean> isDBLocked, defaulting to true. Its finally block released the lock whenever that flag was true, regardless of whether this method had actually acquired it. For updateCount, updateSql, and updateToTag, the pipeline's own LockServiceCommandStep had already taken the lock before this code ran at all (recall Part 2: they never opted out of the auto-wired lock). So the lock got released twice (once here, incorrectly, and once correctly from LockServiceCommandStep.cleanUp()) and the second attempt failed against an already-unlocked row, logging “could not release lock.”
Stop trusting the ThreadLocal's memory of what happened. Check the lock's actual, shared, live state (hasChangeLogLock()) immediately before attempting a release, at both release sites:
} finally {
if (isDBLocked.get()) {
try {
LockServiceFactory.getInstance()
.getLockService(database).releaseLock();
} catch (LockException e) { ... }
}
try {
LockService lockService = LockServiceFactory.getInstance()
.getLockService(database);
if (lockService.hasChangeLogLock()) {
lockService.releaseLock();
}
} catch (LockException e) { ... }
}
AbstractUpdateCommandStep.java: run(), finally block
if (isDBLocked.get()) {
LockServiceFactory.getInstance()
.getLockService(database).releaseLock();
LockService lockService = LockServiceFactory.getInstance()
.getLockService(database);
if (lockService.hasChangeLogLock()) {
lockService.releaseLock();
}
}
LockServiceCommandStep.java: cleanUp()
hasChangeLogLock() flips to false the instant any release succeeds, anywhere. So whichever of the two release-attempts runs first does the real work; the second sees false and quietly no-ops instead of erroring.
A tempting alternative is a second flag, lockAcquiredHere, defaulting to false and set only when this method's own waitForLock() call succeeds. It looks right in isolation. It fails because of the reverse cleanup order from Part 2.
AbstractUpdateCommandStep is added last, so its cleanup fires first, while its own LoggingExecutor is still active. That's the only window in which a release statement gets captured into an updateSql output file; releasing later, from LockServiceCommandStep's cleanup, is too late for the file even though the database state ends up correct either way.The first attempt's lockAcquiredHere flag was only ever set true when this method called waitForLock(), which never happens for updateCount / updateSql / updateToTag, since the pipeline acquires it for them. So the internal release got skipped entirely for those three commands, and the only release left was LockServiceCommandStep.cleanUp()'s, which, per the diagram above, runs after the LoggingExecutor reset. The database still ended up correctly unlocked. The generated SQL file, for updateSql and updateCountSql, silently lost its -- Release Database Lock line. The final hasChangeLogLock()-based fix keeps the release inside AbstractUpdateCommandStep's own finally block whenever the lock is genuinely held (same window, same capture) while making the second attempt a harmless no-op instead of an error.
The pattern underneath both
Neither bug was really about locks or logging. Both were the same shape: a piece of mutable bookkeeping, scoped more broadly than the thing it was supposed to describe, getting asked a question (did I do this?) that only a narrower, more honest source of truth could actually answer.
A shared map keyed by scope-id can't tell threads apart, so key it by thread instead, and let each thread only ever see its own bookkeeping. A ThreadLocal flag can't tell “I acquired this” from “someone else did, and I'm along for the ride”, so stop asking the flag, and ask the lock table itself.
In both cases the fix reads smaller than the bug report. That's usually a sign the root cause was found rather than worked around, the code doesn't need more state to reason about, it needs to be asked the question closer to where the answer actually lives.