Transaction Management
BaseX offers ACID-safe transactions, with multiple readers and a single writer. This applies to the database server, the HTTP server and standalone mode.
Introduction
In a nutshell, a transaction is equal to a command or query. Each command or query sent to the server, each REST or RESTXQ request, and each job becomes a transaction.
Incoming requests are parsed and checked for errors. If the command or query is not correct, the request will not be executed, and the user will receive an error message. Otherwise, the request becomes a transaction and is passed on to the transaction manager.
Please note that:
- Locks cannot be synchronized across BaseX instances that run in different JVMs. If concurrent write operations are to be performed, we generally recommend working with the client/server or the HTTP architecture (see below).
- An unexpected abort of the server during a transaction, caused by a hardware failure or power cut, may lead to an inconsistent database state if a transaction was active at shutdown time. It is advisable to use the
CREATE BACKUPcommand to regularly back up your database. If the worst case occurs, you can try theINSPECTcommand to check if your database has obvious inconsistencies, and useRESTOREto restore the last backed up version of the database.
XQuery Update
Many update operations are triggered by XQuery Update expressions. When executing an updating query, all update operations of the query are stored in a pending update list. They will be executed all at once, so the database is updated atomically. If any of the update sub-operations is erroneous, the overall transaction will be aborted.
Concurrency Control
Updated: Updating queries acquire read locks for databases that are only read.
BaseX supports multiple read and single write operations. All locks that a transaction needs are requested before it is started. This means that:
- Each database has its own queue: An update on database A will not block operations on database B. This requires that the databases accessed by a transaction can be determined before it is evaluated (see below).
- Read transactions on the same database are executed in parallel.
- An updating transaction waits until the running transactions on its databases have finished. Subsequent transactions on these databases wait until the update has completed.
- Read and write locks are distinguished within a single transaction: A database that is only read by an updating transaction will be read-locked, and other read transactions on that database will be executed in parallel (see below).
- Transactions without database access are never locked. They can be executed in parallel with globally locking transactions.
- The maximum number of parallel transactions is limited by the
PARALLELoption. Read transactions are favored, so updates may have to wait longer if many read transactions are running.
Lock Detection
Commands
All commands come with a detector for local locks. Global locking is applied if the glob syntax is used. The following example drops all databases starting with the prefix string new; it results in a global lock:
DROP DB new*
XQuery
Local locks can be applied if it is possible at compile time to associate all database operations with static database names:
| Query | Description |
|---|---|
|
Read lock of the currently opened database |
|
Read lock of the factbook database |
|
Read lock of the documents database |
|
Write lock of the test database |
|
Read lock on the database supplied externally via $db. |
|
Read lock of db1 and db2, as the query is unrolled at compile time. |
|
Write lock of test, as the variable is inlined at compile time. |
|
No lock required |
|
No lock required, as the query is simplified to <doc/> at compile time. |
Updating queries do not write-lock every database they access: A database that is only read will be read-locked. The target of a node update (insert, delete, replace, rename) and of fn:put is supplied by a node, though, and can only be resolved at evaluation time. If a query contains such an expression, its read locks on databases are turned into write locks:
| Query | Description |
|---|---|
|
Read lock of src, write lock of trg |
|
Read lock of src, write lock of trg. Updates in the modify clause of a transform expression are restricted to copied nodes, so they do not affect locking. |
|
Write lock of src and trg, as the target of the insert expression is only known at evaluation time. |
A global lock will be assigned if the static detection fails:
| Query | Description |
|---|---|
|
The name of the database to be opened will only be known at evaluation time. |
|
The loop is too large to be unrolled. The UNROLLLIMIT can be increased to generate 100 db:get function calls and corresponding local locks. |
The functions fn:doc and fn:collection can be used for both accessing database resources and fetching resources at the specified URI (see Access Resources for more details). There are two ways to reduce the number of locks:
- Turn off the
WITHDBoption to prevent the functions from accessing databases; or - use
fetch:docfor fetching resources from URIs, and usedb:getfor accessing databases.
You can consult the query info output (via the Info View of the GUI, via -V on Command-Line or via turning on the QUERYINFO option) to find out which databases are locked by a query, and if local locks or a global lock is applied.
Avoiding Lock Contention
A query keeps its locks until it has finished. If a query accesses a database and then spends time on other work (processing images, calling a web service, …), other transactions on that database have to wait. The following hints help to keep waiting times short:
- Find out what is locked: The query info (see above) shows the locks of a query. The locks of running jobs are returned by
job:list-details. - Keep locks local: Use static database names, as shown in the examples above. With
xquery:eval, a global lock is applied, as the databases of the evaluated query are not known in advance. - Keep locks short: Move database operations to separate jobs.
job:executeevaluates a job and waits for its result, andjob:evalreturns immediately. Each job acquires its own locks and releases them when it has finished. The calling query does not lock the databases accessed by a job, no matter if a query string or a function item is supplied. - Protect other resources: Access to files or Java code can be synchronized with XQuery Locks.
In the following RESTXQ function, the cache database is only locked for looking up and storing a thumbnail. No lock is held while the image is scaled:
declare %rest:path('/thumb/{$key}') function local:thumb($key) {
let $path := $key || '.jpg'
let $cached := job:execute(fn() {
if (db:exists('cache', $path)) { db:get-binary('cache', $path) }
})
return $cached otherwise (
let $thumb := local:scale($key)
return ($thumb, void(job:eval(fn() { db:put-binary('cache', $thumb, $path) })))
)
};
Please note:
- Each job is a transaction of its own. If a job fails, the updates of previous jobs are kept.
- The calling query must not lock a database that is updated by a job evaluated with
job:execute: the job could not start before the calling query has finished, so adeadlockerror is raised instead. The same applies if the job reads a database that is updated by the calling query, or if the calling query is globally locked. - A query must not use the Client Functions to address its own server instance, as this can lead to deadlocks as well.
XQuery Locks
By default, access to external resources (files on hard disk, HTTP requests, …) is not controlled by the transaction manager of BaseX. Custom locks can be assigned via annotations, pragmas or options:
- A lock string may consist of a single key or multiple keys separated with commas.
- Custom locks and database locks are independent: A lock string never conflicts with a database of the same name.
Annotations
In the following module, lock annotations are used to prevent concurrent write operations on the same file:
module namespace config = 'config';
declare %basex:lock('CONFIG') function config:read() as xs:string {
file:read-text('config.txt')
};
declare %updating %basex:lock('CONFIG') function config:write($data as xs:string) {
file:write-text('config.txt', $data)
};
Some explanations:
- If a query calls one of these functions, a lock for the user-defined
CONFIGlock string will be acquired before the query is evaluated. - The lock is a write lock if the query is updating (for example, because it calls
config:write). Otherwise, it is a read lock. - Queries that only read can be evaluated at the same time. A query with a write lock waits until the running queries with the same lock string are finished, and other queries with this lock string wait until it is finished.
Pragmas
Locks can also be declared via pragmas:
(# basex:lock CONFIG #) {
update:output(file:write('config.xml', <config/>))
}
If you enclose code with an update:output function call, it will be treated as updating and a write lock will be assigned.
Options
Locks for the functions of a module can also be assigned via option declarations:
declare option basex:lock 'CONFIG';
update:output(file:write('config.xml', <config/>))
Once again, a write lock is enforced.
Java Modules
Locks can also be acquired on Java functions which are imported and invoked from an XQuery expression. It is advisable to explicitly lock Java code whenever it performs sensitive read and write operations.
File-System Locks
Update Operations
During a database update, a locking file upd.basex will reside in the database directory. If the update fails for some unexpected reason, or if the process is killed ungracefully, this file will not be deleted. In this case, the database cannot be opened anymore, and the message “Database … is being updated, or update was not completed” will be shown instead.
If the locking file is manually removed, you may be able to reopen the database, but you should be aware that the database may have become corrupt due to the interrupted update process, and you should revert to the most recent database backup.
Database Locks
To avoid database corruptions that are caused by accidental write operations from different JVMs, a shared lock is requested on the database table file (tbl.basex) whenever a database is opened. If an update operation is triggered, and if no exclusive lock can be acquired, it will be rejected with an error message, stating that the database is opened by another process.
Please note that you cannot 100% rely on this mechanism. You will be safe when using the client/server or HTTP architecture.
Changelog
Version 13.0- Updated: Updating queries acquire read locks for databases that are only read.
- Removed:
FAIRLOCKoption: Non-fair locking is always used.
- Updated: Lock detection was improved by splitting compilation into multiple steps.
- Updated: Single lock option for reads and writes.
- Updated: Query lock options were moved from
querytobasexnamespace.
- Updated: New
FAIRLOCKoption, improved detection of lock patterns.
- Added: Locks can also be acquired on Java functions.
- Added: database locking introduced, replacing process locking.
- Updated: pin files replaced with shared/exclusive filesystem locking.
- Added: pin files to mark open databases.
- Added: update lock files.