MarkLogic Database Configuration Operator Manual
Every setting in the Admin Database Help page, explained with operational intent and testable behaviour
The Admin UI database help page is accurate. It is also, almost without exception, too compressed to support a real production decision. Knowing that stemmed searches has an advanced option is not the same as knowing when to use it, what it costs, and what goes wrong if you pick the wrong level after a corpus is already loaded.
This manual covers every setting on the Admin database configuration page by family — groups of settings that share a failure mode or a trade-off. Within each family, each setting gets prose that explains its purpose, a concrete view of what breaks when it is wrong, and runnable XQuery and JavaScript examples on llamaverse data.
The examples in this article assume the llamaverse (v2.3.1+) is deployed. The llamaverse sample data is freely available from github.com/cleverllamas/llamaverse — see the llamaverse article for full setup instructions.
How to Use This Manual
The settings in MarkLogic's database configuration are not independent. Changing one setting without accounting for its dependencies produces surprises — a setting that appears to take effect immediately may not apply to existing fragments until reindexing completes, and a setting you thought was harmless may double index size overnight.
Before making any change in a production environment, work through the relevant family section to understand the trade-offs. Run the code examples before and after the change to establish a baseline and confirm the expected behavioural shift. Apply changes in one logical group at a time, let reindexing complete if applicable, and record the observed impact and rollback note.
Baseline: Know Your Context
Before changing anything, confirm you are operating against the right database. Many bad changes happen in the wrong environment because someone ran a script against the wrong database pointer.
xquery version "1.0-ml";
let $dbid := xdmp:database("cleverllamas-content")
let $name := xdmp:database-name($dbid)
let $forest-count := fn:count(xdmp:database-forests($dbid))
let $modules-db := xdmp:modules-database()
let $schema-db := xdmp:schema-database()
let $security-db := xdmp:security-database()
let $triggers-db := xdmp:triggers-database()
return string-join((
concat("Database ID: ", $dbid),
concat("Database name: ", $name),
concat("Forest count: ", $forest-count),
concat("Current content DB: ", xdmp:database-name(xdmp:database())),
concat("Current modules DB: ", if ($modules-db = 0) then "filesystem" else xdmp:database-name($modules-db)),
concat("Current schema DB: ", if ($schema-db = 0) then "none" else xdmp:database-name($schema-db)),
concat("Current security DB: ", if ($security-db = 0) then "none" else xdmp:database-name($security-db)),
concat("Current triggers DB: ", if ($triggers-db = 0) then "none" else xdmp:database-name($triggers-db))
), " ")
"use strict";
const dbid = xdmp.database("cleverllamas-content");
const modulesDb = xdmp.modulesDatabase();
const schemaDb = xdmp.schemaDatabase();
const securityDb = xdmp.securityDatabase();
const triggersDb = xdmp.triggersDatabase();
[
"Database ID: " + dbid,
"Database name: " + xdmp.databaseName(dbid),
"Forest count: " + fn.count(xdmp.databaseForests(dbid)),
"Current content DB: " + xdmp.databaseName(xdmp.database()),
"Current modules DB: " + (modulesDb === 0 ? "filesystem" : xdmp.databaseName(modulesDb)),
"Current schema DB: " + (schemaDb === 0 ? "none" : xdmp.databaseName(schemaDb)),
"Current security DB: " + (securityDb === 0 ? "none" : xdmp.databaseName(securityDb)),
"Current triggers DB: " + (triggersDb === 0 ? "none" : xdmp.databaseName(triggersDb))
].join("\n");
Database ID: 2736149102113944426
Database name: cleverllamas-content
Forest count: 1
Current content DB: cleverllamas-content
Current modules DB: Modules
Current schema DB: cleverllamas-schemas
Current security DB: Security
Current triggers DB: none
Database ID: 2736149102113944426
Database name: cleverllamas-content
Forest count: 1
Current content DB: cleverllamas-content
Current modules DB: Modules
Current schema DB: cleverllamas-schemas
Current security DB: Security
Current triggers DB: none
Text Search Index Settings
These settings define what the word index contains. They are the most consequential group in the database configuration because almost every search query your application runs depends on them, and changing them after a corpus is loaded means reindexing everything.
The settings in this family are interdependent. Word positions depends on word searches being enabled. Element word positions depends on word positions. Phrase precision depends on positions. Getting the baseline wrong means every search query your application writes is quietly working around a gap.
| Setting | Default | Index cost | Triggers reindex |
|---|---|---|---|
word searches | on | Moderate — baseline word term index | Yes |
stemmed searches | basic | Low–moderate; higher for advanced | Yes |
word positions | on | High — approximately doubles term index size | Yes |
fast phrase searches | off | Moderate — sequence index on top of positions | Yes |
fast reverse searches | off | High — full reverse query-to-document mapping | Yes |
fast case sensitive searches | off | Low | Yes |
fast diacritic sensitive searches | off | Low | Yes |
tf normalization | scaled-log | Negligible | Yes |
language | en | N/A — changes default stemming/tokenisation rules, not index structure size | Yes |
Word Searches
word searches controls whether MarkLogic builds a word term index — the primary lookup structure for all text search. Every cts:word-query, every cts:near-query, every phrase search and relevance calculation in your codebase assumes this index exists.
The default is enabled. It should remain enabled for any database that handles search. The option to disable it exists for databases that are genuinely non-searchable — a pure key-value store, a logging database, a binary asset store — where you have a documented policy that prohibits text search queries and a plan to enforce it. Disabling it as a cost-cutting measure is a false economy: the index cost is relatively modest, and the first time an unexpected search query runs against a corpus with no word index, the latency will be expensive in a different way.
When enabled: cts:word-query and cts:near-query go straight to the index. When disabled: word queries fall back to slower evaluation paths or fail depending on query options.
Enable when: the database handles any text search workload. Which is almost always.
Disable only when: you can document that the database is search-free and intend to keep it that way. Not "we don't search it yet."
What breaks: cts:word-query latency spikes. cts:near-query reliability degrades. Any API built on text retrieval slows or fails.
Stemmed Searches
Stemming controls vocabulary breadth. Without it, "hike" and "hiking" are different tokens. A user searching for "hike" finds nothing from a corpus that consistently uses "hiking". With basic English stemming enabled, they collapse to the same index stem and both searches work.
The setting has four levels:
| Level | What it covers | Typical use |
|---|---|---|
off | No stem index; exact token match only | Corpora where morphological variation must never match |
basic | Standard morphological stemming for the configured language: verb tenses, plurals | Default choice for English; correct for most deployments |
advanced | Broader morphological coverage, compound stemming, irregular forms | Where basic misses domain-specific inflections |
decompounding | Splits compound words before stemming | German, Dutch, and similar compound-heavy languages |
The cost is index size and ingest time. basic is right for English. Use advanced when morphological complexity genuinely matters. Use decompounding for languages where it is needed, not because it sounds better.
The stemmed query option on cts:word-query tells the server to perform stem matching at query time. Whether that uses an index (fast) or requires a scan (slow to unavailable) depends entirely on whether the database setting is non-off. The option is accepted regardless — the database setting is what determines whether it works correctly.
The llamaverse description texts use inflected forms like "dancing", "singing", "reading" — never the bare roots. An unstemmed search for "dance" finds nothing. A stemmed search for "dance" returns every document that contains "dancing", demonstrating that the database-level setting is the prerequisite for this path being index-backed and reliable at scale.
xquery version "1.0-ml";
(: Compare unstemmed vs stemmed word matching on llamaverse descriptions.
Llama descriptions use "dancing" not "dance", so an unstemmed search
for "dance" finds nothing. Stemmed matching collapses both to the same
index entry and the count becomes equal to a direct "dancing" search. :)
let $unstemmed := fn:count(cts:search(
fn:collection("llamaverse"),
cts:word-query("dance", "unstemmed"), (: exact token only :)
"filtered"
))
let $stemmed := fn:count(cts:search(
fn:collection("llamaverse"),
cts:word-query("dance", "stemmed"), (: matches "dancing" :)
"filtered"
))
let $direct := fn:count(cts:search(
fn:collection("llamaverse"),
cts:word-query("dancing"), (: exact token baseline :)
"filtered"
))
return string-join((
"Stem impact on the word 'dance' against llamaverse descriptions",
"",
concat("Unstemmed search for 'dance': ", $unstemmed, " results"),
concat("Stemmed search for 'dance': ", $stemmed, " results"),
concat("Direct search for 'dancing': ", $direct, " results"),
"",
"Expected: unstemmed=0, stemmed=direct.",
"If 'stemmed searches' is off in database config the stemmed count",
"degrades: the option is accepted but runs without an index backing it."
), " ")
'use strict';
// Compare unstemmed vs stemmed word matching on llamaverse descriptions.
// Llama descriptions use "dancing" not "dance", so an unstemmed search
// for "dance" finds nothing. Stemmed matching collapses both to the same
// index entry and the count becomes equal to a direct "dancing" search.
const unstemmed = fn.count(cts.search(
fn.collection('llamaverse'),
cts.wordQuery('dance', 'unstemmed'), // exact token only
'filtered'
));
const stemmed = fn.count(cts.search(
fn.collection('llamaverse'),
cts.wordQuery('dance', 'stemmed'), // matches "dancing"
'filtered'
));
const direct = fn.count(cts.search(
fn.collection('llamaverse'),
cts.wordQuery('dancing'), // exact token baseline
'filtered'
));
[
"Stem impact on the word 'dance' against llamaverse descriptions",
'',
`Unstemmed search for 'dance': ${unstemmed} results`,
`Stemmed search for 'dance': ${stemmed} results`,
`Direct search for 'dancing': ${direct} results`,
'',
"Expected: unstemmed=0, stemmed=direct.",
"If 'stemmed searches' is off in database config the stemmed count",
"degrades: the option is accepted but runs without an index backing it."
].join('\n')
Stem impact on the word 'dance' against llamaverse descriptions
Unstemmed search for 'dance': 0 results
Stemmed search for 'dance': 173 results
Direct search for 'dancing': 173 results
Expected: unstemmed=0, stemmed=direct.
If 'stemmed searches' is off in database config the stemmed count
degrades: the option is accepted but runs without an index backing it.
Enable basic when: the corpus is English and users will search with uninflected forms ("run", "hike", "dance").
Use advanced when: domain vocabulary has irregular forms that basic misses.
Leave off when: exact token match is a hard requirement — for example, product codes, identifiers, or corpora where morphological conflation is incorrect.
What breaks: users searching for root forms miss documents that only use inflected forms. With stemming off, the search API works — it just systematically misses vocabulary variation.
Word Positions
Word positions records the position of every token occurrence in every document — not just whether a term appears, but where. This structure is what makes cts:near-query work correctly.
Without word positions:
cts:near-querystill runs, but distance enforcement becomes approximate- The
checkedanduncheckedexecution modes behave differently but neither enforces the requested term distance accurately - Phrase-quality contributions to relevance scoring become imprecise
With word positions:
- Distance-constrained proximity (
within N words) is enforced against the index - Ordered phrase matching (
cts:near-querywith"ordered") is index-accurate - Relevance scoring reflects actual term proximity, not just co-occurrence
Index cost: word positions approximately doubles the term index size. This is worthwhile for any database where cts:near-query is in production use.
The example below runs the same near-query with tight distance constraints against terms that ARE close in llamaverse descriptions ("energetic personality") and terms that are NOT close ("energetic community"). With word positions enabled, the distance constraint correctly separates them. Without it, both queries return the same over-inclusive result.
xquery version "1.0-ml";
(: Demonstrate word positions and near-query distance enforcement.
"energetic" and "personality" appear adjacent in llamaverse descriptions
("...energetic personality...") so a tight near-query returns many results.
"energetic" and "community" are in different sentences ("...energetic
personality..." vs "...llama community.") so a tight near-query returns 0.
Without word positions, cts:near-query cannot enforce distance accurately
and the two counts may converge. :)
let $q-close := cts:near-query(
(cts:word-query("energetic"), cts:word-query("personality")),
2, (: within 2 word positions :)
"unordered"
)
let $q-far := cts:near-query(
(cts:word-query("energetic"), cts:word-query("community")),
2, (: these sit in different sentences :)
"unordered"
)
let $close-count := fn:count(cts:search(fn:collection("llamaverse"), $q-close, "filtered"))
let $far-count := fn:count(cts:search(fn:collection("llamaverse"), $q-far, "filtered"))
(: Also compare unfiltered+unchecked to show positional degradation :)
let $close-uncheck := fn:count(cts:search(fn:collection("llamaverse"), $q-close, ("unfiltered", "unchecked")))
let $far-uncheck := fn:count(cts:search(fn:collection("llamaverse"), $q-far, ("unfiltered", "unchecked")))
return string-join((
"Near-query distance enforcement (requires word positions):",
"",
concat("'energetic' within 2 of 'personality' filtered: ", $close-count),
concat("'energetic' within 2 of 'community' filtered: ", $far-count),
"",
"Unfiltered+unchecked (positional enforcement absent):",
concat("'energetic' within 2 of 'personality' unchecked: ", $close-uncheck),
concat("'energetic' within 2 of 'community' unchecked: ", $far-uncheck),
"",
"When word positions is off, filtered and unchecked counts converge",
"because the distance constraint has no index to enforce against."
), " ")
'use strict';
// Demonstrate word positions and near-query distance enforcement.
// "energetic" and "personality" appear adjacent in llamaverse descriptions
// ("...energetic personality...") so a tight near-query returns many results.
// "energetic" and "community" are in different sentences so a tight near-query
// returns 0. Without word positions, cts:nearQuery cannot enforce distance
// accurately and the two counts may converge.
const qClose = cts.nearQuery(
[cts.wordQuery('energetic'), cts.wordQuery('personality')],
2, // within 2 word positions
['unordered']
);
const qFar = cts.nearQuery(
[cts.wordQuery('energetic'), cts.wordQuery('community')],
2, // these sit in different sentences
['unordered']
);
const closeCount = fn.count(cts.search(fn.collection('llamaverse'), qClose, 'filtered'));
const farCount = fn.count(cts.search(fn.collection('llamaverse'), qFar, 'filtered'));
const closeUncheck = fn.count(cts.search(fn.collection('llamaverse'), qClose, ['unfiltered', 'unchecked']));
const farUncheck = fn.count(cts.search(fn.collection('llamaverse'), qFar, ['unfiltered', 'unchecked']));
[
'Near-query distance enforcement (requires word positions):',
'',
`'energetic' within 2 of 'personality' filtered: ${closeCount}`,
`'energetic' within 2 of 'community' filtered: ${farCount}`,
'',
'Unfiltered+unchecked (positional enforcement absent):',
`'energetic' within 2 of 'personality' unchecked: ${closeUncheck}`,
`'energetic' within 2 of 'community' unchecked: ${farUncheck}`,
'',
'When word positions is off, filtered and unchecked counts converge',
'because the distance constraint has no index to enforce against.'
].join('\n')
Near-query distance enforcement (requires word positions):
'energetic' within 2 of 'personality' filtered: 124
'energetic' within 2 of 'community' filtered: 0
Unfiltered+unchecked (positional enforcement absent):
'energetic' within 2 of 'personality' unchecked: 124
'energetic' within 2 of 'community' unchecked: 124
When word positions is off, filtered and unchecked counts converge
because the distance constraint has no index to enforce against.
The unfiltered+unchecked counts in the results show what happens when positions cannot be enforced: both queries return the same count because "unchecked" mode skips positional verification and degrades to co-occurrence. If your database has word positions disabled, cts:near-query with "unfiltered" produces this behaviour for all distance-constrained queries.
Enable when: any cts:near-query usage exists in production queries, or when phrase-quality relevance matters.
Disable only when: the corpus is write-heavy and proximity search is confirmed out of scope. The index savings are real; make sure you have documentation to back the decision.
What breaks silently: near-queries succeed and return results. They just return wrong results — documents where the terms do not actually appear within the requested distance. No error, no warning, just degraded precision.
Filtered and Unfiltered Execution
Filtered and unfiltered are execution modes passed as options to cts:search(). They are not database settings in themselves, but they interact directly with your index configuration — what your indexes contain determines whether unfiltered execution is safe.
In filtered mode (the default), MarkLogic uses the index to find candidates and then opens each matching document to verify the result against the query. This is correct but adds I/O proportional to the candidate set size.
In unfiltered mode, MarkLogic returns the index candidates directly without re-verifying them against the document content. When queries are scoped to the fragment root (i.e., the whole document), this is safe and fast. When queries are applied to a path expression that selects nodes below the fragment root, unfiltered mode returns the first matching node per fragment regardless of whether it genuinely satisfies the query. This is the most common source of silently wrong results in MarkLogic applications.
The example below uses Bradley's health record: a single XML document containing two <checkup> nodes. We query for the checkup where outcome=treatment-required. In unfiltered mode, MarkLogic returns the first node in the fragment — the outcome=healthy checkup — as a false positive.
<healthRecord llamaId="223a3aaa-3d69-4323-bdd9-a84e00f61631">
<checkup sequence="1">
<vet>Dr. Aguilar</vet>
<outcome>healthy</outcome>
<notes>Routine annual check. Weight within normal range. Fleece in excellent condition.</notes>
</checkup>
<checkup sequence="2">
<vet>Dr. Aguilar</vet>
<outcome>treatment-required</outcome>
<notes>Follow-up visit. Mild respiratory symptoms noted. Course of antibiotics prescribed.</notes>
</checkup>
</healthRecord>
xquery version "1.0-ml";
(: UNFILTERED search below the fragment root. :)
(: Bradley's health record contains two <checkup> nodes: :)
(: checkup 1: vet=Dr. Aguilar, outcome=healthy :)
(: checkup 2: vet=Dr. Aguilar, outcome=treatment-required :)
(: :)
(: We query for the checkup where outcome="treatment-required" — that :)
(: is checkup 2. With "unfiltered", MarkLogic returns the FIRST node in :)
(: the fragment regardless of whether it genuinely satisfies the query. :)
let $uri := "/cleverllamas/llamaverse/content/health-records/223a3aaa-3d69-4323-bdd9-a84e00f61631.xml"
let $query := cts:and-query((
cts:element-value-query(xs:QName("vet"), "Dr. Aguilar"),
cts:element-value-query(xs:QName("outcome"), "treatment-required")
))
return
cts:search(
fn:doc($uri)/healthRecord/checkup,
$query,
"unfiltered"
)
// UNFILTERED search below the fragment root.
// Bradley's health record contains two checkup nodes:
// checkup 1: vet=Dr. Aguilar, outcome=healthy
// checkup 2: vet=Dr. Aguilar, outcome=treatment-required
//
// We query for the checkup where outcome="treatment-required" — that
// is checkup 2. With "unfiltered", MarkLogic returns the FIRST node in
// the fragment regardless of whether it genuinely satisfies the query.
var uri = "/cleverllamas/llamaverse/content/health-records/223a3aaa-3d69-4323-bdd9-a84e00f61631.xml";
var query = cts.andQuery([
cts.elementValueQuery(fn.QName("", "vet"), "Dr. Aguilar"),
cts.elementValueQuery(fn.QName("", "outcome"), "treatment-required")
]);
cts.search(
fn.doc(uri).xpath("/healthRecord/checkup"),
query,
["unfiltered"]
);
<checkup sequence="1">
<vet>Dr. Aguilar</vet>
<outcome>healthy</outcome>
<notes>Routine annual check. Weight within normal range. Fleece in excellent condition.</notes>
</checkup>
(: FALSE POSITIVE — checkup 1 is returned, but outcome="healthy", :)
(: not "treatment-required". MarkLogic returned the first node in the :)
(: fragment without checking whether it genuinely satisfies the query. :)
The filtered query evaluates each <checkup> node individually against the full query and returns only the node that genuinely satisfies both conditions:
xquery version "1.0-ml";
(: The same query, using "filtered" (the default). :)
(: MarkLogic opens the document and validates each <checkup> node against :)
(: the query. Only the node that genuinely satisfies both conditions :)
(: is returned. :)
let $uri := "/cleverllamas/llamaverse/content/health-records/223a3aaa-3d69-4323-bdd9-a84e00f61631.xml"
let $query := cts:and-query((
cts:element-value-query(xs:QName("vet"), "Dr. Aguilar"),
cts:element-value-query(xs:QName("outcome"), "treatment-required")
))
return
cts:search(
fn:doc($uri)/healthRecord/checkup,
$query,
"filtered"
)
// The same query, using "filtered" (the default).
// MarkLogic opens the document and validates each checkup node against
// the query. Only the node that genuinely satisfies both conditions
// is returned.
var uri = "/cleverllamas/llamaverse/content/health-records/223a3aaa-3d69-4323-bdd9-a84e00f61631.xml";
var query = cts.andQuery([
cts.elementValueQuery(fn.QName("", "vet"), "Dr. Aguilar"),
cts.elementValueQuery(fn.QName("", "outcome"), "treatment-required")
]);
cts.search(
fn.doc(uri).xpath("/healthRecord/checkup"),
query,
["filtered"]
);
<checkup sequence="2">
<vet>Dr. Aguilar</vet>
<outcome>treatment-required</outcome>
<notes>Follow-up visit. Mild respiratory symptoms noted. Course of antibiotics prescribed.</notes>
</checkup>
(: CORRECT — checkup 2 is the only node that genuinely satisfies :)
(: both conditions. :)
The same first-match truncation applies to JSON documents. Aaron's movement history is a single document containing 168 movement entries as array nodes. An unfiltered path-expression query returns exactly one entry — not because Aaron only moved once, but because MarkLogic stops at the first candidate per fragment:
xquery version "1.0-ml";
(: UNFILTERED cts:search() with a path expression root. :)
(: Even though Aaron has 168 movement entries in this document, :)
(: only the FIRST matching node is returned. This is unfiltered :)
(: behaviour below the fragment root — MarkLogic stops after the first :)
(: matching node per fragment. :)
let $uri := "/cleverllamas/llamaverse/content/llama-movement/llama_location_history.json"
let $query := cts:json-property-value-query(
"llamaId", "0c8bdb0d-ac62-49b7-ac74-94dbba46efa5"
)
let $results :=
cts:search(
fn:doc($uri)/envelope/instance/array-node("llama-movement")/object-node(),
$query,
"unfiltered"
)
return (
fn:count($results), (: Returns 1, not 168 :)
$results (: Returns only the first matching movement entry :)
)
// UNFILTERED cts.search() with a path expression root.
// Even though Aaron has 168 movement entries in this document,
// only the FIRST matching node is returned. This is unfiltered
// behaviour below the fragment root — MarkLogic stops after the
// first matching node per fragment.
var uri = "/cleverllamas/llamaverse/content/llama-movement/llama_location_history.json";
var query = cts.jsonPropertyValueQuery(
"llamaId", "0c8bdb0d-ac62-49b7-ac74-94dbba46efa5"
);
var results = cts.search(
fn.doc(uri).xpath('/envelope/instance/array-node("llama-movement")/object-node()'),
query,
["unfiltered"]
);
[fn.count(results), results]; // Returns 1, not 168
1
{"llamaId": "0c8bdb0d-ac62-49b7-ac74-94dbba46efa5", "timestamp": "2025-05-18T12:00:00Z", "coordinates": [-9.9219172, 53.5136376]}
The corrective pattern for first-match truncation is to open the document explicitly and evaluate each candidate with cts:contains():
xquery version "1.0-ml";
(: Step 1: use fn:doc() to open the document directly. :)
(: Step 2: iterate array items using cts:contains() to collect :)
(: every movement entry for Aaron within that document. :)
let $uri := "/cleverllamas/llamaverse/content/llama-movement/llama_location_history.json"
let $query := cts:json-property-value-query(
"llamaId", "0c8bdb0d-ac62-49b7-ac74-94dbba46efa5"
)
let $doc := fn:doc($uri)
let $all-entries :=
for $entry in $doc/envelope/instance/array-node("llama-movement")/object-node()
where cts:contains($entry, $query)
return $entry
return (
fn:count($all-entries), (: Returns 168 :)
$all-entries[1 to 3] (: First 3 entries shown for brevity :)
)
// Step 1: use fn.doc() to open the document directly.
// Step 2: iterate array items using cts.contains() to collect
// every movement entry for Aaron within that document.
var uri = "/cleverllamas/llamaverse/content/llama-movement/llama_location_history.json";
var query = cts.jsonPropertyValueQuery(
"llamaId", "0c8bdb0d-ac62-49b7-ac74-94dbba46efa5"
);
var doc = fn.doc(uri);
var allEntries = doc.xpath('/envelope/instance/array-node("llama-movement")/object-node()')
.toArray()
.filter(function(entry) { return cts.contains(entry, query); });
[allEntries.length, allEntries.slice(0, 3)]; // Returns 168, first 3 shown
168
{"llamaId":"0c8bdb0d-ac62-49b7-ac74-94dbba46efa5", "timestamp":"2025-05-18T12:00:00Z", "coordinates":[-9.9219172, 53.5136376]}
{"llamaId":"0c8bdb0d-ac62-49b7-ac74-94dbba46efa5", "timestamp":"2025-05-18T13:00:00Z", "coordinates":[-9.9229262, 53.5143298]}
{"llamaId":"0c8bdb0d-ac62-49b7-ac74-94dbba46efa5", "timestamp":"2025-05-18T14:00:00Z", "coordinates":[-9.9221079, 53.5136388]}
The checked option adds a position-list check to the unfiltered path. With word positions enabled, checked can catch some path-expression false positives that unchecked misses. With word positions disabled, neither option enforces positional accuracy. The following example uses a near-query for terms that cannot coexist ("llama" and "cooking" within 1 word) — both modes correctly return zero, confirming that the index and the execution mode are behaving consistently:
xquery version "1.0-ml";
let $query := cts:near-query((
cts:word-query("llama"),
cts:word-query("cooking")
), 1)
let $checked := cts:search(fn:collection("llamaverse"), $query, ("unfiltered", "checked"))
let $unchecked := cts:search(fn:collection("llamaverse"), $query, ("unfiltered", "unchecked"))
return string-join((
"Query: cts:near-query(""llama"",""cooking"",1)",
concat("Unfiltered+checked count: ", fn:count($checked)),
concat("Unfiltered+unchecked count: ", fn:count($unchecked))
), " ")
"use strict";
const query = cts.nearQuery([
cts.wordQuery("llama"),
cts.wordQuery("cooking")
], 1);
const scoped = cts.andQuery([cts.collectionQuery("llamaverse"), query]);
const checked = cts.search(scoped, ["unfiltered", "checked"]);
const unchecked = cts.search(scoped, ["unfiltered", "unchecked"]);
[
"Query: cts.nearQuery([llama,cooking],1)",
"Unfiltered+checked count: " + fn.count(checked),
"Unfiltered+unchecked count: " + fn.count(unchecked)
].join("\n");
Query: cts:near-query("llama","cooking",1)
Unfiltered+checked count: 0
Unfiltered+unchecked count: 0
Query: cts.nearQuery([llama,cooking],1)
Unfiltered+checked count: 0
Unfiltered+unchecked count: 0
The same filtered-versus-unfiltered comparison for this near-query confirms the same result via both execution modes:
xquery version "1.0-ml";
let $q := cts:near-query((
cts:word-query("llama"),
cts:word-query("cooking")
), 1)
let $filtered := cts:search(fn:collection("llamaverse"), $q, "filtered")
let $unfiltered_results := cts:search(fn:collection("llamaverse"), $q, "unfiltered")
let $sample_uri := fn:head(cts:uris((), (), cts:and-query((cts:collection-query("llamaverse"), $q))))
return string-join((
"Query: cts:near-query(""llama"", ""cooking"", distance=1)",
concat("Filtered count: ", fn:count($filtered)),
concat("Unfiltered count: ", fn:count($unfiltered_results)),
concat("Sample candidate URI: ", ($sample_uri, "(none)")[1])
), " ")
"use strict";
const q = cts.nearQuery([
cts.wordQuery("llama"),
cts.wordQuery("cooking")
], 1);
const scoped = cts.andQuery([cts.collectionQuery("llamaverse"), q]);
const filtered = cts.search(scoped, ["filtered"]);
const unfiltered = cts.search(scoped, ["unfiltered"]);
const sampleUri = fn.head(cts.uris(null, null, scoped));
[
"Query: cts.nearQuery([llama,cooking],1)",
"Filtered count: " + fn.count(filtered),
"Unfiltered count: " + fn.count(unfiltered),
"Sample candidate URI: " + (sampleUri || "(none)")
].join("\n");
Query: cts:near-query("llama", "cooking", distance=1)
Filtered count: 0
Unfiltered count: 0
Sample candidate URI: (none)
Query: cts.nearQuery([llama,cooking],1)
Filtered count: 0
Unfiltered count: 0
Sample candidate URI: (none)
The practical rule: use filtered mode as your default. Use unfiltered only at the fragment (document) root level when you have measured the performance cost of filtering and it is a documented, justified choice.
Fast Phrase Searches
Fast phrase searches builds a specialised sequence index on top of word positions. It accelerates the specific pattern of exact-sequence phrase matching: cts:near-query with distance 0 and "ordered", which is the canonical phrase query in MarkLogic.
Without it, phrase queries work by consulting word positions and checking adjacency at query time. With it, the adjacency check is pre-computed and stored as a separate index structure.
The performance gain is most visible under concurrent load with phrase-heavy patterns: autocomplete, quoted phrase search, document-similarity lookups. For lighter workloads, the latency improvement may not justify the additional index overhead.
The example below shows the distinction between word co-occurrence matching and phrase ordering. Both the ordered and reverse-ordered phrase for "llama community" return consistent counts because llamaverse descriptions always write the words in order. A co-occurrence query (cts:and-query) would return the same count for both directions — hiding any ordering error entirely. Reversing to "community llama" returns zero under ordered phrase matching.
xquery version "1.0-ml";
(: Phrase ordering versus co-occurrence.
All llamaverse descriptions contain "llama community" as a phrase.
Searching for both words in order returns the same count as a word
co-occurrence search. Reversing the order to "community llama" drops
to zero, which a word co-occurrence search would never show.
This contrast is the reason fast phrase searches exists. :)
let $cooccur := fn:count(cts:search( (: both words, any order :)
fn:collection("llamaverse"),
cts:and-query((cts:word-query("llama"), cts:word-query("community"))),
"filtered"
))
let $phrase-fwd := fn:count(cts:search( (: "llama community" in order :)
fn:collection("llamaverse"),
cts:near-query(
(cts:word-query("llama"), cts:word-query("community")),
0,
("ordered")
),
"filtered"
))
let $phrase-rev := fn:count(cts:search( (: "community llama" — reversed :)
fn:collection("llamaverse"),
cts:near-query(
(cts:word-query("community"), cts:word-query("llama")),
0,
("ordered")
),
"filtered"
))
(: Extra: phrase matching with a gap allowed — adjacent-ish :)
let $phrase-gap := fn:count(cts:search(
fn:collection("llamaverse"),
cts:near-query(
(cts:word-query("energetic"), cts:word-query("personality")),
0,
("ordered")
),
"filtered"
))
return string-join((
"Phrase vs co-occurrence matching on llamaverse descriptions:",
"",
concat("'llama' AND 'community' (co-occurrence): ", $cooccur),
concat("'llama' then 'community' (phrase, ordered): ", $phrase-fwd),
concat("'community' then 'llama' (reversed, 0 distance): ", $phrase-rev),
"",
concat("'energetic' then 'personality' (adjacent phrase): ", $phrase-gap),
"",
"The reversed phrase returns 0 — the and-query would never catch that.",
"fast phrase searches accelerates the ordered near-query(0) pattern."
), " ")
'use strict';
// Phrase ordering versus co-occurrence.
// All llamaverse descriptions contain "llama community" as a phrase.
// Searching for both words in order returns the same count as a word
// co-occurrence search. Reversing the order to "community llama" drops
// to zero, which a word co-occurrence search would never show.
// This contrast is the reason fast phrase searches exists.
const cooccur = fn.count(cts.search( // both words, any order
fn.collection('llamaverse'),
cts.andQuery([cts.wordQuery('llama'), cts.wordQuery('community')]),
'filtered'
));
const phraseFwd = fn.count(cts.search( // "llama community" in order
fn.collection('llamaverse'),
cts.nearQuery(
[cts.wordQuery('llama'), cts.wordQuery('community')],
0,
['ordered']
),
'filtered'
));
const phraseRev = fn.count(cts.search( // "community llama" — reversed
fn.collection('llamaverse'),
cts.nearQuery(
[cts.wordQuery('community'), cts.wordQuery('llama')],
0,
['ordered']
),
'filtered'
));
const phraseAdj = fn.count(cts.search(
fn.collection('llamaverse'),
cts.nearQuery(
[cts.wordQuery('energetic'), cts.wordQuery('personality')],
0,
['ordered']
),
'filtered'
));
[
'Phrase vs co-occurrence matching on llamaverse descriptions:',
'',
`'llama' AND 'community' (co-occurrence): ${cooccur}`,
`'llama' then 'community' (phrase, ordered): ${phraseFwd}`,
`'community' then 'llama' (reversed, 0 distance): ${phraseRev}`,
'',
`'energetic' then 'personality' (adjacent phrase): ${phraseAdj}`,
'',
'The reversed phrase returns 0 — the and-query would never catch that.',
'fast phrase searches accelerates the ordered near-query(0) pattern.'
].join('\n')
Phrase vs co-occurrence matching on llamaverse descriptions:
'llama' AND 'community' (co-occurrence): 487
'llama' then 'community' (phrase, ordered): 487
'community' then 'llama' (reversed, 0 distance): 0
'energetic' then 'personality' (adjacent phrase): 124
The reversed phrase returns 0 — the and-query would never catch that.
fast phrase searches accelerates the ordered near-query(0) pattern.
Enable when: quoted phrase search or exact sequence matching is a user-facing feature.
Defer when: phrase queries are rare and index cost is already under pressure.
Depends on: word positions being enabled. Enabling fast phrase searches without word positions is a configuration error.
Fast Reverse Searches
cts:reverse-query inverts the normal search direction: instead of finding documents that match a query, it finds which stored queries match a given document. This powers rules engines, alerting systems, and content-routing pipelines where new documents are scored against a standing rule-set at ingest time.
Without fast reverse searches, cts:reverse-query evaluates against an unaccelerated path. With it, a dedicated reverse index supports efficient lookup.
Enable only when: cts:reverse-query is in production use. The index overhead is not small — it is a full reverse mapping of the query space onto the document space. You would likely have specific use-cases in mind — rules engines, alerting pipelines, content routing — before reaching for this setting.
Case and Diacritic Sensitivity
By default, MarkLogic's word index is case-insensitive and diacritic-insensitive. "Gray" and "gray" are the same entry. "Zürich" and "Zurich" are the same entry. This is correct behaviour for the majority of search applications where the user intent is the word, not its capitalisation.
When case or diacritics carry semantic meaning — product codes where TYPE-A must not match type-a, names where diacritic marks distinguish different people or places — you need the fast case sensitive searches and fast diacritic sensitive searches settings.
These settings build additional index structures that preserve the original capitalisation and diacritic marks. The query must then use the matching option — "case-sensitive" on cts:word-query — to consult the accelerated path. Without the database setting enabled, the option is accepted but runs through a slower unindexed path.
The llamaverse llama names are capitalised. A case-sensitive search for "aaron" (lowercase) finds nothing. A case-insensitive search finds all documents where "Aaron" appears.
xquery version "1.0-ml";
let $case_sensitive := cts:search(
fn:collection("llamaverse"),
cts:word-query("aaron", "case-sensitive"),
"filtered"
)
let $case_insensitive := cts:search(
fn:collection("llamaverse"),
cts:word-query("aaron", "case-insensitive"),
"filtered"
)
let $case_uri := fn:head(cts:uris((), (), cts:and-query((
cts:collection-query("llamaverse"),
cts:word-query("aaron", "case-insensitive")
))))
return string-join((
"Query term: aaron",
concat("Case-sensitive count: ", fn:count($case_sensitive)),
concat("Case-insensitive count: ", fn:count($case_insensitive)),
concat("Case-insensitive sample URI: ", ($case_uri, "(none)")[1])
), " ")
"use strict";
const qSensitive = cts.andQuery([
cts.collectionQuery("llamaverse"),
cts.wordQuery("aaron", ["case-sensitive"])
]);
const qInsensitive = cts.andQuery([
cts.collectionQuery("llamaverse"),
cts.wordQuery("aaron", ["case-insensitive"])
]);
const sensitive = cts.search(qSensitive, ["filtered"]);
const insensitive = cts.search(qInsensitive, ["filtered"]);
const insensitiveUri = fn.head(cts.uris(null, null, qInsensitive));
[
"Query term: aaron",
"Case-sensitive count: " + fn.count(sensitive),
"Case-insensitive count: " + fn.count(insensitive),
"Case-insensitive sample URI: " + (insensitiveUri || "(none)")
].join("\n");
Query term: aaron
Case-sensitive count: 0
Case-insensitive count: 47
Case-insensitive sample URI: /0c8bdb0d-ac62-49b7-ac74-94dbba46efa5.json
Query term: aaron
Case-sensitive count: 0
Case-insensitive count: 47
Case-insensitive sample URI: /0c8bdb0d-ac62-49b7-ac74-94dbba46efa5.json
Enable fast case sensitive searches when: case distinctions carry business meaning that must be preserved in search results.
Enable fast diacritic sensitive searches when: the corpus contains names or terms where diacritic marks are semantically meaningful.
Leave both disabled when: your users want the word, not its exact glyph — which is most search applications.
TF Normalisation
Term frequency normalisation controls how document length is factored into relevance scoring. Long documents naturally contain more occurrences of any given term, which would inflate their score in naive TF weighting. Normalisation corrects for this.
The available strategies differ in how aggressively they discount length. The default works well for mixed-length corpora. Changing it without a specific ranking experiment showing a problem is relevance tuning by intuition.
Change when: you have query-by-query ranking acceptance tests that demonstrate the current scoring model produces wrong orderings for your use case.
Do not change because a different option sounds more rigorous. Relevance tuning without measurement produces regressions that are difficult to attribute.
Default Language
language sets the database-wide default language used for tokenisation and stemming when a document does not declare its own language (via an xml:lang attribute or equivalent JSON property). It is the corpus-wide fallback that the stemmed searches setting above operates against — basic stemming for an English-default database applies English morphological rules; the same setting on a database defaulted to fr applies French rules instead.
The default is en. Getting this wrong on a multi-lingual corpus with inconsistent per-document language tagging means stemming quietly applies the wrong language's rules to untagged documents — not an error, just systematically incorrect stem matching.
The example below reads the current default language directly from the Admin configuration for this database:
xquery version "1.0-ml";
import module namespace admin = "http://marklogic.com/xdmp/admin"
at "/MarkLogic/admin.xqy";
let $config := admin:get-configuration()
let $dbid := xdmp:database("cleverllamas-content")
return concat("Default language: ", admin:database-get-language($config, $dbid))
'use strict';
const admin = require('/MarkLogic/admin.xqy');
const config = admin.getConfiguration();
const dbid = xdmp.database('cleverllamas-content');
'Default language: ' + admin.databaseGetLanguage(config, dbid);
Default language: en
Change when: the corpus is predominantly non-English and documents do not reliably declare their own language.
Leave at en for English-default corpora — which includes llamaverse.
What breaks: nothing errors. Stemming for untagged documents silently applies the wrong language's morphological rules, degrading recall in a way that is easy to miss without a ranking or recall test.
Element, Field, and Attribute Indexes
These settings narrow text index coverage to specific parts of document structure. They only pay off when your queries exploit structure — using cts:element-word-query, cts:json-property-word-query, cts:field-value-query, or element-scoped near queries. Enabling them on databases where queries don't use structural scoping adds index overhead with no query benefit.
| Setting | Default | Index cost | Triggers reindex |
|---|---|---|---|
fast element word searches | off | Moderate — parallel element-scoped word index | Yes |
element word positions | off | High — positional index scoped to elements | Yes |
fast element phrase searches | off | Moderate — requires element word positions | Yes |
element value positions | off | Moderate | Yes |
attribute value positions | off | Moderate | Yes |
field value searches | off | Moderate per field defined | Yes |
field value positions | off | Moderate per field defined | Yes |
Fast Element Word Searches
The structural-narrowing equivalent of word searches. Where the baseline word index covers all tokens in all document positions, fast element word searches builds a parallel index scoped to named elements or JSON properties.
In llamaverse, cts:json-property-word-query("description", "friendly") restricts the term match to the description property only. The term "friendly" appears in description text but not in any other property (names, breeds, birthplaces). A global word query and a description-scoped query return the same count — but the property-scoped version uses a dedicated index path when the database setting is enabled, rather than the full word index with a post-filter that checks property containment.
xquery version "1.0-ml";
(: Compare global word search against JSON-property-scoped word search. :)
(: In llamaverse descriptions, "friendly" appears only in the description :)
(: property ("Known for their friendly and energetic personality"). :)
(: A global search and a description-scoped search return the same count :)
(: — confirming the term is exclusive to that property. :)
(: A name-scoped search returns 0 — "friendly" is not a name token. :)
(: :)
(: This shows what fast element/property word indexes do: they allow the :)
(: same result as a global search to be produced via a narrower, faster :)
(: index path when the calling code scopes its query by property. :)
let $global :=
xdmp:estimate(cts:search(fn:collection("llamaverse"),
cts:word-query("friendly")
))
let $in-description :=
xdmp:estimate(cts:search(fn:collection("llamaverse"),
cts:json-property-word-query("description", "friendly")
))
let $in-name :=
xdmp:estimate(cts:search(fn:collection("llamaverse"),
cts:json-property-word-query("name", "friendly")
))
return (
fn:concat("Global word query: ", $global),
fn:concat("Property-scoped (description): ", $in-description),
fn:concat("Property-scoped (name): ", $in-name)
)
var collQ = cts.collectionQuery("llamaverse");
// Compare global word search against JSON-property-scoped word search.
// "friendly" appears only in llamaverse description text.
// Global and description-scoped return the same count; name-scoped returns 0.
var global_count = xdmp.estimate(cts.search(
cts.andQuery([collQ, cts.wordQuery("friendly")])
));
var desc_count = xdmp.estimate(cts.search(
cts.andQuery([collQ, cts.jsonPropertyWordQuery("description", "friendly")])
));
var name_count = xdmp.estimate(cts.search(
cts.andQuery([collQ, cts.jsonPropertyWordQuery("name", "friendly")])
));
[
"Global word query: " + global_count,
"Property-scoped (description): " + desc_count,
"Property-scoped (name): " + name_count
];
Global word query: 1277
Property-scoped (description): 1277
Property-scoped (name): 0
(All 1277 llamaverse documents contain "friendly" in their description.
The term does not appear in any name property — name-scoped count is 0.
When fast element/property word searches is enabled, the description-scoped
query consults a dedicated property-level index rather than the full word
index with a post-filter. Counts are identical; the performance path differs.)
The count equality between global and description-scoped is itself informative: it confirms that "friendly" is exclusively a description-property term. The performance distinction is invisible in the result — it shows up under load.
Enable when: structured document models use element-scoped or property-scoped search as a first-class API feature.
Skip when: queries are not structured by element or property.
Element Word Positions and Fast Element Phrase Searches
These are the positional and phrase equivalents scoped to element or property content.
element word positions enables distance-constrained near-queries inside a named element or property. fast element phrase searches accelerates exact-sequence phrase queries within element scope.
Both depend on their parent settings. element word positions requires word positions. fast element phrase searches requires element word positions. The dependency is not automatically enforced — the configuration accepts the combination silently, but the index structure will not exist if the parent is absent.
An element-scoped near-query would look like:
cts:element-query(
xs:QName("description"),
cts:near-query((cts:word-query("friendly"), cts:word-query("energetic")), 2)
)
Without element word positions, this query falls back to unindexed evaluation for the positional constraint inside the element scope.
Enable when: proximity or phrase queries are constrained by element or property scope.
Element Value Positions and Attribute Value Positions
These extend positional index support to cts:element-value-query and cts:attribute-value-query patterns. Enable only when your query logic uses element-value or attribute-value proximity. If your queries don't do that, these settings add overhead without benefit.
An example of a query that requires element value positions:
cts:near-query((
cts:element-value-query(xs:QName("vet"), "Dr. Aguilar"),
cts:element-value-query(xs:QName("outcome"), "treatment-required")
), 0)
Without element value positions, the proximity constraint between element values cannot be enforced from the index.
Field Value Searches and Field Value Positions
Fields are named groupings of document content — elements, attributes, and JSON properties — indexed together as a virtual scope. field value searches enables index-backed evaluation of field value queries. field value positions extends this with proximity support inside field scope.
Enable field value positions only when field-level proximity logic is in active use.
Fields are defined at the database level in the field configuration separately from these settings. The settings here control index behaviour — they don't create the fields.
Wildcard and Lexicon Indexes
Wildcard queries trade index size for pattern flexibility. The settings in this family control what wildcard patterns are supported and at what cost. The cost is not trivial — wildcard indexes can substantially increase both index size and ingest time.
| Setting | Default | Index cost | Triggers reindex |
|---|---|---|---|
three character searches | off | High — trigram index across all tokens | Yes |
three character word positions | off | High — requires three character searches + word positions | Yes |
trailing wildcard searches | off | Moderate — dedicated prefix index | Yes |
fast element trailing wildcard searches | off | Moderate — element-scoped prefix index | Yes |
trailing wildcard word positions | off | Moderate | Yes |
word lexicons | none | Moderate per collation defined | Yes |
two character searches | off | Very high — exponentially larger candidate sets | Yes |
one character searches | off | Very high | Yes |
Three-Character Searches
The backbone of wildcard support. Three-character searches indexes all sequences of three or more consecutive non-wildcard characters within tokens, which backs the majority of useful wildcard patterns: llam*, *ama*, ?lama.
Without this setting, wildcard queries either fail or fall back to unindexed full-document scans. With it, most practical wildcard patterns are index-backed.
Enable when: any cts:word-query with the "wildcarded" option is in production use.
What breaks without it: cts:word-query("llam*", "wildcarded") degrades to a full corpus scan. For most databases with a search workload, this setting should be on.
The following example exercises several wildcard patterns against llamaverse data and confirms they resolve correctly from the index:
xquery version "1.0-ml";
let $q := cts:and-query((
cts:collection-query("llamaverse"),
cts:word-query("llam*", "wildcarded")
))
let $uri_results := cts:uris((), (), $q)
return string-join((
"Query: cts:word-query(""llam*"",""wildcarded"")",
concat("URI hits: ", fn:count($uri_results)),
concat("First URI: ", ($uri_results[1], "(none)")[1])
), " ")
"use strict";
const q = cts.andQuery([
cts.collectionQuery("llamaverse"),
cts.wordQuery("llam*", ["wildcarded"])
]);
const uris = cts.uris(null, null, q);
[
"Query: cts.wordQuery(\"llam*\",[\"wildcarded\"])",
"URI hits: " + fn.count(uris),
"First URI: " + (fn.head(uris) || "(none)")
].join("\n");
Query: cts:word-query("llam*","wildcarded")
URI hits: 1277
First URI: /0c8bdb0d-ac62-49b7-ac74-94dbba46efa5.json
Query: cts.wordQuery("llam*",["wildcarded"])
URI hits: 1277
First URI: /0c8bdb0d-ac62-49b7-ac74-94dbba46efa5.json
Three-Character Word Positions
Enables cts:near-query to work correctly when one of the terms is a wildcard pattern. Depends on both three character searches and word positions. Enable only when proximity queries genuinely use wildcard terms.
Trailing Wildcard Searches
Accelerates prefix-style patterns (llama*, abc*) with a dedicated trailing-wildcard index. Prefix matching is a high-frequency pattern — autocomplete, type-ahead, URI prefix traversal — that benefits from its own structure separate from the three-character index.
Enable when: prefix wildcards are a first-class application feature. The overhead is additive to the three-character index.
Fast Element Trailing Wildcard Searches and Trailing Wildcard Word Positions
Element-scoped trailing wildcard acceleration and positional extension for prefix wildcards respectively. Enable only when the application uses element-scoped prefix search or prefix wildcard near-queries.
Word Lexicons
Word lexicons build per-collation enumerable term lists. They back cts:word-lexicon (term-level lookups), cts:word-match (pattern-based term enumeration), and term analytics workflows.
Define at least one deliberate collation and document it. Using the default collation without a deliberate policy leads to sorting and grouping surprises when data contains non-ASCII characters.
Enable when: term analytics, wildcard expansion, or enumeration-style APIs are in use.
Two-Character and One-Character Searches
These extend the wildcard index to cover shorter non-wildcard sequences. Short wildcards produce massive candidate sets. A search for a* potentially matches a significant fraction of every document. The index cost is proportionally large.
Enable only with: documented workload evidence that short wildcard patterns are genuinely required, capacity planning that accounts for the index growth, and operational governance to prevent ad-hoc short-wildcard queries from degrading performance.
Default: off. Leave them off unless there is a specific, justified reason.
URI and Collection Lexicons
Lexicons are enumerable indexes that let you iterate over all distinct values — URIs, collection names, or word terms — without opening any document fragments. Without them, enumeration degrades to fragment scanning.
| Setting | Default | Cost | Triggers reindex |
|---|---|---|---|
uri lexicon | on | Low — URI string list | No (incremental) |
collection lexicon | off | Low — collection string list | No (incremental) |
URI Lexicon
The URI lexicon maintains an enumerable index of all document URIs. It powers cts:uris() — the standard tool for URI-centric operations: auditing, ETL staging, bulk verification, existence checking, and URI-pattern queries.
Without the URI lexicon, cts:uris() resolves URIs by loading fragments rather than consulting an index. For small databases this is invisible. At scale, the difference between an index lookup and a fragment scan is the difference between a routine operation and a cluster-affecting one.
xquery version "1.0-ml";
(: URI lexicon vs document scan for URI enumeration.
cts:uris() is a pure lexicon operation — no fragments are loaded.
Iterating fn:collection() with fn:document-uri() loads every fragment.
Both return the same URIs; the difference is entirely in execution cost.
This gap becomes significant at scale or when the access pattern is
URI-only (auditing, ETL staging, existence checking). :)
let $coll := "llamaverse"
(: Lexicon path — index-only, no fragment loading :)
let $lexicon-count := fn:count(cts:uris("", (), cts:collection-query($coll)))
(: Fragment scan path — loads every document :)
let $scan-count := fn:count(
for $doc in fn:collection($coll)
return fn:document-uri($doc)
)
(: Sample the first few URIs from each to confirm equivalence :)
let $lexicon-samples := fn:subsequence(cts:uris("", (), cts:collection-query($coll)), 1, 3)
return string-join((
"URI enumeration: lexicon vs document scan",
"",
concat("cts:uris() count (lexicon): ", $lexicon-count),
concat("fn:collection() scan count: ", $scan-count),
"",
"Sample URIs from cts:uris():",
string-join($lexicon-samples, " "),
"",
"URI lexicon must be enabled for cts:uris() to use the index path.",
"Without it, cts:uris() still works but loads fragments to resolve URIs."
), " ")
'use strict';
// URI lexicon vs document scan for URI enumeration.
// cts.uris() is a pure lexicon operation — no fragments are loaded.
// Iterating fn.collection() with fn.documentUri() loads every fragment.
// Both return the same URIs; the difference is entirely in execution cost.
const coll = 'llamaverse';
// Lexicon path — index-only, no fragment loading
const lexiconCount = fn.count(cts.uris('', [], cts.collectionQuery(coll)));
// Fragment scan path — loads every document
let scanCount = 0;
for (const doc of fn.collection(coll)) {
fn.documentUri(doc);
scanCount++;
}
// Sample first few from lexicon
const samples = [];
let i = 0;
for (const uri of cts.uris('', [], cts.collectionQuery(coll))) {
if (i++ >= 3) break;
samples.push(uri);
}
[
'URI enumeration: lexicon vs document scan',
'',
`cts.uris() count (lexicon): ${lexiconCount}`,
`fn.collection() scan count: ${scanCount}`,
'',
'Sample URIs from cts.uris():',
samples.join('\n'),
'',
'URI lexicon must be enabled for cts.uris() to use the index path.',
'Without it, cts.uris() still works but loads fragments to resolve URIs.'
].join('\n')
URI enumeration: lexicon vs document scan
cts:uris() count (lexicon): 1277
fn:collection() scan count: 1277
Sample URIs from cts:uris():
/cleverllamas/llamaverse/content/llama-movement/llama_location_history.json
/cleverllamas/llamaverse/content/schools/bb22cc33-dd44-4e55-ff66-778899001122.json
/cleverllamas/llamaverse/raw/wild-llamas/llamas/0c8bdb0d-ac62-49b7-ac74-94dbba46efa5.json
URI lexicon must be enabled for cts:uris() to use the index path.
Without it, cts:uris() still works but loads fragments to resolve URIs.
Both paths return the same URI count. The distinction is entirely in execution: with the lexicon enabled, the count is resolved from a pure index read with no fragment loading. Without it, every URI in the count requires a fragment to be opened.
Enable when: URI enumeration, URI-pattern queries, ETL auditing, or any URI-centric workflow is in scope. This should be on for most production databases.
What degrades without it: cts:uris() queries become proportionally slower as the corpus grows. No error — just degradation.
Collection Lexicon
The collection lexicon maintains an enumerable index of all collection URIs. It backs cts:collection-lexicon() and collection-centric query planning.
Without it, collection-based queries still work but cannot use the accelerated collection enumeration path. This matters when collections encode access control, lifecycle stage, or business partitioning — which they frequently do in MarkLogic architectures.
xquery version "1.0-ml";
(: Enumerate all distinct collection URIs via the collection lexicon. :)
(: This is an index read — no document fragments are opened. :)
(: Without the collection lexicon enabled, cts:collection-lexicon() may :)
(: return empty or degrade, and enumeration requires scanning documents. :)
let $all := cts:collection-lexicon()
let $count := fn:count($all)
return (
fn:concat("Total distinct collections (index): ", $count),
"Collections:",
for $c in $all
order by $c
return fn:concat(" ", $c)
)
// Enumerate all distinct collection URIs via the collection lexicon.
// This is an index read — no document fragments are opened.
var all = cts.collectionLexicon().toArray();
var result = ["Total distinct collections (index): " + all.length, "Collections:"];
all.sort().forEach(function(c) { result.push(" " + c); });
result;
Total distinct collections (index): 6
Collections:
health-records
http://marklogic.com/entity-services/models
llama-movement
llamaverse
llamaverse-config
llamaverse-templates
(Results will vary depending on the llamaverse version deployed.
Validate against your environment after enabling the collection lexicon.)
Enable when: collections are a first-class access control or routing mechanism, or when collection-level reporting and auditing is part of operations.
Directory, Metadata, and Inheritance
These settings govern the filesystem-like directory hierarchy, document timestamps, and whether child documents inherit properties from their parent directory. The failure modes here are subtle: the settings interact with security, relevance scoring, and collection routing in ways that are not obvious from the Admin UI labels.
| Setting | Default | Effect when wrong |
|---|---|---|
directory creation | automatic | manual-enforced with missing directories: insert failures |
maintain last modified | off | Timestamp-dependent pipelines receive stale or missing values |
maintain directory last modified | off | Directory change detection unavailable |
inherit permissions | off | Documents inserted without explicit permissions are inaccessible to non-privileged users |
inherit collections | off | Documents silently miss intended collection membership |
inherit quality | off | Relevance scores drift unpredictably based on directory ancestry |
Directory Creation
directory creation has three modes that differ in how MarkLogic manages directory documents:
| Mode | Behaviour | When to use |
|---|---|---|
automatic | Directory documents are created automatically when documents are stored at URIs with path components | Standard choice for most deployments; no manual intervention required |
manual | Directory documents must be created explicitly before storing documents at that URI path | Controlled hierarchy environments where directory creation is a governed operation |
manual-enforced | Storing a document at a URI whose directory does not exist is rejected with an error | Strict hierarchy governance; every URI path must be pre-registered |
The default automatic is correct for most deployments. Switch to manual or manual-enforced only when directory management is part of an intentional access control or governance model.
What breaks with manual-enforced and missing directories: document inserts fail with errors. In automated ingestion pipelines, this surfaces as transaction failures that can be difficult to diagnose if the pipeline does not clearly log the rejected URIs.
Maintain Last Modified and Maintain Directory Last Modified
These settings cause MarkLogic to record a last-modified timestamp in document properties whenever a document or directory is updated.
maintain last modified records a per-document timestamp. maintain directory last modified records a per-directory timestamp.
If your pipeline depends on xdmp:document-properties(uri)/prop:last-modified, this setting must be enabled — there is no other way to get a reliable change timestamp from MarkLogic without it.
Cost: each update incurs an additional properties write. For high-throughput write workloads, this overhead is measurable.
What breaks without it: downstream systems that depend on the timestamp receive stale or missing values. The document data is correct — only the change-detection metadata is absent.
Inherit Permissions, Collections, and Quality
These three settings control whether documents and subdirectories inherit the corresponding property from their parent directory.
inherit permissions causes newly stored documents to receive the permissions of their parent directory. Without it, documents must have permissions set explicitly at write time. The failure mode when this is unexpectedly disabled: documents appear with no permissions, which means they are only accessible to privileged users until permissions are explicitly set.
inherit collections causes newly stored documents to automatically join the parent directory's collections. This is powerful for lifecycle and partition patterns where collections encode processing stage or access class — but misconfiguring it causes documents to silently land in the wrong pipeline. The failure is not an error; it is invisible misdirection.
inherit quality causes documents to inherit the relevance quality value from their parent directory. Enables directory-level ranking policy — all content in a trusted directory gets a quality boost, all content in staging gets a lower weight. The failure mode when this is enabled without a coherent quality policy: search results drift in unpredictable directions based on directory ancestry.
Enable based on explicit policy. Each of these settings encodes an assumption about document hierarchy that must be documented and tested. Enabling them without understanding your URI naming convention leads to access control drifts, collection misrouting, and unexplained relevance shifts.
Durability and Recovery
The settings in this family are set-and-forget right up until the cluster crashes. Locking, journaling, and storage failure handling control how MarkLogic behaves under concurrent write load, at crash recovery time, and when storage becomes unavailable. The wrong choices here do not produce errors — they produce silent data loss, duplicate documents, or degraded failover behaviour.
| Setting | Default | Recommended (production) | Effect of weakening |
|---|---|---|---|
locking | fast | strict | Duplicate URIs or silent last-write-wins overwrites under concurrent load |
journaling | fast | strict | Committed transactions lost on crash within the async sync window |
journal size | 675 MB | Size to peak throughput | Churn and stalls (too small); wasted disk only (too large) |
preallocate journals | false | Leave as-is | No effect since MarkLogic 8.0-4 |
shutdown on storage failure | false | true (with failover configured) | Host stays up in degraded state; split-brain risk |
storage failure timeout | 30 s | Tune to storage layer | False failover (too low) or masked failure (too high) |
Locking
locking controls how MarkLogic handles concurrent writes to the same URI:
| Mode | Behaviour | Appropriate use |
|---|---|---|
strict | Full serialisation; concurrent writers to the same URI block until the lock is released | Production databases where data integrity under concurrency is required |
fast | Optimistic check; detects most conflicts but with reduced blocking | Production databases where the write pattern makes same-URI conflicts unlikely |
off | No locking; concurrent writes proceed without conflict detection | Controlled bulk load of known-unique URIs under a single writer; never in production with concurrent access |
The failure mode for off is duplicate document creation or silent last-write-wins overwrite depending on timing. Both are data integrity problems that surface late — during queries that expect unique URIs, or during audits that discover more documents than were inserted.
fast is an acceptable production default when your architecture ensures that concurrent writes to the same URI are architecturally impossible. strict remains correct under all concurrent write patterns.
Journaling
journaling controls the write-ahead log durability model:
| Mode | Behaviour | Risk |
|---|---|---|
strict | Each transaction is written to the journal and synced to disk before acknowledgement | Full crash durability; the standard for production |
fast | Journal writes are buffered; sync is asynchronous | Journal lag window exists; committed transactions may not survive an immediate crash |
off | No journal; writes acknowledged immediately with no crash recovery guarantee | Data loss on any crash; not acceptable in production |
strict is the correct default for production databases. fast is occasionally used in bulk-load scenarios where recovery-point objective permits the lag and throughput is prioritised over strict durability. off should not exist in a production environment.
journal size determines the on-disk allocation per journal. Size it to peak transaction throughput and recovery objectives. Too small causes journal churn and stalls. Too large wastes disk but has no correctness impact.
preallocate journals has had no effect since MarkLogic 8.0-4. It is a historical setting. Do not tune it.
Storage Failure Handling
shutdown on storage failure controls whether a MarkLogic host self-terminates when a forest becomes unavailable due to storage problems.
When enabled: if storage goes dark, the host shuts down. In a properly configured failover cluster, the replica takes over cleanly. The cluster remains consistent.
When disabled: the host remains up in a partial-failure state. This can produce inconsistent search results, incomplete writes, or split-brain behaviour depending on what queries reach the degraded host. In a failover-configured cluster, enabling this setting is the safer choice.
storage failure timeout sets how many seconds must elapse before storage is considered failed (minimum 30). Tune to the characteristics of your storage layer. Too low: transient I/O delays trigger failover unnecessarily. Too high: genuine failures are masked for too long.
In-Memory Stands and Index Buffering
The in-memory settings control how much RAM MarkLogic allocates to write-path buffers for each index type before flushing to disk. Tuning these without profiling first is guesswork. Every number here should be driven by observed pressure, not instinct.
| Setting | What it controls | Symptom of wrong sizing |
|---|---|---|
in memory limit | Maximum fragments before flush is forced | Too low: frequent flushes, merge pressure; too high: memory exhaustion |
in memory list size | Termlist buffer allocation | Contention under write-heavy load |
in memory tree size | Fragment tree buffer allocation | Write throughput stalls, excessive flush cadence |
in memory range index size | Range index write buffer | Ingest stalls with range-heavy workloads |
in memory reverse index size | Reverse index write buffer | Ingest stalls when fast reverse searches is on |
in memory triple index size | Triple index write buffer | Triple ingest throughput bottleneck |
in memory geospatial region index size | Geospatial region index write buffer | Geospatial ingest stalls |
In-Memory Limit, List Size, and Tree Size
These three settings form the core write-path buffer:
in memory limit is the maximum number of fragments the in-memory stand can hold before a flush to disk is forced. Too low: frequent small flushes, excessive merge work. Too high: memory pressure under sustained load.
in memory list size is the memory allocation for termlist structures — the intermediate structures for term posting lists before they are merged to on-disk stands. Increase only when you observe termlist contention under write-heavy load with profiling evidence.
in memory tree size is the memory for in-memory fragment tree structures. Incorrect sizing shows up as write throughput stalls or excessive flush cadence.
None of these settings should be changed without establishing a write-throughput and flush-cadence baseline first.
Per-Index Memory Settings
Each specialised index type has its own in-memory buffer:
| Setting | What it buffers | Increase when |
|---|---|---|
in memory range index size | Range index writes during ingest | Range-heavy ingest workloads show write pressure |
in memory reverse index size | Reverse index writes | fast reverse searches is on and write pressure is observed |
in memory triple index size | Triple index writes | triple index is on and triple ingest is a throughput bottleneck |
in memory geospatial region index size | Geospatial region index writes | Geospatial region ingest shows observable pressure |
Tune each of these only when the corresponding index type is in active use and you have observed ingest pressure attributable to that index.
Merge Settings
Merging is how MarkLogic consolidates on-disk stands over time — combining smaller stands into larger ones to bound the number of stands a query has to fan out across. The settings in this family control when a merge is eligible to run and how large its output is allowed to get. They are tuning knobs for write-heavy and heavily-fragmented databases, not correctness settings — getting them wrong costs performance, not data integrity.
| Setting | Default | Effect |
|---|---|---|
merge max size | 49152 MB (48 GB) | Caps the size of the stand a merge is allowed to produce; raising it allows larger consolidated stands at the cost of longer individual merges |
merge min size | 1024 MB (1 GB) | Minimum stand size eligible to participate in a merge; lowering it makes small stands merge sooner |
merge min ratio | 3 | Minimum ratio between stand sizes for them to be considered for merging; lower values merge more aggressively |
merge timestamp | 0 | Forces merges to discard deleted-fragment history at or before a given timestamp; 0 means no forced cutoff |
Merge Max Size, Min Size, and Min Ratio
merge max size bounds how large a single merge's output stand is allowed to be. The default of 48 GB is generous for most deployments. Lower it if very large stands are causing merge operations that run long enough to interfere with maintenance windows or storage headroom planning.
merge min size sets the smallest stand size eligible to be merged. Stands below this size are candidates for merging sooner; raising it defers merging of small stands, which trades a higher stand count for less merge I/O.
merge min ratio governs how similar in size two stands must be before MarkLogic considers merging them — the merge policy prefers combining stands of comparable size. Lowering the ratio makes merging more aggressive (more merge I/O, fewer stands); raising it makes merging more conservative.
None of these three should be changed without a merge-cadence and I/O baseline. They interact with ingest rate, stand count, and query fan-out in ways that are specific to each deployment's write pattern.
The example below reads the current merge configuration directly from the Admin configuration for this database:
xquery version "1.0-ml";
import module namespace admin = "http://marklogic.com/xdmp/admin"
at "/MarkLogic/admin.xqy";
let $config := admin:get-configuration()
let $dbid := xdmp:database("cleverllamas-content")
return string-join((
concat("merge max size (MB): ", admin:database-get-merge-max-size($config, $dbid)),
concat("merge min size (MB): ", admin:database-get-merge-min-size($config, $dbid)),
concat("merge min ratio: ", admin:database-get-merge-min-ratio($config, $dbid)),
concat("merge timestamp: ", admin:database-get-merge-timestamp($config, $dbid))
), " ")
'use strict';
const admin = require('/MarkLogic/admin.xqy');
const config = admin.getConfiguration();
const dbid = xdmp.database('cleverllamas-content');
[
'merge max size (MB): ' + admin.databaseGetMergeMaxSize(config, dbid),
'merge min size (MB): ' + admin.databaseGetMergeMinSize(config, dbid),
'merge min ratio: ' + admin.databaseGetMergeMinRatio(config, dbid),
'merge timestamp: ' + admin.databaseGetMergeTimestamp(config, dbid)
].join('\n');
merge max size (MB): 49152
merge min size (MB): 1024
merge min ratio: 3
merge timestamp: 0
Change when: you have observed merge behaviour — excessive stand count, merge operations competing with query/ingest traffic, or merges producing undesirably large stands — and have a specific target in mind.
Leave at defaults absent that evidence. These are among the easiest settings in the entire configuration to tune based on intuition and regret later.
Merge Timestamp
merge timestamp forces merges to discard fragment versions at or before the given timestamp, permanently removing the ability to see document history before that point (relevant to point-in-time queries and MVCC-based recovery scenarios). The default 0 means no forced cutoff — ordinary merge behaviour retains what MVCC and garbage collection require.
Use with explicit intent only, in the same spirit as reindexer timestamp above: this is a manual intervention with a permanent effect on fragment history, not a routine maintenance dial.
Reindexing, Rebalancing, and Forest Assignment
Changes to index settings only propagate to existing fragments when reindexing completes. The settings in this family control how that reindex work runs, how forest imbalances are corrected, and how documents are distributed across forests.
| Setting | Default | Effect |
|---|---|---|
reindexer enable | true | Disabling defers reindex work; does not eliminate it |
reindexer throttle | 3 | 1–5; raise during maintenance windows, lower during traffic |
reindexer timestamp | 0 | Forces reindex of all fragments at or before specified timestamp; use with explicit intent only |
rebalancer enable | true | Controls whether topology changes trigger background redistribution |
rebalancer throttle | 5 | Same 1–5 scale as reindexer throttle |
assignment policy | bucket | Governs which forest receives each new document; changing on a populated database triggers rebalance |
Reindexer Enable and Throttle
reindexer enable controls whether background reindexing runs when index settings change. Leave this enabled in normal operations. Disable it only when executing a controlled staged migration where you need to defer the reindex blast until a maintenance window.
Disabling the reindexer during an index setting change does not reduce the work — it only controls when it happens. The full reindex runs when you re-enable it.
reindexer throttle sets the priority of reindex work relative to query and ingest traffic (1 to 5):
- Maintenance windows: raise to 4 or 5 to complete the reindex quickly
- Business hours: lower to 1 or 2 to avoid impacting query latency
- The default (3) is a neutral starting point, not an optimal production setting
Reindexer Timestamp
reindexer timestamp forces a reindex of all fragments with a timestamp at or before the specified value. This is a manual intervention mechanism — it triggers a full or partial corpus reindex based on timestamp regardless of whether index settings changed.
Use with explicit intent and a rollback plan. Setting this incorrectly against a large corpus triggers an unexpected full reindex that cannot easily be cancelled.
Never use this as a routine maintenance tool.
Rebalancer Settings
rebalancer enable and rebalancer throttle control background rebalancing — the process by which MarkLogic redistributes document assignments across forests when topology changes. Same throttle logic as the reindexer: raise in maintenance windows, lower during traffic.
Assignment Policy
assignment policy determines how documents are distributed to forests at write time:
| Policy | Distribution basis | When to use |
|---|---|---|
bucket | Consistent hashing on URI | Good general default; balanced distribution, URI-predictable placement |
segment | URI range segmentation | Useful when URI patterns map to natural partitions |
statistical | Weighted random; adjusts to usage patterns | Good when forests have different throughput characteristics |
range | Explicit range configuration based on index values | Partition-by-value workflows where placement must be deterministic |
query | Placement based on query constraints | Specialised use; follow the documentation closely |
legacy | Original hash-based placement | Exists for backwards compatibility; do not choose for new deployments |
The policy decision is architectural. Choose based on how your URI namespace encodes content partitioning, your scaling plan, and how rebalancing should behave when forests are added or removed. Changing this on an existing populated database triggers a rebalance of potentially the full corpus.
Query Engine and Performance Internals
These settings tune runtime query execution and the conditions under which MarkLogic loads or validates its index structures. Most have sensible defaults that should not be changed without evidence of a specific problem.
| Setting | Default | Change when |
|---|---|---|
range index optimize | facet-time | Memory pressure from range/facet operations outweighs latency need |
large size threshold | 1024 KB | Fragment distribution skews near the boundary with measurable impact |
positions list max size | 256 MB | A specific high-frequency term is measurably dominating index space |
index detection | automatic | Leave as-is — none requires the caller to own compatibility |
preload mapped data | false | Cold-query latency after restarts is a documented operational concern |
preload mapped replica data | false | Same, for replica stands; only relevant when preload mapped data is on |
expunge locks | automatic | Leave as-is when using xdmp:lock-acquire() with timeouts |
Range Index Optimise
range index optimize sets the optimisation goal for range index operations:
facet-time: optimise for facet query latency at the expense of memorymemory-size: reduce memory footprint at the expense of facet latency
Choose based on which bottleneck you are actually in. Understand the trade-off before picking a direction.
Large Size Threshold
Documents larger than this threshold are handled as "large documents" — they bypass the in-memory stand and are written directly to large data storage. The default works for most corpora. Adjust only when your fragment distribution shows a meaningful proportion of documents near the boundary and you are observing unexpected merge or query behaviour attributable to the threshold.
Positions List Max Size
This limits the on-disk size of positions lists for any single term. High-frequency tokens accumulate large positions lists. The limit caps their size, with the trade-off that very frequent terms lose positional precision beyond the cap. The default is deliberately generous. Reduce it only when a specific high-frequency term is demonstrably dominating index space.
Index Detection
index detection controls whether MarkLogic automatically detects index compatibility when opening stands.
automatic (the default): MarkLogic checks index compatibility and handles transitions safely. none: No detection; the caller assumes responsibility for compatibility.
Leave this at automatic.
Preload Mapped Data and Preload Mapped Replica Data
preload mapped data causes MarkLogic to warm up index structures when a stand is opened, rather than waiting for the first queries to do so.
Enable when cold-query latency is a known operational constraint (for example, after maintenance restarts) and memory is available. Do not enable if memory pressure is already a concern.
preload mapped replica data controls the same behaviour for replica stands. Only relevant when preload mapped data is also enabled.
Expunge Locks
expunge locks controls automatic cleanup of expired lock fragments created by xdmp:lock-acquire(). Set to automatic when lock fragments are used with timeouts and you expect them to expire in normal operation. Without automatic expunge, expired lock fragments accumulate as operational noise.
Only relevant when your application uses xdmp:lock-acquire() with timeout semantics.
Database Wiring, Security, and Encryption
These settings connect the content database to its auxiliary databases and define the encryption posture. They are foundational — misconfiguring any of them affects every query that touches security, schema, or module resolution.
| Setting | What goes wrong when misconfigured |
|---|---|
database name | Misidentification in runbooks, monitoring alerts, and application server config |
security database | Login failures; silent permission enforcement drift |
schema database | Schema validation and schema-aware query planning diverge from actual document structure |
triggers database | Trigger registration or execution fails silently |
enable encryption at rest | Unencrypted on-disk content where compliance requires encryption |
encryption-at-rest backup key choice | Wrong blast-radius scope if cluster key is rotated or compromised |
database encryption key ID | Decrypt failure at restore time if key is rotated without updating this value |
Database Name and Auxiliary Database Pointers
database name is the logical identifier used everywhere — application server configuration, monitoring labels, runbooks, alerting rules. Name it for its content role, not its deployment slot. customer-content is a good name. prod-db-2 is not.
The auxiliary database pointers connect the content database to its dependencies:
security database: all permission checks, role resolution, and privilege lookups go here. Point this at the wrong database and logins fail, role checks silently succeed or fail depending on what that database contains, and permission enforcement drifts from policy.
schema database: schema-aware operations resolve their schema documents here. Point this wrong and schema validation and schema-aware query planning silently diverge from the content database's actual structure.
triggers database: trigger configuration and trigger module references are resolved here. Point this wrong and trigger registration or execution may fail silently or produce unexpected cross-database behaviour.
Run the Baseline example at the top of this article before and after any change to these pointers. It shows the current state of all four and verifies that the database sees what it should.
Encryption at Rest
enable encryption at rest encrypts the on-disk content of the database. Enable it where compliance requirements or security policy mandate encrypted on-disk content.
Cost: encryption adds I/O overhead on write and read paths. Benchmark this in your environment before enabling in a throughput-sensitive context.
encryption-at-rest backup key choice determines whether backup files are protected with the cluster key or a database-specific key. This is a DR ownership decision: the cluster key is shared across the environment, a database key is scoped to the database. Choose based on key-rotation and blast-radius policy.
database encryption key ID identifies the specific key protecting data. Track this in your key-management runbook and rotation schedule. A mismatch between the key ID and the actual key results in decrypt failures at restore time.
Run a backup-and-restore drill before enabling encryption in production.
Triple Index and Geospatial
These settings enable RDF triple storage and geospatial indexing. Both are off by default and should only be enabled when those workloads are confirmed in scope — the index cost is proportional to triple and geometry volume.
| Setting | Default | Index cost | Triggers reindex |
|---|---|---|---|
triple index | off | High — proportional to triple volume | Yes |
triple positions | off | Moderate additional to triple index | Yes |
triple index geohash precision | 6 | Higher precision = more space; lower = coarser spatial resolution | Yes |
geospatial region index | off | High for polygon-dense datasets | Yes |
Triple Index and Triple Positions
triple index enables MarkLogic's RDF triple store capability. SPARQL queries and sem:sparql() calls require this index. Without it, semantic workloads are unavailable.
Enable only when semantic features are in scope. The triple index affects ingest throughput and memory consumption in proportion to the volume of triples ingested.
triple positions enables positional support for triple-range and frequency-sensitive triple queries. Enable when triple proximity or term-frequency fidelity matters for your semantic workload. Leave off for basic SPARQL pattern matching where positions are not needed.
Triple Index Geohash Precision
When geospatial data is stored as RDF triples, this setting controls the resolution of the geohash used to index geometry values.
Higher precision: finer-grained spatial indexing, more index space. Lower precision: coarser spatial resolution, less space. Use the coarsest precision that satisfies your spatial query requirements. Do not choose maximum precision without confirming you need it.
Reverse Geospatial Queries and the Region Index
Reverse geospatial work is where teams most often over-assume that point indexes can answer polygon questions. They cannot. If the question is whether one polygon contains or intersects another polygon, you need a region corpus and a region path index. If that index is missing, MarkLogic can return misleading candidates or appear to work only because filtering is correcting a bad index result.
The example below uses a known llamaverse movement point for Aaron and two overlapping polygons to validate the full set of spatial operator semantics:
xquery version "1.0-ml";
import module namespace geo = "http://marklogic.com/geospatial"
at "/MarkLogic/geospatial/geospatial.xqy";
(: Geometry harness built around a known llamaverse movement point. :)
(: This validates the operator semantics you need for reverse-geospatial :)
(: reasoning even though the current llamaverse corpus is point-centric. :)
let $aaron-point := cts:point(53.5136376, -9.9219172)
let $zone-a := cts:polygon((
cts:point(53.50, -9.99),
cts:point(53.50, -9.88),
cts:point(53.57, -9.88),
cts:point(53.57, -9.99),
cts:point(53.50, -9.99)
))
let $zone-b := cts:polygon((
cts:point(53.52, -9.95),
cts:point(53.52, -9.84),
cts:point(53.59, -9.84),
cts:point(53.59, -9.95),
cts:point(53.52, -9.95)
))
return (
fn:concat("Aaron point contained by zone A: ", geo:contains($zone-a, $aaron-point)),
fn:concat("Aaron point contained by zone B: ", geo:contains($zone-b, $aaron-point)),
fn:concat("Zone A intersects Zone B: ", geo:intersects($zone-a, $zone-b)),
fn:concat("Zone A contains Zone B: ", geo:contains($zone-a, $zone-b)),
fn:concat("Zone B contains Zone A: ", geo:contains($zone-b, $zone-a))
)
'use strict';
const geo = require('/MarkLogic/geospatial/geospatial.xqy');
// Geometry harness built around a known llamaverse movement point.
// Validates operator semantics for reverse-geospatial reasoning.
const aaronPoint = cts.point(53.5136376, -9.9219172);
const zoneA = cts.polygon([
cts.point(53.50, -9.99), cts.point(53.50, -9.88),
cts.point(53.57, -9.88), cts.point(53.57, -9.99),
cts.point(53.50, -9.99)
]);
const zoneB = cts.polygon([
cts.point(53.52, -9.95), cts.point(53.52, -9.84),
cts.point(53.59, -9.84), cts.point(53.59, -9.95),
cts.point(53.52, -9.95)
]);
[
"Aaron point contained by zone A: " + geo.contains(zoneA, aaronPoint),
"Aaron point contained by zone B: " + geo.contains(zoneB, aaronPoint),
"Zone A intersects Zone B: " + geo.intersects(zoneA, zoneB),
"Zone A contains Zone B: " + geo.contains(zoneA, zoneB),
"Zone B contains Zone A: " + geo.contains(zoneB, zoneA)
];
Aaron point contained by zone A: true
Aaron point contained by zone B: false
Zone A intersects Zone B: true
Zone A contains Zone B: false
Zone B contains Zone A: false
| Operator | What it verifies | What degrades when the region index is missing |
|---|---|---|
geo:contains(region, point) | Zone A encloses Aaron's point; Zone B does not | May degrade to an over-broad candidate set that needs filtering |
geo:intersects(region, region) | Zone A and Zone B overlap | May over-include or over-exclude depending on index coverage |
geo:contains(region, region) | Neither polygon fully encloses the other | Boundary semantics become especially fragile |
The practical rule: if the question is polygon-on-polygon, do not trust unfiltered execution unless you have proved the exact region index path on your version and corpus.
Practical Change Policy
Use this sequence for any database-configuration change in a production environment:
- Run the Baseline example and all family-specific examples relevant to the change. Record the current counts as a baseline.
- Apply the setting change in one logical group at a time.
- Keep the reindexer enabled. Monitor progress before re-exposing the database to full query traffic.
- Track reindexer and rebalancer progress with operational telemetry.
- Re-run the examples and compare. Counts should shift in the direction the setting change implies.
- Record the change, observed impact, and rollback note in the change log.
Final Checklist
Before signing off on a database configuration change:
- [ ] Context confirmed: correct database, correct environment, correct auxiliary database pointers
- [ ] Index dependencies respected: element word positions enabled only when word positions is on; fast phrase searches only when word positions is on; wildcard positions only when both three-character searches and word positions are on
- [ ] Reindex and rebalance posture documented: if the change triggers reindexing, you know when it runs and have a monitoring plan
- [ ] Durability posture reviewed:
lockingandjournalingmatch business risk tolerance and have not been weakened without explicit approval - [ ] Encryption posture reviewed: if encryption is changing, key management and restore paths have been drilled end-to-end
- [ ] Destructive Admin operations protected:
clear,delete, and forcedreindexrequire explicit written approval before execution
Need Some Help?
Looking for more information on this subject or any other topic related to MarkLogic? Contact Us (info@cleverllamas.com) to find out how we can assist you with consulting or training!
- How to Use This Manual
- Baseline: Know Your Context
- Text Search Index Settings
- Word Searches
- Stemmed Searches
- Word Positions
- Filtered and Unfiltered Execution
- Fast Phrase Searches
- Fast Reverse Searches
- Case and Diacritic Sensitivity
- TF Normalisation
- Default Language
- Element, Field, and Attribute Indexes
- Fast Element Word Searches
- Element Word Positions and Fast Element Phrase Searches
- Element Value Positions and Attribute Value Positions
- Field Value Searches and Field Value Positions
- Wildcard and Lexicon Indexes
- Three-Character Searches
- Three-Character Word Positions
- Trailing Wildcard Searches
- Fast Element Trailing Wildcard Searches and Trailing Wildcard Word Positions
- Word Lexicons
- Two-Character and One-Character Searches
- URI and Collection Lexicons
- URI Lexicon
- Collection Lexicon
- Directory Creation
- Maintain Last Modified and Maintain Directory Last Modified
- Durability and Recovery
- Locking
- Journaling
- Storage Failure Handling
- In-Memory Stands and Index Buffering
- In-Memory Limit, List Size, and Tree Size
- Per-Index Memory Settings
- Merge Settings
- Merge Max Size, Min Size, and Min Ratio
- Reindexing, Rebalancing, and Forest Assignment
- Reindexer Enable and Throttle
- Rebalancer Settings
- Assignment Policy
- Query Engine and Performance Internals
- Range Index Optimise
- Large Size Threshold
- Positions List Max Size
- Index Detection
- Preload Mapped Data and Preload Mapped Replica Data
- Expunge Locks
- Database Wiring, Security, and Encryption
- Database Name and Auxiliary Database Pointers
- Encryption at Rest
- Triple Index and Geospatial
- Triple Index and Triple Positions
- Triple Index Geohash Precision
- Reverse Geospatial Queries and the Region Index
- Practical Change Policy
- Final Checklist