Series overview
Part 6 of 1833% complete
2026-05-12•13 min read

Mapping and analysis in depth

Chapter 05 declared three name analysers and two normalisers and promised to explain them. This chapter does that with the _analyze API and the million profiles in user-profile-v1. You watch each stage of analysis transform José Fernandes-D'Souza, see why a term query for Prashant Kumar finds nothing while a match query finds 81,676 profiles, and reproduce the two mapping failures that exhaust clusters: fielddata on a text field and mapping explosion.

You need the lab with user-profile-v1 loaded (chapter 05). Every count in this chapter was checked against the equivalent PostgreSQL query. The chapter takes about 40 minutes; the scratch indices it creates are deleted as you go.

Text and keyword: two ways to store a string

What it is. A keyword field indexes the whole value as one term, exactly as given (after an optional normaliser). A text field runs the value through an analyser, which splits it into terms, usually words, and transforms them. Queries match terms, not the original string. See the text and keyword references.

Why it matters at 300 million profiles. Every string field in the projection must be one or the other, or both through a multi-field. Choose text for a field that is really an identifier and exact lookups stop working; choose keyword for a name and nobody can find Kumar inside Prashant Kumar.

Example. The _analyze API shows exactly which terms a piece of text produces. Compare the standard analyser, which text fields use by default, with the keyword analyser:

GET _analyze?filter_path=tokens.token
{ "analyzer": "standard", "text": "José Fernandes-D'Souza" }
{"tokens":[{"token":"josé"},{"token":"fernandes"},{"token":"d'souza"}]}
GET _analyze?filter_path=tokens.token
{ "analyzer": "keyword", "text": "José Fernandes-D'Souza" }
{"tokens":[{"token":"José Fernandes-D'Souza"}]}

The standard analyser produced three lowercase terms; the hyphen split two of them, and the apostrophe stayed inside d'souza. The keyword analyser produced one term, unchanged.

Common mistake. Thinking of a field as “a string” and reasoning about queries as if they compared strings. They compare terms. The next section shows the consequence.

Why a term query finds nothing and a match query finds too much

What it is. A term query looks up its value in the inverted index as given: no analysis for text fields. A match query analyses its value with the field’s search analyser first, then looks up the resulting terms. For keyword fields, both apply the field’s normaliser, if it has one.

Why it matters at 300 million profiles. Mixing these up produces the most common bug report about search: “the profile exists but search can’t find it.” The data is there; the query and the index disagree about terms.

Example. Run each of these as GET user-profile-read/_count with the query shown. Chapter 05 mapped fullName as text with a fullName.keyword sub-field.

QueryCountWhy
{"term": {"fullName": "Prashant Kumar"}}0No single indexed term equals Prashant Kumar; the index holds prashant and kumar
{"term": {"fullName": "prashant"}}33,333Exactly the indexed term
{"match": {"fullName": "Prashant Kumar"}}81,676Analysed into prashant OR kumar; any profile with either word matches
{"match": {"fullName": {"query": "Prashant Kumar", "operator": "and"}}}1,666Both words required
{"term": {"fullName.keyword": "Prashant Kumar"}}1,666Whole-value match; the lowercase_ascii normaliser lowercases the query value too
{"match": {"fullName.keyword": "prashant"}}0A keyword field matches whole values only

The counts match PostgreSQL: 33,333 profiles have the first name Prashant, 1,666 are named exactly Prashant Kumar, and 81,676 have first name Prashant or last name Kumar.

Read the third row carefully. A match query’s default operator is OR, so a search box sending Prashant Kumar returns every Prashant and every Kumar, ranked by relevance. That is often the right behaviour for a ranked list, and always the wrong behaviour for a count or a filter. Chapter 11 decides when to use and, minimum_should_match, and filters.

Common mistake. Using term on a text field because “I want an exact match.” Exact matching belongs on a keyword field or sub-field. On a text field, term works only if you pass an already-analysed term, which means duplicating the analyser in application code.

The analysis pipeline, one stage at a time

What it is. An analyser is three stages, applied in order:

  1. Character filters transform the raw text before it is split, for example replacing & with and or stripping HTML.
  2. One tokeniser splits the text into tokens and records each token’s position and character offsets.
  3. Token filters transform, remove, or add tokens: lowercase them, fold accents, expand synonyms, generate prefixes.

Tata & Sons

Character filter

& to and

Tokeniser

standard

Token filter

lowercase

tata | and | sons

Tata & Sons

Character filter

& to and

Tokeniser

standard

Token filter

lowercase

tata | and | sons

The diagram follows Tata & Sons through a mapping character filter that replaces & with and, the standard tokeniser that splits the text into words, and a lowercase token filter, producing three terms: tata, and, and sons.

Why it matters at 300 million profiles. Every stage runs on every document at index time and on every query at search time. Each choice decides what users can find, and expensive stages, such as n-gram generation, multiply index size.

Example. Try each stage in isolation. You can define an analyser inline in _analyze without creating an index.

A mapping character filter:

GET _analyze?filter_path=tokens.token
{
"char_filter": [ { "type": "mapping", "mappings": ["& => and"] } ],
"tokenizer": "standard",
"filter": ["lowercase"],
"text": "Tata & Sons"
}
{"tokens":[{"token":"tata"},{"token":"and"},{"token":"sons"}]}

Two tokenisers on the same name. The standard tokeniser splits on Unicode word boundaries; the whitespace tokeniser splits only on spaces:

GET _analyze?filter_path=tokens.token
{ "tokenizer": "whitespace", "text": "José Fernandes-D'Souza" }
{"tokens":[{"token":"José"},{"token":"Fernandes-D'Souza"}]}

With the whitespace tokeniser, a search for Fernandes would not match this profile, because the hyphenated surname is one token. For names, the standard tokeniser’s behaviour is what users expect.

Common mistake. Adding stemming or stop-word filters to a name analyser because a generic “English” analyser has them. Names are not English prose. An English stemmer can reduce different names to the same stem, and a stop-word list can delete a real name that happens to be a common word.

The name analyser, and the price of accent folding

What it is. Chapter 05’s name_index analyser is the standard tokeniser followed by lowercase and asciifolding, which converts characters such as é to their closest ASCII equivalent.

GET user-profile-v1/_analyze?filter_path=tokens.token
{ "analyzer": "name_index", "text": "José Fernandes" }
{"tokens":[{"token":"jose"},{"token":"fernandes"}]}

Why it matters at 300 million profiles. Users type names the way their keyboard lets them. A support agent searching jose must find José, and a user who types JOSÉ must find the same profiles:

GET user-profile-read/_count?filter_path=count
{ "query": { "match": { "fullName": "jose" } } }
{"count":33333}

{"match": {"fullName": "JOSÉ"}} returns the same 33,333: both query forms become the term jose.

Trade-off. Folding throws information away. After folding, a search cannot tell José from Jose, and names that differ only by a diacritic rank identically. The preserve_original option keeps both forms:

GET _analyze?filter_path=tokens.token
{
"tokenizer": "standard",
"filter": [ "lowercase", { "type": "asciifolding", "preserve_original": true } ],
"text": "José"
}
{"tokens":[{"token":"jose"},{"token":"josé"}]}

With preserve_original, an accented query matches both terms and scores exact-accent matches higher, at the cost of more terms in the index. This series folds without preserving, because the query shapes treat José and Jose as the same person’s name. Needs validation: if your users routinely search in scripts other than Latin, such as Devanagari, folding does nothing for them; transliteration is a separate problem this series does not solve.

Common mistake. Folding at index time but not at search time. Any analyser used for search must produce terms the index can contain. Chapter 05 keeps name_index and name_search identical except for synonyms.

Index-time and search-time analysers

What it is. A text field can use a different analyser for queries than for documents, through search_analyzer. See index and search analysis.

Why it matters at 300 million profiles. Two features in this index depend on it. Synonyms are expanded only at search time, so the synonym list can change without reindexing. Autocomplete prefixes are generated only at index time, so a query for pras is not itself split into pr, pra, and pras.

Example. The search analyser expands synonyms from the profile-name-synonyms set:

GET user-profile-v1/_analyze?filter_path=tokens.token
{ "analyzer": "name_search", "text": "Mohd Imran" }
{"tokens":[{"token":"mohammed"},{"token":"mohammad"},{"token":"muhammad"},{"token":"mohd"},{"token":"imran"}]}

So {"match": {"fullName": "Mohd"}} counts 33,333 profiles, the same as a search for Mohammed, even though no profile contains the word Mohd. The prefix analyser runs only when documents are indexed:

GET user-profile-v1/_analyze?filter_path=tokens.token
{ "analyzer": "name_prefix_index", "text": "Prashant Kumar" }
{"tokens":[{"token":"pr"},{"token":"pra"},{"token":"pras"},{"token":"prash"},{"token":"prasha"},
{"token":"prashan"},{"token":"prashant"},{"token":"ku"},{"token":"kum"},{"token":"kuma"},{"token":"kumar"}]}

A query for pras against fullName.prefix is analysed with plain name_index, becomes the single term pras, and matches the stored fragment: {"match": {"fullName.prefix": "pras"}} counts 33,333, all the Prashants. Chapter 07 compares this with the other autocomplete options.

Common mistake. Using the prefix analyser at search time as well. The query pras would become pr OR pra OR pras, and pr matches Priya, Prashant, and every other name starting with those letters.

Normalisers: analysis for keyword fields

What it is. A normaliser is an analyser without a tokeniser. It can use only filters that work on single characters, such as lowercasing and folding, and always produces exactly one term.

Why it matters at 300 million profiles. Exact filters and lookups on city, state, and email must survive differences in case from users and upstream systems, without turning those fields into word-matched text.

Example.

GET user-profile-v1/_analyze?filter_path=tokens.token
{ "normalizer": "lowercase_ascii", "text": "Bengaluru URBAN Née" }
{"tokens":[{"token":"bengaluru urban nee"}]}

You can also analyse as a field, which applies whatever that field’s mapping specifies:

GET user-profile-v1/_analyze?filter_path=tokens.token
{ "field": "email", "text": "José.F@Example.com" }
{"tokens":[{"token":"josé.f@example.com"}]}

The email normaliser lowercased the value and left the accent alone, which is chapter 05’s lowercase_only decision.

Common mistake. Normalising at index time only in the application, by lowercasing values before indexing. Queries built elsewhere will not lowercase, and the two sides drift. A normaliser applies to both sides automatically.

See what a document actually indexed

The term vectors API shows the terms stored for one document, per field. It is the quickest way to answer “why doesn’t this profile match?”

GET user-profile-read/_termvectors/1?fields=fullName,fullName.prefix,fullName.keyword&filter_path=term_vectors.*.terms
fullName priya, kumar
fullName.prefix pr, pri, priy, priya, ku, kum, kuma, kumar
fullName.keyword priya kumar

The response is condensed here to the term lists. The full response also includes each term’s frequency, position, and character offsets, which is what highlighting uses in chapter 11. One source value, Priya Kumar, produced three independent sets of terms, one per multi-field.

Fielddata: why text fields refuse to sort and aggregate

What it is. Sorting and aggregations need to go from a document to its values. For keyword, numeric, and date fields, Elasticsearch stores that direction on disk at index time, as doc values. text fields have no doc values. The only way to sort or aggregate on one is fielddata: an in-memory structure built on the JVM heap by un-inverting the inverted index. It is disabled by default, which is why chapter 05’s aggregation on fullName failed.

Why it matters at 300 million profiles. Fielddata is loaded per shard, stays on the heap until evicted, and grows with the number of documents and distinct terms. On 300 million profiles with real, diverse names, one careless aggregation can load an enormous structure onto every data node’s heap.

Example. Build a scratch copy of the names with fielddata enabled, then aggregate on it:

PUT ch06-fielddata
{
"settings": { "number_of_shards": 1, "number_of_replicas": 0 },
"mappings": { "properties": { "fullName": { "type": "text", "fielddata": true } } }
}
POST _reindex?refresh=true&filter_path=total,created,failures
{
"source": { "index": "user-profile-read", "_source": ["fullName"] },
"dest": { "index": "ch06-fielddata" }
}
{"total":1000000,"created":1000000,"failures":[]}
GET ch06-fielddata/_search?filter_path=aggregations
{ "size": 0, "aggs": { "top_names": { "terms": { "field": "fullName", "size": 4 } } } }
{"aggregations":{"top_names":{"doc_count_error_upper_bound":0,"sum_other_doc_count":1799959,
"buckets":[{"key":"khan","doc_count":50011},{"key":"das","doc_count":50010},
{"key":"fernandes","doc_count":50010},{"key":"gupta","doc_count":50010}]}}}

The first problem is the answer itself. The buckets are words, not names: khan, not Zoya Khan. The same aggregation on fullName.keyword, which uses doc values, returns whole names:

{"buckets":[{"key":"zoya khan","doc_count":1668},{"key":"aditi das","doc_count":1667},
{"key":"aditi fernandes","doc_count":1667},{"key":"aditi gupta","doc_count":1667}]}

The second problem is memory. After the aggregation:

GET _cat/fielddata?v&h=node,field,size
node field size
4483bb3f14e5 fullName 1.5mb
...

The lab’s names use only 50 distinct words, so 1.5 MB is a small number. Real names have far more distinct terms, and the structure grows with both terms and documents. The fielddata circuit breaker rejects requests that would exceed its limit, 40% of the heap by default; in the lab, GET _nodes/stats/breaker reports that limit as 409.5 MB. A tripped breaker fails the query, which is better than an out-of-memory node, but either way the aggregation does not work. Delete the scratch index:

DELETE ch06-fielddata

Common mistake. Following the error message’s own suggestion, “set fielddata=true”, to make an aggregation work. The fix is a keyword sub-field, which chapter 05’s mapping already has. The text field reference documents fielddata and its memory cost.

Mapping explosion: when documents define the schema

What it is. With dynamic mapping, every new field name in any document adds a field to the index mapping, permanently. When documents carry data as field names, such as user IDs, campaign names, or arbitrary tags as keys, the mapping grows without bound. That is mapping explosion. The mapping is part of the cluster state, which every node holds and the master publishes on every change.

Why it matters at 300 million profiles. A profile projection tends to attract “just one more” nested object: marketing preferences, feature flags, consent records. If those arrive as open-ended key-value maps under dynamic mapping, the mapping grows with the business, not with the design.

Example. Create a scratch index with default dynamic mapping, and index a profile whose preferences object has 600 campaign keys, campaign_0_opt_in to campaign_599_opt_in:

PUT ch06-explosion
{ "settings": { "number_of_shards": 1, "number_of_replicas": 0 } }
PUT ch06-explosion/_doc/1
{ "userId": "1", "preferences": { "campaign_0_opt_in": true, "campaign_1_opt_in": true, ... } }

The ... stands for the remaining 598 keys; generate the document with a short script rather than typing it. Afterwards, the mapping holds 600 new boolean fields. A second profile with 500 different campaign keys is rejected:

document_parsing_exception: [1:11569] failed to parse: Limit of total fields [1000] has been exceeded while adding new fields [398]

That limit is index.mapping.total_fields.limit, 1,000 by default; see mapping limit settings. It is a safety net, not a design. Raising it is rarely the right fix.

The design fix depends on whether you need to query the data:

NeedMappingEffect
No unknown fields at all"dynamic": "strict"Documents with unknown fields are rejected (chapter 05)
Keep unknown fields, never query them"dynamic": falseStored in _source, not indexed, not added to the mapping
Query arbitrary keys as exact values"type": "flattened"The whole object is one field; every leaf value is indexed as a keyword

The flattened type handles the same two documents without growing the mapping:

PUT ch06-flat
{
"settings": { "number_of_shards": 1, "number_of_replicas": 0 },
"mappings": {
"dynamic": "strict",
"properties": {
"userId": { "type": "keyword" },
"preferences": { "type": "flattened" }
}
}
}

After indexing both documents, {"term": {"preferences.campaign_700_opt_in": "true"}} counts 1, and the mapping still shows only "preferences": {"type": "flattened"}. The trade-off is that every value is treated as a keyword: no numeric ranges, no full-text analysis. Delete both scratch indices:

DELETE ch06-explosion
DELETE ch06-flat

Common mistake. Treating the field limit error as the problem and raising the limit. The problem is that document data is being used as schema.

Common mistakes with mapping and analysis

  • Mapping everything as text. Filters on status and city then behave like word searches, cannot be aggregated, and match partial values.
  • Searching keyword fields with match, or text fields with term, without knowing what each does. Decide per field which query types are allowed, and enforce that in the query builder (chapter 10).
  • Enabling fielddata on a text field. Add a keyword sub-field instead.
  • Letting documents define fields. Use strict, false, or flattened deliberately.
  • Changing analysers on a live index. Analysis settings of an existing index cannot be changed while it is open, and even when they can, already-indexed terms keep the old analysis. A new analyser means a new index version (chapter 14).

What you now know, and what comes next

You can predict which terms any field produces with _analyze, explain from the analysis chain why a query matches or misses, and choose between term and match per field. You have seen fielddata answer the wrong question at a memory cost, and a dynamic mapping hit its field limit.

This chapter used the lab’s small name vocabulary. Memory figures for fielddata here are illustrations of the mechanism, not estimates for your data.

Chapter 07 builds the name search users actually type into: autocomplete compared across edge_ngram, search_as_you_type, and the completion suggester; fuzzy matching for typos; synonyms updated without reindexing; and the _explain API for understanding why one profile ranks above another.

Elasticsearch

Type to search the site.

↑↓ navigate⏎ openPowered by Pagefind