Series overview
Part 7 of 1839% complete
2026-05-14•10 min read

Autocomplete, typos, synonyms, and name relevance

Users type names into a search box, one letter at a time, with typos, in spellings that differ from the stored record. This chapter makes that work on user-profile-v1. You build three scratch indices over the same million names to compare Elasticsearch’s autocomplete options by size and behaviour, confirm the choice made in chapter 05, add typo tolerance, change the synonym list on a live index without reindexing, and read a relevance score line by line with the _explain API.

You need the lab with user-profile-v1 loaded (chapter 05) and the analysis concepts from chapter 06. The chapter takes about 45 minutes. The scratch indices are deleted at the end.

Three ways to autocomplete, built side by side

Elasticsearch offers three mechanisms for prefix search. They differ in what they index, what queries they support, and whether they can be combined with filters.

MechanismWhat it indexesHow you query it
edge_ngram on a text sub-fieldEvery leading fragment of every word: pr, pra, … prashantAn ordinary match query
search_as_you_type fieldWords, word pairs and triples (shingles), and their prefixes, in generated sub-fieldsA multi_match query of type bool_prefix
completion fieldThe whole value in an in-memory prefix structure (a finite-state transducer)The completion suggester, not a query

Why it matters at 300 million profiles. Autocomplete runs on every keystroke, so it is usually the highest-rate query the Search API serves, and its index structures are among the largest in the mapping. Choosing wrongly is expensive twice: once in disk and heap, and again in a reindex to change it.

Example. Create four scratch indices. They share a small analysis block, a plain name analyser n and a prefix analyser np, and differ only in how fullName is mapped. The baseline has no autocomplete structure at all, so it shows what each option adds.

PUT ch07-baseline
{
"settings": {
"number_of_shards": 1, "number_of_replicas": 0,
"analysis": {
"filter": { "eg": { "type": "edge_ngram", "min_gram": 2, "max_gram": 15 } },
"analyzer": {
"n": { "type": "custom", "tokenizer": "standard", "filter": ["lowercase", "asciifolding"] },
"np": { "type": "custom", "tokenizer": "standard", "filter": ["lowercase", "asciifolding", "eg"] }
}
}
},
"mappings": { "properties": {
"fullName": { "type": "text", "analyzer": "n" },
"state": { "type": "keyword" }
} }
}

Create ch07-edge, ch07-sayt, and ch07-completion with the same settings block, changing only the fullName mapping:

fullName in ch07-edge
{ "type": "text", "analyzer": "n",
"fields": { "prefix": { "type": "text", "analyzer": "np", "search_analyzer": "n" } } }
fullName in ch07-sayt
{ "type": "search_as_you_type", "analyzer": "n" }
fullName in ch07-completion
{ "type": "text", "analyzer": "n",
"fields": { "suggest": { "type": "completion", "analyzer": "n" } } }

Copy the names and states into each index, then merge each to one segment so that sizes are comparable:

POST _reindex?refresh=true&filter_path=created,failures
{ "source": { "index": "user-profile-read", "_source": ["fullName", "state"] },
"dest": { "index": "ch07-baseline" } }
POST ch07-baseline/_forcemerge?max_num_segments=1

Repeat both requests for ch07-edge, ch07-sayt, and ch07-completion. Each reindex reports {"created":1000000,"failures":[]}. Then compare:

GET _cat/indices/ch07-*?v&h=index,docs.count,store.size&s=index
index docs.count store.size
ch07-baseline 1000000 24.2mb
ch07-completion 1000000 84.5mb
ch07-edge 1000000 32.2mb
ch07-sayt 1000000 47mb

These are lab measurements on one million short synthetic names, each index merged to one segment, _source included. Over the baseline, edge_ngram added about 8 MB, search_as_you_type about 23 MB, and completion about 60 MB. Needs validation: absolute sizes depend on name length and diversity, so measure your own data as chapter 08 describes. The order is what these numbers show, not the ratios at scale.

How each option behaves

Run the same user intentions against each index. A user in Bihar has typed prashant ku.

edge_ngram, with the query analysed by the plain analyser and every word required:

GET ch07-edge/_count?filter_path=count
{ "query": { "bool": {
"must": { "match": { "fullName.prefix": { "query": "prashant ku", "operator": "and" } } },
"filter": { "term": { "state": "Bihar" } }
} } }
{"count":345}

search_as_you_type, with a bool_prefix query across the generated sub-fields:

GET ch07-sayt/_count?filter_path=count
{ "query": { "bool": {
"must": { "multi_match": {
"query": "prashant ku", "type": "bool_prefix", "operator": "and",
"fields": ["fullName", "fullName._2gram", "fullName._3gram"]
} },
"filter": { "term": { "state": "Bihar" } }
} } }
{"count":345}

Both agree with PostgreSQL, where first_name = 'Prashant' AND last_name LIKE 'Ku%' AND state = 'Bihar' counts 345. Without "operator": "and", the bool_prefix query returns 16,378 profiles, because any word may match. That is fine for a ranked list, and misleading for a count.

completion, through the suggest API:

GET ch07-completion/_search?filter_path=suggest.names.options.text
{ "suggest": { "names": {
"prefix": "prashant k",
"completion": { "field": "fullName.suggest", "size": 5, "skip_duplicates": true }
} } }
{"suggest":{"names":[{"options":[{"text":"Prashant Khan"},{"text":"Prashant Kumar"}]}]}}

The suggester returns distinct values, not documents. That makes it good for “which names exist?” and useless for “which profiles match?”. Now try the two things a profile search needs. First, a match at the start of the surname:

GET ch07-completion/_search?filter_path=suggest.names.options.text
{ "suggest": { "names": { "prefix": "kum", "completion": { "field": "fullName.suggest", "size": 5 } } } }
{}

No suggestions: the completion structure matches from the start of the whole value only. Both ch07-edge and ch07-sayt return 50,009 profiles for kum, every Kumar. Second, a filter. Send a query that matches nothing, alongside the suggester:

GET ch07-completion/_search?filter_path=hits.total,suggest.names.options.text
{
"query": { "term": { "state": "Atlantis" } },
"suggest": { "names": { "prefix": "pras", "completion": { "field": "fullName.suggest", "size": 3, "skip_duplicates": true } } }
}
{"hits":{"total":{"value":0,"relation":"eq"}},
"suggest":{"names":[{"options":[{"text":"Prashant Chatterjee"},{"text":"Prashant Das"},{"text":"Prashant Fernandes"}]}]}}

Zero hits, and suggestions anyway. The suggester ignores the query entirely. The only built-in way to restrict it is context fields declared in the mapping, which must be chosen in advance and cannot express arbitrary combinations of state, city, status, and gender.

The recommendation for this index

edge_ngram sub-fieldsearch_as_you_typecompletion
Combines with any filterYesYesNo; fixed contexts only
Matches the start of any wordYesYesNo; start of value only
Returns documents, with highlightingYesYesNo; returns values
Extra index size in the labSmallestMediumLargest
Scores phrase order (prashant ku above ku… prashant)NoYes, through shinglesNot applicable
QueryOrdinary matchmulti_match of type bool_prefixSuggest API

Recommendation. Use an edge_ngram sub-field, which is what user-profile-v1 already has as fullName.prefix. The Search API filters every autocomplete request by at least accountStatus, and often by state or city, so the completion suggester is ruled out. Between the other two, edge_ngram was smaller in the lab, needs only a plain match query, and gives complete control over the analysis chain, including the accent folding from chapter 06.

Trade-off. search_as_you_type is a credible alternative. It requires no custom analyser, and its shingle sub-fields rank matches whose word order follows the input higher, which helps with longer, multi-word values. For two- and three-word personal names, that advantage is small. The completion suggester is the right tool for a different feature: a global, unfiltered “names in our system” suggestion list where memory-resident speed matters and filters do not.

Common mistake. Setting min_gram to 1. Every name then produces one-letter terms, and a single typed character matches a large share of the index. With min_gram: 2, autocomplete starts on the second keystroke, which users rarely notice.

Autocomplete on the real index

The same query works on user-profile-v1, with the status filter every Search API request will carry and highlighting to show what matched:

GET user-profile-read/_search?filter_path=hits.total,hits.hits._source,hits.hits.highlight
{
"size": 2,
"_source": ["userId", "fullName", "city"],
"query": { "bool": {
"must": { "match": { "fullName.prefix": { "query": "prashant ku", "operator": "and" } } },
"filter": [ { "term": { "state": "bihar" } }, { "term": { "accountStatus": "ACTIVE" } } ]
} },
"highlight": { "fields": { "fullName.prefix": {} } }
}
{"hits":{"total":{"value":237,"relation":"eq"},"hits":[
{"_source":{"userId":"6000","fullName":"Prashant Kumar","city":"Gaya"},
"highlight":{"fullName.prefix":["<em>Prashant</em> <em>Kumar</em>"]}},
{"_source":{"userId":"11400","fullName":"Prashant Kumar","city":"Gaya"},
"highlight":{"fullName.prefix":["<em>Prashant</em> <em>Kumar</em>"]}}]}}

237 active profiles, matching PostgreSQL. The state filter used bihar in lowercase, and chapter 05’s normaliser made it match Bihar. The whole word is highlighted, not just the typed ku, because the edge n-grams share the word’s start and end offsets. If you need to highlight only the typed characters, the client can compute that from the input; the highlighter cannot.

Clean up the scratch indices:

DELETE ch07-baseline,ch07-edge,ch07-sayt,ch07-completion

Typo tolerance with fuzziness

What it is. A match query with fuzziness also accepts terms within a small edit distance of each query term: one or two insertions, deletions, substitutions, or transpositions. "fuzziness": "AUTO" allows 0 edits for terms of 1–2 characters, 1 edit for 3–5, and 2 for longer terms, as the match query reference describes.

Why it matters at 300 million profiles. Support agents search while on the phone, and they mistype. Fuzziness recovers those searches, but each fuzzy term expands into every matching variant in the index, so its cost grows with the size of the name vocabulary, which is large at 300 million real profiles.

Example. prashnat swaps two letters:

Query on fullNameCount
{"query": "prashnat", "operator": "and"}0
{"query": "prashnat", "fuzziness": "AUTO", "operator": "and"}33,333
{"query": "prashnat kumr", "fuzziness": "AUTO", "operator": "and"}1,666

The last row finds exactly the Prashant Kumars despite two typos. Short names are protected by AUTO: rao allows one edit and still counts 49,980 profiles, the same as the exact query, because no other surname in the data is one edit away.

Trade-off. Fuzzy matches score like exact ones unless you separate them. Chapter 11 puts an exact match and a fuzzy match in the same bool query, so exact spellings rank first and typo matches follow. For large vocabularies, prefix_length, the number of leading characters that must match exactly, and max_expansions cap the cost.

Common mistake. Adding fuzziness to the fullName.prefix field. Every edge n-gram becomes a fuzzy term, so pr with one edit matches almost every two-letter fragment in the index. Keep autocomplete exact and apply fuzziness to the whole-word field only.

Synonyms you can change on a live index

What it is. A synonyms set is stored in the cluster and referenced by a synonym_graph token filter. Because chapter 05 used it only in the search analyser and marked it updateable, changing a rule reloads the analyser immediately, with no reindex.

Why it matters at 300 million profiles. Name variants surface continuously from support tickets: Fatima and Fathima, Lakshmi and Laxmi. If adding one required rebuilding a 300-million-document index, the list would never be maintained.

Example. The set currently has no rule for Fathima:

GET user-profile-read/_count?filter_path=count
{ "query": { "match": { "fullName": "fathima" } } }
{"count":0}

Add one rule with the synonym rule API:

PUT _synonyms/profile-name-synonyms/fatima
{ "synonyms": "fatima, fathima" }
{"result":"created","reload_analyzers_details":{"_shards":{"total":2,"successful":2,"failed":0},
"reload_details":[{"index":"user-profile-v1","reloaded_analyzers":["name_search"],"reloaded_node_ids":["..."]}]}}

The response confirms that name_search was reloaded on user-profile-v1. _analyze with name_search on Fathima now returns fatima and fathima.

The request-cache trap

Now repeat the exact same count request:

{"count":0}

Still zero. Change the request text in any way, for example FATHIMA instead of fathima, and the count is 33,333. The first request was answered from the shard request cache, which caches size: 0 results keyed by a hash of the request body. Elasticsearch invalidates that cache when a refresh picks up changed documents or when the mapping changes. A synonym reload is neither, so the cached answer from before the change survived.

Clear the request cache after every synonym change:

POST user-profile-v1/_cache/clear?request=true

After that, the original request returns {"count":33333}. Search requests that return hits (size greater than 0) are not cached by default, so they were correct immediately. Counts and aggregations are the ones at risk. This behaviour was observed in the Elasticsearch 9.5.4 lab; treat the cache clear as part of the synonym-change procedure.

Common mistake. Putting synonyms in the index analyser. Every rule change then needs a reindex, and the expanded terms inflate the index. Keep synonyms at search time unless you have measured a reason not to.

Why one profile outranks another: BM25 and _explain

What it is. Elasticsearch scores text matches with BM25 by default. For each query term that matches, the score grows with how rare the term is across the index (inverse document frequency, IDF) and with how often it appears in this document’s field, adjusted for field length. The explain API shows the full calculation for one document.

Why it matters at 300 million profiles. Relevance is corpus-wide. Chapter 01 quoted PostgreSQL’s documentation that its ranking functions do not use global information. Elasticsearch’s does, and that is visible in every explanation.

Example. Search for prashant kumar. The top profiles are named Prashant Kumar with a score of 6.3967; profiles named Prashant Jha score 3.4012. Ask why, for profile 600, a Prashant Kumar:

GET user-profile-read/_explain/600
{ "query": { "match": { "fullName": "prashant kumar" } } }
6.396736 sum of:
3.4011931 weight(fullName:prashant), boost * idf * tf from:
2.2 boost
3.4011934 idf = log(1 + (N - n + 0.5) / (n + 0.5)), with n = 33333, N = 1000000
0.45454544 tf = freq / (freq + k1 * (1 - b + b * dl / avgdl)), with freq = 1, k1 = 1.2, b = 0.75, dl = 2, avgdl = 2
2.995543 weight(fullName:kumar), boost * idf * tf from:
2.2 boost
2.9955432 idf, with n = 50009, N = 1000000
0.45454544 tf, as above

The response is condensed here; it nests each line as a JSON object with value, description, and details. Three things are visible:

  • The score is a sum over matched terms. A Prashant Jha matches only prashant, so it scores 3.4012, the first half.
  • Rarer terms weigh more. prashant appears in 33,333 profiles and kumar in 50,009, so prashant contributes more. On real data, a rare surname in the query counts for much more than a common one.
  • Field length matters. dl / avgdl compares this name’s length in terms with the average. A longer name that contains the same words scores a little lower. The 2.2 boost is BM25’s k1 + 1, folded in by Lucene.

Common mistake. Treating identical scores as a ranking. All 1,666 Prashant Kumars score exactly 6.396736, so their order is arbitrary unless the query adds a tiebreaker sort. Chapter 11 adds updatedAt and userId after the score, and chapter 13 depends on that for stable pagination.

What you built, and what comes next

You compared the three autocomplete mechanisms on the same million names, by size and by behaviour, and confirmed fullName.prefix with edge n-grams as the right choice for a filtered profile search. You added typo tolerance, added a synonym rule to a live index, and found that counts need a request-cache clear after synonym changes. You can read a BM25 score line by line.

The sizes in this chapter come from short synthetic names on one node. They show the relative cost of each option, not what your index will need.

Chapter 08 turns measurements like these into a capacity plan: shard counts, replicas, refresh and translog settings, and a sizing method for 300 million profiles, built from sample data and validated by load tests instead of by rules of thumb.

Elasticsearch

Type to search the site.

↑↓ navigate⏎ openPowered by Pagefind