Series overview
Part 5 of 1828% complete
2026-05-11•11 min read

Designing the user-profile index

This chapter creates the index the rest of the series searches. By the end you have user-profile-v1 behind two aliases, user-profile-read and user-profile-write, with a strict mapping derived field by field from chapter 01’s query shapes, and the lab’s million profiles loaded into it and checked against PostgreSQL.

The chapter makes the decisions: which fields exist, what each one must support, and how the index is named and addressed. Chapter 06 explains the analysis machinery behind the text fields, and chapter 07 tunes autocomplete and name relevance. You need the lab from chapter 01 with chapter 04’s row_version column applied.

Why the index name has a version

What it is. The physical index is called user-profile-v1. Applications never use that name. They use two aliases, which are names that point at one or more indices: user-profile-read for searches and user-profile-write for the indexer.

Why it matters at 300 million profiles. Chapter 02 showed that primary shard count and existing field types cannot change in place. Sooner or later you need a new mapping or a different shard count, and that means building user-profile-v2 next to v1, filling it, and switching traffic. With aliases, the switch is one atomic _aliases request, and neither the Search API nor the indexer is redeployed. Without them, every application that hard-codes user-profile must change at the same moment, and there is no clean way back.

Two aliases instead of one is a deliberate choice. During a reindex (chapter 14), the indexer must write to the new index while searches still read the old one. Separate aliases let you move them at different times. They also give you separate security boundaries: chapter 16 grants the Search API read access to user-profile-read only.

Example. The index creation request in Stage 2 declares both aliases. is_write_index: true marks the index that receives writes when the write alias points to more than one index.

Common mistake. Creating the index as user-profile “for now”. An alias cannot have the same name as an existing index, so the first reindex turns into a rename under pressure.

Decide what every field must do

Before writing a mapping, list what each field has to support. Chapter 01’s query shapes Q1–Q7 decide it. Everything the mapping does follows from this table.

FieldFull-text searchExact filter or lookupSortAggregateReturnedMapping
userIdYes (Q1)Tiebreaker (Q6)Yeskeyword
fullNameYes (Q3)Yes, alphabeticalYestext + .prefix + .keyword
firstName, lastNameYes, field-specificYestext
emailYes (Q1), case-insensitiveOn exact lookup onlykeyword, lowercase normalizer
mobileNumberYes (Q1)On exact lookup onlykeyword
city, stateYes, as part of a search boxYes (Q4), case-insensitiveYes (chapter 12)Yeskeyword, normalizer, + .text
countryYesYesYeskeyword
pincodeYes (Q4), and prefixYesYeskeyword
genderYes (Q4)YesYeskeyword
accountStatusYes (Q4)YesYeskeyword
createdAtRangeYesDate histogramYesdate
updatedAtRangeYes (Q5)Date histogramYesdate
dateOfBirthYesdate, not indexed, no doc values

Three decisions in this table are worth stating explicitly.

Identifiers are keywords, even when they look like numbers. userId, mobileNumber, and pincode are never compared as numbers, summed, or searched by range. As keyword, they are matched exactly and support prefix queries, such as all pincodes starting 8000. Numeric types are optimised for range queries, which these fields never need.

dateOfBirth is stored, but not searchable. No query shape filters or sorts by date of birth. Mapped with "index": false and "doc_values": false, it is kept in _source for display and costs nothing in the inverted index or on disk for sorting. If a later requirement needs age filters, that is a mapping change and a reindex, a cost you should pay only when there is a real need. The stronger option, leaving it out of the projection altogether, is often right for personal data. Chapter 16 returns to data minimisation.

Nothing else exists. dynamic: strict rejects any document that carries a field not in this table. A field added upstream in PostgreSQL does not quietly start appearing in search results or growing the mapping. Someone has to decide to add it.

Stage 1 — Create the name synonyms set

The name analysers refer to a synonyms set stored in the cluster, so it must exist before the index. Synonyms let a search for Mohd find Mohammed. Chapter 07 explains how they work and how to change them without reindexing.

PUT _synonyms/profile-name-synonyms
{
"synonyms_set": [
{ "id": "mohammed", "synonyms": "mohammed, mohammad, muhammad, mohd" },
{ "id": "lakshmi", "synonyms": "lakshmi, laxmi" }
]
}
{"result":"created","reload_analyzers_details":{"_shards":{"total":0,"successful":0,"failed":0},"reload_details":[]}}

Stage 2 — Write the index definition

Keep the definition in a file under version control, next to the code that depends on it. Chapter 10’s Spring Boot service loads this same file, so the index in production is created from exactly what was reviewed.

Create es/user-profile-v1.json in the lab directory:

es/user-profile-v1.json
{
"settings": {
"index": {
"number_of_shards": 1,
"number_of_replicas": 0,
"refresh_interval": "1s"
},
"analysis": {
"filter": {
"name_edge_ngram": {
"type": "edge_ngram",
"min_gram": 2,
"max_gram": 15
},
"name_synonyms": {
"type": "synonym_graph",
"synonyms_set": "profile-name-synonyms",
"updateable": true
}
},
"analyzer": {
"name_index": {
"type": "custom",
"tokenizer": "standard",
"filter": ["lowercase", "asciifolding"]
},
"name_search": {
"type": "custom",
"tokenizer": "standard",
"filter": ["lowercase", "asciifolding", "name_synonyms"]
},
"name_prefix_index": {
"type": "custom",
"tokenizer": "standard",
"filter": ["lowercase", "asciifolding", "name_edge_ngram"]
}
},
"normalizer": {
"lowercase_ascii": {
"type": "custom",
"filter": ["lowercase", "asciifolding"]
},
"lowercase_only": {
"type": "custom",
"filter": ["lowercase"]
}
}
}
},
"mappings": {
"dynamic": "strict",
"properties": {
"userId": { "type": "keyword" },
"fullName": {
"type": "text",
"analyzer": "name_index",
"search_analyzer": "name_search",
"fields": {
"prefix": { "type": "text", "analyzer": "name_prefix_index", "search_analyzer": "name_index" },
"keyword": { "type": "keyword", "normalizer": "lowercase_ascii", "ignore_above": 256 }
}
},
"firstName": { "type": "text", "analyzer": "name_index", "search_analyzer": "name_search" },
"lastName": { "type": "text", "analyzer": "name_index", "search_analyzer": "name_search" },
"email": { "type": "keyword", "normalizer": "lowercase_only" },
"mobileNumber": { "type": "keyword" },
"city": {
"type": "keyword",
"normalizer": "lowercase_ascii",
"fields": { "text": { "type": "text", "analyzer": "name_index" } }
},
"state": {
"type": "keyword",
"normalizer": "lowercase_ascii",
"fields": { "text": { "type": "text", "analyzer": "name_index" } }
},
"country": { "type": "keyword" },
"pincode": { "type": "keyword" },
"dateOfBirth": { "type": "date", "index": false, "doc_values": false },
"gender": { "type": "keyword" },
"accountStatus": { "type": "keyword" },
"createdAt": { "type": "date" },
"updatedAt": { "type": "date" }
}
},
"aliases": {
"user-profile-read": {},
"user-profile-write": { "is_write_index": true }
}
}

The decisions in this file, grouped by what they protect:

Settings are lab values, and say so. One primary and zero replicas fit a single-node laptop. They are not a recommendation for 300 million profiles; chapter 08 derives production values. refresh_interval is set explicitly to one second, so the search-idle behaviour from chapter 02 does not apply.

Three analysers for one name. name_index splits a name into words, lowercases them, and folds accents, so José is stored as jose. name_search does the same to the query and then expands synonyms. name_prefix_index also emits leading fragments of each word, pr, pra, pras, and so on, for autocomplete. Chapter 06 shows each one working with the _analyze API.

Synonyms apply at search time only. The name_synonyms filter is marked updateable, which Elasticsearch allows only in search analysers. The payoff is that changing the synonyms set takes effect without reindexing 300 million documents. Chapter 07 demonstrates it.

fullName is one value indexed three ways. A multi-field indexes the same source value into several sub-fields: fullName for word matching, fullName.prefix for autocomplete, and fullName.keyword for sorting and exact comparisons. The document still contains fullName once.

Normalisers make exact filters forgiving where users type. A normaliser is an analyser for keyword fields that produces a single token. city and state use lowercase_ascii, so a filter for patna or PATNA matches Patna. email uses lowercase_only, because email addresses are case-insensitive in practice, but folding accents in them would change the address. accountStatus and gender have no normaliser: their values come from the application’s enums, and an exact, case-sensitive match catches a client sending the wrong constant.

city and state are keywords with a text sub-field. Filtering needs the exact value, and the search box needs word matching and highlighting. The main field serves the more common use, filtering.

ignore_above: 256 on fullName.keyword. A value longer than 256 characters is not indexed in that sub-field, rather than producing an oversized term. Such a name is almost certainly bad data. See the ignore_above reference.

Stage 3 — Create the index and check the aliases

Create the index from the file. From a shell in the lab directory:

Terminal window
set -a; source .env; set +a
curl -s -u "elastic:$ELASTIC_PASSWORD" -X PUT "localhost:9200/user-profile-v1" \
-H 'Content-Type: application/json' --data-binary @es/user-profile-v1.json
{"acknowledged":true,"shards_acknowledged":true,"index":"user-profile-v1"}

Then check where the aliases point:

GET _cat/aliases/user-profile-*?v&h=alias,index,is_write_index
alias index is_write_index
user-profile-read user-profile-v1 -
user-profile-write user-profile-v1 true

Stage 4 — Load the lab’s profiles

The real ingestion path is the outbox indexer and the backfill job, both built in Kotlin in chapter 14. Until then, a short shell loader fills the index so chapters 06 to 13 have data to search. It exports every profile from PostgreSQL as bulk NDJSON, with row_version as the external version, and sends it in chunks.

Create es/export-profiles.sql:

es/export-profiles.sql
-- One bulk action line and one document line per profile, as NDJSON.
-- The external version is the row version, so re-running the load is idempotent.
SELECT json_build_object(
'index', json_build_object('_id', user_id::text,
'version', row_version,
'version_type', 'external'))::text
|| E'\n' ||
json_build_object(
'userId', user_id::text,
'fullName', full_name,
'firstName', first_name,
'lastName', last_name,
'email', email,
'mobileNumber', mobile_number,
'city', city,
'state', state,
'country', country,
'pincode', pincode,
'dateOfBirth', date_of_birth,
'gender', gender,
'accountStatus', account_status,
'createdAt', created_at,
'updatedAt', updated_at)::text
FROM user_profile
ORDER BY user_id;

Create es/load-v1.sh:

es/load-v1.sh
#!/usr/bin/env bash
# Throwaway loader for the lab. Chapter 14 replaces it with a Kotlin backfill job.
set -euo pipefail
set -a; source .env; set +a
ES="http://localhost:9200"
AUTH="elastic:${ELASTIC_PASSWORD}"
mkdir -p data
rm -f data/chunk-*
echo "Exporting profiles from PostgreSQL..."
docker compose exec -T postgres psql -U profiles -d profiles -At -f - \
< es/export-profiles.sql > data/profiles.ndjson
split -l 20000 data/profiles.ndjson data/chunk-
echo "Disabling refresh for the load..."
curl -s -u "$AUTH" -X PUT "$ES/user-profile-write/_settings" \
-H 'Content-Type: application/json' -d '{"index":{"refresh_interval":"-1"}}' > /dev/null
failed=0
for chunk in data/chunk-*; do
response=$(curl -s -u "$AUTH" -X POST "$ES/user-profile-write/_bulk?filter_path=errors,items.*.status" \
-H 'Content-Type: application/x-ndjson' --data-binary "@$chunk")
if [[ "$response" == *'"errors":true'* ]]; then
# 409 means Elasticsearch already holds this version or a newer one: not a failure.
bad=$(grep -o '"status":[0-9]*' <<< "$response" | grep -Evc '"status":(200|201|409)' || true)
if (( bad > 0 )); then
echo "$chunk: $bad failed items"
failed=$((failed + bad))
fi
fi
done
echo "Restoring refresh and refreshing..."
curl -s -u "$AUTH" -X PUT "$ES/user-profile-write/_settings" \
-H 'Content-Type: application/json' -d '{"index":{"refresh_interval":"1s"}}' > /dev/null
curl -s -u "$AUTH" -X POST "$ES/user-profile-write/_refresh" > /dev/null
echo "Done. Failed items: $failed"

The script applies three rules from earlier chapters. It writes through the user-profile-write alias, never the index name. It turns off refresh for the duration of the load and restores it at the end, as the indexing-speed guidance cited in chapter 02 recommends. And it classifies bulk items with chapter 03’s table, so a 409 is not a failure.

Run it:

Terminal window
chmod +x es/load-v1.sh
./es/load-v1.sh
Exporting profiles from PostgreSQL...
Disabling refresh for the load...
Restoring refresh and refreshing...
Done. Failed items: 0

In the lab this took about 45 seconds. Run it a second time: every item comes back 409, because every document already has that version, and the script still reports Failed items: 0. That is the idempotency chapter 04 designed for.

Security note — data/profiles.ndjson is a full export of the profile table, about 470 MB for the lab. For synthetic data that is harmless. For real data, an export like this is exactly the kind of file that leaks: keep it out of version control, and never produce one from production outside a controlled process.

Checkpoint: the index matches PostgreSQL

GET user-profile-read/_count?filter_path=count
{"count":1000000}
Terminal window
docker compose exec postgres psql -U profiles -d profiles -At -c "SELECT count(*) FROM user_profile;"
1000000

The counts match: chapter 04 deleted one profile and inserted one. Check those two changes arrived:

GET user-profile-read/_doc/42?filter_path=_version,_source.city,_source.accountStatus
{"_version":4,"_source":{"city":"Gaya","accountStatus":"SUSPENDED"}}

GET user-profile-read/_doc/43 returns "found":false. The Elasticsearch _version equals PostgreSQL’s row_version, which is what makes later writes from the indexer safe.

GET _cat/indices/user-profile-v1?v&h=index,docs.count,docs.deleted,store.size
index docs.count docs.deleted store.size
user-profile-v1 1000000 0 157.6mb

Chapter 08 uses that store size, measured on your own data, as the starting point for capacity planning.

Stage 5 — Check the matrix behaves as designed

A mapping is a set of promises. Test the important ones now, while they are cheap to fix.

Email lookup ignores case. The normaliser is applied to the query value too:

GET user-profile-read/_search?filter_path=hits.total,hits.hits._source
{
"query": { "term": { "email": "PRIYA.KUMAR.1@EXAMPLE.COM" } },
"_source": ["userId", "fullName", "email"]
}
{"hits":{"total":{"value":1,"relation":"eq"},
"hits":[{"_source":{"userId":"1","fullName":"Priya Kumar","email":"priya.kumar.1@example.com"}}]}}

City filters ignore case; status filters do not.

GET user-profile-read/_count?filter_path=count
{ "query": { "bool": { "filter": [
{ "term": { "city": "patna" } },
{ "term": { "accountStatus": "ACTIVE" } }
] } } }
{"count":70214}

That is exactly SELECT count(*) FROM user_profile WHERE city = 'Patna' AND account_status = 'ACTIVE' in PostgreSQL. The same query with "accountStatus": "active" returns 0, as intended.

Unknown fields are rejected.

PUT user-profile-write/_doc/999
{ "userId": "999", "fullName": "Test", "nickname": "x" }
strict_dynamic_mapping_exception: [1:46] mapping set to strict, dynamic introduction of [nickname] within [_doc] is not allowed

Text fields refuse to aggregate. Try to count profiles by fullName:

GET user-profile-read/_search
{ "size": 0, "aggs": { "names": { "terms": { "field": "fullName" } } } }
illegal_argument_exception: Fielddata is disabled on [fullName] in [user-profile-v1]. Text fields are not
optimised for operations that require per-document field data like aggregations and sorting, so these
operations are disabled by default. Please use a keyword field instead. ...

That is the protection working. Use fullName.keyword for this. Chapter 06 explains why turning on fielddata instead is dangerous.

A returned-only field does not always fail loudly. A term query on dateOfBirth fails:

query_shard_exception: failed to create query: Cannot search on field [dateOfBirth] since it is not indexed nor has doc values.

But a range query on the same field returns zero hits and no error:

GET user-profile-read/_search?size=0
{ "query": { "range": { "dateOfBirth": { "gte": "1990-01-01" } } } }
{"took":0,"timed_out":false,"_shards":{"total":1,"successful":1,"skipped":0,"failed":0},
"hits":{"total":{"value":0,"relation":"eq"},"max_score":null,"hits":[]}}

A caller who builds a date-of-birth filter against this index gets an empty result, not an error. The protection against that is the Search API’s request contract (chapter 10), which exposes only the filters the matrix allows.

When index templates help, and why this series does not use one

An index template applies settings, mappings, and aliases automatically to any new index whose name matches a pattern. Composable templates are built from reusable component templates. You could put this chapter’s settings and mappings into a component template, and have an index template apply it to every user-profile-v* index:

PUT _component_template/user-profile-schema
{ "template": { "settings": { ... }, "mappings": { ... } } }
PUT _index_template/user-profile
{ "index_patterns": ["user-profile-v*"], "priority": 100, "composed_of": ["user-profile-schema"] }

The simulate API shows what a future index would receive:

POST _index_template/_simulate_index/user-profile-v2?filter_path=template.settings.index.number_of_shards,template.mappings.dynamic
{"template":{"settings":{"index":{"number_of_shards":"1"}},"mappings":{"dynamic":"strict"}}}

Trade-off. Templates earn their place when indices are created automatically and often, as with time-based data streams for logs. A search projection gets a new index rarely, deliberately, and usually because the mapping changed. A file per version, user-profile-v1.json and later user-profile-v2.json, makes each version’s definition reviewable and diffable. A template, by contrast, is shared state in the cluster that silently applies to the next matching index, including one created by a typo. This series uses files. If you tried the template requests, delete both before continuing:

DELETE _index_template/user-profile
DELETE _component_template/user-profile-schema

Common mistakes with index design

  • Designing the mapping from the table instead of from the queries. Mirroring every PostgreSQL column produces a large index full of fields nobody queries, and puts more personal data at risk.
  • Mapping identifiers as numbers. Long IDs and phone numbers as long lose leading zeros and prefix queries, and gain nothing.
  • One alias for reading and writing. It works until the first reindex, when the indexer and the search traffic need to point at different indices.
  • Leaving dynamic at its default. The first unexpected field from upstream changes the mapping permanently, as chapter 02 showed.

What you built, and what comes next

You have user-profile-v1 holding one million profiles, reachable only through user-profile-read and user-profile-write, with a strict mapping in which every field exists because a query shape needs it. You checked the count against PostgreSQL and tested each promise of the capability matrix.

The settings are for a laptop. Shard and replica counts for 300 million profiles are chapter 08’s subject, and they need measurements this chapter cannot provide.

Chapter 06 opens up the text fields: how analysers, tokenisers, token filters, character filters, and normalisers turn José Fernandes into searchable terms, and why fielddata and uncontrolled mappings are the two fastest ways to exhaust a cluster’s memory.

ElasticsearchPostgres

Type to search the site.

↑↓ navigate⏎ openPowered by Pagefind