Series overview
Part 11 of 1861% complete
2026-05-21•11 min read

Query design for profile search

This chapter designs the query behind the Search API and replaces chapter 10’s placeholder. By the end, UserQueryBuilder produces one request shape for every search: a single scored clause for the name, which tolerates typos and combines name, city, and state; every other constraint as a cached, unscored filter; a sort that is always total; highlighting; and a source filter. You build each piece in Kibana Dev Tools first, check its counts against PostgreSQL, and then write it in Kotlin.

The chapter’s centre is the request chapter 01 used as its example: find active users named Prashant in Bihar, updated recently. You write it as JSON and as Kotlin, and confirm that both produce the same request and the same seven profiles.

You need the lab over TLS and the search-api module from chapter 10. Requests use Kibana Dev Tools syntax. The chapter takes about 60 minutes.

Query context and filter context

What it is. Every clause in a search runs in one of two contexts. In query context, a clause answers “how well does this document match?” and contributes to the relevance score. In filter context, it answers only “does it match?”, yes or no. Filter context is used by filter and must_not clauses in a bool query. Elastic’s query and filter context guide states that frequently used filters are cached automatically.

Why it matters at 300 million profiles. Most constraints in profile search are exact: status, state, city, gender, pincode, a date range. Scoring them is wasted work on every matching document, and scored clauses cannot use the filter cache. Keeping exact constraints in filter context makes them cheaper and reusable across requests.

Example. The same term on state, first in must, then in filter:

GET user-profile-read/_search?filter_path=hits.total,hits.hits._score&size=2
{ "query": { "bool": { "must": [ { "term": { "state": "bihar" } } ] } } }
{"hits":{"total":{"value":10000,"relation":"gte"},"hits":[{"_score":1.6106621},{"_score":1.6106621}]}}
GET user-profile-read/_search?filter_path=hits.total,hits.hits._score&size=2
{ "query": { "bool": { "filter": [ { "term": { "state": "bihar" } } ] } } }
{"hits":{"total":{"value":10000,"relation":"gte"},"hits":[{"_score":0.0},{"_score":0.0}]}}

Same documents; in must every one received the same computed score of 1.61, which carries no information, because every match has the same state.

Common mistake. Putting every clause in must because “they all have to match”. They do, and filter also means “has to match”. Use must only for the clauses whose degree of match should order the results.

The bool query, clause by clause

A bool query combines clauses with four occurrence types:

ClauseMust match?Scores?Cached?Use for
mustYes, all of themYesNoThe name clause: how well a profile matches the search text
filterYes, all of themNoConsideredStatus, state, city, gender, pincode, date ranges
shouldDepends on minimum_should_matchYesNoAlternatives, and optional clauses that raise the score
must_notMust not match anyNoConsideredExclusions

minimum_should_match is where should becomes subtle. The bool query reference is precise: when a bool has at least one should clause and no must or filter clauses, the default is 1; otherwise it is 0, and should clauses only add score. The name clause in this chapter sets it explicitly, so its meaning does not change when someone adds a filter next to it.

must_not is filter context too. Excluding two statuses in Bihar:

GET user-profile-read/_count?filter_path=count
{ "query": { "bool": {
"filter": [ { "term": { "state": "bihar" } } ],
"must_not": [ { "terms": { "accountStatus": ["DELETED", "SUSPENDED"] } } ]
} } }
{"count":159861}

PostgreSQL’s state = 'Bihar' AND account_status NOT IN ('DELETED', 'SUSPENDED') agrees. The Search API expresses the same intent positively, with accountStatus as a filter, which is easier to read and to cache.

Exact lookups use term queries

What it is. A term query matches one exact value in a keyword field, after the field’s normaliser. Chapter 10’s lookups by email and mobile number use it, and a lookup by userId uses get.

Why it matters at 300 million profiles. An exact lookup should cost the same at 300 million profiles as at one million: a direct term lookup, not a scan. It also must never match a different profile, which rules out analysed fields.

Example. The prefix query is the exact-value relative of term for keyword fields. All profiles in pincodes starting 8000:

GET user-profile-read/_count?filter_path=count
{ "query": { "prefix": { "pincode": "8000" } } }
{"count":100109}

That matches PostgreSQL’s pincode LIKE '8000%' exactly.

Common mistake. Sending exact lookups through the full-text search endpoint. A match query for an email address analyses it, and a mistyped address may still match something. Exact lookups get their own endpoints and term queries, as chapter 10 built.

Name search across name, city, and state

What it is. A multi_match query runs one text against several fields. Its cross_fields type treats those fields as if they were one big field, so each word may match in a different field. That is what a search box needs for input like prashant patna: prashant is in the name, patna is in the city.

Why it matters at 300 million profiles. Users do not say which field a word belongs to. Requiring every word in one field finds nothing for prashant patna; matching any word in any field returns every Prashant and everyone in Patna.

Example. Search the name, city, and state, with the name weighted three times higher:

GET user-profile-read/_count?filter_path=count
{ "query": { "multi_match": {
"query": "prashant patna", "type": "cross_fields", "operator": "and",
"fields": ["fullName^3", "city.text", "state.text"]
} } }
{"count":0}

Zero, although PostgreSQL has 3,343 Prashants in Patna. The validate API with explain=true shows the query Elasticsearch actually built:

GET user-profile-read/_validate/query?explain=true&filter_path=explanations.explanation
{ "query": { "multi_match": {
"query": "prashant patna", "type": "cross_fields", "operator": "and",
"fields": ["fullName^3", "city.text", "state.text"]
} } }
{"explanations":[{"explanation":"((+blended(terms:[state.text:prashant, city.text:prashant]) +blended(terms:[state.text:patna, city.text:patna])) | (+fullName:prashant +fullName:patna)^3.0)"}]}

cross_fields blends only fields that share a search analyser. fullName uses name_search, with synonyms, and city.text and state.text use name_index, so Elasticsearch built two groups and applied operator: and inside each: all words in the name, or all words in the place. Neither group can contain both words.

Force one analyser for the whole query:

GET user-profile-read/_validate/query?explain=true&filter_path=explanations.explanation
{ "query": { "multi_match": {
"query": "prashant patna", "type": "cross_fields", "operator": "and",
"analyzer": "name_search",
"fields": ["fullName^3", "city.text", "state.text"]
} } }
{"explanations":[{"explanation":"+blended(terms:[state.text:prashant, fullName:prashant^3.0, city.text:prashant]) +blended(terms:[state.text:patna, fullName:patna^3.0, city.text:patna])"}]}

Now each word must appear in some field, and the count is 3,343. Because the forced analyser is name_search, synonyms apply across all three fields: mohd patna counts 3,347, exactly PostgreSQL’s Mohammeds in Patna.

Common mistake. Trusting cross_fields without checking the groups. When fields have different analysers, the query silently becomes something else. _validate/query?explain=true is the fastest way to see what a text query really is.

Tolerating typos without letting them win

cross_fields does not support fuzziness, and a support agent will type prashnat kumr. The name clause therefore offers two ways to match, and a profile needs only one:

{ "bool": {
"should": [
{ "multi_match": { "query": "prashnat kumr", "type": "cross_fields", "operator": "and",
"analyzer": "name_search", "fields": ["fullName^3", "city.text", "state.text"] } },
{ "match": { "fullName": { "query": "prashnat kumr", "operator": "and",
"fuzziness": "AUTO", "prefix_length": 1 } } }
],
"minimum_should_match": 1
} }

Run as a search, it finds the 1,666 Prashant Kumars through the fuzzy clause alone, each scoring 5.22. The correctly spelled prashant kumar matches both clauses and scores 25.59. Exact spellings rank above typo matches without any extra rule, because a document that satisfies both alternatives earns both scores. prefix_length: 1 requires the first letter to be right, which keeps fuzzy expansion cheap and avoids matching names that differ in their first letter.

Autocomplete is a separate mode

Chapter 07 chose an edge n-gram sub-field for autocomplete. It gets its own mode rather than being mixed into the full search, because its semantics differ: every word is a prefix, and there is no typo tolerance.

{ "match": { "fullName.prefix": { "query": "pras", "operator": "and" } } }

Scores, sorting, and when to ignore relevance

What it is. Results are sorted by _score unless you say otherwise. Sorting by a field, such as updatedAt, skips score computation for that ordering. In chapter 10’s filter-only search sorted by UPDATED_AT, every hit had "score": null. The sort reference describes the options.

Why it matters at 300 million profiles. Relevance is meaningful only when there is text to match. An admin list of “profiles in Kochi, newest first” should not pay for scoring, and its order should not depend on BM25 statistics that shift as the index changes.

Example. The Search API’s sort is always one of two, and both end with tiebreakers:

sortOrder
RELEVANCE_score descending, updatedAt descending, userId ascending
UPDATED_ATupdatedAt descending, userId ascending

Chapter 07 showed that all 1,666 Prashant Kumars score exactly the same. Without updatedAt and userId after the score, their order would be whatever the shards returned, and it could change between requests. The final userId makes the order total, because it is unique, which chapter 13’s cursor pagination requires.

Common mistake. Adding track_scores or scripting a custom score to rescue a sort that should simply not use relevance. Decide per use case: text search sorts by relevance; lists and exports sort by fields.

Highlighting and source filtering

Highlighting returns the matched parts of each field with markers. Request it for the fields the name clause searches; Elasticsearch returns fragments only for fields that actually matched. See the highlighting reference.

Source filtering limits _source to the listed fields. Chapter 10’s RESULT_FIELDS keeps email, mobile number, and date of birth out of every result list. That is a privacy control, not only an optimisation. See retrieving selected fields.

Security note — Highlighting reads the stored source of the fields it highlights. Highlight only fields that are safe to return. A highlight on email would return the address inside the highlight section even with email excluded from _source.

The composed request: active users named Prashant in Bihar, updated recently

Everything in this chapter combines into one request. “Recently” here means within 90 days. now-90d/d rounds to the start of the day, so the filter value is identical for every request on the same day, and can be cached, instead of changing every millisecond.

GET user-profile-read/_search
{
"size": 3,
"track_total_hits": true,
"_source": ["userId", "fullName", "city", "state", "accountStatus", "updatedAt"],
"query": {
"bool": {
"must": [
{
"bool": {
"should": [
{
"multi_match": {
"query": "prashant",
"type": "cross_fields",
"operator": "and",
"analyzer": "name_search",
"fields": ["fullName^3", "city.text", "state.text"]
}
},
{
"match": {
"fullName": { "query": "prashant", "operator": "and", "fuzziness": "AUTO", "prefix_length": 1 }
}
}
],
"minimum_should_match": 1
}
}
],
"filter": [
{ "term": { "accountStatus": "ACTIVE" } },
{ "term": { "state": "bihar" } },
{ "range": { "updatedAt": { "gte": "now-90d/d" } } }
]
}
},
"sort": [
{ "_score": "desc" },
{ "updatedAt": "desc" },
{ "userId": "asc" }
],
"highlight": {
"fields": { "fullName": {}, "city.text": {}, "state.text": {} }
}
}

The response, condensed to the first hit:

{
"hits": {
"total": { "value": 7, "relation": "eq" },
"hits": [
{
"_id": "857970",
"_score": 13.604773,
"_source": { "userId": "857970", "fullName": "Prashant Jha", "city": "Patna", "state": "Bihar",
"accountStatus": "ACTIVE", "updatedAt": "2026-08-01T00:00:00+00:00" },
"highlight": { "fullName": ["<em>Prashant</em> Jha"] },
"sort": [13.604773, 1785542400000, "857970"]
},
...
]
}
}

Seven profiles, on the day this was run. Because the date window is relative, your count depends on the day you run it. Check it against the source of truth:

SELECT count(*) FROM user_profile
WHERE first_name = 'Prashant' AND state = 'Bihar' AND account_status = 'ACTIVE'
AND updated_at >= date_trunc('day', now() - interval '90 days');

Each hit’s sort array holds its values for the three sort keys. Chapter 13 passes the last hit’s array back as a cursor.

Stage 1 — The query in Kotlin

The request contract gains two fields: a nameMode that selects full search or autocomplete, and updatedWithinDays for relative date windows. Replace SearchDtos.kt:

search-api/src/main/kotlin/in/o612/eng/usersearch/api/web/SearchDtos.kt
package `in`.o612.eng.usersearch.api.web
import jakarta.validation.constraints.Max
import jakarta.validation.constraints.Min
import jakarta.validation.constraints.Pattern
import jakarta.validation.constraints.Size
import java.time.Instant
import java.time.LocalDate
enum class Gender { MALE, FEMALE, OTHER }
enum class AccountStatus { ACTIVE, INACTIVE, SUSPENDED, DELETED }
enum class SortOrder { RELEVANCE, UPDATED_AT }
/** FULL: whole words, typo-tolerant, across name, city, and state. PREFIX: autocomplete on the name. */
enum class NameMode { FULL, PREFIX }
/** Query parameters of GET /api/users/search. Only the filters chapter 05's capability matrix allows exist. */
data class UserSearchRequest(
@field:Size(min = 2, max = 100) val name: String? = null,
val nameMode: NameMode = NameMode.FULL,
@field:Size(max = 60) val city: String? = null,
@field:Size(max = 60) val state: String? = null,
@field:Pattern(regexp = "\\d{6}") val pincode: String? = null,
val gender: Gender? = null,
val accountStatus: AccountStatus = AccountStatus.ACTIVE,
val updatedSince: Instant? = null,
@field:Min(1) @field:Max(3650) val updatedWithinDays: Int? = null,
val sort: SortOrder = SortOrder.RELEVANCE,
@field:Min(1) @field:Max(50) val size: Int = 20,
)
data class UserSearchResponse(
val total: TotalHits,
val hits: List<UserSearchHit>,
)
/** `exact = false` means "at least [value]": Elasticsearch stopped counting (chapter 09). */
data class TotalHits(val value: Long, val exact: Boolean)
/** A search result. Deliberately without email, mobile number, or date of birth. */
data class UserSearchHit(
val userId: String,
val fullName: String,
val city: String,
val state: String,
val accountStatus: String,
val updatedAt: Instant,
val score: Double?,
val highlights: Map<String, List<String>>,
)
/** A single profile, returned only by exact lookups. */
data class ProfileDetail(
val userId: String,
val fullName: String,
val email: String,
val mobileNumber: String,
val city: String,
val state: String,
val pincode: String,
val dateOfBirth: LocalDate?,
val gender: String,
val accountStatus: String,
val updatedAt: Instant,
)

Replace UserQueryBuilder.kt with the full design:

search-api/src/main/kotlin/in/o612/eng/usersearch/api/search/UserQueryBuilder.kt
package `in`.o612.eng.usersearch.api.search
import co.elastic.clients.elasticsearch._types.FieldValue
import co.elastic.clients.elasticsearch._types.SortOptions
import co.elastic.clients.elasticsearch._types.SortOrder as EsSortOrder
import co.elastic.clients.elasticsearch._types.query_dsl.Operator
import co.elastic.clients.elasticsearch._types.query_dsl.Query
import co.elastic.clients.elasticsearch._types.query_dsl.TextQueryType
import co.elastic.clients.elasticsearch.core.SearchRequest
import co.elastic.clients.elasticsearch.core.search.Highlight
import co.elastic.clients.elasticsearch.core.search.HighlightField
import co.elastic.clients.json.JsonData
import co.elastic.clients.util.NamedValue
import `in`.o612.eng.usersearch.api.web.NameMode
import `in`.o612.eng.usersearch.api.web.SortOrder
import `in`.o612.eng.usersearch.api.web.UserSearchRequest
import `in`.o612.eng.usersearch.index.UserProfileIndex
/**
* Turns a validated search request into an Elasticsearch request. Pure: no I/O, so it is unit-tested directly.
*
* Scoring happens only in the name clause. Every other constraint is a filter.
*/
object UserQueryBuilder {
/** The only fields a result list may return. */
val RESULT_FIELDS = listOf("userId", "fullName", "city", "state", "accountStatus", "updatedAt")
/** Fields a free-text name search covers; the name counts three times as much as the place. */
private val TEXT_FIELDS = listOf("fullName^3", "city.text", "state.text")
fun build(request: UserSearchRequest): SearchRequest = SearchRequest.of { s ->
s.index(UserProfileIndex.READ_ALIAS)
.size(request.size)
.source { src -> src.filter { f -> f.includes(RESULT_FIELDS) } }
.query(query(request))
.sort(sort(request.sort))
request.name?.let { s.highlight(highlight(request.nameMode)) }
s
}
fun query(request: UserSearchRequest): Query = Query.of { q ->
q.bool { b ->
request.name?.let { b.must(nameQuery(it, request.nameMode)) }
filters(request).forEach { b.filter(it) }
b
}
}
/** The only scored clause. */
fun nameQuery(name: String, mode: NameMode): Query = when (mode) {
NameMode.FULL -> Query.of { q ->
q.bool { b ->
b.should(crossFields(name))
.should(fuzzyName(name))
.minimumShouldMatch("1")
}
}
NameMode.PREFIX -> Query.of { q ->
q.match { m -> m.field("fullName.prefix").query(name).operator(Operator.And) }
}
}
/**
* Every word must appear in the name, city, or state, in any combination: "prashant patna".
* One analyser for all fields keeps them in one blended group; without it, operator AND applies per field group.
*/
private fun crossFields(text: String): Query = Query.of { q ->
q.multiMatch { mm ->
mm.query(text)
.type(TextQueryType.CrossFields)
.operator(Operator.And)
.analyzer("name_search")
.fields(TEXT_FIELDS)
}
}
/** Typo tolerance on the name alone. Scores below exact matches, so exact spellings rank first. */
private fun fuzzyName(text: String): Query = Query.of { q ->
q.match { m ->
m.field("fullName").query(text).operator(Operator.And).fuzziness("AUTO").prefixLength(1)
}
}
/** Exact constraints: filter context, no scoring, cacheable. */
fun filters(request: UserSearchRequest): List<Query> = buildList {
add(term("accountStatus", request.accountStatus.name))
request.city?.let { add(term("city", it)) }
request.state?.let { add(term("state", it)) }
request.pincode?.let { add(term("pincode", it)) }
request.gender?.let { add(term("gender", it.name)) }
request.updatedSince?.let { add(updatedAtFrom(it.toString())) }
// Rounded to the day, so the same filter is reused, and cached, for a whole day.
request.updatedWithinDays?.let { add(updatedAtFrom("now-${it}d/d")) }
}
/** Relevance first when requested, then newest, then userId so every order is total. */
fun sort(order: SortOrder): List<SortOptions> = buildList {
if (order == SortOrder.RELEVANCE) add(SortOptions.of { it.score { sc -> sc.order(EsSortOrder.Desc) } })
add(SortOptions.of { it.field { f -> f.field("updatedAt").order(EsSortOrder.Desc) } })
add(SortOptions.of { it.field { f -> f.field("userId").order(EsSortOrder.Asc) } })
}
fun highlight(mode: NameMode): Highlight {
val fields = when (mode) {
NameMode.FULL -> listOf("fullName", "city.text", "state.text")
NameMode.PREFIX -> listOf("fullName.prefix")
}
return Highlight.of { h -> h.fields(fields.map { NamedValue.of(it, HighlightField.of { f -> f }) }) }
}
private fun term(field: String, value: String): Query =
Query.of { q -> q.term { t -> t.field(field).value(FieldValue.of(value)) } }
private fun updatedAtFrom(from: String): Query =
Query.of { q -> q.range { r -> r.untyped { u -> u.field("updatedAt").gte(JsonData.of(from)) } } }
}

Three details map directly to the JSON. nameQuery builds the same two-way should with minimumShouldMatch("1") for the FULL mode, and the edge n-gram match for PREFIX. filters puts every exact constraint in filter context, including the rounded now-Nd/d range. And highlight passes its fields as NamedValue entries: in client 9.x, highlight fields are an ordered list, not a map.

In UserSearchService.kt, replace the highlights = hit.highlight(), line in search so that callers see field names from the contract, not index internals such as fullName.prefix or city.text:

search-api/src/main/kotlin/in/o612/eng/usersearch/api/search/UserSearchService.kt
// "fullName.prefix" and "city.text" are index details; callers see "fullName" and "city".
highlights = hit.highlight().mapKeys { (field, _) -> field.substringBefore('.') },

Check that Kotlin and JSON agree

The Java API Client’s request objects print their JSON body with toString(). For UserSearchRequest(name = "prashant", state = "bihar", updatedWithinDays = 90, size = 3), UserQueryBuilder.build(request).toString() produces, reformatted:

{
"_source": { "includes": ["userId", "fullName", "city", "state", "accountStatus", "updatedAt"] },
"highlight": { "fields": [ { "fullName": {} }, { "city.text": {} }, { "state.text": {} } ] },
"query": { "bool": {
"filter": [
{ "term": { "accountStatus": { "value": "ACTIVE" } } },
{ "term": { "state": { "value": "bihar" } } },
{ "range": { "updatedAt": { "gte": "now-90d/d" } } }
],
"must": [ { "bool": {
"minimum_should_match": "1",
"should": [
{ "multi_match": { "analyzer": "name_search", "fields": ["fullName^3", "city.text", "state.text"],
"operator": "and", "query": "prashant", "type": "cross_fields" } },
{ "match": { "fullName": { "fuzziness": "AUTO", "operator": "and", "prefix_length": 1, "query": "prashant" } } }
]
} } ]
} },
"size": 3,
"sort": [ { "_score": { "order": "desc" } }, { "updatedAt": { "order": "desc" } }, { "userId": { "order": "asc" } } ]
}

It is the Dev Tools request, with the long form of term and sort, and highlight fields as a list. The only intended difference is track_total_hits: the API keeps the default and reports exact: false above 10,000. Chapter 17 turns this comparison into a unit test.

Stage 2 — Run it through the API

Restart search-api with chapter 10’s environment variables, and run the composed search:

Terminal window
curl -s 'localhost:8080/api/users/search?name=prashant&state=bihar&updatedWithinDays=90&size=3'
total {value: 7, exact: true}
857970 Prashant Jha Patna 2026-08-01T00:00:00Z 13.604773 {fullName: [<em>Prashant</em> Jha]}
518970 Prashant Jha Patna 2026-07-24T00:00:00Z 13.604773 {fullName: [<em>Prashant</em> Jha]}
598470 Prashant Mishra Patna 2026-07-23T00:00:00Z 13.604773 {fullName: [<em>Prashant</em> Mishra]}

The JSON response is condensed here to one line per hit. The same seven profiles, in the same order, as the Dev Tools request. The other modes, each checked against PostgreSQL for active profiles:

RequestTotalPostgreSQLWhat it shows
name=pras&nameMode=PREFIX&state=bihar4,7154,715Autocomplete on the start of any word, highlighted as <em>Prashant</em> Jha
name=prashnat%20kumr1,1911,191Typos recovered by the fuzzy clause, scoring 5.22
name=prashant%20patna2,3632,363Words split across name and city; highlights include {"city": ["<em>Patna</em>"]}
name=mohd&city=patna2,3662,366Search-time synonyms, highlighted as <em>Mohammed</em> Mishra

An unknown nameMode value is rejected with 400 before any request reaches Elasticsearch.

Checkpoint

The composed search returns the same total as the PostgreSQL query from the previous section, for the day you run it, and its first hit carries a highlights.fullName fragment. If name=prashant%20patna returns 0, the analyzer("name_search") line is missing from crossFields.

Common mistakes in query design

  • Scoring exact constraints. Status, place, and date belong in filter.
  • Leaving minimum_should_match implicit. Its default changes when a filter is added to the same bool.
  • Trusting cross_fields across fields with different analysers. Check with _validate/query?explain=true, or force one analyser.
  • Sorting by _score alone. Ties are common in names; add a field sort and a unique tiebreaker.
  • Highlighting or returning fields the caller may not see. Source filtering and highlight fields are part of the privacy design.
  • Unrounded now in date filters. now-90d changes every millisecond and defeats caching; now-90d/d changes once a day.

What you built, and what comes next

The Search API now runs one designed query: a single scored name clause that combines cross-field matching with typo tolerance, filters for every exact constraint, a total sort order, highlighting, and source filtering. You verified the composed request in Dev Tools, in Kotlin, and through the API, and every count against PostgreSQL.

The API still returns only the first page. size stops at 50, and there is no way to ask for the next 50.

Chapter 12 adds aggregations: facet counts by state, city, and status for the same search, with the accuracy limits that come with them. Chapter 13 then adds cursor pagination with search_after and point-in-time.

ElasticsearchSpring BootKotlin

Type to search the site.

↑↓ navigate⏎ openPowered by Pagefind