TL;DR — In large‑scale search, BM25’s k1 and b parameters control term‑frequency saturation and document‑length bias; tuning them for your corpus can improve nDCG by 15‑30 % while keeping latency sub‑second. The key is to measure the average document length, inspect term‑frequency distributions, and iterate with A/B tests on real query traffic. When done right, BM25 becomes a low‑overhead, high‑quality ranking function that scales horizontally across shards and replicas.

Search relevance is the differentiator between a product users love and one they abandon. BM25, the default ranking function in many open‑source and commercial search engines, provides a small but powerful set of knobs that directly affect how terms are weighted relative to document length and collection statistics. Because it is based on a probabilistic model of term relevance, small, data‑driven adjustments often yield measurable gains without requiring a full learning‑to‑rank pipeline. In this post we will explore the mechanics of BM25, walk through concrete tuning steps for large‑scale indexes, and share patterns that production teams at companies such as Elastic, Netflix, and Airbnb use to keep search fast and relevant.

BM25 Fundamentals

BM25 (Okapi BM25) is the de‑facto standard ranking function used by Elasticsearch, OpenSearch, Solr, and many other search platforms. It scores a document d for a query q as

score(d,q) = Σ_{t∈q} IDF(t) * (freq(t,d) * (k1 + 1) / (freq(t,d) + k1 * (1 - b + b * |d| / avdl)))
  • IDF(t) inverts the document frequency of term t across the whole collection, giving higher weight to rarer terms.
  • freq(t,d) is the raw term frequency of t in document d.
  • |d| is the length of document d (usually measured in tokens).
  • avdl is the average document length across the index.
  • k1 controls term‑frequency saturation: larger values let high frequencies contribute more before diminishing returns set in.
  • b controls document‑length normalization: b = 0 disables length normalization, b = 1 normalizes fully to the average length.

In practice, most deployments start with the classic defaults k1 = 1.5 and b = 0.75, but these values were tuned on English newswire corpora. When your data differs—shorter documents, multi‑lingual content, or heavy technical jargon—those defaults may be sub‑optimal.

Term‑frequency saturation

Increasing k1 flattens the saturation curve, allowing documents with many occurrences of a query term to retain a higher share of the IDF boost. Decreasing k1 makes the score drop off more quickly, which can be useful when you want to penalize keyword stuffing or when the term frequency distribution is heavily skewed.

Document‑length normalization

The b parameter stretches or compresses the length‑adjustment term. A b close to 0 means long documents are treated almost the same as short ones; a b close to 1 forces long documents to receive a proportionally lower score. Choosing b is often a trade‑off between recall (favoring longer documents) and precision (favoring concise, relevant passages).

Key Parameters k1 and b

Choosing k1

A practical starting point is to inspect the term‑frequency histogram of your corpus. If you see many documents with frequency ≥ 5, a higher k1 (≈ 2.0) may prevent over‑penalisation. Conversely, if most documents contain the term once or twice, k1 ≈ 1.2 can give more discriminative power. A common approach is to run a quick grid search:

  1. Set k1 ∈ {1.0, 1.2, 1.5, 1.8, 2.0}.
  2. Keep b fixed at 0.75.
  3. Evaluate nDCG@10 on a held‑out validation set of queries.
  4. Pick the k1 that maximises nDCG without a large latency hit.

Adjusting b

The b parameter is more data‑dependent than k1. If your index contains many long technical manuals, a lower b (≈ 0.5) can keep those documents competitive. If you want the search to favour concise answers (e.g., FAQ entries), raise b toward 0.8 or 0.9. As with k1, validate with A/B tests on real query logs.

Document Length Normalization

The average document length avdl is computed automatically by the search engine, but it’s worth monitoring. A sudden shift—e.g., after a bulk import of longer articles—will change the normalization baseline and can silently degrade relevance. Most engines let you override avdl per‑field or per‑index if you need a custom baseline.

In Elasticsearch you can view the current avdl:

GET /my_index/_stats/indexing?filter_path=**.avg_doc_len

If you detect a systematic offset, you can re‑index with a adjusted field length or apply a search‑time boost to compensate.

Scaling BM25 in Production

Large‑scale indexes—tens of millions of documents across multiple shards—introduce additional considerations:

  • Shard‑local relevance: BM25 scores are computed independently on each shard. After scoring, the coordinating node merges results, applying a reciprocal rank fusion or simple top‑N selection. Ensure that the same k1/b values are deployed across all shards; version drift can cause inconsistent rankings.

  • Cache‑friendly parameters: Because k1 and b are static per index, the scoring logic can be baked into the query execution plan. Elasticsearch caches IDF values per shard; changing these parameters triggers an IDF recompute, which may cause a brief spike in query latency.

  • Horizontal scaling: When you add new shards, re‑balance the cluster and recompute avdl. Use the _cluster/reroute API to force a re‑balance if you notice score drift after adding nodes.

  • Failure mode: A common pitfall is forgetting to redeploy the new k1/b after a rolling restart. If the old parameters remain in the segment metadata, queries may still use the old values, leading to “stale relevance”. Always verify the active BM25 settings with GET /_cat/indices?v&h=index,store.analysis.

Architecture pattern: “parameter‑per‑collection”

Some organisations maintain a tiny metadata service that stores the current BM25 tuple per index. Deployment pipelines read this service, render the values into the index settings, and trigger a rolling restart. This pattern decouples parameter changes from code deployments and enables rapid A/B testing via feature flags.

Practical Tuning Workflow

  1. Establish a baseline – Deploy the default k1 = 1.5, b = 0.75 and capture nDCG@10 on a representative query log (ideally > 10 k queries).
  2. Measure avdl and term‑frequency distribution – Use engine‑specific APIs or a quick Python script to compute the histogram.
  3. Run a coarse grid search – Vary k1 over {1.0, 1.2, 1.5, 1.8, 2.0} while keeping b fixed, then vary b over {0.5, 0.75, 0.9} with the best k1.
  4. Evaluate with A/B testing – Route a percentage of traffic (e.g., 5 %) to the new configuration and compare lift in click‑through rate, conversion, or whatever metric matters for your product.
  5. Iterate – If a statistically significant improvement appears, promote the new parameters to all traffic; otherwise, adjust the search space and repeat.
  6. Monitor – Set up alerts for sudden drops in nDCG or latency spikes after a parameter change. A rolling restart that forgets to update the settings is the most common cause of regressions.

Example: Updating BM25 in Elasticsearch

PUT /my_index/_settings
{
  "index": {
    "similarity": {
      "bm25": {
        "type": "BM25",
        "k1": 1.8,
        "b": 0.6
      }
    }
  }
}

The request above updates the similarity model without re‑indexing. After the change, Elasticsearch recomputes IDF on the next refresh interval (usually 1 s) and applies the new k1/b to all subsequent searches.

Key Takeaways

  • BM25’s two knobs, k1 (term‑frequency saturation) and b (document‑length normalization), have an outsized impact on relevance when your corpus deviates from the original newswire assumptions.
  • Always measure avdl and term‑frequency distributions before tuning; a data‑driven starting point saves time and avoids guesswork.
  • Grid‑search combined with A/B testing on real query logs is the most reliable way to discover optimal values.
  • Keep BM25 parameters cluster‑wide consistent; shard‑level drift silently degrades ranking quality.
  • Cache‑friendly parameter updates (e.g., Elasticsearch _settings API) let you iterate quickly without full re‑indexing.
  • Monitor nDCG, latency, and business metrics after each change; a regression in one area often signals a mis‑configured parameter.

Further Reading