ELI RESEARCH

AI-search recommendation stability: how repeatable is the answer?

Repeat the same buyer question and the shortlist may change.

Prompt variants × repeated runs — incomplete
Research in progress

Stability study collecting repeated observations. Do not derive stability from one run per prompt.

RELEASE STANDARD

Methodology is publishable before findings are.

Eli can publish the question universe, schema, definitions and quality controls before a benchmark release. It will only publish an aggregate when the underlying observations are complete, reviewable and described with their coverage and limitations. Until then, this page makes no performance, ranking or customer-outcome claim.

Two kinds of stability

Measure repetition stability for same prompt and paraphrase stability for controlled wording variants.

Observation design

Store intent ID, variant, engine, date/time, brand set, recommendation order and citations.

Overlap metrics

Potential metrics include brand-set overlap, top recommendation consistency, Jaccard overlap and rank/context agreement; publish only used formulas.

Why stability matters

High rate can be fragile; moderate rate can be consistent. Track volatility and whether changes survive paraphrases.

FAQ

Questions about the research design and release boundary.

Why repeats?

Generative answers vary.

How many?

Publish threshold after methodology defines it.

Paraphrases keywords?

No, controlled same-intent variants.

Stable but rare?

Yes.

Means correctness?

No.

CONNECT YOUR WEBSITE

Measure your recommendation stability

Recommendation stability measures how consistently the same brands appear across repeated observations and controlled paraphrases of the same buyer intent. It separates durable visibility from one-off output variance.

Measure your recommendation stability