ELI RESEARCH

Eli AI Search Benchmark methodology

A benchmark is useful only when someone can understand what was measured.

Prompt design → version lock → capture → normalize → QA → release
Research in progress

Methodology is public. No benchmark aggregate will be published until raw observations are complete, normalized and quality-checked.

RELEASE STANDARD

Methodology is publishable before findings are.

Eli can publish the question universe, schema, definitions and quality controls before a benchmark release. It will only publish an aggregate when the underlying observations are complete, reviewable and described with their coverage and limitations. Until then, this page makes no performance, ranking or customer-outcome claim.

Research questions and prompt construction

Measure recommendation/mention/citation frequency, cited domains, engine overlap and stability. Prompts belong to commercial intent and controlled paraphrases retain stable IDs.

Observation record

Every record has prompt ID, intent, variant, engine/model where known, date, raw answer, mentioned/recommended brands, citations, market, language and notes.

Metrics

Recommendation and mention rates use recommended/mentioned eligible observations divided by eligible universe. Citation denominator must be explicit. Publish exact stability formula when implemented.

Normalization and extraction

Normalize spelling variants, product versus parent brand and aliases while retaining original answer. Capture URL/domain/source type and do not infer citations not exposed.

Quality control

Verify prompt version, missing answers, extraction errors, ambiguous recommendations, entity normalization and citation parsing before locking release.

FAQ

Questions about the research design and release boundary.

Why fix prompt set?

So question changes do not cause result changes.

Why paraphrases?

To measure sensitivity to wording.

Why raw answers?

For auditable classifications/aggregates.

Ambiguous recommendations?

Flag for review, not forced classification.

Can methodology change?

Yes, with versioned methodology/release.

CONNECT YOUR WEBSITE

Inspect the data schema

Eli's AI-search benchmark uses a fixed, versioned buyer-question universe, controlled paraphrases, explicit engine/date metadata, saved raw answers, normalized brand entities and transparent metric formulas. Changing the prompt universe creates a new benchmark version.

Inspect the data schema