ArticleRetrieval & Evaluation

Cosine Similarity Is Not an Answer Confidence Score

A passage can resemble your question without containing the fact needed to answer it. Choose an admission rule using evidence labels, then test the answer separately.

Last reviewed: 2026-10-09

Glass vessels of different shapes on black platforms, illuminated by cyan and amber light on a reflective surface.
Similarity compares representations; evidence sufficiency requires checking the facts needed for the answer.
TL;DR

Use cosine similarity to compare embeddings. Use labelled cases to decide whether a retrieval score should admit context. Use an answer rubric to judge the generated response. These are three different tasks. A threshold is a decision rule, and a score that passes it is still not a calibrated probability that the answer is supported.

Similarity, Support and Correctness

The Sentence Transformers similarity documentation describes comparing text embeddings: higher similarity identifies semantically closer text pairs. That quantity concerns the texts being compared. Reading it as the probability that a later generated answer is correct adds a claim the similarity calculation does not establish.

Keep three questions visible in your logs. Does the retrieved text concern the question? Does the supplied context contain sufficient evidence for the required answer? Does the generated answer correctly use that evidence? A topical passage can fail the second check. An answer can fail the third even when the context passes the second, for example by assigning a documented setting to the wrong software version.

The scikit-learn calibration documentation defines calibration through agreement between predicted probabilities and observed label frequencies. A reliability diagram compares those quantities in groups. To build an answer-support probability, you must first define the support label and validate a learned mapping on independent data. Applying a cutoff to raw similarity does not perform that work.

A Similar Passage Without the Required Fact

Consider this invented support question: “What is the default request timeout in Client Library v4?” The retrieved passage says: “Client Library v4 supports request timeouts. Set the request_timeout option to override the default.” It shares the product, version and topic with the question, but omits the default value. This illustration has no measured similarity score.

If those are the only supplied excerpts, the evidence label is insufficient. The expected response is to state that the default value is missing, rather than invent a duration. A second passage documenting a v3 default is also insufficient for this v4 question unless evidence establishes that the setting carries over. Repeated keywords and plausible history do not fill the missing fact.

Now add the actual v4 settings table containing the default. Corpus answerability changes when that document becomes available. Retrieved-context sufficiency changes only if retrieval and packing include the required table. This distinction helps diagnose whether you need better retrieval, a different admission rule or changes to generation. The Retrieval and RAG guide covers the whole pipeline; this worksheet isolates the threshold decision.

Label Questions Before Choosing a Threshold

Freeze a corpus snapshot and create questions representative of your feature. Include answerable questions with named evidence and unanswerable questions whose required facts are absent. Include near misses: the wrong version, a similar product, an outdated policy and a passage that mentions the setting without its value. For multi-part questions, list every required fact; one matching paragraph should not count as a complete answer.

For each case, record the question ID, corpus-answerable label, required facts, evidence IDs and expected behaviour when evidence is absent. Resolve disputed labels by inspecting the source text. Keep ambiguous cases visible instead of silently labelling them negative. Separate the tuning questions from held-out validation questions before inspecting score distributions, grouping paraphrases of the same question together so they do not appear in both sets.

Run retrieval once with the frozen embedding model, query processing, filters, candidate limit and context packing rule. Save the candidate IDs, scores and exact text sent to generation. Add a second label: whether this retrieved context contains all required evidence. An answerable question whose evidence was missed is a retrieval failure; admitting its insufficient context is a false acceptance by the support gate. The AI Evals guide explains how evaluation cases connect to product behaviour.

Define the Rule and Count Both Errors

For a simple first experiment, define one scalar score per question: the maximum cosine similarity among its retrieved candidates. Admit the frozen context when that score is at least the selected threshold; otherwise abstain. Treat an empty candidate list as abstention explicitly. This rule is deliberately limited: a maximum score does not express coverage of several required facts. If you use a different aggregation or per-chunk filtering, document it and label the resulting context instead.

The scikit-learn threshold guide separates score prediction from taking action and makes tuning depend on the chosen metric. Here, “positive” means sufficient retrieved context, and “accept” means passing the gate. Use that definition consistently in your confusion table.

Evidence labelGate acceptsGate abstains
Sufficient contextTrue acceptance: generation may proceedFalse rejection: usable evidence was withheld
Insufficient contextFalse acceptance: generation proceeds without required supportTrue rejection: missing support is recognised

Calculate false acceptance rate as false acceptances divided by all insufficient-context cases. Calculate false rejection rate as false rejections divided by all sufficient-context cases. Also report insufficient-context acceptances divided by all accepted cases: that answers a different operational question, how often admitted context lacks support. Show counts beside rates and mark a rate undefined when its denominator is zero. Break out corpus-unanswerable questions and retrieval misses so their causes remain visible.

Under this fixed maximum-score rule, raising the threshold can only remove acceptances. It can remove both unsafe and useful ones; it does not establish that generated answers improve. Decide which errors the product can tolerate before selecting a threshold. A support tool that can hand off a question may choose differently from a feature where unnecessary abstention blocks a routine workflow. These are product tradeoffs, not a universal numerical cutoff.

Use a Held-Out Decision Worksheet

Sweep candidate thresholds on the tuning set, then lock the selected rule and run the separate validation set. The threshold guide warns about overfitting when training and tuning reuse data. Reserve validation for the final decision rather than repeatedly changing the threshold after every validation miss. If you tune again, obtain fresh held-out cases before making a new performance claim.

Fill the following worksheet from your own runs. No retrieval or generation experiment was performed for this article, and the rows contain no benchmark results.

Decision fieldRecord before adoption
ConfigurationCorpus snapshot, embedding model, metric, filters, candidate count, packing rule and score aggregation
Labels and splitQuestion-set version, sufficient/insufficient counts, corpus-answerability breakdown and disputed labels
Chosen thresholdNot measured; fill the value and equality rule after tuning
Validation errorsNot measured; false acceptance and rejection counts, denominators and rates
Accepted-context riskNot measured; insufficient accepted contexts / all accepted contexts
Answer outcomeNot measured; supported correct answers, unsupported claims and appropriate abstentions
Acceptance limitsProduct limits for both error types, owner, decision date and next review trigger

Evaluate generation with the same contract you request: answer from supplied excerpts, cite evidence for factual claims and state when a required fact is missing. Score factual accuracy, completeness and citation support separately. Record unsupported answers among accepted cases as well as successful answers among all questions. Otherwise a restrictive threshold can make accepted-answer quality look good simply by declining useful questions. A correct answer supplied from elsewhere still fails an evidence-only support contract.

Revalidate When the System Changes

Attach the threshold to its recorded configuration. Re-run the labelled comparison when you change embedding models, corpus contents, chunk boundaries, filters or query processing. A document update can change whether a question is answerable; review the labels before treating changed decisions as regressions. The embedding migration guide gives a broader procedure for comparing retrieval across an index change.

If the answer model or prompt changes while retrieval stays fixed, repeat the answer-support evaluation even if gate decisions are identical. If real traffic introduces a new question category, add labelled cases for it and review the threshold decision against the original product limits. Retain the previous configuration and decision record so a failed validation has a concrete rollback target.

If you later display a probability, specify the event it predicts: sufficient context, fully supported answer or correct answer under a particular rubric. Keep calibration data separate from fitting data, check reliability on held-out cases and report group sizes. A model calibrated for one event does not certify another. Until that process is validated, show the score as a similarity signal and explain the admission rule without turning it into a confidence percentage.

Primary sources checked 2026-10-09

The example, label scheme, error denominators, worksheet and review process are this article's proposed method. No universal threshold or measured improvement is claimed.