Category: AI Search & Retrieval
Definition
Recall@k is a retrieval evaluation metric that measures how many of the relevant results are successfully retrieved within the top k results.
The k represents the number of results being evaluated.
For example:
- Recall@5 measures retrieval within the top 5 results.
- Recall@10 measures retrieval within the top 10.
- Recall@100 measures retrieval within the top 100.
Recall@k is especially useful for evaluating first-stage retrieval, where the goal is to make sure important candidates are not missed.
Formula
A simplified formula is:
Recall@k = Relevant results retrieved in top k ÷ Total relevant results
For example, suppose there are 10 relevant documents for a query.
If the first 20 retrieved results contain 8 of them:
Recall@20 = 8 ÷ 10 = 0.80
So the system has an 80% Recall@20.
Why It Matters
A retrieval system can only rank or use information that it successfully retrieves.
If an important document never enters the candidate set, a later reranker cannot select it.
This makes recall particularly important in multi-stage retrieval systems.
A common pipeline might look like:
Large document collection → First-stage retrieval → Candidate set → Reranking → Final results
If the first stage has poor recall, valuable information may be lost before the more sophisticated stages have a chance to evaluate it.
Example
Imagine a search query has 20 genuinely relevant documents.
A retrieval system returns its top 100 candidates.
Among those 100 candidates, 18 are genuinely relevant.
The calculation is:
Recall@100 = 18 ÷ 20 = 90%
This means the retrieval system successfully captured 90% of the known relevant documents within its first 100 results.
Recall@k vs. Precision@k
Recall@k and Precision@k measure different things.
Recall@k asks:
How many of the relevant documents did we retrieve?
Precision@k asks:
How many of the retrieved documents were relevant?
For example, suppose 10 relevant documents exist and the system retrieves 20 results containing 8 relevant documents.
Then:
Recall@20 = 8 ÷ 10 = 80%
Precision@20 = 8 ÷ 20 = 40%
A retrieval system may therefore have high recall while still returning many irrelevant candidates.
Why Different Values of k Matter
The value of k depends on where the metric is being used.
A first-stage retrieval system may evaluate Recall@100 or Recall@1000 because it needs to capture enough candidates for downstream ranking.
A user-facing search system may care more about very small values of k, where the top few results matter most.
Comparing multiple values of k can reveal where relevant information is being lost.
Recall@k in AI Search
Recall@k can be used to evaluate systems that retrieve:
- Documents
- Passages
- Knowledge-base entries
- Product records
- Web pages
- Other searchable information
It is particularly useful when retrieval is followed by a more expensive ranking or generation stage.
A system may deliberately retrieve a relatively large candidate set to maximize recall before applying a sophisticated reranker.
Why Recall@k Matters for AI Visibility
Recall@k provides a useful framework for understanding one important part of AI visibility:
Was relevant information retrieved at all?
If a source is not included in the candidate set, later ranking and answer-generation stages cannot normally select it.
For publishers, this does not mean there is a specific Recall@k number to optimize for on external AI search engines. Those systems use their own retrieval architectures and evaluation methods.
Instead, the concept reinforces the importance of creating content that is clearly relevant, well-structured, and easy for retrieval systems to identify.
Related Terms
- Retrieval Recall
- Retrieval Precision
- Retrieval F1 Score
- Precision@k
- First-Stage Retrieval
- Second-Stage Retrieval
- Candidate Generation
- Retrieval Evaluation
- Retrieval Quality
- Passage Retrieval
- Re-Ranking
In Simple Terms
Recall@k measures how much of the relevant information a retrieval system successfully finds within its top k results.
