Retrieval
Does the retriever find the lemmas a proof needs?
Measured on its own, before any proving. The ground truth for a theorem is the set of indexed Mathlib theorems its reference proof uses. Premises from held-out modules are removed from the index when it is built, so no retriever can return the target or anything downstream of it.
Recall and mean reciprocal rank
164 theorems have at least one reachable ground-truth premise. 95% bootstrap intervals over theorems.
| Retriever | Theorems | n | MRR | R@1 | R@5 | R@8 | R@10 | R@20 | R@50 |
|---|---|---|---|---|---|---|---|---|---|
| BM25 | All | 164 | 0.117 | 2% | 8% | 9% | 10% | 14% | 17% |
| BM25 | Mathlib held-out | 119 | 0.136 | 2% | 7% | 9% | 10% | 12% | 15% |
| BM25 | Authored | 45 | 0.065 | 4% | 9% | 9% | 9% | 18% | 20% |
| Dense (bge-small) | All | 164 | 0.088 | 2% | 5% | 7% | 8% | 11% | 16% |
| Dense (bge-small) | Mathlib held-out | 119 | 0.094 | 1% | 5% | 5% | 7% | 9% | 13% |
| Dense (bge-small) | Authored | 45 | 0.074 | 4% | 7% | 12% | 12% | 17% | 23% |
| Hybrid (RRF) | All | 164 | 0.125 | 3% | 7% | 10% | 10% | 13% | 19% |
| Hybrid (RRF) | Mathlib held-out | 119 | 0.129 | 1% | 6% | 8% | 9% | 12% | 18% |
| Hybrid (RRF) | Authored | 45 | 0.116 | 9% | 11% | 16% | 16% | 18% | 24% |
Recall inside the agent runs
Mean share of a theorem's ground-truth premises that appeared in the eight lemmas actually shown to the model.
| Configuration | Mean recall@8 | Verified |
|---|---|---|
| Full (plan + retrieval + repair) | 11% | 7% |
| Full − compiler feedback | 11% | 6% |
| Full − memory | 11% | 6% |
| Full − skeleton | 11% | 7% |
| Hybrid retrieval | 11% | 2% |
| Hybrid retrieval + repair | 11% | 5% |
| BM25 retrieval + repair | 8% | 7% |
| Dense retrieval + repair | 8% | 5% |