Repair loop
Does feeding Lean's errors back help?
A repair round shows the model its previous proof and exactly what Lean printed, nothing more. The fair comparison is against spending the same number of calls on independent fresh drafts (Direct ×4), not against a single draft.
Repair success
Among theorems whose first draft failed, the share later verified within the same run.
| Configuration | First draft failed | Later verified |
|---|---|---|
| Full (plan + retrieval + repair) | 167 | 4% |
| Full − compiler feedback | 167 | 2% |
| Full − memory | 167 | 2% |
| Full − retrieval | 166 | 1% |
| Full − skeleton | 165 | 2% |
| Hybrid retrieval + repair | 171 | 3% |
| BM25 retrieval + repair | 168 | 4% |
| Dense retrieval + repair | 171 | 4% |
| Repair | 173 | 6% |
| Repair (paraphrased prompt) | 169 | 4% |
Which errors get repaired
For each class of first error, how often the very next round in the same run produced a verified proof. Classes with fewer than five cases are omitted; 95% Wilson intervals.