Card 1 — Cover
Card 2 — Prover Leaderboard
| Prover | pass@1 | pass@32 |
|---|---|---|
| Goedel-Architect | 99.2% | 100%* |
| OProver-32B | — | 93.3% |
| DeepSeek-Prover-V2 (671B) | 61.9% | 88.9% |
- Goedel-Architect: 75.6% (no NL seed) / 88.8% (with NL seed, 597/672)
- OProver-32B: 11.3% @pass@32
- DeepSeek-Prover-V2: ~7.4% @pass@1024 (prior result)
- Goedel-Architect uses DeepSeek-V4-Flash API; pass@1 reported on standard MiniF2F-test split.
- OProver reports pass@32 (unbiased estimate over n=64 rollouts); no independent external replication yet for either system.
- Kimina-Prover (not yet reported on these specific splits this week).
Card 3 — New Benchmarks & Harnesses
Card 4 — Signal vs Hype
| Item | Verdict |
|---|---|
| Goedel-Architect 99.2% pass@1 MiniF2F | ✓ Lean kernel-verified, open-source pipeline |
| OProver 32B beats DeepSeek 671B | ✓ Open weights + OProofs corpus public |
| LeanMarathon 0 sorries, 7 theorems | ✓ GitHub artifacts, runnable |
| TheoremBench coverage finding | ⚠ No external replication yet |
| LLM Eval Gemini/Claude scores | ⚠ refine@32 budget, not pass@1 |
- PutnamBench gains for Goedel-Architect depend on NL seeding (human-provided proof outlines). Strong result, but pass@1 without seeding (75.6%) is the more conservative baseline.
- OProver's training relies on OProofs corpus bootstrapped from public Lean resources; contamination with miniF2F training split not explicitly audited.
- FLT formalization: no new sorry closures this week. Workshop Jul 6–10 is the next milestone.
- Mathlib Initiative: no new public Q2 metrics this week; last known trajectory on track for <1-week PR review cycles.


Comments
Sign in to comment.