|
| 1 | +# Phase 3E Final Execution Report |
| 2 | + |
| 3 | +## Final status |
| 4 | + |
| 5 | +- Final verdict: **COMPLETE** |
| 6 | +- Repository status: **MERGED TO `main`** |
| 7 | +- Baseline commit: `c4313f6ad3a28b5901255259c2b39b67ee836413` |
| 8 | +- Phase 3E implementation commit: `bcfe625` |
| 9 | +- Merge commit on `main`: `69313bd` |
| 10 | +- Feature branch: `benchmark/phase3e-experiment` |
| 11 | +- Pull request: `#2 — Finalize Phase 3E benchmark study` |
| 12 | + |
| 13 | +## Gates |
| 14 | + |
| 15 | +| Gate | Result | |
| 16 | +|---|---| |
| 17 | +| Clean working tree | PASS | |
| 18 | +| Clean install | PASS | |
| 19 | +| Ruff lint | PASS | |
| 20 | +| Ruff format | PASS | |
| 21 | +| Tests | PASS — 57 tests | |
| 22 | +| Cross-batch conflict semantics | PASS — A/B/C | |
| 23 | +| Stateful rerun/update/stale semantics | PASS | |
| 24 | +| Oracle equivalence | PASS — 100k and 1m | |
| 25 | +| Process interruption recovery | PASS | |
| 26 | +| 100k cohort | PASS — 90 runs, 11,250 query samples | |
| 27 | +| 1m cohort | PASS — 48 runs, 4,500 query samples | |
| 28 | +| Q1–Q5 equality | PASS | |
| 29 | +| Analysis regeneration | PASS | |
| 30 | +| GitHub Actions push CI | PASS — Python 3.11 and 3.12 | |
| 31 | +| GitHub Actions PR CI | PASS — Python 3.11 and 3.12 | |
| 32 | +| Pull request merge | PASS — PR #2 | |
| 33 | + |
| 34 | +The 1m resource-bounded design uses one complete stateful correctness sequence per strategy and five measured `initial_load` / `exact_rerun` repetitions per strategy. The correctness-only repetition is excluded from performance aggregation by `scripts/analyze_benchmark.py`. |
| 35 | + |
| 36 | +## Corrected 1m analysis |
| 37 | + |
| 38 | +The original generated report accidentally included the correctness-only `repetition == 0` run in the regenerated performance summary. The merged implementation fixes this by excluding repetition 0 from performance aggregation while preserving it for correctness evidence. |
| 39 | + |
| 40 | +Corrected median initial-load wall times: |
| 41 | + |
| 42 | +| Strategy | Median (ms) | |
| 43 | +|---|---:| |
| 44 | +| A — atomic full replacement | 39,050.38 | |
| 45 | +| B — incremental upsert | 21,173.67 | |
| 46 | +| C — append history + materialized current state | 22,539.42 | |
| 47 | + |
| 48 | +For this recorded 1m cohort, Strategy B's median initial-load wall time was **45.8% lower** than Strategy A's. Equivalently, Strategy A took about **1.84×** as long as Strategy B. |
| 49 | + |
| 50 | +The 100k result is unchanged: Strategy B's median initial-load wall time was **49.4% lower** than Strategy A's. |
| 51 | + |
| 52 | +## Canonical source |
| 53 | + |
| 54 | +The canonical project state is the merged Git repository at: |
| 55 | + |
| 56 | +- Phase 3E implementation commit: `bcfe625` |
| 57 | +- Main merge commit: `69313bd` |
| 58 | + |
| 59 | +Earlier Manus-generated ZIP/diff/format-patch artifacts predate the local analysis correction and should not be treated as the canonical final source unless regenerated from the merged repository. |
| 60 | + |
| 61 | +## CI evidence |
| 62 | + |
| 63 | +Remote GitHub Actions was executed after the corrected Phase 3E implementation was pushed: |
| 64 | + |
| 65 | +- Push workflow run: `35874939037` — PASS |
| 66 | +- Pull-request workflow run: `35875147446` — PASS |
| 67 | +- Python 3.11 correctness job — PASS |
| 68 | +- Python 3.12 correctness job — PASS |
| 69 | + |
| 70 | +## Main remaining limitation |
| 71 | + |
| 72 | +Evidence is single-machine, synthetic, single-process, and SQLite-specific. No universal performance claim, distributed scalability claim, publication claim, or production-readiness claim is made. |
0 commit comments