Skip to content

Commit ca3a3c1

Browse files
committed
Add corrected Phase 3E execution report
1 parent 69313bd commit ca3a3c1

1 file changed

Lines changed: 72 additions & 0 deletions

File tree

Lines changed: 72 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,72 @@
1+
# Phase 3E Final Execution Report
2+
3+
## Final status
4+
5+
- Final verdict: **COMPLETE**
6+
- Repository status: **MERGED TO `main`**
7+
- Baseline commit: `c4313f6ad3a28b5901255259c2b39b67ee836413`
8+
- Phase 3E implementation commit: `bcfe625`
9+
- Merge commit on `main`: `69313bd`
10+
- Feature branch: `benchmark/phase3e-experiment`
11+
- Pull request: `#2 — Finalize Phase 3E benchmark study`
12+
13+
## Gates
14+
15+
| Gate | Result |
16+
|---|---|
17+
| Clean working tree | PASS |
18+
| Clean install | PASS |
19+
| Ruff lint | PASS |
20+
| Ruff format | PASS |
21+
| Tests | PASS — 57 tests |
22+
| Cross-batch conflict semantics | PASS — A/B/C |
23+
| Stateful rerun/update/stale semantics | PASS |
24+
| Oracle equivalence | PASS — 100k and 1m |
25+
| Process interruption recovery | PASS |
26+
| 100k cohort | PASS — 90 runs, 11,250 query samples |
27+
| 1m cohort | PASS — 48 runs, 4,500 query samples |
28+
| Q1–Q5 equality | PASS |
29+
| Analysis regeneration | PASS |
30+
| GitHub Actions push CI | PASS — Python 3.11 and 3.12 |
31+
| GitHub Actions PR CI | PASS — Python 3.11 and 3.12 |
32+
| Pull request merge | PASS — PR #2 |
33+
34+
The 1m resource-bounded design uses one complete stateful correctness sequence per strategy and five measured `initial_load` / `exact_rerun` repetitions per strategy. The correctness-only repetition is excluded from performance aggregation by `scripts/analyze_benchmark.py`.
35+
36+
## Corrected 1m analysis
37+
38+
The original generated report accidentally included the correctness-only `repetition == 0` run in the regenerated performance summary. The merged implementation fixes this by excluding repetition 0 from performance aggregation while preserving it for correctness evidence.
39+
40+
Corrected median initial-load wall times:
41+
42+
| Strategy | Median (ms) |
43+
|---|---:|
44+
| A — atomic full replacement | 39,050.38 |
45+
| B — incremental upsert | 21,173.67 |
46+
| C — append history + materialized current state | 22,539.42 |
47+
48+
For this recorded 1m cohort, Strategy B's median initial-load wall time was **45.8% lower** than Strategy A's. Equivalently, Strategy A took about **1.84×** as long as Strategy B.
49+
50+
The 100k result is unchanged: Strategy B's median initial-load wall time was **49.4% lower** than Strategy A's.
51+
52+
## Canonical source
53+
54+
The canonical project state is the merged Git repository at:
55+
56+
- Phase 3E implementation commit: `bcfe625`
57+
- Main merge commit: `69313bd`
58+
59+
Earlier Manus-generated ZIP/diff/format-patch artifacts predate the local analysis correction and should not be treated as the canonical final source unless regenerated from the merged repository.
60+
61+
## CI evidence
62+
63+
Remote GitHub Actions was executed after the corrected Phase 3E implementation was pushed:
64+
65+
- Push workflow run: `35874939037` — PASS
66+
- Pull-request workflow run: `35875147446` — PASS
67+
- Python 3.11 correctness job — PASS
68+
- Python 3.12 correctness job — PASS
69+
70+
## Main remaining limitation
71+
72+
Evidence is single-machine, synthetic, single-process, and SQLite-specific. No universal performance claim, distributed scalability claim, publication claim, or production-readiness claim is made.

0 commit comments

Comments
 (0)