You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Commit f90f30e
Browse filesBrowse the repository at this point in the historyBrowse files
The repository includes a reproducible benchmark harness. Its results are cohort-specific and
10
-
are not a claim of production readiness, universal performance, scalability, or research novelty.
5
+
A correctness-first, reproducible time-series ETL and SQLite benchmark for comparing three loading strategies under stateful updates, duplicates, conflicts, stale revisions, and interruption recovery.
6
+
7
+
## Why this project exists
8
+
9
+
The project studies a practical systems question:
10
+
11
+
> How do different SQLite loading strategies behave when they must preserve the same logical state under reruns, updates, duplicates, stale revisions, conflicts, and failures?
12
+
13
+
The benchmark compares:
14
+
15
+
-**Strategy A — Atomic full replacement**
16
+
-**Strategy B — Transactional incremental upsert**
17
+
-**Strategy C — Append-only history with a materialized current-state projection**
18
+
19
+
All three strategies are validated against the same independent logical-state oracle before performance results are summarized.
20
+
21
+
## Highlights
22
+
23
+
- Deterministic synthetic time-series generator
24
+
- Strict normalization and revision semantics
25
+
- Independent logical-state oracle with stable SHA-256 checksums
26
+
- Stateful workload execution
27
+
- Cross-batch same-revision conflict rejection
28
+
- Stale-revision protection
29
+
- Process-interruption and transactional recovery tests
30
+
- Five fixed query workloads with warm-up and timed samples
31
+
- Reproducible 100k and 1m benchmark cohorts
32
+
- GitHub Actions validation on Python 3.11 and 3.12
33
+
34
+
## Benchmark results
35
+
36
+
Median initial-load wall time from the corrected Phase 3E analysis:
37
+
38
+
| Cohort | Strategy A | Strategy B | Strategy C |
39
+
|---|---:|---:|---:|
40
+
| 100k rows | 3,429.25 ms | 1,733.58 ms | 1,868.53 ms |
41
+
| 1m rows | 39,050.38 ms | 21,173.67 ms | 22,539.42 ms |
42
+
43
+
For the recorded environment and workloads:
44
+
45
+
- At **100k rows**, Strategy B had **49.4% lower** median initial-load wall time than Strategy A.
46
+
- At **1m rows**, Strategy B had **45.8% lower** median initial-load wall time than Strategy A.
47
+
- At **1m rows**, Strategy A took about **1.84×** as long as Strategy B.
48
+
49
+
These are cohort-specific observations, not universal performance claims.
CLI values override a YAML file. Only `source`, `seed`, and `rows` are accepted:
153
+
The full benchmark is intentionally separate from ordinary CI.
86
154
87
-
```bash
88
-
fdp run-all --config config.yaml --output-dir /tmp/fdp-run
155
+
## Stateful semantics
156
+
157
+
The logical key is:
158
+
159
+
```text
160
+
(series_id, event_ts_utc)
89
161
```
90
162
91
-
The package never searches the source checkout for configuration. It does not use `.env`, so
92
-
no `.env.example` is required.
163
+
Exact event identity is:
164
+
165
+
```text
166
+
(series_id, event_ts_utc, revision)
167
+
```
168
+
169
+
Rules:
170
+
171
+
- exact duplicate events are safe no-ops
172
+
- same-revision payload conflicts reject the entire batch
173
+
- higher revisions update current state
174
+
- lower revisions cannot regress current state
175
+
- empty replacement is rejected by default
176
+
177
+
Timestamps are UTC Unix seconds aligned to one minute. Prices and optional volume use fixed-point integers with scale `10^-6`.
93
178
94
-
## Logical model
179
+
## Query suite
95
180
96
-
The logical key is `(series_id, event_ts_utc)` and exact event identity is
97
-
`(series_id, event_ts_utc, revision)`. Timestamps are UTC Unix seconds aligned to one minute.
98
-
Prices and optional volume use fixed-point integers with scale `10^-6`.
181
+
Each accepted measured state uses:
99
182
100
-
Duplicate/update rules:
183
+
-**Q1** — point lookup
184
+
-**Q2** — range lookup
185
+
-**Q3** — period aggregate
186
+
-**Q4** — latest-state lookup
187
+
-**Q5** — filtered-return query
101
188
102
-
- an exact duplicate event is a safe no-op;
103
-
- different payloads for the same key and revision reject the entire batch, including across
104
-
previously committed batches;
105
-
- the highest revision is current;
106
-
- a lower revision cannot regress the current state;
107
-
- an empty replacement is rejected by default.
189
+
Each query uses two warm-ups and 30 timed warm-cache samples.
108
190
109
-
The SQLite current-state table has a composite primary key, required checks, and an
110
-
`event_ts_utc` index. Loads use WAL, `synchronous=FULL`, foreign keys, a 5-second busy timeout,
111
-
and one explicit transaction. The staging table is validated against the independent oracle
112
-
before publication. A pre-commit failure rolls back to the prior committed table.
191
+
Result signatures are checked for equality across strategies before timing summaries are used.
113
192
114
-
## Tests and local CI-equivalent checks
193
+
## Correctness and recovery
194
+
195
+
The benchmark verifies:
196
+
197
+
- independent-oracle equivalence
198
+
- stable canonical checksums
199
+
- duplicate idempotency
200
+
- cross-batch conflict rejection
201
+
- stale-revision protection
202
+
- rollback on injected failure
203
+
- process interruption before commit
204
+
- database integrity after recovery
205
+
- deterministic rerun to the expected state
206
+
207
+
Process interruption testing is not presented as exhaustive OS- or power-loss testing.
208
+
209
+
## Tests and CI
210
+
211
+
Local checks:
115
212
116
213
```bash
117
214
ruff check .
118
215
ruff format --check .
119
216
pytest -q
120
-
121
-
tmp_dir="$(mktemp -d)"
122
-
fdp run-all --source synthetic --seed 20270916 --rows 1000 \
123
-
--output-dir "$tmp_dir/output"
124
217
```
125
218
126
-
Tests use temporary directories and make no live network calls.
219
+
The GitHub Actions workflow runs correctness checks on:
220
+
221
+
- Python 3.11
222
+
- Python 3.12
127
223
128
224
## Docker
129
225
130
-
The image installs the package, runs as a non-root user, and defaults to the deterministic
131
-
offline 1,000-row flow:
226
+
The image installs the package, runs as a non-root user, and defaults to the deterministic offline 1,000-row flow:
132
227
133
228
```bash
134
-
docker build -t finance-data-pipelines:phase3d.
229
+
docker build -t finance-data-pipelines:phase3e.
135
230
docker run --rm -v "$PWD/docker-output:/work/output" \
136
-
finance-data-pipelines:phase3d
231
+
finance-data-pipelines:phase3e
137
232
```
138
233
139
-
Docker support is part of the baseline, but a particular release should be called verified
140
-
only when build and run commands have actually executed in that release environment.
141
-
142
-
## Current maturity and limitations
143
-
144
-
- Three strategies with shared logical semantics; physical storage trade-offs differ.
145
-
- Single-process, single-writer SQLite baseline.
146
-
- Synthetic correctness fixture only; no market-behavior claims.
147
-
- No concurrency or distributed-system evaluation.
148
-
- Benchmark results are machine-specific and do not establish universal rankings.
149
-
- Deterministic exception injection and process-interruption tests cover transactional recovery;
150
-
they are not a claim that every OS/power-loss mode has been tested.
234
+
Docker support is part of the repository, but a release should only be called Docker-verified when those commands have actually been executed in that release environment.
151
235
152
236
## Project layout
153
237
154
238
```text
155
239
src/fdp/
156
240
cli.py installed Click interface
157
-
pipeline.py ordinary Python orchestration
241
+
pipeline.py orchestration
158
242
extract.py deterministic synthetic generator
159
243
transform.py strict normalization
160
-
validation.py loader-side validation and revision resolution
244
+
validation.py validation and revision resolution
161
245
oracle.py independent current-state oracle
162
246
load.py atomic SQLite full replacement
163
-
encoding.py canonical binary checksum encoding
247
+
encoding.py canonical checksum encoding
164
248
manifest.py deterministic JSON manifests
165
249
parquet_io.py frozen Parquet schema
166
-
strategies.py alternative SQLite implementations with shared logical semantics
167
-
scripts/ benchmark.py and analyze_benchmark.py
168
-
tests/ isolated correctness and CLI tests
250
+
strategies.py benchmark loading strategies
251
+
252
+
scripts/
253
+
benchmark.py
254
+
analyze_benchmark.py
255
+
256
+
tests/
257
+
correctness, CLI, strategy, and interruption tests
258
+
259
+
docs/phase3e/
260
+
corrected execution report
261
+
benchmark charts
262
+
169
263
.github/workflows/ci.yml
170
264
```
265
+
266
+
## Reproducibility
267
+
268
+
The Phase 3E benchmark uses:
269
+
270
+
- seed `20270916`
271
+
- deterministic synthetic data
272
+
- fixed SQLite settings
273
+
- rotated strategy order
274
+
- five measured repetitions
275
+
- two query warm-ups
276
+
- 30 timed samples per query
277
+
278
+
The 1m correctness-only `repetition == 0` run is preserved for correctness evidence but excluded from performance aggregation.
279
+
280
+
## Scope and limitations
281
+
282
+
This repository is a controlled systems benchmark, not a production trading system or a market-analysis project.
283
+
284
+
Current limitations:
285
+
286
+
- single-machine evidence
287
+
- synthetic data
288
+
- single-process execution
289
+
- single-writer SQLite baseline
290
+
- no distributed-system evaluation
291
+
- no concurrent-writer benchmark
292
+
- no universal performance ranking claim
293
+
- no publication or research-novelty claim
294
+
295
+
## Phase 3E repository state
296
+
297
+
- Phase 3E implementation commit: `bcfe625`
298
+
- Merge commit on `main`: `69313bd`
299
+
- Corrected execution report commit: `ca3a3c1`
300
+
- Pull request: `#2`
301
+
302
+
The canonical source is the merged Git repository.
0 commit comments