Skip to content

Commit f90f30e

Browse files
authored
Merge pull request #3 from SoheilGtex/docs/phase3e-release-polish
Polish README and add Phase 3E benchmark charts
2 parents ca3a3c1 + fb5701d commit f90f30e

7 files changed

Lines changed: 443 additions & 75 deletions

File tree

‎README.md‎

Lines changed: 207 additions & 75 deletions
Original file line numberDiff line numberDiff line change
@@ -1,19 +1,78 @@
11
# Finance Data Pipelines
22

3-
Finance Data Pipelines is a small, correctness-first time-series ETL and SQLite loading
4-
project. The current release provides a deterministic offline generator, strict validation,
5-
an independent logical-state oracle, and three comparable SQLite strategies: **A: atomic
6-
full replacement**, **B: incremental upsert**, and **C: append-only history with a current
7-
projection**.
3+
[![CI](https://github.com/SoheilGtex/finance-data-pipelines/actions/workflows/ci.yml/badge.svg)](https://github.com/SoheilGtex/finance-data-pipelines/actions/workflows/ci.yml)
84

9-
The repository includes a reproducible benchmark harness. Its results are cohort-specific and
10-
are not a claim of production readiness, universal performance, scalability, or research novelty.
5+
A correctness-first, reproducible time-series ETL and SQLite benchmark for comparing three loading strategies under stateful updates, duplicates, conflicts, stale revisions, and interruption recovery.
6+
7+
## Why this project exists
8+
9+
The project studies a practical systems question:
10+
11+
> How do different SQLite loading strategies behave when they must preserve the same logical state under reruns, updates, duplicates, stale revisions, conflicts, and failures?
12+
13+
The benchmark compares:
14+
15+
- **Strategy A — Atomic full replacement**
16+
- **Strategy B — Transactional incremental upsert**
17+
- **Strategy C — Append-only history with a materialized current-state projection**
18+
19+
All three strategies are validated against the same independent logical-state oracle before performance results are summarized.
20+
21+
## Highlights
22+
23+
- Deterministic synthetic time-series generator
24+
- Strict normalization and revision semantics
25+
- Independent logical-state oracle with stable SHA-256 checksums
26+
- Stateful workload execution
27+
- Cross-batch same-revision conflict rejection
28+
- Stale-revision protection
29+
- Process-interruption and transactional recovery tests
30+
- Five fixed query workloads with warm-up and timed samples
31+
- Reproducible 100k and 1m benchmark cohorts
32+
- GitHub Actions validation on Python 3.11 and 3.12
33+
34+
## Benchmark results
35+
36+
Median initial-load wall time from the corrected Phase 3E analysis:
37+
38+
| Cohort | Strategy A | Strategy B | Strategy C |
39+
|---|---:|---:|---:|
40+
| 100k rows | 3,429.25 ms | 1,733.58 ms | 1,868.53 ms |
41+
| 1m rows | 39,050.38 ms | 21,173.67 ms | 22,539.42 ms |
42+
43+
For the recorded environment and workloads:
44+
45+
- At **100k rows**, Strategy B had **49.4% lower** median initial-load wall time than Strategy A.
46+
- At **1m rows**, Strategy B had **45.8% lower** median initial-load wall time than Strategy A.
47+
- At **1m rows**, Strategy A took about **1.84×** as long as Strategy B.
48+
49+
These are cohort-specific observations, not universal performance claims.
50+
51+
### 100k cohort
52+
53+
![100k ingestion wall time](docs/phase3e/charts/100k/ingestion_wall_time.svg)
54+
55+
![100k query latency](docs/phase3e/charts/100k/query_latency.svg)
56+
57+
![100k storage footprint](docs/phase3e/charts/100k/storage_footprint.svg)
58+
59+
### 1m cohort
60+
61+
![1m ingestion wall time](docs/phase3e/charts/1m/ingestion_wall_time.svg)
62+
63+
![1m query latency](docs/phase3e/charts/1m/query_latency.svg)
64+
65+
![1m storage footprint](docs/phase3e/charts/1m/storage_footprint.svg)
66+
67+
See the corrected execution report:
68+
69+
[`docs/phase3e/phase3e-final-execution-report-corrected.md`](docs/phase3e/phase3e-final-execution-report-corrected.md)
1170

1271
## Requirements
1372

1473
- Python 3.11 or newer
1574
- SQLite supplied by Python
16-
- No API key or network data source
75+
- No API key or network data source required
1776

1877
## Installation
1978

@@ -31,13 +90,11 @@ For development:
3190
python -m pip install -e '.[dev]'
3291
```
3392

34-
`pyproject.toml` is the authoritative dependency definition. `requirements*.txt` are
35-
compatibility wrappers only.
93+
`pyproject.toml` is the authoritative dependency definition. `requirements*.txt` are compatibility wrappers only.
3694

3795
## Deterministic offline quickstart
3896

39-
Choose the output directory explicitly. All generated Parquet, JSON, SQLite, WAL, and SHM
40-
files remain under that directory unless `--db-path` explicitly selects another location.
97+
Choose the output directory explicitly:
4198

4299
```bash
43100
fdp run-all \
@@ -47,8 +104,9 @@ fdp run-all \
47104
--output-dir /tmp/fdp-run
48105
```
49106

50-
The command performs two identical atomic full replacements so exact-rerun idempotency is
51-
verified. It prints correctness fields as JSON and writes:
107+
The command performs two identical atomic full replacements so exact-rerun idempotency is verified.
108+
109+
It writes:
52110

53111
```text
54112
/tmp/fdp-run/
@@ -59,112 +117,186 @@ verified. It prints correctness fields as JSON and writes:
59117
warehouse.db
60118
```
61119

62-
The quickstart produces correctness evidence only. Comparative results are generated separately
63-
by the controlled experiment harness described below.
120+
The quickstart produces correctness evidence only. Comparative benchmark results are generated separately.
64121

65122
## Controlled experiment
66123

67-
The benchmark harness is `scripts/benchmark.py`. It freezes SQLite WAL and
68-
`synchronous=FULL`, runs stateful workload sequences from one fresh database per strategy and
69-
repetition, rotates strategy order, validates every transition against the independent oracle,
70-
and writes `environment.json`, `dataset_manifest.json`, `runs.jsonl`, `query_samples.jsonl`,
71-
`summary.csv`, `correctness.csv`, and `recovery.csv`. Each accepted condition uses five fixed
72-
queries with two warm-ups and 30 timed warm-cache samples. Analysis is regenerated from raw files
73-
with `scripts/analyze_benchmark.py`. A small offline smoke is reproducible with:
124+
The benchmark harness is:
125+
126+
```text
127+
scripts/benchmark.py
128+
```
129+
130+
It:
131+
132+
- freezes SQLite WAL and `synchronous=FULL`
133+
- uses one fresh database per strategy and repetition
134+
- rotates strategy order
135+
- validates every accepted state against the independent oracle
136+
- records environment and dataset metadata
137+
- records state transitions and recovery evidence
138+
- benchmarks five fixed query workloads
139+
- regenerates analysis from raw files only
140+
141+
A small offline smoke run is available with:
74142

75143
```bash
76144
make benchmark-small
77145
```
78146

79-
The full experiment is intentionally separate from ordinary CI. It uses synthetic data only;
80-
no financial or market conclusion is supported. Any report must identify the exact machine,
81-
source hash, workload definitions, and limitations of its cohort.
147+
Analysis is regenerated with:
82148

83-
## Optional configuration file
149+
```bash
150+
python scripts/analyze_benchmark.py <raw_dir> <output_dir>
151+
```
84152

85-
CLI values override a YAML file. Only `source`, `seed`, and `rows` are accepted:
153+
The full benchmark is intentionally separate from ordinary CI.
86154

87-
```bash
88-
fdp run-all --config config.yaml --output-dir /tmp/fdp-run
155+
## Stateful semantics
156+
157+
The logical key is:
158+
159+
```text
160+
(series_id, event_ts_utc)
89161
```
90162

91-
The package never searches the source checkout for configuration. It does not use `.env`, so
92-
no `.env.example` is required.
163+
Exact event identity is:
164+
165+
```text
166+
(series_id, event_ts_utc, revision)
167+
```
168+
169+
Rules:
170+
171+
- exact duplicate events are safe no-ops
172+
- same-revision payload conflicts reject the entire batch
173+
- higher revisions update current state
174+
- lower revisions cannot regress current state
175+
- empty replacement is rejected by default
176+
177+
Timestamps are UTC Unix seconds aligned to one minute. Prices and optional volume use fixed-point integers with scale `10^-6`.
93178

94-
## Logical model
179+
## Query suite
95180

96-
The logical key is `(series_id, event_ts_utc)` and exact event identity is
97-
`(series_id, event_ts_utc, revision)`. Timestamps are UTC Unix seconds aligned to one minute.
98-
Prices and optional volume use fixed-point integers with scale `10^-6`.
181+
Each accepted measured state uses:
99182

100-
Duplicate/update rules:
183+
- **Q1** — point lookup
184+
- **Q2** — range lookup
185+
- **Q3** — period aggregate
186+
- **Q4** — latest-state lookup
187+
- **Q5** — filtered-return query
101188

102-
- an exact duplicate event is a safe no-op;
103-
- different payloads for the same key and revision reject the entire batch, including across
104-
previously committed batches;
105-
- the highest revision is current;
106-
- a lower revision cannot regress the current state;
107-
- an empty replacement is rejected by default.
189+
Each query uses two warm-ups and 30 timed warm-cache samples.
108190

109-
The SQLite current-state table has a composite primary key, required checks, and an
110-
`event_ts_utc` index. Loads use WAL, `synchronous=FULL`, foreign keys, a 5-second busy timeout,
111-
and one explicit transaction. The staging table is validated against the independent oracle
112-
before publication. A pre-commit failure rolls back to the prior committed table.
191+
Result signatures are checked for equality across strategies before timing summaries are used.
113192

114-
## Tests and local CI-equivalent checks
193+
## Correctness and recovery
194+
195+
The benchmark verifies:
196+
197+
- independent-oracle equivalence
198+
- stable canonical checksums
199+
- duplicate idempotency
200+
- cross-batch conflict rejection
201+
- stale-revision protection
202+
- rollback on injected failure
203+
- process interruption before commit
204+
- database integrity after recovery
205+
- deterministic rerun to the expected state
206+
207+
Process interruption testing is not presented as exhaustive OS- or power-loss testing.
208+
209+
## Tests and CI
210+
211+
Local checks:
115212

116213
```bash
117214
ruff check .
118215
ruff format --check .
119216
pytest -q
120-
121-
tmp_dir="$(mktemp -d)"
122-
fdp run-all --source synthetic --seed 20270916 --rows 1000 \
123-
--output-dir "$tmp_dir/output"
124217
```
125218

126-
Tests use temporary directories and make no live network calls.
219+
The GitHub Actions workflow runs correctness checks on:
220+
221+
- Python 3.11
222+
- Python 3.12
127223

128224
## Docker
129225

130-
The image installs the package, runs as a non-root user, and defaults to the deterministic
131-
offline 1,000-row flow:
226+
The image installs the package, runs as a non-root user, and defaults to the deterministic offline 1,000-row flow:
132227

133228
```bash
134-
docker build -t finance-data-pipelines:phase3d .
229+
docker build -t finance-data-pipelines:phase3e .
135230
docker run --rm -v "$PWD/docker-output:/work/output" \
136-
finance-data-pipelines:phase3d
231+
finance-data-pipelines:phase3e
137232
```
138233

139-
Docker support is part of the baseline, but a particular release should be called verified
140-
only when build and run commands have actually executed in that release environment.
141-
142-
## Current maturity and limitations
143-
144-
- Three strategies with shared logical semantics; physical storage trade-offs differ.
145-
- Single-process, single-writer SQLite baseline.
146-
- Synthetic correctness fixture only; no market-behavior claims.
147-
- No concurrency or distributed-system evaluation.
148-
- Benchmark results are machine-specific and do not establish universal rankings.
149-
- Deterministic exception injection and process-interruption tests cover transactional recovery;
150-
they are not a claim that every OS/power-loss mode has been tested.
234+
Docker support is part of the repository, but a release should only be called Docker-verified when those commands have actually been executed in that release environment.
151235

152236
## Project layout
153237

154238
```text
155239
src/fdp/
156240
cli.py installed Click interface
157-
pipeline.py ordinary Python orchestration
241+
pipeline.py orchestration
158242
extract.py deterministic synthetic generator
159243
transform.py strict normalization
160-
validation.py loader-side validation and revision resolution
244+
validation.py validation and revision resolution
161245
oracle.py independent current-state oracle
162246
load.py atomic SQLite full replacement
163-
encoding.py canonical binary checksum encoding
247+
encoding.py canonical checksum encoding
164248
manifest.py deterministic JSON manifests
165249
parquet_io.py frozen Parquet schema
166-
strategies.py alternative SQLite implementations with shared logical semantics
167-
scripts/ benchmark.py and analyze_benchmark.py
168-
tests/ isolated correctness and CLI tests
250+
strategies.py benchmark loading strategies
251+
252+
scripts/
253+
benchmark.py
254+
analyze_benchmark.py
255+
256+
tests/
257+
correctness, CLI, strategy, and interruption tests
258+
259+
docs/phase3e/
260+
corrected execution report
261+
benchmark charts
262+
169263
.github/workflows/ci.yml
170264
```
265+
266+
## Reproducibility
267+
268+
The Phase 3E benchmark uses:
269+
270+
- seed `20270916`
271+
- deterministic synthetic data
272+
- fixed SQLite settings
273+
- rotated strategy order
274+
- five measured repetitions
275+
- two query warm-ups
276+
- 30 timed samples per query
277+
278+
The 1m correctness-only `repetition == 0` run is preserved for correctness evidence but excluded from performance aggregation.
279+
280+
## Scope and limitations
281+
282+
This repository is a controlled systems benchmark, not a production trading system or a market-analysis project.
283+
284+
Current limitations:
285+
286+
- single-machine evidence
287+
- synthetic data
288+
- single-process execution
289+
- single-writer SQLite baseline
290+
- no distributed-system evaluation
291+
- no concurrent-writer benchmark
292+
- no universal performance ranking claim
293+
- no publication or research-novelty claim
294+
295+
## Phase 3E repository state
296+
297+
- Phase 3E implementation commit: `bcfe625`
298+
- Merge commit on `main`: `69313bd`
299+
- Corrected execution report commit: `ca3a3c1`
300+
- Pull request: `#2`
301+
302+
The canonical source is the merged Git repository.

0 commit comments

Comments
 (0)