This project originated from a prediction problem posed by Dr. Jamshid Saeidian during a Numerical Analysis class at Kharazmi University:
Can a student's final Numerical Analysis grade be predicted from early-semester information?
The subsequent mathematical formulation, experiments, software implementation, and analysis were developed independently. This wording does not imply supervision, endorsement, coauthorship, or research collaboration.
Can a student's final Numerical Analysis grade be predicted from early-semester information? This repository studies the question as a controlled numerical-modeling problem rather than as an empirical student-record analysis.
The project originated from the classroom prediction question described above. The subsequent mathematical formulation, experiments, software implementation, and analysis were developed independently.
Scientific Phases 1–5 are complete. Phase 6 packages the accepted work as a reproducible research-software release with a technical report, provenance metadata, artifact manifests, automated verification, and CI. Phase 1 provides the synthetic least-squares baseline; Phase 2 provides explicit numerical least-squares solvers; Phase 3 provides conditioning and perturbation analysis; Phase 4 provides Tikhonov/Ridge regularization, lambda selection, bias–variance–stability analysis, and classical OLS uncertainty quantification; Phase 5 provides early-semester held-out forecasting, ablation, and contamination robustness.
The model is
where the design matrix contains an explicit intercept and five early-semester predictors. The baseline estimates parameters by minimizing numpy.linalg.lstsq.
- deterministic synthetic modeling with known ground-truth coefficients;
- ordinary least-squares estimation and prediction;
- parameter-recovery evaluation;
- reproducible train/test experiments;
- MAE, RMSE,
$R^2$ , and relative coefficient-error metrics.
Python 3.11 or newer is required. From the repository root:
python -m venv .venv
. .venv/bin/activate
python -m pip install -e ".[dev]"
python scripts/run_baseline.py
pytest
ruff check .The complete Phase 1–5 reproduction sequence is:
python scripts/run_baseline.py
python scripts/run_solver_comparison.py
python scripts/run_conditioning_study.py
python scripts/run_perturbation_study.py
python scripts/run_regularization_study.py
python scripts/run_uncertainty_study.py
python scripts/run_forecasting_study.pyThe Phase 6 packaging checks are:
python scripts/capture_environment.py
python scripts/verify_repository.pyFor exact release reproduction, use the recorded Python version and requirements-release.txt, then install the project with python -m pip install -e . --no-deps. Normal CI intentionally uses the flexible pyproject.toml ranges for compatibility testing. Deliberately ill-conditioned experiments can materially amplify cross-environment floating-point differences at coefficient level; exact values quoted here refer to the recorded release environment, while the packaged CSV files remain the canonical release record.
The project progresses from least squares to solver stability, rank deficiency, conditioning, perturbation sensitivity, Tikhonov regularization, bias–variance, uncertainty, held-out forecasting, ablation, and robustness.
Source code is under src/numerical_learning/; experiment entry points are under scripts/; tests are under tests/; mathematical and release documentation is under docs/; generated tables, figures, and reproducibility metadata are under results/.
Run ruff check ., pytest, and python scripts/verify_repository.py. GitHub Actions repeats installation, Ruff, pytest, and repository verification. The experiment manifest, environment snapshot, and SHA-256 artifact manifest provide release provenance.
The observations are mathematically controlled synthetic data, not records from Kharazmi University or any other students. Therefore, the results support software and numerical verification only; they do not support educational conclusions, causal inference, or claims of statistical significance.
Completed analyses include the reproducible synthetic least-squares baseline, solver comparison, conditioning, perturbation, Ridge/Tikhonov regularization, lambda selection, bias–variance/stability analysis, classical OLS uncertainty verification, early-semester held-out forecasting, feature ablation, and specified response/feature contamination robustness. Still future work includes bootstrap uncertainty, conformal prediction, nonlinear models, and state-space/Kalman methods.
Phase 2 extends the baseline with explicit Normal Equations, reduced unpivoted QR, an SVD/pseudoinverse solver, and the existing numpy.linalg.lstsq reference. It adds controlled well-conditioned, Phase 1, near-rank-deficient, and exactly rank-deficient cases, together with solver diagnostics and a small raw-versus-standardized conditioning diagnostic. Failures are retained as results rather than hidden. The implementation does not claim that any solver is universally superior; conclusions are limited to the executed synthetic cases.
In the exact rank-deficient case, SVD and numpy.linalg.lstsq produce minimum-norm solutions with nearly zero residuals, while the chosen generating coefficient vector is a different valid representation. This demonstrates that excellent fitted values do not imply unique parameter recovery: rank-deficient designs admit multiple coefficient vectors with the same observations.
Phase 3 studies a compact deterministic sweep from well-conditioned to near-rank-deficient designs in exact and modest-noise modes. It measures condition numbers, Gram-matrix conditioning, solver-dependent coefficient and fitted-value behavior, response perturbation amplification, design-matrix sensitivity, and raw-versus-standardized Phase 1 coordinates. The conditioning-sweep RMSE is an in-sample fitted-value error because the same observations are used for fitting and evaluation; Phase 3 does not test generalization or out-of-sample forecasting. All observations remain synthetic; results are numerical-analysis evidence, not educational findings.
The executed sweep used nine values of lstsq had coefficient errors near
For every nonzero lstsq. Using the spectral matrix 2-norm, the corresponding design-perturbation study increased from approximately
Phase 4 uses centered SVD Ridge with an unpenalized intercept, training-only standardization, a dimension-aware
The bias–variance experiment used 30 deterministic Gaussian-noise replicates. For fixed
The uncertainty experiment selects existing design rows by measured leverage: median (
Phase 5 uses a new controlled synthetic generator with four nested information checkpoints: pre-semester, early semester, mid semester, and later semester. The 20-seed clean study reuses each seed's train/validation/test partition across checkpoints and evaluates OLS, validation-selected Ridge, and transparent Huber IRLS. These checkpoints represent feature availability, not time-series dependence.
Mean held-out RMSE was:
| Checkpoint | OLS | Ridge | Huber |
|---|---|---|---|
| Pre-semester | 1.839638 | 1.841547 | 1.839715 |
| Early semester | 1.568967 | 1.558670 | 1.568399 |
| Mid semester | 1.384252 | 1.379206 | 1.385090 |
| Later semester | 1.301504 | 1.296868 | 1.297964 |
Ridge had the lowest mean RMSE at the early-, mid-, and later-semester checkpoints. At the pre-semester checkpoint, OLS had the lowest mean RMSE by a very small numerical margin; all three estimators were close. The adjacent Ridge RMSE reductions were 0.283, 0.179, and 0.082, respectively. These are finite-sample synthetic results, not significance claims.
The strongest replicated Ridge ablation result was removal of the combined assessment group, which increased RMSE by 0.195 on average. This is a predictive contribution measure conditional on the other features, not a causal or educational-importance claim. In the 20-seed robustness study, 20% response contamination of training targets produced mean clean-test RMSE degradation of 0.315 for OLS, 0.156 for Ridge, and 0.072 for Huber, with degradation standard deviations of 0.177, 0.107, and 0.085, respectively. Under 20% feature contamination applied to raw training predictors before preprocessing, mean degradation was 0.089 for OLS, 0.103 for Ridge, and 0.082 for Huber, with standard deviations of 0.081, 0.067, and 0.074. Huber was more resistant to the specified response contamination, but it did not universally dominate under feature contamination; it does not automatically solve high-leverage predictor corruption.
All Phase 5 results remain controlled synthetic evidence. They do not establish actual student forecasting performance, empirical realism, causality, deployment readiness, or universal superiority of any model.
See docs/problem_formulation.md and docs/mathematical_notes.md for the formal model, derivations, assumptions, and numerical interpretation.
The main technical report, reproducibility guide, and results index provide detailed interpretation, exact commands, and artifact provenance. The project summary is a concise reviewer-oriented overview. The research-integrity statement records the synthetic-data and claim boundaries, and CITATION.cff contains the software citation metadata.