The project originated from a prediction problem posed by Dr. Jamshid Saeidian during a Numerical Analysis class at Kharazmi University: whether a student's final Numerical Analysis grade could be predicted from early-semester information. The extension documented here was developed independently. No authorized real student dataset is currently available, so Phase 1 uses controlled synthetic observations and makes no claims about actual students.
The research motivation is numerical rather than merely predictive. A transparent linear model provides a small setting in which the design matrix, least-squares solution, known generating parameters, and numerical error can all be inspected before later work on stability, inverse problems, regularization, and uncertainty.
For observation
The intercept is therefore included exactly once. The parameter vector is
The data-generating model is
where noise_std. The implemented coefficient vector is
The five features are sampled independently using a local NumPy generator. Prior GPA and prerequisite grade lie in
The random generator is numpy.random.default_rng(seed), so the same configuration and seed produce identical arrays without relying on global random state.
The baseline solves
using numpy.linalg.lstsq. The implementation deliberately does not manually implement normal equations, QR, SVD, pseudoinverse analysis, conditioning experiments, or regularization in Phase 1.
Because the synthetic generator exposes
For test targets
root mean squared error
and
The implementation rejects constant targets for which the
A zero ground-truth coefficient vector is rejected for this relative metric.
The reproducible baseline uses the first 75% of the deterministic sample as training data and the remaining 25% as test data. This split verifies data generation, fitting, prediction, metrics, and ground-truth recovery. It is not a final statistically rigorous validation protocol, and no educational or causal conclusion can be drawn from the synthetic experiment.
Later phases may study numerical solver comparisons, conditioning, perturbation and stability, regularization, uncertainty quantification, robustness, feature ablation, temporal forecasting, and nonlinear dynamics. Those topics are intentionally outside Phase 1.