Skip to content
Kudos AI

Statistical Learning Toolkit

Least squares, logistic regression, ridge and lasso, and k-fold cross-validation implemented from their estimating equations and checked against scikit-learn.

Python (NumPy, scikit-learn for validation)active

A compact library implementing the core of statistical learning directly from the mathematics: the normal equations for least squares, iteratively reweighted least squares for logistic regression, the closed-form ridge solution, and coordinate descent for the lasso. Each estimator is validated numerically against its scikit-learn counterpart, so a reader can see that a derivation carried out on paper produces the same coefficients as the production implementation. On top of the estimators sits a resampling layer - k-fold and leave-one-out cross-validation - used to draw the bias-variance curve empirically: as model flexibility rises, training error falls monotonically while validation error turns upward, and the toolkit plots exactly where that turn happens. Implementation is in progress and no source repository has been published yet.

Highlights

  • Normal equations, IRLS, closed-form ridge, and coordinate-descent lasso written from the estimating equations
  • Every estimator checked coefficient-by-coefficient against scikit-learn
  • k-fold and leave-one-out cross-validation implemented over a shared splitting interface
  • Bias-variance curve produced empirically, showing where validation error turns upward

Related articles

7 min readStatistical Learning Foundations

What Is Statistical Learning?

The setup behind every predictive model: estimating an unknown function f from data, the split between reducible and irreducible error, and why prediction and inference pull in different directions.

StatisticsMachine LearningMathematics
7 min readSupervised Learning

Linear Regression from First Principles

Derive the least-squares coefficients by differentiating the residual sum of squares, then work a complete five-observation fit by hand: coefficients, fitted values, residuals, RSS, and R-squared, each verified numerically.

StatisticsMachine LearningMathematics
7 min readSupervised Learning

Logistic Regression and Classification

Why a straight line cannot model a probability, how the logistic function fixes it, and what the coefficients mean in log-odds, with a gradient-ascent step and a converged fit computed and checked numerically.

StatisticsMachine LearningOptimization
7 min readSupervised Learning

Regularization: Ridge and Lasso

Adding a penalty on coefficient size to trade a little bias for a large reduction in variance, and why the L1 penalty sets coefficients exactly to zero while L2 only shrinks them, with both fitted numerically.

StatisticsMachine LearningOptimization
7 min readStatistical Learning Foundations

Cross-Validation and Resampling

Why training error is a biased estimate of test error, and how the validation set, leave-one-out, and k-fold approaches fix it, with a five-fold LOOCV computation worked out observation by observation.

StatisticsMachine Learning
7 min readStatistical Learning Foundations

The Bias-Variance Tradeoff

The exact decomposition of expected test error into squared bias, variance, and irreducible noise, demonstrated numerically with a 2,000-run simulation where all three terms are measured separately and shown to add up.

StatisticsMachine LearningMathematics