Aller au contenu
Kudos AI

Statistical Learning Toolkit

Moindres carrés, régression logistique, ridge et lasso, et validation croisée k-fold, implémentés depuis leurs équations d’estimation et vérifiés face à scikit-learn.

Python (NumPy, scikit-learn for validation)actif

Une bibliothèque compacte implémentant le cœur de l’apprentissage statistique directement à partir des mathématiques : équations normales pour les moindres carrés, moindres carrés repondérés itérativement pour la régression logistique, solution en forme close pour la ridge, et descente par coordonnées pour le lasso. Chaque estimateur est validé numériquement face à son équivalent scikit-learn, de sorte qu’un lecteur constate qu’une dérivation menée sur papier produit les mêmes coefficients que l’implémentation de production. Au-dessus des estimateurs se trouve une couche de rééchantillonnage - validation croisée k-fold et leave-one-out - utilisée pour tracer empiriquement la courbe biais-variance : à mesure que la flexibilité du modèle augmente, l’erreur d’entraînement décroît de façon monotone tandis que l’erreur de validation finit par remonter, et la bibliothèque trace exactement où ce retournement se produit. L’implémentation est en cours et aucun dépôt source n’a encore été publié.

Points forts

  • Équations normales, IRLS, ridge en forme close et lasso par descente de coordonnées, écrits depuis les équations d’estimation
  • Chaque estimateur vérifié coefficient par coefficient face à scikit-learn
  • Validation croisée k-fold et leave-one-out implémentées sur une interface de découpage commune
  • Courbe biais-variance produite empiriquement, montrant où l’erreur de validation remonte

Articles associés

7 min de lectureStatistical Learning Foundations

What Is Statistical Learning?

The setup behind every predictive model: estimating an unknown function f from data, the split between reducible and irreducible error, and why prediction and inference pull in different directions.

StatistiqueApprentissage automatiqueMathématiques
7 min de lectureSupervised Learning

Linear Regression from First Principles

Derive the least-squares coefficients by differentiating the residual sum of squares, then work a complete five-observation fit by hand: coefficients, fitted values, residuals, RSS, and R-squared, each verified numerically.

StatistiqueApprentissage automatiqueMathématiques
7 min de lectureSupervised Learning

Logistic Regression and Classification

Why a straight line cannot model a probability, how the logistic function fixes it, and what the coefficients mean in log-odds, with a gradient-ascent step and a converged fit computed and checked numerically.

StatistiqueApprentissage automatiqueOptimisation
7 min de lectureSupervised Learning

Regularization: Ridge and Lasso

Adding a penalty on coefficient size to trade a little bias for a large reduction in variance, and why the L1 penalty sets coefficients exactly to zero while L2 only shrinks them, with both fitted numerically.

StatistiqueApprentissage automatiqueOptimisation
7 min de lectureStatistical Learning Foundations

Cross-Validation and Resampling

Why training error is a biased estimate of test error, and how the validation set, leave-one-out, and k-fold approaches fix it, with a five-fold LOOCV computation worked out observation by observation.

StatistiqueApprentissage automatique
7 min de lectureStatistical Learning Foundations

The Bias-Variance Tradeoff

The exact decomposition of expected test error into squared bias, variance, and irreducible noise, demonstrated numerically with a 2,000-run simulation where all three terms are measured separately and shown to add up.

StatistiqueApprentissage automatiqueMathématiques