Skip to content
Kudos AI

Encyclopedia

A concise, cross-linked reference. Each entry connects to related concepts and the articles that go deeper.

Browse by topic

A

B

Backpropagation

The algorithm that computes the gradient of a neural network’s loss with respect to every weight, by applying the chain rule backwards through the network.

Deep LearningOptimizationMathematics

Bagging and Random Forests

Ensemble methods that reduce variance by averaging many models fitted to bootstrap resamples, with random forests additionally decorrelating the trees by restricting the features available at each split.

Machine LearningStatistics

Bayes’ Theorem

A rule for updating the probability of a hypothesis in light of new evidence, by inverting a conditional probability.

ProbabilityStatistics

Bayesian Network

A directed acyclic graph whose nodes are random variables and whose edges express direct influence, with a conditional probability table at each node, that together define a full joint distribution as a product of local factors.

ProbabilityArtificial Intelligence

Belief State

The probability distribution an agent holds over the states it might be in, given everything it has done and perceived - the thing it can act on when the state itself is hidden.

Artificial IntelligenceProbability

Bellman Equation

The self-consistency condition that the utility of a state equals its immediate reward plus the discounted value of the best action available from it, averaged over the outcomes that action cannot control.

Artificial IntelligenceMathematicsOptimization

Bias-Variance Trade-off

The decomposition of a model’s expected prediction error into bias, variance, and irreducible noise, and the tension that reducing one of the first two typically increases the other.

StatisticsMachine Learning

C

Causal Graph

A drawing of assumed cause-and-effect relationships as arrows between variables, used to decide which variables must be adjusted for and which must not - a question the data alone cannot answer.

StatisticsMachine Learning

Condition Number

The ratio of the largest to the smallest curvature of a loss surface, which alone determines how fast gradient descent can converge on it.

OptimizationMathematicsMachine Learning

Confidence Interval

A range computed from data by a procedure that, repeated over many samples, contains the true value a stated proportion of the time. The stated proportion is a property of the procedure, not of any particular interval it produces.

StatisticsProbability

Confounding

A variable that influences both the treatment and the outcome, so that a comparison between the treated and the untreated measures the difference between the groups as well as the effect of the treatment.

StatisticsMachine Learning

Conjugate Prior

A prior chosen so that the posterior belongs to the same family, which turns Bayesian updating into arithmetic on the parameters and makes the prior readable as a number of imagined observations.

ProbabilityMachine LearningMathematics

Constraint Satisfaction Problem

A problem stated as a set of variables, a domain of permitted values for each, and constraints restricting which combinations of values may be taken simultaneously, so that a general solver can reason about its structure without any domain knowledge.

Artificial IntelligenceSearch & Planning

Convolutional Neural Network

A neural network that applies learned filters across an input’s spatial extent, sharing weights so the same pattern is detected wherever it occurs.

Deep LearningComputer Vision

Cross-Entropy

A measure of the difference between two probability distributions, used as the standard loss function for classification.

Information TheoryDeep LearningMachine Learning

Cross-Validation

A resampling method that estimates a model’s test error by repeatedly fitting it on part of the data and evaluating it on the part held out.

StatisticsMachine Learning

D

E

F

G

H

K

L

M

Markov Decision Process

A formal model of sequential decision-making in which outcomes are partly random, defined by states, actions, transition probabilities, and rewards.

Reinforcement LearningProbabilityArtificial Intelligence

Matrix Factorisation

A model that explains a sparse table of interactions as the product of two small matrices, giving every user and every item a short vector of learned traits whose dot product predicts the missing entries.

Machine LearningStatistics

Maximum Likelihood Estimation

A method of fitting a model by choosing the parameter values that make the observed data most probable.

StatisticsProbability

Minimax

A decision rule for two-player zero-sum games in which each player chooses the move maximizing their own worst-case outcome against optimal opposition.

Game TheorySearch & PlanningArtificial Intelligence

Mixed Strategy

A strategy that selects among the available actions according to a probability distribution rather than choosing one deterministically.

Game Theory

Multiple Comparisons

The inflation of false positives that occurs whenever more than one test, metric, segment or stopping point is allowed to produce the headline. Each additional chance raises the probability that something crosses the threshold by luck alone.

StatisticsMachine Learning

Mutual Information

How many bits observing one variable tells you about another. It is zero exactly when the two are independent, it catches dependence of any shape rather than linear dependence only, and nothing computed downstream can increase it.

MathematicsMachine Learning

N

O

P

p-value

The probability of observing data at least as extreme as the data in hand, computed under the assumption that the null hypothesis is true. It measures how unusual the sample would be in a world where the effect is absent, and nothing else.

StatisticsProbability

PAC Learning

A definition of learnability in which an algorithm must return, with high probability, a hypothesis whose true error is within a chosen tolerance - using a number of samples that is bounded in advance rather than discovered afterwards.

Machine LearningMathematics

Partially Observable MDP

A Markov decision process in which the agent cannot observe its state directly, only noisy percepts of it - solved in principle by treating the distribution over states as the state of an ordinary, fully observable MDP.

Artificial IntelligenceProbability

Perplexity

The exponential of a model’s average cross entropy, read as the number of equally likely options it is effectively choosing between at each step.

Information TheoryMachine LearningArtificial Intelligence

Planning Graph

A layered structure alternating literal levels and action levels, annotated with mutual-exclusion links, that bounds in polynomial time what a planning problem can achieve by a given step.

Artificial Intelligence

Precision and Recall

Two rates that split what accuracy hides: precision is the share of predicted positives that are real, and recall is the share of real positives that were found.

Machine LearningStatistics

Pretraining and Fine-Tuning

The two-stage recipe of first training a model on a large generic corpus, then adapting it to a specific task with a much smaller labelled dataset.

Generative AIDeep LearningMachine Learning

Principal Component Analysis

A technique that re-expresses data in new uncorrelated coordinates ordered by how much variance each explains, allowing dimension reduction by keeping only the first few.

Machine LearningStatisticsMathematics

Prisoner’s Dilemma

A game in which each player has a dominant strategy, yet both playing it produces an outcome worse for both than mutual cooperation would have been.

Game Theory

Q

R

S

Self-Attention

A mechanism that lets every position in a sequence attend to every other, computing each output as a weighted sum of values whose weights come from query-key similarity.

Generative AIDeep LearningNatural Language Processing

Softmax

A function that turns a vector of real scores into a probability distribution by exponentiating each score and dividing by the total, preserving their order while making them positive and summing to one.

Machine LearningMathematicsArtificial Intelligence

Spline

A piecewise polynomial joined at chosen points called knots, constrained so that the function and its lower derivatives stay continuous there, giving local flexibility without the wild behaviour of a high-degree polynomial.

StatisticsMachine LearningMathematics

Stationarity

A property of a series whose statistical behaviour does not depend on when you look at it: the mean, the variance and the correlation structure are the same in every window. Almost every classical method assumes it, and most real series lack it.

StatisticsMachine Learning

Statistical Power

The probability that a test rejects the null hypothesis when a specified alternative is true. It is the chance of finding an effect that is genuinely there, and it is fixed by the design before any data are collected.

StatisticsProbability

Stochastic Gradient Descent

Gradient descent in which each step uses the gradient of a small random sample of the data rather than all of it, trading an exact direction for far more steps per unit of compute.

OptimizationMachine LearningMathematics

Support Vector Machine

A classifier that separates classes with the boundary leaving the widest possible margin, determined only by the closest training points.

Machine LearningOptimizationMathematics

T

V