Skip to content
Kudos AI

Support Vector Machine

A classifier that separates classes with the boundary leaving the widest possible margin, determined only by the closest training points.

Also known as: SVM, Support vector network

Understanding Support Vector Machine

When two classes are separable, infinitely many boundaries achieve zero training error, and they are not equally good. A boundary passing very close to some training points is fragile: a slightly different sample would place those points on the wrong side. The support vector machine makes robustness explicit by choosing the boundary that maximizes the margin, the distance to the nearest point of either class.

The resulting solution has an unusual and useful property. Only the points sitting on the margin boundary, the support vectors, influence where the boundary falls. Every other training point could be moved, or deleted entirely, without changing the fit. This is a sharp contrast with logistic regression, where every observation contributes to the likelihood and therefore to the estimate.

Real data is rarely cleanly separable, so the practical formulation uses a soft margin that permits violations at a cost. A parameter, conventionally C, sets the exchange rate: a large C penalizes violations heavily and yields a narrow margin that fits the training data closely, while a small C tolerates more violations for a wider, more stable boundary. This is the bias-variance trade-off appearing again with a different name on the dial.

Non-linear boundaries come from the kernel trick. The optimization depends on the data only through inner products between pairs of points, so replacing that inner product with a kernel function computes what the inner product would be in some higher-dimensional space, without ever constructing coordinates in that space. The radial basis function kernel corresponds to an infinite-dimensional space and is the usual default.

How to Calculate

minimize ½‖w‖² + C Σᵢ ξᵢ subject to yᵢ(w·xᵢ + b) ≥ 1 − ξᵢ, ξᵢ ≥ 0

where

w, b
the weight vector and offset defining the separating hyperplane
‖w‖
the norm of w; the margin width is 2/‖w‖, so minimizing ‖w‖ maximizes the margin
yᵢ ∈ {−1, +1}
the class label of observation i
ξᵢ
slack: how far observation i intrudes into or across the margin
C
the cost of margin violations, trading margin width against training error

Example of Support Vector Machine

Take two clearly separated clusters of points. Many straight lines separate them, but the SVM selects the one running down the middle of the empty corridor between the classes, as far as possible from both. The points touching the edges of that corridor are the support vectors.

Now consider data arranged as one class in a central blob surrounded by a ring of the other class. No straight line can separate these. A linear SVM fails on this problem in principle, not merely in practice.

An SVM with a radial basis function kernel separates it cleanly. The kernel measures similarity that decays with distance, so nearby points influence one another strongly and distant points barely at all, producing a closed curved boundary around the central blob. No explicit coordinates in the higher-dimensional space are ever computed.

Advantages and Disadvantages

Pros

  • Effective in high-dimensional settings, including when features outnumber observations.
  • The solution depends only on the support vectors, so it is memory-efficient at prediction time.
  • Kernels provide flexible non-linear boundaries within a single, convex optimization.

Cons

  • Training scales poorly with the number of observations, limiting it on very large datasets.
  • Produces no probability estimate natively; calibration requires an extra fitting step.
  • Requires feature scaling and careful tuning of C and any kernel parameters.

Frequently Asked Questions

Why does only a subset of the data determine the boundary?

It falls out of the optimization. Points comfortably on the correct side of the margin satisfy their constraint with room to spare, so they exert no force on the solution. Only points on or inside the margin have active constraints, and those are the support vectors.

What does the C parameter control?

The cost of allowing a point to violate the margin. Large C means violations are expensive, producing a tight boundary that risks overfitting; small C permits violations for a wider, more robust margin. It is chosen by cross-validation.

When should a linear kernel be preferred over RBF?

When the number of features is large relative to the number of observations, as in text classification, the classes are frequently close to linearly separable already, and a linear kernel is faster and less prone to overfitting. RBF is the reasonable default otherwise.

The Bottom Line

A support vector machine picks the widest-margin boundary, depends only on the closest points, and reaches non-linear decision surfaces through kernels rather than explicit feature construction. Its cost is poor scaling in sample size and the absence of native probabilities.