Skip to content
PostMachine Learning / ML-Cheat-Sheet.html

ML-Cheat-Sheet

2024-12-25
Back to Blog

ML-Cheat-Sheet ​

Basic Rules of Differentiation ​

Basic Rules

  • Constant Rule: ddxC=0

  • Power Rule: ddxxn=nxn−1

  • Linear Combination: ddx[af(x)+bg(x)]=af′(x)+bg′(x)

  • Product Rule: ddx[f(x)g(x)]=f′(x)g(x)+f(x)g′(x)

  • Quotient Rule: ddx[f(x)g(x)]=f′(x)g(x)−f(x)g′(x)[g(x)]2

  • Chain Rule: ddxf(g(x))=f′(g(x))g′(x)

  • Exponential:ddxex=ex | |ddxax=axln⁡(a)

  • Logarithmic ddxln⁡(x)=1x || ddxloga⁡(x)=1xln⁡(a)

Linear Regression ​

1. Hypothesis ​

hθ(x)=θTx=θ0+θ1x1+⋯+θnxn

2. Cost Function ​

Mean Squared Error (MSE): J(θ)=12m∑i=1m(hθ(x(i))−y(i))2

3. Optimization ​

  • Gradient Descent: θj:=θj−α1m∑i=1m(hθ(x(i))−y(i))xj(i)
  • Normal Equation: θ=(XTX)−1XTy

Logistic Regression ​

1. Hypothesis ​

hθ(x)=11+e−θTx

  • Prediction Rule:
    • Predict y=1 if hθ(x)≥0.5, otherwise y=0.

2. Cost Function ​

Log Loss: J(θ)=1m∑i=1m[−y(i)log⁡(hθ(x(i)))−(1−y(i))log⁡(1−hθ(x(i)))]

3. Optimization ​

4. Sigmoid Properties ​

  • Output: g(z)∈[0,1]
  • Derivative: g′(z)=g(z)(1−g(z))

Ridge Regression ​

Loss Function ​

Adds L2 regularization to prevent overfitting: J(θ)=12m∑i=1m(hθ(x(i))−y(i))2+λ∑j=1nθj2

  • λ: Regularization parameter. Higher values shrink θj.

Optimization ​

  • Closed-form Solution: θ=(XTX+λI)−1XTy
  • Gradient Descent: θj:=θj−α(1m∑i=1m(hθ(x(i))−y(i))xj(i)+2λθj)

Bayesian Classification ​

Dataset ​

  • T={(x1,y1),(x2,y2),…,(xN,yN)}
  • xi=(x1,…,xn), yi∈{c1,…,cK}

Posterior Probability ​

The probability of class ck given input x: P(y=ck∣x)∝P(y=ck)P(x∣y=ck)

If features are conditionally independent: P(y=ck∣x)∝P(y=ck)∏jP(xj∣y=ck)

SVM ​

Hard SVMHyperplane: H={w|wTx+b=0}Constraint: yi(wTxi+b)≥1 ∀iGoal: min12||w||2 s.t. yi(wtxi+b)≥1Lagrangian:L(w,b,α)=12||w||2−∑iαi(yi(wTxi+b)−1),αi≥0Partial derivative: ∂L∂w=w−∑iαiyixi=0 ∂L∂b=−∑iαiyi=0Solution: ||w||2=(∑iαiyixi)T(∑iαiyixi)=∑i∑jαiαjyiyjxiTxjLagrangian becomes: L=∑iαi−12∑i∑jαiαjyiyjxiTxjs.t. ∑iαiyi=0 and αi≥0∀iWeight vector: w∗=∑iαiyixiBias: b∗=yi−∑iαiyixiTxj

Soft SVMHyperplane: H={w|wTx+b=0}Constraint: yi(wTxi+b)≥1−ξi,ξi≥0,∀iGoal: min12||w||2+C∑i=1nξi,s.t.yi(wTxi+b)≥1−ξi,ξi≥0
Lagrangian: L(w,b,α,ξ)=12||w||2+C∑i=1nξi−∑i=1nαi(yi(wTxi+b)−1+ξi)−∑i=1nμiξi,αi,μi≥0Partial Derivative: ∂L∂w=w−∑i=1nαiyixi=0,∂L∂b=−∑i=1nαiyi=0,∂L∂ξi=C−αi−μi=0Solution: ||w||2=∑i=1n∑j=1nαiαjyiyjxiTxjDual Problem: L=maxα∑i=1nαi−12∑i=1n∑j=1nαiαjyiyjxiTxjs.t. ∑i=1nαiyi=0,0≤αi≤C
Weight vector: w∗=∑i=1nαiyixiBias: b∗=yk−∑i=1nαiyixiTxkfor any 0<αk<CThe reason that ξ disappears: The slack variables ξi disappear in the dual problem because they are implicitly handled through the Lagrange multipliers αi. By taking the derivative of the Lagrangian with respect to ξi, we obtain:∂L∂ξi=C−αi−μi=0 This relationship ensures that αi is bounded by 0≤αi≤C. Consequently, the slack variables αi do not explicitly appear in the dual formulation. Instead, the dual problem balances maximizing the margin and allowing for misclassification through the constraint on αi.

Kernel SVMHyperplane: H={w|wTϕ(x)+b=0}Constraint: yi(wTϕ(xi)+b)≥1−ξi,ξi≥0,∀iGoal: min12||w||2+C∑i=1nξi,s.t.yi(wTϕ(xi)+b)≥1−ξiLagrangian (Dual): L(α)=∑i=1nαi−12∑i=1n∑j=1nαiαjyiyjK(xi,xj)s.t. ∑i=1nαiyi=0,0≤αi≤C,∀iWeight vector: w=∑i=1nαiyiϕ(xi)Decision Function: f(x)=sign(∑i=1nαiyiK(xi,x)+b)Bias: b=yk−∑i=1nαiyiK(xi,xk)∀sup vec 0<αk<CKernel Functions:
Linear: K(xi,xj)=xiTxj
Polynomial: K(xi,xj)=(xiTxj+c)d
Gaussian (RBF): K(xi,xj)=exp⁡(−||xi−xj||22σ2)
Sigmoid: K(xi,xj)=tanh⁡(κxiTxj+c)

MLE and MAP ​

MLE ​

构建似然函数:联合分布 L(θ)=∏i=1nP(Xi|θ)。 取对数简化计算:ln⁡L(θ)=∑i=1nln⁡P(Xi|θ)。 求导并设为 0:ddθln⁡L(θ)=0,解得 θ^MLE。 验证极值:通过二阶导数等方式确保是最大值。

MAP ​

结合先验构建后验概率:P(θ|X)∝P(X|θ)P(θ)。 取对数后验函数:ln⁡P(θ|X)∝ln⁡P(X|θ)+ln⁡P(θ)。 求导并设为 0:ddθln⁡P(θ|X)=0,解得 θ^MAP。 验证极值:确保找到最大值。