Skip to content
PostMachine Learning / Assignment

ML-As-2

2024-10-16
Back to Blog

ML-As-2 ​

Point Estimation ​

The Poisson distribution is a useful discrete distribution which can be used to model the number of occurrences of something per unit time. For example, in networking, packet arrival density is often modeled with the Poisson distribution. If X is Poisson distributed, i.e., X Poisson(λ), its probability mass function takes the following form:

P(X|λ)=λXe−λX!

It can be shown that E(X)=λ. Assume now we have n i.i.d. data points from Poisson(λ):D=X1,…,Xn . (For the purpose of this problem, you can only use the knowledge about the Poisson and Gamma distributions provided in this problem.)

(a) ​

Show that the sample mean λ^=1nΣi=1nXi is the maximum likelihood estimate (MLE) of λ and it is unbiased (Eλ^=λ).

Finding the MLE

L(λ)=∏i=1nP(Xi|λ)=∏i=1nλXie−λXi!ln⁡L(λ)=∑i=1n(Xiln⁡λ−λ−ln⁡(Xi!))ddλln⁡L(λ)=∑i=1n(Xiλ−1)=0∑i=1nXi=nλλ^=1n∑i=1nXi

Unbiasedness

E(λ^)=E(1n∑i=1nXi)

Since  Xi  are i.i.d., we can take the expectation inside the sum:

E(λ^)=1n∑i=1nE(Xi)=1n∑i=1nλ=nλn=λ

Therefore,  E(λ^)=λ, confirming that λ^  is an unbiased estimator of  λ . of λ^

(b) ​

Now let's be Bayesian and put a prior distribution over λ. Assuming that λ follows a Gamma distribution with the parameters (α,β) , its probability density function:

p(λ|α,β)=βαΓ(α)λα−1e−βλ

Where Γ(α)=(α−1)! (here we assume α is a positive integer). Compute the posterior distribution λ.

P(θ|λ)=P(X|λ)P(λ|α,β)P(X)P(θ|λ)∝P(X|λ)P(λ|α,β)=λXe−λX!βαΓ(α)λα−1e−βλP(θ|λ)∝λX+α−1e−λ(β+1)

Let α′=X+α , β′=β+1 Then the distribution is still a Gamma distribution

(c) ​

Derive an analytic expression for the maximum a posterior (MAP) of λ under Gamma(α,β) prior.

MAP(λ)=∏i=1nP(Xi|λ)=∏i=1nP(Xi|λ)P(λ)P(X)∝∏i=1nP(Xi|λ)P(λ)∏i=1nP(Xi|λ)P(λ)∝log∏i=1nP(Xi|λ)P(λ)logP(λ∣X)∝log(∏ni=1P(Xi​∣λ)P(λ))∝∑i=1n​logP(Xi​∣λ)+logP(λ)∑i=1n​logP(Xi​∣λ)+logP(λ)

Prior Distribution P(λ)

P(λ|α,β)=βαΓ(α)λα−1e−βλlogP(λ|α,β)∝(α−1)logλ−βλ

Likelihood function P(Xi|λ)

P(X|λ)=λXe−λX!​logP(Xi​∣λ)∝Xilogλ−λ∑i=1n​logP(Xi​∣λ)+logP(λ)
MAP(λ)∝∑i=1nXilogλ−nλ+(α−1)logλ−βλ=logλ(∑i=1nXi+α−1)−λ(n+β)ddλlogP(λ|X)=∑i=1nXi+α−1λ−(n+β)=0λMAP=∑i=1nXi+α−1n+β

Source of Error: Part 1 ​

(a) ​

The bias of an estimator is defined as E[μ^]−μ

The bias is 1−μ

The variance of an estimator is defined as Var(μ^)=E[(μ^−E[μ^2])]

∴Var(μ^)=0

This is not a good estimator, since the bias is large when the true value of μ is not 1. Usually we don’t have any information about the true value of μ, so it is unreasonable to assume it is equal to 1.

(b) ​

E(μ^)=μ the bias is 0. This is an unbiased estimator. The variance of this estimator is Var(μ^)=Var(y1)=1

This is not a good estimator since its variability does not decrease with the sample size.

(c) ​

−2∑i=1n(yi−μ)+2λμ=0μ^=1n+λ∑iyi=nn+λy¯E[μ^]=1n+λE[∑iyi]=nn+λμ

Bias of the estimator :

bias=−λμn+λ

Variance of the estimator :

Var(μ^)=Var(1n+λ∑iyi)=1(n+λ)2∑iVar(yi)=n(n+λ)2σ2

Source of Error: Part 2 ​

(a) ​

(b) ​

The error is equal to 0.

Because p(X|Y=0) and p(X|Y=1) do not overlap.

Just check whether it is in the interval [-4,-1] or in the interval [1,4]

(c) ​

P[error]=P[x∈[0,1]]×P[error|x∈[0,1]]=(P[x∈[0,1]|y=0]P[y=0]+P[x∈[0,1]|y=1]P[y=1])×P[error|x∈[0,1]]=(14×12+14×12)×12=18

(d) ​

  • E[X|Y=0]=−2.5 and Var[X|Y=0]=34 (using the variance formula for the uniform distribution),
  • E[X|Y=1]=2.5 and Var[X|Y=1]=34.

Since we are approximating p(X|Y) using a normal distribution, we have:

  • p^(X|Y=0)=N(−2.5,0.75),
  • p^(X|Y=1)=N(2.5,0.75).

Using these, for x<0, we find p^(X|Y=0)>p^(X|Y=1), and for x>0, p^(X|Y=0)<p^(X|Y=1). Therefore, the classifier will make no error in classifying new points.

(e) ​

Given a finite amount of data, we will not learn the mean and variance of p(X|Y) perfectly. Therefore, the classifier's error will increase due to the limited data. In this scenario, we would have both bias and error in our model.

Gaussian (Naïve) Bayes and Logistic Regression ​

No, the new P(Y|X) is no longer the form used by logistic regression.

P(Y=1|X)=P(Y=1)P(X|Y=1)P(Y=1)P(X|Y=1)+P(Y=0)P(X|Y=0)=11+P(Y=0)P(X|Y=0)P(Y=1)P(X|Y=1)=11+exp⁡(ln⁡P(Y=0)P(X|Y=0)P(Y=1)P(X|Y=1))=11+exp⁡(ln⁡1−ππ+ln⁡P(X|Y=0)P(X|Y=1))=11+exp⁡(ln⁡1−ππ+∑iln⁡P(Xi|Y=0)P(Xi|Y=1))

The log ratio of class-conditional probabilities:

∑iln⁡P(Xi|Y=0)P(Xi|Y=1)=∑iln⁡12πσi0exp⁡(−(Xi−μi0)22σi02)12πσi1exp⁡(−(Xi−μi1)22σi12)

Simplifies to:

=∑iln⁡σi1σi0+∑i((Xi−μi1)22σi12−(Xi−μi0)22σi02)=∑iln⁡σi1σi0+∑iσi02−σi122σi02σi12Xi2+2(μi0σi12−μi1σi02σi02σi12)Xi+μi12σi02−μi02σi122σi02σi12

Probability of P(Y=1|X):

P(Y=1|X)=11+exp⁡(ln⁡1−ππ+∑iln⁡P(Xi|Y=0)P(Xi|Y=1))

Simplifies to:

P(Y=1|X)=11+exp⁡(w0+∑iwiXi+∑iviXi2)w0=ln⁡1−ππ+∑i(ln⁡σi1σi0+μi12σi02−μi02σi122σi02σi12)wi=μi0σi12−μi1σi02σi02σi12vi=σi02−σi122σi02σi12