Skip to content
PostDeep Learning / 00-Basic

Activation Function

2024-04-06
Back to Blog

Activation Function

Requirement**

  • 增加非线性表达 使得神经网络可以拟合任意函数
  • 连续可导的函数 可以使用梯度下降法进行参数更新
  • 定义域是 R 可以映射所有实数
  • 单调递增的函数 不改变输入的响应状态

saturation function

饱和函数

+ Def

x  f(x)0
  • 导致梯度消失
    • 参数不会被更新
  • Sigmoid
  • Tanh

Unsaturated function

非饱和函数

  • Rectified Linear Unit 修正线性单元 RELU
    • 解决梯度消失问题
  • RELU Leaky ReLU, Parametric ReLU, ...

Sigmoid

σ=11+ey   σ=σ(1σ)

[!failure]+ Cons

  • 非零均值函数

    • 导致参数同时(正向/反向)更新,不利于收敛
  • 导数最大值 14

    • 导致每层梯度被动缩小 4 倍
      • 导致开始的几层梯度几乎不变
        • 就是梯度消失现象 gradient vanishing problem

-

Sigmoid

tanh

[!abstract]+

tanh=1ey1+ey   tanh=1tanh2

[!success]+ Pros

  • 零均值函数
    • 比 Sigmoid 更快收敛

[!failure]+ Cons

  • 饱和函数
    • 梯度消失

-

tanh

ReLU (Rectified Linear Unit)

[!abstract]+

ReLU(x)={x,x>00,x0

[!success]+ Pros

  • 非零均值函数
    • 收敛速度快
  • 非饱和函数
    • 避免梯度消失问题