本章目录
第三章 神经网络基础
3.1 引言:从逻辑回归到神经网络
在上一章中,我们深入探讨了逻辑回归,这是一种强大的线性分类器。然而,现实世界中的许多问题都具有非线性特性,这就需要我们探索更复杂的模型。本章将介绍神经网络,这是一类受生物神经系统启发的强大算法,能够学习复杂的非线性关系。
神经网络不仅在理论上很有吸引力,在实践中也已经在各个领域取得了巨大成功,从计算机视觉到自然语言处理,再到游戏AI。下面从基本组件出发,探索神经网络和深度学习的基础知识。
3.2 神经网络的历史与发展
历史小知识: 神经网络的概念可以追溯到20世纪40年代。1943年,沃伦·麦卡洛克(Warren McCulloch)和沃尔特·皮茨(Walter Pitts)提出了人工神经元的数学模型。反向传播思想在更早的研究中已有发展;1986年,David Rumelhart、Geoffrey Hinton和Ronald Williams的工作使其在多层神经网络训练中得到广泛关注。
神经网络的发展大致可以分为以下几个阶段:
- 感知器时代(1950s-1960s):Frank Rosenblatt在1958年发明了感知器,这是最早的神经网络模型之一。
- 研究低潮(1970s):Marvin Minsky和Seymour Papert的著作《感知器》指出了单层感知器的局限性;技术能力、研究预期和资助环境等多重因素使神经网络研究一度降温。
- 反向传播与复兴(1980s-1990s):反向传播的推广和应用重新点燃了神经网络研究的热情。
- 深度学习革命(2000s-至今):得益于大数据和强大的计算能力,深度学习在各个领域取得了突破性进展。
3.3 从生物到人工神经元
人工神经网络的灵感来源于生物神经系统。让我们简单比较一下生物神经元和人工神经元:
生物神经元:
- 树突接收输入信号
- 细胞体处理信号
- 轴突传输输出信号
人工神经元:
- 输入特征 (x1, x2, ..., xn)
- 权重 (w1, w2, ..., wn) 和偏置 (b)
- 激活函数 (如sigmoid, ReLU)
人工神经元的数学表示:
y = f(Σ(wi * xi) + b)
其中 f 是激活函数,Σ(wi * xi) + b 是加权和加偏置。
3.4 单神经元分类器:从逻辑回归到神经网络
经典感知器使用阈值激活和感知器学习规则;上一章的逻辑回归则使用sigmoid函数和对数损失。下面的Keras代码实现的是单神经元逻辑分类器,它与感知器结构相似,但训练目标并不相同。
让我们用TensorFlow/Keras实现一个简单的感知器:
import tensorflow as tf
import numpy as np
from tensorflow.keras.models import Sequential
from tensorflow.keras.layers import Dense
# 创建一个单神经元逻辑分类模型
model = Sequential([
Dense(1, activation='sigmoid', input_shape=(2,))
])
# 编译模型
model.compile(optimizer='sgd', loss='binary_crossentropy', metrics=['accuracy'])
# 准备一些示例数据(例如,实现AND逻辑门)
X = np.array([[0, 0], [0, 1], [1, 0], [1, 1]], dtype=np.float32)
y = np.array([0, 0, 0, 1], dtype=np.float32)
# 训练模型
model.fit(X, y, epochs=1000, verbose=0)
# 测试模型
print(model.predict(X))
这个简单的例子展示了如何使用TensorFlow/Keras创建和训练单神经元逻辑分类器。在接下来的部分,我们将加入隐藏层和非线性激活函数,构建更复杂的神经网络。
在下一节中,我们将探讨多层感知器(MLP),这是向更复杂神经网络结构迈出的第一步。
3.5 计算能力的飞跃:神经网络的推动力
神经网络的概念虽然早在20世纪40年代就被提出,但直到近年来才真正得到广泛应用。这种戏剧性的发展很大程度上归功于计算能力的飞跃。让我们回顾一下推动神经网络发展的关键计算里程碑:
- CPU的进步(1970s-2000s) * 摩尔定律的体现:处理器速度和晶体管数量的指数级增长。 * 影响:使得更复杂的神经网络模型成为可能,但训练大型网络仍然耗时。
- GPU计算的兴起(2000s初) * NVIDIA在2006-2007年推出CUDA,使得GPU可以用于通用计算。 * 影响:显著加速了神经网络的训练过程,特别是在处理图像数据时。
- 分布式计算和大数据(2000s中期) * Hadoop(2006)和Spark(2014)等框架的出现。 * 影响:使得在大规模数据集上训练神经网络成为可能。
- 云计算的普及(2010s) * Amazon EC2(2006)、Google Cloud Platform(2008)、Microsoft Azure(2010)的推出。 * 影响:降低了进行大规模机器学习实验的硬件门槛。
- 专用AI硬件(2016年至今) * Google的TPU(Tensor Processing Unit)、NVIDIA的Tesla V100等。 * 影响:进一步加速了深度学习模型的训练和推理过程。
- 开源深度学习框架(2015年至今) * TensorFlow(2015)、PyTorch(2016)等的出现。 * 影响:大大降低了开发和部署神经网络的技术门槛。
技术小知识: 以图像识别为例,2012年的AlexNet使用两块GTX 580 GPU训练了数天。2015年的ResNet显著提高了网络深度和识别准确率,但训练成本仍然较高。具体耗时取决于模型、数据、硬件和软件实现,不能脱离实验条件直接比较。
这些计算能力的进步不仅加速了神经网络的训练过程,还使得更深、更复杂的网络结构成为可能。例如,从2012年8层的AlexNet,到2015年152层的ResNet,再到如今拥有数十亿参数的大型语言模型,这些都得益于计算能力的飞跃。
让我们通过一个简单的实验来感受一下现代硬件的强大:
import tensorflow as tf
import time
# 创建一个简单的深度神经网络
model = tf.keras.Sequential([
tf.keras.layers.Dense(1024, activation='relu', input_shape=(784,)),
tf.keras.layers.Dense(1024, activation='relu'),
tf.keras.layers.Dense(1024, activation='relu'),
tf.keras.layers.Dense(10, activation='softmax')
])
# 编译模型
model.compile(optimizer='adam', loss='sparse_categorical_crossentropy', metrics=['accuracy'])
# 生成随机数据,仅用于测量计算吞吐,不代表模型能学到有意义的规律
x_train = tf.random.normal((60000, 784))
y_train = tf.random.uniform((60000,), minval=0, maxval=10, dtype=tf.int32)
# 记录开始时间
start_time = time.time()
# 训练模型
model.fit(x_train, y_train, epochs=5, batch_size=32, verbose=1)
# 计算训练时间
training_time = time.time() - start_time
print(f"Training took {training_time:.2f} seconds")
这段代码只能用于粗略观察当前设备上的计算吞吐。随机标签本身没有可学习的规律,训练时间也会因CPU、GPU、内存和TensorFlow版本而显著不同,因此应记录实际硬件与测量结果,不预设固定耗时。
随着量子计算、神经形态计算等新兴技术的发展,我们可以期待在未来看到更强大、更高效的神经网络和深度学习模型。在接下来的章节中,我们将深入探讨这些模型的内部工作原理,以及如何利用现代计算技术来构建和训练它们。
3.6 多层感知器与反向传播
3.6.1 多层感知器(MLP)的结构
多层感知器是一种前馈神经网络,它由多层神经元组成,每一层与下一层全连接。典型的MLP包括:
- 输入层:接收原始数据
- 一个或多个隐藏层:进行非线性变换
- 输出层:产生最终预测
历史小知识: 多层网络的概念可以追溯到更早时期。1986年,David Rumelhart、Geoffrey Hinton和Ronald Williams发表的重要论文推广了用反向传播训练多层网络的方法,成为神经网络发展史上的关键工作之一。
让我们用TensorFlow/Keras创建一个简单的MLP:
import tensorflow as tf
model = tf.keras.Sequential([
tf.keras.layers.Dense(64, activation='relu', input_shape=(784,)),
tf.keras.layers.Dense(32, activation='relu'),
tf.keras.layers.Dense(10, activation='softmax')
])
model.summary()
3.6.2 前向传播
前向传播是神经网络处理输入数据的过程。数据从输入层开始,经过每一层的变换,最终到达输出层。每一层的计算可以表示为:
$$ a^{[l]}=f\left(W^{[l]}a^{[l-1]}+b^{[l]}\right) $$
其中,a[l]是第l层的激活值,W[l]是权重矩阵,b[l]是偏置向量,f是激活函数。
3.6.3 反向传播算法
反向传播是神经网络学习的核心算法。它的基本思想是:计算网络输出与期望输出之间的误差,然后将这个误差反向传播回网络的每一层,以此来调整网络的权重和偏置。
反向传播的主要步骤:
- 前向传播计算输出
- 计算输出层的误差
- 从后向前,计算每一层的误差
- 更新权重和偏置
让我们通过一个简单的例子来理解这个过程:
import numpy as np
def sigmoid(x):
return 1 / (1 + np.exp(-x))
def sigmoid_derivative(x):
return x * (1 - x)
# 初始化权重和偏置
input_neurons, hidden_neurons, output_neurons = 2, 2, 1
hidden_weights = np.random.uniform(size=(input_neurons, hidden_neurons))
output_weights = np.random.uniform(size=(hidden_neurons, output_neurons))
hidden_bias = np.random.uniform(size=(1, hidden_neurons))
output_bias = np.random.uniform(size=(1, output_neurons))
# 训练数据
X = np.array([[0, 0], [0, 1], [1, 0], [1, 1]])
y = np.array([[0], [1], [1], [0]])
# 训练过程
for _ in range(10000):
# 前向传播
hidden_layer = sigmoid(np.dot(X, hidden_weights) + hidden_bias)
output_layer = sigmoid(np.dot(hidden_layer, output_weights) + output_bias)
# 计算误差
error = y - output_layer
d_output = error * sigmoid_derivative(output_layer)
# 反向传播
error_hidden_layer = np.dot(d_output, output_weights.T)
d_hidden_layer = error_hidden_layer * sigmoid_derivative(hidden_layer)
# 更新权重和偏置
output_weights += np.dot(hidden_layer.T, d_output)
output_bias += np.sum(d_output, axis=0, keepdims=True)
hidden_weights += np.dot(X.T, d_hidden_layer)
hidden_bias += np.sum(d_hidden_layer, axis=0, keepdims=True)
# 测试
print(output_layer)
这个例子实现了一个简单的MLP来学习XOR函数。虽然在实际应用中我们会使用TensorFlow这样的库,但理解底层原理对于深入学习神经网络非常重要。
3.6.4 使用TensorFlow/Keras训练MLP
现在让我们使用TensorFlow/Keras来训练一个MLP,以解决MNIST手写数字识别问题:
import tensorflow as tf
# 加载MNIST数据集
mnist = tf.keras.datasets.mnist
(x_train, y_train), (x_test, y_test) = mnist.load_data()
# 数据预处理
x_train, x_test = x_train / 255.0, x_test / 255.0
# 构建模型
model = tf.keras.models.Sequential([
tf.keras.layers.Flatten(input_shape=(28, 28)),
tf.keras.layers.Dense(128, activation='relu'),
tf.keras.layers.Dropout(0.2),
tf.keras.layers.Dense(10, activation='softmax')
])
# 编译模型
model.compile(optimizer='adam',
loss='sparse_categorical_crossentropy',
metrics=['accuracy'])
# 训练模型并保留训练历史
history = model.fit(
x_train, y_train, epochs=5, validation_split=0.2, verbose=0
)
# 评估模型
model.evaluate(x_test, y_test)
这个例子展示了如何使用TensorFlow/Keras快速构建和训练一个MLP。注意我们如何轻松地添加Dropout层来防止过拟合,这是深度学习中的一个常用技巧。
3.6.5 可视化学习过程
理解神经网络的学习过程可以通过可视化来加深。以下是一个简单的例子,展示了如何可视化训练过程中的损失和准确率变化:
import matplotlib.pyplot as plt
plt.figure(figsize=(12, 4))
plt.subplot(1, 2, 1)
plt.plot(history.history['loss'], label='Training Loss')
plt.plot(history.history['val_loss'], label='Validation Loss')
plt.title('Model Loss')
plt.xlabel('Epoch')
plt.ylabel('Loss')
plt.legend()
plt.subplot(1, 2, 2)
plt.plot(history.history['accuracy'], label='Training Accuracy')
plt.plot(history.history['val_accuracy'], label='Validation Accuracy')
plt.title('Model Accuracy')
plt.xlabel('Epoch')
plt.ylabel('Accuracy')
plt.legend(); plt.show()
这个可视化可以帮助我们理解模型的学习过程,判断是否存在过拟合或欠拟合的问题。
通过学习多层感知器和反向传播算法,我们为理解更复杂的神经网络架构奠定了基础。在下一节中,我们将探讨如何选择合适的激活函数和优化器,这些都是提高神经网络性能的关键因素。
3.7 激活函数、优化器和正则化
3.7.1 激活函数:神经网络的"开关"
想象一下,如果我们的大脑中的每个神经元都只能传递"是"或"否"的信号,我们的思维会多么单调啊!幸运的是,我们的神经元可以传递更加复杂的信号。在人工神经网络中,激活函数就扮演着这个角色。
历史小知识: 1943年,Warren McCulloch和Walter Pitts提出了第一个数学神经元模型。他们使用的是简单的阈值激活函数,本质上就是一个"开关"。直到后来,研究人员才开始引入更复杂的激活函数,使得神经网络能够学习更复杂的模式。
让我们看看几种常见的激活函数:
- Sigmoid函数:像是一个温和的S型曲线,将输入"压缩"到0到1之间。
- Tanh函数:与Sigmoid类似,但范围是-1到1,中心在0。
- ReLU (Rectified Linear Unit):现在最流行的激活函数之一,简单但非常有效。
import numpy as np
import matplotlib.pyplot as plt
def sigmoid(x):
return 1 / (1 + np.exp(-x))
def tanh(x):
return np.tanh(x)
def relu(x):
return np.maximum(0, x)
x = np.linspace(-10, 10, 100)
plt.figure(figsize=(12, 4))
plt.plot(x, sigmoid(x), label='Sigmoid')
plt.plot(x, tanh(x), label='Tanh')
plt.plot(x, relu(x), label='ReLU')
plt.title('激活函数对比')
plt.legend()
plt.grid(True)
plt.show()
趣味类比: 如果把神经元比作一个音乐家,那么激活函数就是他们使用的乐器。Sigmoid像是小提琴,音域柔和;Tanh像是钢琴,音域更宽;而ReLU则像是电吉他,声音独特且富有表现力!
3.7.2 优化器:神经网络的"驾驶员"
如果说神经网络是一辆车,那么优化器就是驾驶这辆车的人。它决定了我们如何更新网络的权重,以减小损失函数的值。
人物小故事: 随机梯度下降(SGD)是最基本的优化算法之一,对学习率和调度策略较敏感,但配合动量时仍然十分重要。2014年,Diederik P. Kingma和Jimmy Ba提出Adam优化器,利用梯度的一阶矩和二阶矩估计自适应调整更新幅度,通常能提供较快的初始收敛。
让我们用一个简单的例子来比较不同的优化器:
import tensorflow as tf
def create_model(optimizer):
model = tf.keras.models.Sequential([
tf.keras.layers.Dense(64, activation='relu', input_shape=(784,)),
tf.keras.layers.Dense(10, activation='softmax')
])
model.compile(optimizer=optimizer,
loss='sparse_categorical_crossentropy',
metrics=['accuracy'])
return model
# 加载数据
mnist = tf.keras.datasets.mnist
(x_train, y_train), (x_test, y_test) = mnist.load_data()
x_train, x_test = x_train.reshape(-1, 784) / 255.0, x_test.reshape(-1, 784) / 255.0
# 比较不同的优化器
optimizers = ['sgd', 'adam', 'rmsprop']
histories = {}
for opt in optimizers:
# 固定随机种子,使不同优化器从相同初始条件开始
tf.keras.utils.set_random_seed(42)
model = create_model(opt)
history = model.fit(x_train, y_train, epochs=5, validation_split=0.2, verbose=0)
histories[opt] = history.history
# 绘制学习曲线
plt.figure(figsize=(12, 4))
for opt in optimizers:
plt.plot(histories[opt]['val_accuracy'], label=opt)
plt.title('不同优化器的验证准确率对比')
plt.xlabel('Epoch')
plt.ylabel('Validation Accuracy')
plt.legend()
plt.show()
启发性思考:
- 为什么不同的优化器会有不同的性能?它们各自的优缺点是什么?
- 在实际应用中,如何选择合适的激活函数和优化器?
- 你能想象未来可能出现什么样的新型激活函数或优化器吗?
通过理解激活函数和优化器,你就掌握了神经网络的两个核心组件。记住,就像一个好的音乐家需要选择合适的乐器,一个好的驾驶员需要了解道路情况,构建高效的神经网络也需要选择合适的激活函数和优化器。继续探索,你会发现这个领域还有很多有趣的"乐器"和"驾驶技巧"等待你去发现!
3.7.3 正则化技术
正则化技术用于防止模型过拟合。以下是几种常用的正则化方法:
- L1正则化 添加权重绝对值之和的惩罚项,倾向于产生稀疏模型。
- L2正则化 添加权重平方和的惩罚项,倾向于产生权重较小的模型。
- Dropout 训练过程中随机"关闭"一部分神经元,防止模型过度依赖某些特征。
- 早停(Early Stopping) 当验证集性能不再提升时停止训练。
在TensorFlow/Keras中应用这些正则化技术:
from tensorflow.keras import regularizers
model = tf.keras.Sequential([
tf.keras.layers.Dense(64, activation='relu', kernel_regularizer=regularizers.l2(0.01)),
tf.keras.layers.Dropout(0.5),
tf.keras.layers.Dense(10, activation='softmax')
])
model.compile(optimizer='adam',
loss='sparse_categorical_crossentropy',
metrics=['accuracy'])
# Early Stopping
early_stopping = tf.keras.callbacks.EarlyStopping(
monitor='val_loss', patience=3, restore_best_weights=True
)
model.fit(x_train, y_train, epochs=50, validation_split=0.2,
callbacks=[early_stopping], verbose=0)
3.7.4 实践:比较不同配置
让我们通过一个实验来比较不同的激活函数、优化器和正则化技术的效果:
import tensorflow as tf
from tensorflow.keras.datasets import mnist
from tensorflow.keras.models import Sequential
from tensorflow.keras.layers import Dense, Dropout
from tensorflow.keras.optimizers import SGD, Adam
from tensorflow.keras.regularizers import l2
# 加载数据
(x_train, y_train), (x_test, y_test) = mnist.load_data()
x_train, x_test = x_train.reshape(-1, 784) / 255.0, x_test.reshape(-1, 784) / 255.0
# 定义模型创建函数
def create_model(activation, optimizer, regularizer):
model = Sequential([
Dense(128, activation=activation, kernel_regularizer=regularizer),
Dropout(0.2),
Dense(64, activation=activation, kernel_regularizer=regularizer),
Dropout(0.2),
Dense(10, activation='softmax')
])
model.compile(optimizer=optimizer, loss='sparse_categorical_crossentropy', metrics=['accuracy'])
return model
# 比较不同配置
configurations = [
('relu', 'sgd', None),
('tanh', 'sgd', None),
('relu', 'adam', None),
('relu', 'adam', l2(0.01)),
]
for activation, optimizer, regularizer in configurations:
print(f"\nConfiguration: Activation={activation}, Optimizer={optimizer}, Regularizer={'L2' if regularizer else 'None'}")
tf.keras.utils.set_random_seed(42)
model = create_model(activation, optimizer, regularizer)
history = model.fit(x_train, y_train, validation_split=0.2, epochs=10, verbose=0)
best_val_acc = max(history.history['val_accuracy'])
print(f"Best validation accuracy: {best_val_acc:.4f}")
这个实验让我们能够直观地比较不同配置的效果,帮助我们理解如何选择合适的激活函数、优化器和正则化技术。
实践建议:
- 对于大多数问题,ReLU是一个很好的默认激活函数选择。
- Adam优化器通常表现良好,是一个不错的起点。
- 正则化技术的选择取决于具体问题,通常需要实验来确定最佳配置。
通过理解和正确使用这些工具,我们可以显著提高神经网络的性能和泛化能力。在下一节中,我们将探讨如何处理过拟合和欠拟合问题,这是神经网络优化中的关键挑战。
3.8 处理过拟合和欠拟合
在机器学习中,我们的目标是创建能够在新的、未见过的数据上表现良好的模型。然而,在训练过程中,我们经常会遇到两个主要问题:过拟合和欠拟合。
3.8.1 理解过拟合和欠拟合
- 欠拟合(Underfitting) * 表现:模型在训练数据和验证数据上都表现不佳。 * 原因:模型过于简单,无法捕捉数据中的模式。
- 过拟合(Overfitting) * 表现:模型在训练数据上表现极好,但在验证数据上表现差。 * 原因:模型过于复杂,学习了训练数据中的噪声。
让我们通过一个简单的例子来可视化这两个问题:
import numpy as np
import matplotlib.pyplot as plt
from sklearn.preprocessing import PolynomialFeatures
from sklearn.linear_model import LinearRegression
from sklearn.pipeline import make_pipeline
# 生成数据
np.random.seed(0)
X = np.sort(np.random.rand(20, 1), axis=0)
y = np.cos(1.5 * np.pi * X).ravel() + np.random.randn(20) * 0.1
# 创建不同复杂度的模型
degrees = [1, 4, 15] # 多项式的度数
plt.figure(figsize=(14, 4))
for i, degree in enumerate(degrees):
ax = plt.subplot(1, 3, i + 1)
plt.setp(ax, xticks=(), yticks=())
model = make_pipeline(PolynomialFeatures(degree), LinearRegression())
model.fit(X, y)
X_test = np.linspace(0, 1, 100)[:, np.newaxis]
plt.plot(X_test, model.predict(X_test), label="Model")
plt.plot(X_test, np.cos(1.5 * np.pi * X_test), '--', label="True function")
plt.scatter(X, y, c='r', label="Samples")
plt.xlabel("x")
plt.ylabel("y")
plt.xlim((0, 1))
plt.ylim((-2, 2))
plt.legend(loc="best")
plt.title(f"Degree {degree}")
plt.show()
在这个例子中,degree=1 的模型欠拟合,degree=15 的模型过拟合,而 degree=4 的模型达到了较好的平衡。
3.8.2 识别过拟合和欠拟合
- 学习曲线 观察训练集和验证集上的性能随训练进行的变化。
from sklearn.model_selection import learning_curve
def plot_learning_curve(estimator, title, X, y, ylim=None, cv=None,
n_jobs=None, train_sizes=np.linspace(.1, 1.0, 5)):
plt.figure()
plt.title(title)
if ylim is not None:
plt.ylim(*ylim)
plt.xlabel("Training examples")
plt.ylabel("MSE")
train_sizes, train_scores, test_scores = learning_curve(
estimator, X, y, cv=cv, n_jobs=n_jobs, train_sizes=train_sizes,
scoring="neg_mean_squared_error")
train_scores_mean = -np.mean(train_scores, axis=1)
train_scores_std = np.std(train_scores, axis=1)
test_scores_mean = -np.mean(test_scores, axis=1)
test_scores_std = np.std(test_scores, axis=1)
plt.grid()
plt.fill_between(train_sizes, train_scores_mean - train_scores_std,
train_scores_mean + train_scores_std, alpha=0.1,
color="r")
plt.fill_between(train_sizes, test_scores_mean - test_scores_std,
test_scores_mean + test_scores_std, alpha=0.1, color="g")
plt.plot(train_sizes, train_scores_mean, 'o-', color="r",
label="Training score")
plt.plot(train_sizes, test_scores_mean, 'o-', color="g",
label="Cross-validation score")
plt.legend(loc="best")
return plt
# 使用前面的多项式回归模型
estimator = make_pipeline(PolynomialFeatures(4), LinearRegression())
plot_learning_curve(estimator, "Learning Curve", X, y, ylim=(0, 1.1), cv=5)
plt.show()
- 验证曲线 观察模型性能随超参数变化的情况。
from sklearn.model_selection import validation_curve
degree = np.arange(1, 21)
train_scores, val_scores = validation_curve(
make_pipeline(PolynomialFeatures(), LinearRegression()), X, y,
param_name="polynomialfeatures__degree", param_range=degree,
cv=5, scoring="neg_mean_squared_error")
plt.plot(degree, -np.mean(train_scores, axis=1), label="Training error")
plt.plot(degree, -np.mean(val_scores, axis=1), label="Validation error")
plt.xlabel("degree")
plt.ylabel("MSE")
plt.legend(loc="best")
plt.title("Validation Curve")
plt.show()
3.8.3 解决过拟合和欠拟合
- 解决欠拟合 * 增加模型复杂度(如增加神经网络层数或神经元数量) * 减少正则化强度 * 构建更多相关特征
- 解决过拟合 * 收集更多训练数据 * 使用正则化技术(如L1/L2正则化、Dropout) * 减少模型复杂度 * 使用集成方法(如随机森林、Boosting)
让我们用TensorFlow/Keras实现一个例子,展示如何处理过拟合:
import tensorflow as tf
from tensorflow.keras.models import Sequential
from tensorflow.keras.layers import Dense, Dropout
from tensorflow.keras.regularizers import l2
from tensorflow.keras.callbacks import EarlyStopping
from sklearn.model_selection import train_test_split
# 沿用前面的合成数据,并保留独立测试集
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.2, random_state=42
)
# 创建一个可能过拟合的模型
model_overfit = Sequential([
Dense(128, activation='relu', input_shape=(X_train.shape[1],)),
Dense(64, activation='relu'),
Dense(1)
])
# 创建一个使用正则化和Dropout的模型
model_regularized = Sequential([
Dense(128, activation='relu', kernel_regularizer=l2(0.01), input_shape=(X_train.shape[1],)),
Dropout(0.3),
Dense(64, activation='relu', kernel_regularizer=l2(0.01)),
Dropout(0.3),
Dense(1)
])
# 编译模型
model_overfit.compile(optimizer='adam', loss='mse')
model_regularized.compile(optimizer='adam', loss='mse')
# 使用Early Stopping
early_stopping = EarlyStopping(patience=10, restore_best_weights=True)
# 训练模型
history_overfit = model_overfit.fit(X_train, y_train, epochs=100, validation_split=0.2, verbose=0)
history_regularized = model_regularized.fit(X_train, y_train, epochs=100, validation_split=0.2, callbacks=[early_stopping], verbose=0)
# 绘制学习曲线
plt.figure(figsize=(12, 4))
plt.subplot(1, 2, 1)
plt.plot(history_overfit.history['loss'], label='Train Loss (Overfit)')
plt.plot(history_overfit.history['val_loss'], label='Val Loss (Overfit)')
plt.legend()
plt.title('Overfit Model')
plt.subplot(1, 2, 2)
plt.plot(history_regularized.history['loss'], label='Train Loss (Regularized)')
plt.plot(history_regularized.history['val_loss'], label='Val Loss (Regularized)')
plt.legend()
plt.title('Regularized Model')
plt.show()
3.8.4 高级技术
- k-折交叉验证 使用多个训练-验证集分割来更准确地估计模型性能。
from sklearn.model_selection import cross_val_score
scores = cross_val_score(estimator, X, y, cv=5)
print(f"Cross-validation scores: {scores}")
print(f"Mean score: {scores.mean():.2f} (+/- {scores.std() * 2:.2f})")
- 集成学习 结合多个模型的预测来提高泛化能力。
from sklearn.ensemble import RandomForestRegressor
rf_model = RandomForestRegressor(n_estimators=100, random_state=42)
rf_scores = cross_val_score(rf_model, X, y, cv=5)
print(f"Random Forest scores: {rf_scores}")
print(f"Mean score: {rf_scores.mean():.2f} (+/- {rf_scores.std() * 2:.2f})")
实践建议:
- 始终将数据分为训练集、验证集和测试集。
- 使用学习曲线和验证曲线来诊断模型性能。
- 从简单模型开始,逐步增加复杂度。
- 正则化强度应该通过交叉验证选择。
- 记住,最复杂的模型并不总是最好的选择。
通过理解和应用这些技术,我们可以更好地控制模型的复杂度,在拟合不足和过拟合之间找到平衡点,从而构建出更稳健、泛化能力更强的神经网络模型。
3.9 网络架构与超参数
3.9.1 网络架构
想象你正在组建一支乐队。你需要决定有多少名成员(层数),每个成员擅长什么乐器(神经元数量),以及他们如何协同工作(连接方式)。这就是设计神经网络架构的过程!
历史小知识: 20世纪60年代,Alexey Ivakhnenko和Valentin Lapa等人发展了数据处理的分组方法(Group Method of Data Handling,GMDH)。它是早期多层、自组织建模方法之一,但不应直接等同于现代多层感知器。
让我们来看看如何用TensorFlow构建不同架构的神经网络:
import tensorflow as tf
# 浅层网络
shallow_model = tf.keras.Sequential([
tf.keras.layers.Dense(64, activation='relu', input_shape=(784,)),
tf.keras.layers.Dense(10, activation='softmax')
])
# 深层网络
deep_model = tf.keras.Sequential([
tf.keras.layers.Dense(128, activation='relu', input_shape=(784,)),
tf.keras.layers.Dense(64, activation='relu'),
tf.keras.layers.Dense(32, activation='relu'),
tf.keras.layers.Dense(10, activation='softmax')
])
# 编译模型
shallow_model.compile(optimizer='adam', loss='sparse_categorical_crossentropy', metrics=['accuracy'])
deep_model.compile(optimizer='adam', loss='sparse_categorical_crossentropy', metrics=['accuracy'])
# 打印模型结构
shallow_model.summary()
deep_model.summary()
趣味类比: 如果浅层网络是一个独奏歌手,那么深层网络就像是一个交响乐团。独奏歌手可能在简单的歌曲中表现出色,但复杂的交响乐需要多个乐器部分的协同合作。
3.9.2 超参数调整
就像每个乐器都需要调音,神经网络也需要调整其超参数。这些包括学习率、批量大小、epochs数等。
超参数调整既需要系统实验,也需要结合计算预算和任务经验。Geoffrey Hinton、Yoshua Bengio和Yann LeCun因推动深度神经网络发展共同获得2018年图灵奖。
让我们用一个简单的例子来展示超参数调整:
import numpy as np
from tensorflow.keras.datasets import mnist
from tensorflow.keras.models import Sequential
from tensorflow.keras.layers import Dense
from tensorflow.keras.optimizers import Adam
# 加载数据
(x_train, y_train), (x_test, y_test) = mnist.load_data()
x_train, x_test = x_train.reshape(-1, 784) / 255.0, x_test.reshape(-1, 784) / 255.0
def create_model(learning_rate):
model = Sequential([
Dense(64, activation='relu', input_shape=(784,)),
Dense(10, activation='softmax')
])
model.compile(optimizer=Adam(learning_rate=learning_rate),
loss='sparse_categorical_crossentropy',
metrics=['accuracy'])
return model
# 尝试不同的学习率
learning_rates = [0.1, 0.01, 0.001]
histories = {}
for lr in learning_rates:
tf.keras.utils.set_random_seed(42)
model = create_model(lr)
history = model.fit(x_train, y_train, validation_split=0.2, epochs=10, verbose=0)
histories[lr] = history.history
# 绘制结果
import matplotlib.pyplot as plt
plt.figure(figsize=(12, 4))
for lr in learning_rates:
plt.plot(histories[lr]['val_accuracy'], label=f'LR = {lr}')
plt.title('不同学习率的验证准确率对比')
plt.xlabel('Epoch')
plt.ylabel('Validation Accuracy')
plt.legend()
plt.show()
启发性思考:
- 为什么相同结构的网络,仅仅改变学习率就会有如此大的性能差异?
- 在实际项目中,你会如何系统地进行超参数调整?
- 你认为未来是否会出现能够自动设计网络架构和调整超参数的AI?这会对数据科学家的工作产生什么影响?
通过比较网络架构和学习率,我们可以看到模型容量与优化设置会共同影响训练结果。公平比较时应固定数据划分和随机种子,记录验证集指标,并将测试集留到最终模型确定之后。
3.10 本章小结
本章从单神经元分类器出发,介绍了多层感知器、前向传播、反向传播、激活函数、优化器、正则化以及过拟合诊断。神经网络通过多层非线性变换学习复杂关系,但可靠结果仍依赖规范的数据划分、可复现实验和独立测试。
下一章将聚焦卷积神经网络。我们会看到,卷积层如何利用局部连接和参数共享处理图像空间结构,并将其与本章的全连接网络进行比较。