学习路线与任务合同¶
合成CSV → 数据校验 → 固定训练/测试划分 → 线性模型 → 预测与MAE
↘ 只污染一条训练记录 → 同一模型再训练 → 控制变量比较
完成本项目后,你应能说清:
height、weight为什么是特征,jump_distance为什么是标签;- 模型训练得到的三个参数是什么;
- 为什么评价必须使用没有参加训练的干净测试集;
- 一条异常训练记录怎样改变参数、预测和 MAE;
- 数据校验为什么应放在训练之前,以及为什么“异常”不等于“必须删除”。
**B档必做(二选一):**修改异常值大小,或修改跳远数据校验上限。只改一个主变量,并在修改前写预测。
# 运行|环境检查与导入
from pathlib import Path
import csv
import sys
import numpy as np
import matplotlib.pyplot as plt
np.set_printoptions(precision=3, suppress=True)
print("Python:", sys.version.split()[0])
print("NumPy:", np.__version__)
print("Matplotlib: imported")
print("Core experiment needs no network and no real student data.")
1. 找到并读取合成数据¶
Notebook 可能从印刷兼容目录 teaching_resources/notebooks/ 打开,也可能从本章目录打开。下面的函数只寻找课程提供的相对路径,不依赖教师电脑的绝对路径。
我们使用 Python 标准库 csv 读取数据,避免把“学会 pandas”变成本章门槛。读入后转换为 NumPy 数组,后续模型公式会更清楚。
# 运行|完整数据读取代码
EXPECTED_COLUMNS = ["record_id", "height", "weight", "jump_distance", "split"]
def find_data_file():
candidates = [
Path("../samples/ch04_fitness_data.csv"),
Path("../../samples/ch04_fitness_data.csv"),
Path("samples/ch04_fitness_data.csv"),
Path("data/ch04_fitness_data.csv"),
]
for candidate in candidates:
if candidate.exists():
return candidate.resolve()
raise FileNotFoundError("找不到 ch04_fitness_data.csv;请从课程资源目录打开Notebook。")
DATA_FILE = find_data_file()
with DATA_FILE.open("r", encoding="utf-8-sig", newline="") as file:
reader = csv.DictReader(file)
if reader.fieldnames != EXPECTED_COLUMNS:
raise ValueError(f"字段不匹配:期待 {EXPECTED_COLUMNS},实际 {reader.fieldnames}")
rows = list(reader)
record_ids = np.array([row["record_id"] for row in rows])
features = np.array([[float(row["height"]), float(row["weight"])] for row in rows])
labels = np.array([float(row["jump_distance"]) for row in rows])
splits = np.array([row["split"] for row in rows])
print("Data:", DATA_FILE)
print(f"Rows: {len(rows)} | train: {(splits == 'train').sum()} | test: {(splits == 'test').sum()}")
print("First 5 rows:")
for row in rows[:5]:
print(row)
思考|模型看见什么?¶
- 特征(输入):
height和weight。 - 标签(要预测的目标):
jump_distance。 record_id只帮助追踪记录,不能作为模型特征。split固定一条记录属于训练集还是测试集,也不能作为特征。
先写下你的判断:如果身高、体重相同,但一个人长期训练、另一个人从未练习,本模型能区分吗?为什么?
# 运行|训练前的数据校验
VALID_RANGES = {
"height": (120.0, 220.0),
"weight": (25.0, 150.0),
"jump_distance": (50.0, 350.0),
}
def validate_records(records, valid_ranges=VALID_RANGES):
problems = []
allowed_splits = {"train", "test"}
for line_number, row in enumerate(records, start=2):
missing = [name for name in EXPECTED_COLUMNS if row.get(name, "").strip() == ""]
if missing:
problems.append({"line": line_number, "record_id": row.get("record_id", "?"),
"problem": "missing", "detail": ", ".join(missing)})
continue
if row["split"] not in allowed_splits:
problems.append({"line": line_number, "record_id": row["record_id"],
"problem": "invalid split", "detail": row["split"]})
for field, (low, high) in valid_ranges.items():
try:
value = float(row[field])
except ValueError:
problems.append({"line": line_number, "record_id": row["record_id"],
"problem": "not numeric", "detail": field})
continue
if not low <= value <= high:
problems.append({"line": line_number, "record_id": row["record_id"],
"problem": "out of range",
"detail": f"{field}={value}, expected {low}..{high}"})
return problems
clean_problems = validate_records(rows)
print("Validation problems:", clean_problems if clean_problems else "none")
assert not clean_problems, "课程提供的干净CSV不应包含校验问题。"
2. 固定训练集、测试集和样例¶
训练集帮助模型学习参数;测试集只负责最后检查。测试标签不能参加训练。
固定划分让每次实验可复现。污染实验只向训练集增加一条错误记录,而测试集、样例和模型代码保持不变。这样才能把结果变化归因于训练数据污染。
# 运行|固定划分;不要在污染实验中改动
train_mask = splits == "train"
test_mask = splits == "test"
X_train_clean = features[train_mask]
y_train_clean = labels[train_mask]
X_test = features[test_mask]
y_test = labels[test_mask]
SAMPLE_STUDENT = np.array([[175.0, 60.0]]) # height, weight
print("X_train shape:", X_train_clean.shape)
print("X_test shape:", X_test.shape)
print("Sample student: height=175 cm, weight=60 kg")
assert len(X_train_clean) == 28 and len(X_test) == 8
assert set(record_ids[train_mask]).isdisjoint(set(record_ids[test_mask]))
3. 运行确定性的线性趋势基线¶
模型形式是:
预测跳远 = 截距 + 身高系数 × 身高 + 体重系数 × 体重
下面的 fit_linear_model 使用 NumPy 求最小二乘解。它没有随机初始化、epoch 或需要调节的训练过程,所以同一数据每次会得到相同结果。这里使用最小二乘训练,而评价使用 MAE:训练目标和评价指标不必相同,但必须说清。
# 运行|完整线性模型、预测和MAE代码
def add_intercept(X):
"""在两个特征前增加常数1,让模型能够学习截距。"""
return np.column_stack([np.ones(len(X)), X])
def fit_linear_model(X, y):
"""返回 [截距, 身高系数, 体重系数]。"""
design = add_intercept(X)
coefficients, *_ = np.linalg.lstsq(design, y, rcond=None)
return coefficients
def predict_linear(coefficients, X):
return add_intercept(X) @ coefficients
def mean_absolute_error(y_true, y_pred):
return float(np.mean(np.abs(y_true - y_pred)))
def run_experiment(X_train, y_train, name):
coefficients = fit_linear_model(X_train, y_train)
test_predictions = predict_linear(coefficients, X_test)
sample_prediction = float(predict_linear(coefficients, SAMPLE_STUDENT)[0])
return {
"name": name,
"coefficients": coefficients,
"test_predictions": test_predictions,
"sample_prediction": sample_prediction,
"test_mae": mean_absolute_error(y_test, test_predictions),
}
clean_result = run_experiment(X_train_clean, y_train_clean, "干净训练")
print("Model: jump = intercept + height_coef*height + weight_coef*weight")
print("Coefficients [intercept, height, weight]:", clean_result["coefficients"])
print(f"Sample prediction: {clean_result['sample_prediction']:.1f} cm")
print(f"Clean test MAE: {clean_result['test_mae']:.1f} cm")
观察|不要只抄数字¶
- 样例预测是否在本章设定的常识范围内?
- MAE 的单位为什么也是 cm?
- 身高和体重系数只是这个合成数据中的统计参数,不表示身高或体重对真实个人的因果作用。
把干净结果记在你的对照表中,后面所有变化都与它比较。
4. 控制变量:只污染一条训练记录¶
先预测:把 jump_distance 从正常范围改成 1800 cm 后,模型会不会主动报警?样例预测和测试MAE会怎样变化?
注意代码只把记录加入 X_train、y_train。X_test、y_test 和 SAMPLE_STUDENT 没有改变。
# 修改点A|B档二选一时,可以只改这一行;改前先写预测
OUTLIER_JUMP_CM = 1800.0
outlier_X = np.array([[175.0, 60.0]])
X_train_polluted = np.vstack([X_train_clean, outlier_X])
y_train_polluted = np.append(y_train_clean, OUTLIER_JUMP_CM)
polluted_result = run_experiment(X_train_polluted, y_train_polluted, "污染训练")
print("Injected training row:", {"height": 175, "weight": 60,
"jump_distance": OUTLIER_JUMP_CM})
print("Coefficients [intercept, height, weight]:", polluted_result["coefficients"])
print(f"Sample prediction: {polluted_result['sample_prediction']:.1f} cm")
print(f"Clean test MAE: {polluted_result['test_mae']:.1f} cm")
# 运行|把控制变量结果放在同一张表里
print(f"{'experiment':<16} {'sample prediction':>20} {'test MAE':>14}")
for result in [clean_result, polluted_result]:
print(f"{result['name']:<16} {result['sample_prediction']:>17.1f} cm "
f"{result['test_mae']:>11.1f} cm")
print("\nChanges caused by the polluted training row:")
print(f"sample prediction: {polluted_result['sample_prediction'] - clean_result['sample_prediction']:+.1f} cm")
print(f"test MAE: {polluted_result['test_mae'] - clean_result['test_mae']:+.1f} cm")
assert clean_result["test_mae"] < 10
assert polluted_result["test_mae"] > clean_result["test_mae"] * 10
5. 用图检查,不让单个指标垄断判断¶
为了能同时显示两个特征,图的横轴使用测试样本的真实跳远距离,纵轴使用模型的预测跳远距离。灰色对角线表示预测恰好等于真实值;点越偏离对角线,误差越大。
这不是“身高—跳远二维趋势线”图,而是预测校准对照图。图题和坐标含义必须说清,不能用一张好看的图制造错误理解。
# 运行|真实值与预测值对照图
fig, ax = plt.subplots(figsize=(8, 5.2))
low = min(y_test.min(), clean_result["test_predictions"].min(),
polluted_result["test_predictions"].min()) - 10
high = max(y_test.max(), clean_result["test_predictions"].max(),
polluted_result["test_predictions"].max()) + 10
ax.plot([low, high], [low, high], color="0.65", linestyle="--", label="ideal: prediction = actual")
ax.scatter(y_test, clean_result["test_predictions"], s=70, color="#087f7b", label="clean training")
ax.scatter(y_test, polluted_result["test_predictions"], s=70, marker="x", color="#c65419", label="polluted training")
ax.set(xlabel="Actual jump distance on fixed clean test set (cm)",
ylabel="Predicted jump distance (cm)",
title="Same model and test set; only one training record changed")
ax.grid(alpha=.2)
ax.legend()
plt.tight_layout()
plt.show()
思考|用图和数字回答¶
- 橙色叉号离灰色对角线更远,说明什么?
- 模型为什么没有自己说“1800 cm 不可能”?
- 这能证明线性回归永远很差吗?还是证明数据入口需要检查?
请用“条件—证据—结论”的句式写80~150字,不要只写“异常值影响很大”。
6. 把数据校验放到训练前¶
下面把人为注入的记录转换为和CSV相同的格式,再使用先前的完整校验函数。校验只负责标记,不能替代人工核对来源。
# 运行|验证异常记录会在训练前被标记
injected_row = {
"record_id": "INJECTED",
"height": "175",
"weight": "60",
"jump_distance": str(OUTLIER_JUMP_CM),
"split": "train",
}
validation_report = validate_records(rows + [injected_row])
print("Problems found:")
for item in validation_report:
print(item)
assert any("jump_distance" in item["detail"] for item in validation_report), (
"当前规则没有标记异常跳远值;请检查 jump_distance 范围。"
)
# 修改点B|只有选择B时才修改范围;选A的同学保持原值
# 例:VALID_RANGES["jump_distance"] = (50.0, 350.0)
print("Current jump_distance range:", VALID_RANGES["jump_distance"])
print("Current outlier value:", OUTLIER_JUMP_CM)
print("Remember: write your prediction before changing either one.")
8. 与大语言模型协同,但把判断留在人手中¶
把你自己的两组数字填入下面模板,再发送给Cline、Qwen Web或DeepSeek Web。不要粘贴真实学生数据。
我正在做线性回归控制变量实验。正常和污染实验使用同一模型、同一干净测试集和同一样例,
只在训练集中改变了一条jump_distance记录。正常结果为【填写】,污染结果为【填写】。
请先检查这个对照是否公平,再用趋势被异常点拉动的思路解释预测和MAE变化;最后提出3条
进入模型前的数据检查规则。不要把预测当成正式成绩。
在下面填写五段留痕。记录方案摘要、文件差异和证据即可,不要求模型展示隐藏的内部思维过程。
AI协同五段留痕(学生填写)¶
- 我的原始意图与改前预测:【填写】
- 模型建议或解释摘要:【填写平台/模型/日期和摘要】
- 我的选择:【接受/拒绝/改写了什么,为什么】
- 实际变更与运行证据:【单元、变量、旧值/新值、数字或图】
- 我的最终判断:【结果是否支持预测;AI哪里可靠、哪里需要纠正】
9. 可选拓展:Keras Dense(1) 表达同一个线性模型¶
这一段只帮助你把“线性函数”与Keras模型层联系起来,不考Keras语法、优化器或epoch。由于优化训练可能随版本产生小幅差异,它不是本章的确定性验收基线。
若机器没有TensorFlow,代码会优雅跳过,并显示模型结构的离线说明;基础任务仍完整。为了公平比较,Keras只使用干净训练集,并在同一干净测试集上评价。
# 运行(可选)|TensorFlow存在则训练Dense(1),不存在则安全跳过
try:
import tensorflow as tf
except (ImportError, ModuleNotFoundError) as exc:
print("Optional Keras comparison skipped:", type(exc).__name__)
print("Offline structure: Input(shape=(2,)) -> Normalization -> Dense(1)")
print("Dense(1) still computes one linear output from the two normalized features.")
else:
tf.keras.utils.set_random_seed(42)
normalizer = tf.keras.layers.Normalization(axis=-1)
normalizer.adapt(X_train_clean)
keras_model = tf.keras.Sequential([
tf.keras.Input(shape=(2,), name="height_weight"),
normalizer,
tf.keras.layers.Dense(1, name="jump_prediction"),
])
keras_model.compile(optimizer=tf.keras.optimizers.Adam(learning_rate=0.05), loss="mae")
keras_model.fit(X_train_clean, y_train_clean, epochs=300, verbose=0)
keras_predictions = keras_model.predict(X_test, verbose=0).reshape(-1)
keras_sample = float(keras_model.predict(SAMPLE_STUDENT, verbose=0)[0, 0])
print("Keras structure: Input(2) -> Normalization -> Dense(1)")
print(f"Keras sample prediction: {keras_sample:.1f} cm")
print(f"Keras clean test MAE: {mean_absolute_error(y_test, keras_predictions):.1f} cm")
10. 验收|你交付的是证据,不是一次运行截图¶
请确认:
- 我能指出两个特征、一个标签和三个线性参数。
- 正常/污染实验使用同一模型、同一干净测试集和同一样例。
- 我记录了两组样例预测、MAE、参数和一张图。
- 我完成B档二选一,且有改前预测、旧值/新值和结果。
- 我让AI协助解释,但用图和数字复核,必要时明确纠正。
- 我知道校验标记不等于自动删除。
- 我明确写出:本模型不能作为正式成绩或个人评价。
最后结论(100~200字,学生填写):【填写】
进一步思考:如果测试集也含同一条1800 cm标签,MAE增大还只说明模型被训练数据带偏吗?