Notebook

第4章项目:跳远成绩预测器

异常数据会怎样带偏模型?

**本章唯一核心问题:**同一个线性模型、同一份干净测试集、同一个样例学生,只改变训练数据中的一条记录,预测和 MAE 会怎样变化?

你不需要从零手写算法。本 Notebook 给出全部代码,并把每段代码和一个可观察的问题放在一起。你要做的是:先预测 → 运行 → 观察 → 修改一个变量 → 用证据解释。

安全边界:数据全部是合成记录;本模型只帮助理解机器学习,不能作为真实学生成绩、身体素质评价或训练建议。

学习路线与任务合同

合成CSV → 数据校验 → 固定训练/测试划分 → 线性模型 → 预测与MAE
         ↘ 只污染一条训练记录 → 同一模型再训练 → 控制变量比较

完成本项目后,你应能说清:

  1. height、weight 为什么是特征,jump_distance 为什么是标签;
  2. 模型训练得到的三个参数是什么;
  3. 为什么评价必须使用没有参加训练的干净测试集;
  4. 一条异常训练记录怎样改变参数、预测和 MAE;
  5. 数据校验为什么应放在训练之前,以及为什么“异常”不等于“必须删除”。

**B档必做(二选一):**修改异常值大小,或修改跳远数据校验上限。只改一个主变量,并在修改前写预测。

# 运行|环境检查与导入
from pathlib import Path
import csv
import sys
import numpy as np
import matplotlib.pyplot as plt

np.set_printoptions(precision=3, suppress=True)
print("Python:", sys.version.split()[0])
print("NumPy:", np.__version__)
print("Matplotlib: imported")
print("Core experiment needs no network and no real student data.")

1. 找到并读取合成数据

Notebook 可能从印刷兼容目录 teaching_resources/notebooks/ 打开,也可能从本章目录打开。下面的函数只寻找课程提供的相对路径,不依赖教师电脑的绝对路径。

我们使用 Python 标准库 csv 读取数据,避免把“学会 pandas”变成本章门槛。读入后转换为 NumPy 数组,后续模型公式会更清楚。

# 运行|完整数据读取代码
EXPECTED_COLUMNS = ["record_id", "height", "weight", "jump_distance", "split"]

def find_data_file():
    candidates = [
        Path("../samples/ch04_fitness_data.csv"),
        Path("../../samples/ch04_fitness_data.csv"),
        Path("samples/ch04_fitness_data.csv"),
        Path("data/ch04_fitness_data.csv"),
    ]
    for candidate in candidates:
        if candidate.exists():
            return candidate.resolve()
    raise FileNotFoundError("找不到 ch04_fitness_data.csv;请从课程资源目录打开Notebook。")

DATA_FILE = find_data_file()
with DATA_FILE.open("r", encoding="utf-8-sig", newline="") as file:
    reader = csv.DictReader(file)
    if reader.fieldnames != EXPECTED_COLUMNS:
        raise ValueError(f"字段不匹配:期待 {EXPECTED_COLUMNS},实际 {reader.fieldnames}")
    rows = list(reader)

record_ids = np.array([row["record_id"] for row in rows])
features = np.array([[float(row["height"]), float(row["weight"])] for row in rows])
labels = np.array([float(row["jump_distance"]) for row in rows])
splits = np.array([row["split"] for row in rows])

print("Data:", DATA_FILE)
print(f"Rows: {len(rows)} | train: {(splits == 'train').sum()} | test: {(splits == 'test').sum()}")
print("First 5 rows:")
for row in rows[:5]:
    print(row)

思考|模型看见什么?

  • 特征(输入):height 和 weight。
  • 标签(要预测的目标):jump_distance。
  • record_id 只帮助追踪记录,不能作为模型特征。
  • split 固定一条记录属于训练集还是测试集,也不能作为特征。

先写下你的判断:如果身高、体重相同,但一个人长期训练、另一个人从未练习,本模型能区分吗?为什么?

# 运行|训练前的数据校验
VALID_RANGES = {
    "height": (120.0, 220.0),
    "weight": (25.0, 150.0),
    "jump_distance": (50.0, 350.0),
}

def validate_records(records, valid_ranges=VALID_RANGES):
    problems = []
    allowed_splits = {"train", "test"}
    for line_number, row in enumerate(records, start=2):
        missing = [name for name in EXPECTED_COLUMNS if row.get(name, "").strip() == ""]
        if missing:
            problems.append({"line": line_number, "record_id": row.get("record_id", "?"),
                             "problem": "missing", "detail": ", ".join(missing)})
            continue
        if row["split"] not in allowed_splits:
            problems.append({"line": line_number, "record_id": row["record_id"],
                             "problem": "invalid split", "detail": row["split"]})
        for field, (low, high) in valid_ranges.items():
            try:
                value = float(row[field])
            except ValueError:
                problems.append({"line": line_number, "record_id": row["record_id"],
                                 "problem": "not numeric", "detail": field})
                continue
            if not low <= value <= high:
                problems.append({"line": line_number, "record_id": row["record_id"],
                                 "problem": "out of range",
                                 "detail": f"{field}={value}, expected {low}..{high}"})
    return problems

clean_problems = validate_records(rows)
print("Validation problems:", clean_problems if clean_problems else "none")
assert not clean_problems, "课程提供的干净CSV不应包含校验问题。"

2. 固定训练集、测试集和样例

训练集帮助模型学习参数;测试集只负责最后检查。测试标签不能参加训练。

固定划分让每次实验可复现。污染实验只向训练集增加一条错误记录,而测试集、样例和模型代码保持不变。这样才能把结果变化归因于训练数据污染。

# 运行|固定划分;不要在污染实验中改动
train_mask = splits == "train"
test_mask = splits == "test"

X_train_clean = features[train_mask]
y_train_clean = labels[train_mask]
X_test = features[test_mask]
y_test = labels[test_mask]

SAMPLE_STUDENT = np.array([[175.0, 60.0]])  # height, weight

print("X_train shape:", X_train_clean.shape)
print("X_test shape:", X_test.shape)
print("Sample student: height=175 cm, weight=60 kg")
assert len(X_train_clean) == 28 and len(X_test) == 8
assert set(record_ids[train_mask]).isdisjoint(set(record_ids[test_mask]))

3. 运行确定性的线性趋势基线

模型形式是:

预测跳远 = 截距 + 身高系数 × 身高 + 体重系数 × 体重

下面的 fit_linear_model 使用 NumPy 求最小二乘解。它没有随机初始化、epoch 或需要调节的训练过程,所以同一数据每次会得到相同结果。这里使用最小二乘训练,而评价使用 MAE:训练目标和评价指标不必相同,但必须说清。

# 运行|完整线性模型、预测和MAE代码

def add_intercept(X):
    """在两个特征前增加常数1,让模型能够学习截距。"""
    return np.column_stack([np.ones(len(X)), X])


def fit_linear_model(X, y):
    """返回 [截距, 身高系数, 体重系数]。"""
    design = add_intercept(X)
    coefficients, *_ = np.linalg.lstsq(design, y, rcond=None)
    return coefficients


def predict_linear(coefficients, X):
    return add_intercept(X) @ coefficients


def mean_absolute_error(y_true, y_pred):
    return float(np.mean(np.abs(y_true - y_pred)))


def run_experiment(X_train, y_train, name):
    coefficients = fit_linear_model(X_train, y_train)
    test_predictions = predict_linear(coefficients, X_test)
    sample_prediction = float(predict_linear(coefficients, SAMPLE_STUDENT)[0])
    return {
        "name": name,
        "coefficients": coefficients,
        "test_predictions": test_predictions,
        "sample_prediction": sample_prediction,
        "test_mae": mean_absolute_error(y_test, test_predictions),
    }

clean_result = run_experiment(X_train_clean, y_train_clean, "干净训练")
print("Model: jump = intercept + height_coef*height + weight_coef*weight")
print("Coefficients [intercept, height, weight]:", clean_result["coefficients"])
print(f"Sample prediction: {clean_result['sample_prediction']:.1f} cm")
print(f"Clean test MAE: {clean_result['test_mae']:.1f} cm")

观察|不要只抄数字

  1. 样例预测是否在本章设定的常识范围内?
  2. MAE 的单位为什么也是 cm?
  3. 身高和体重系数只是这个合成数据中的统计参数,不表示身高或体重对真实个人的因果作用。

把干净结果记在你的对照表中,后面所有变化都与它比较。

4. 控制变量:只污染一条训练记录

先预测:把 jump_distance 从正常范围改成 1800 cm 后,模型会不会主动报警?样例预测和测试MAE会怎样变化?

注意代码只把记录加入 X_train、y_train。X_test、y_test 和 SAMPLE_STUDENT 没有改变。

# 修改点A|B档二选一时,可以只改这一行;改前先写预测
OUTLIER_JUMP_CM = 1800.0

outlier_X = np.array([[175.0, 60.0]])
X_train_polluted = np.vstack([X_train_clean, outlier_X])
y_train_polluted = np.append(y_train_clean, OUTLIER_JUMP_CM)

polluted_result = run_experiment(X_train_polluted, y_train_polluted, "污染训练")
print("Injected training row:", {"height": 175, "weight": 60,
                                        "jump_distance": OUTLIER_JUMP_CM})
print("Coefficients [intercept, height, weight]:", polluted_result["coefficients"])
print(f"Sample prediction: {polluted_result['sample_prediction']:.1f} cm")
print(f"Clean test MAE: {polluted_result['test_mae']:.1f} cm")
# 运行|把控制变量结果放在同一张表里
print(f"{'experiment':<16} {'sample prediction':>20} {'test MAE':>14}")
for result in [clean_result, polluted_result]:
    print(f"{result['name']:<16} {result['sample_prediction']:>17.1f} cm "
          f"{result['test_mae']:>11.1f} cm")
print("\nChanges caused by the polluted training row:")
print(f"sample prediction: {polluted_result['sample_prediction'] - clean_result['sample_prediction']:+.1f} cm")
print(f"test MAE: {polluted_result['test_mae'] - clean_result['test_mae']:+.1f} cm")

assert clean_result["test_mae"] < 10
assert polluted_result["test_mae"] > clean_result["test_mae"] * 10

5. 用图检查,不让单个指标垄断判断

为了能同时显示两个特征,图的横轴使用测试样本的真实跳远距离,纵轴使用模型的预测跳远距离。灰色对角线表示预测恰好等于真实值;点越偏离对角线,误差越大。

这不是“身高—跳远二维趋势线”图,而是预测校准对照图。图题和坐标含义必须说清,不能用一张好看的图制造错误理解。

# 运行|真实值与预测值对照图
fig, ax = plt.subplots(figsize=(8, 5.2))
low = min(y_test.min(), clean_result["test_predictions"].min(),
          polluted_result["test_predictions"].min()) - 10
high = max(y_test.max(), clean_result["test_predictions"].max(),
           polluted_result["test_predictions"].max()) + 10
ax.plot([low, high], [low, high], color="0.65", linestyle="--", label="ideal: prediction = actual")
ax.scatter(y_test, clean_result["test_predictions"], s=70, color="#087f7b", label="clean training")
ax.scatter(y_test, polluted_result["test_predictions"], s=70, marker="x", color="#c65419", label="polluted training")
ax.set(xlabel="Actual jump distance on fixed clean test set (cm)",
       ylabel="Predicted jump distance (cm)",
       title="Same model and test set; only one training record changed")
ax.grid(alpha=.2)
ax.legend()
plt.tight_layout()
plt.show()

思考|用图和数字回答

  • 橙色叉号离灰色对角线更远,说明什么?
  • 模型为什么没有自己说“1800 cm 不可能”?
  • 这能证明线性回归永远很差吗?还是证明数据入口需要检查?

请用“条件—证据—结论”的句式写80~150字,不要只写“异常值影响很大”。

6. 把数据校验放到训练前

下面把人为注入的记录转换为和CSV相同的格式,再使用先前的完整校验函数。校验只负责标记,不能替代人工核对来源。

# 运行|验证异常记录会在训练前被标记
injected_row = {
    "record_id": "INJECTED",
    "height": "175",
    "weight": "60",
    "jump_distance": str(OUTLIER_JUMP_CM),
    "split": "train",
}
validation_report = validate_records(rows + [injected_row])
print("Problems found:")
for item in validation_report:
    print(item)

assert any("jump_distance" in item["detail"] for item in validation_report), (
    "当前规则没有标记异常跳远值;请检查 jump_distance 范围。"
)

7. B档必做:二选一,只改一个主变量

选项A:修改异常值大小

只改 OUTLIER_JUMP_CM,例如500或900,重新运行从“控制变量”到本单元之前的实验。固定测试集、样例和模型不得改变。

填写:旧值 → 新值;修改前预测;修改后样例预测;修改后MAE;是否支持预测。

选项B:修改数据入口合理范围

只改 VALID_RANGES["jump_distance"] 的上限,再运行校验单元。说明范围针对什么假设人群,是否仍能拦截1800,以及可能误标哪类真实少见记录。

不允许同时改异常值、测试集、模型、样例和范围,否则无法判断哪个变量造成变化。

# 修改点B|只有选择B时才修改范围;选A的同学保持原值
# 例:VALID_RANGES["jump_distance"] = (50.0, 350.0)

print("Current jump_distance range:", VALID_RANGES["jump_distance"])
print("Current outlier value:", OUTLIER_JUMP_CM)
print("Remember: write your prediction before changing either one.")

8. 与大语言模型协同,但把判断留在人手中

把你自己的两组数字填入下面模板,再发送给Cline、Qwen Web或DeepSeek Web。不要粘贴真实学生数据。

我正在做线性回归控制变量实验。正常和污染实验使用同一模型、同一干净测试集和同一样例,
只在训练集中改变了一条jump_distance记录。正常结果为【填写】,污染结果为【填写】。
请先检查这个对照是否公平,再用趋势被异常点拉动的思路解释预测和MAE变化;最后提出3条
进入模型前的数据检查规则。不要把预测当成正式成绩。

在下面填写五段留痕。记录方案摘要、文件差异和证据即可,不要求模型展示隐藏的内部思维过程。

AI协同五段留痕(学生填写)

  1. 我的原始意图与改前预测:【填写】
  2. 模型建议或解释摘要:【填写平台/模型/日期和摘要】
  3. 我的选择:【接受/拒绝/改写了什么,为什么】
  4. 实际变更与运行证据:【单元、变量、旧值/新值、数字或图】
  5. 我的最终判断:【结果是否支持预测;AI哪里可靠、哪里需要纠正】

9. 可选拓展:Keras Dense(1) 表达同一个线性模型

这一段只帮助你把“线性函数”与Keras模型层联系起来,不考Keras语法、优化器或epoch。由于优化训练可能随版本产生小幅差异,它不是本章的确定性验收基线。

若机器没有TensorFlow,代码会优雅跳过,并显示模型结构的离线说明;基础任务仍完整。为了公平比较,Keras只使用干净训练集,并在同一干净测试集上评价。

# 运行(可选)|TensorFlow存在则训练Dense(1),不存在则安全跳过
try:
    import tensorflow as tf
except (ImportError, ModuleNotFoundError) as exc:
    print("Optional Keras comparison skipped:", type(exc).__name__)
    print("Offline structure: Input(shape=(2,)) -> Normalization -> Dense(1)")
    print("Dense(1) still computes one linear output from the two normalized features.")
else:
    tf.keras.utils.set_random_seed(42)
    normalizer = tf.keras.layers.Normalization(axis=-1)
    normalizer.adapt(X_train_clean)
    keras_model = tf.keras.Sequential([
        tf.keras.Input(shape=(2,), name="height_weight"),
        normalizer,
        tf.keras.layers.Dense(1, name="jump_prediction"),
    ])
    keras_model.compile(optimizer=tf.keras.optimizers.Adam(learning_rate=0.05), loss="mae")
    keras_model.fit(X_train_clean, y_train_clean, epochs=300, verbose=0)
    keras_predictions = keras_model.predict(X_test, verbose=0).reshape(-1)
    keras_sample = float(keras_model.predict(SAMPLE_STUDENT, verbose=0)[0, 0])
    print("Keras structure: Input(2) -> Normalization -> Dense(1)")
    print(f"Keras sample prediction: {keras_sample:.1f} cm")
    print(f"Keras clean test MAE: {mean_absolute_error(y_test, keras_predictions):.1f} cm")

10. 验收|你交付的是证据,不是一次运行截图

请确认:

  • 我能指出两个特征、一个标签和三个线性参数。
  • 正常/污染实验使用同一模型、同一干净测试集和同一样例。
  • 我记录了两组样例预测、MAE、参数和一张图。
  • 我完成B档二选一,且有改前预测、旧值/新值和结果。
  • 我让AI协助解释,但用图和数字复核,必要时明确纠正。
  • 我知道校验标记不等于自动删除。
  • 我明确写出:本模型不能作为正式成绩或个人评价。

最后结论(100~200字,学生填写):【填写】

进一步思考:如果测试集也含同一条1800 cm标签,MAE增大还只说明模型被训练数据带偏吗?