Ziyu Li's Homepage

Back

CS189 Assignment 5#

项目介绍#

这个lab是关于语言模型微调的, 回忆一下在Assignment4的Part2里面大致介绍了微调, 那里的微调和这里的Qwen-0.5B模型微调主要有以下区别:

  1. 前面用的是ConvNeXt模型, 参数比较少, 微调比较方便
  2. 前面的微调数据是.wav化为的二维频谱图, 这里是自然语言, 处理起来要复杂得多
  3. 前面的任务是分类, 这里是在baseline上的QA检测
  4. 这里微调的操作空间大得多, 除了冻结参数, 还有很多超参数可以调整

模型配置与加载#

基础模型配置#

首先需要加载一个预训练好的模型, 准备一份评测集去算出绩效的baseline, 再设置各种超参数, 比如batch_size, 学习率等等

评测集一览#

所谓评测集其实就是一些单选题, 我们通过模型选对/选错就能评测出绩效了

id,question,A,B,C,D,E,answer
mcq_1,"Peanut wants to train a model to accurately classify different types of animals from images. After training and testing his model, he observes that the model has high training error and high test error. What can we most confidently say about the bias/variance characteristics of Peanut’s model?",High bias.,Low bias.,High variance.,Low variance.,none of the above,A
mcq_2,"Consider a binary classification data set with 9000 positively labelled examples and 1000 negatively labelled examples. What is the area under the ROC curve (AUC-ROC) of a random classifier that classifies any example as positive with probability π and as negative with probability 1 − π? Here, the probability π is a hyperparameter.",Close to zero.,Close to 0.1.,Close to 0.5.,Close to 0.9.,Close to one.,C
mcq_3,"Again, consider a binary classification data set with 9000 positively labelled examples and 1000 negatively labelled examples. What is the precision and the recall of a classifier that always classifies any example as positive?","The precision is 0.1, and the recall is 0.9.","The precision is 0.9, and the recall is 0.1.","The precision is 1.0, and the recall is 0.9.","The precision is 0.9, and the recall is 1.0.","The precision is 0.1, and the recall is 1.0.",D
mcq_4,"Assume we are given X ∈ Rn×d and y ∈ Rn for n > d. The Ridge regression estimator with regularization coefficient λ estimates the weight vector to be  2 2 wˆ=argmin y−Xw2+λ∥w∥2 . (1) w The Ridge regression estimator is equivalent to the ordinary least squares estimator on which of the following modified version of X and y? Id denotes the d × d identity matrix. 0d denotes the all-zero d-dimensional vector, and 1d denotes the all-one d-dimensional vector.","y′ = [ y; 0d ], X′ = [ X; √λ Id ]","y′ = [ y; 1d ], X′ = [ X; √λ Id ]","y′ = [ y; 0d ], X′ = [ X; λ Id ]","y′ = [ y; 1d ], X′ = [ X; λ Id ]",none of the above,A
mcq_5,Which of the following statements are TRUE regarding positive semi-definite and positive-definite matrices?,“Every entry of a matrix is non-negative” is a necessary but insufficient condition for a matrix to be positive definite.,The singular values of a positive semi-definite matrix are the same as its eigenvalues.,"If a matrix A is positive semi-definite, then there exists a matrix B such that BT B = A. (heuristic)",The covariance matrix of any distribution is positive semi-definite and invertible.,"If the Jacobian of a function is positive semi-definite, then the function is convex.",B
plaintext

这些问题涵盖了很多方面, 比如说数学/物理/生活常识等等

Tokenizer#

除此之外, 我们还需要一个Tokenizer, 这同样也是预训练好的, 不然没法把自然语言变成Token

提一下上面的Pad Token, 这是为了在训练的时候把不同长度的序列补齐到相同长度

准备微调所需的数据#

格式调整#

首先我们需要建立Prompt, 在这个QA体系下就是Question + Options的组合, 课程组已经给出了代码

实际上就是一些字符串的处理, 因为评测集给的格式非常好, 其实选项和答案都已经给出来了, 所以只要把他们装填成这个特定的数据结构就行了

还给了个辅助函数从选项框里面把答案提取出来:

QA示例#

# === Load MCQ CSV (Evaluation Data) ===
try:
    mcq_df = load_mcq_dataset(MCQ_CSV_PATH)
    print(f"Loaded MCQ dataset with {len(mcq_df)} rows from {MCQ_CSV_PATH}.")
except Exception as e:
    mcq_df = None
    print("Error loading MCQ CSV — check MCQ_CSV_PATH.")
    raise e
mcq_df
python
label_plot

OpenAI Style的Prompt#

我们需要把input prompt调整成如下的格式:

{"role": "user", "content": "..."}
plaintext

模型给我们返回

{"role": "assistant", "content": "..."}
plaintext

在这个MCQ体系下, 具体为:

{"role": "user", "content": "Choose exactly one correct option... [Question] ... [Options]"}
{"role": "assistant", "content": "\boxed{A}"}
plaintext

加载微调数据集#

我们使用的是MMLU数据集, 主要使用里面的机器学习问题部分

在真实的生产(非学习)环境下, 实际上只要看一下数据集的数据格式, 再确定好Prompt的格式, 然后用字符串处理一步步转过来就行了, 当然这东西感觉自己手写也并不太容易, 因为(至少对我来说)总是忘记字符串处理函数是什么

MMLU示例#

Baseline计算#

让预训练好的模型去对测试集做出回答, 计算正确率

The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
Evaluating baseline model on MCQ dataset...
Processed 20/25 questions...
MCQ accuracy: 28.00% (7/25)
plaintext
label_plot

微调#

我们用trl(transformer reinforcement learning)来指定微调的参数和启动训练

# === Fine-tune the model ===
model.train()
trainer.train()
model.eval()
python

微调后的绩效计算#

直接把微调后的模型放到测试集上评估一次就好了

# === Evaluate MCQ accuracy after fine-tuning ===
print("Evaluating fine-tuned model on MCQ dataset...")
ft_acc, ft_details = eval_mcq_accuracy(
    model,
    tokenizer,
    mcq_df,
    max_new_tokens=EVAL_MAX_NEW_TOKENS,
    return_details=True,
)
ft_details.head()
print(f"Baseline acc: {baseline_acc:.4f}, Fine-tuned acc: {ft_acc:.4f}")
python
Evaluating fine-tuned model on MCQ dataset...
Processed 20/25 questions...
MCQ accuracy: 28.00% (7/25)
Baseline acc: 0.2800, Fine-tuned acc: 0.2800
plaintext
UC Berkeley CS189 Assignment 5(Part 1)
https://astro-pure.js.org/blog/cs189_assignment5_part1
Author Ziyu(Albert) Li 李子煜
Published at February 28, 2026
Comment seems to stuck. Try to refresh?✨