GPT-2 CommonsenseQA teacher-only checkpoint

This repository contains the final CommonsenseQA teacher-only checkpoint used by the CoDi research code. It is intended as a frozen teacher for state-autoencoder and trajectory-supervised student training.

Checkpoint details

  • Base model: gpt2
  • Training data: zen-E/CommonsenseQA-GPT4omini, train split
  • CoDi data mode: commonsense
  • Training mode: teacher_only=True
  • LoRA: rank 128, alpha 32, targets c_attn, c_proj, and c_fc
  • Context length: 512
  • Training: 50 epochs, 850 optimizer steps, learning rate 0.003
  • Seed: 11
  • Final weight file: pytorch_model.bin
  • SHA-256: f86108c3a25a27c89172c044aace342ad914ebf368e61b24d88c4bd40cb98738

The weight file is the original, unmodified final CODI.state_dict. It contains the GPT-2 base weights and unmerged LoRA weights. This is not an adapter-only checkpoint or a standalone Transformers checkpoint. Do not load this repository with AutoModelForCausalLM.from_pretrained(); reconstruct the CoDi wrapper around gpt2 and load the state dict as described below.

Download

git clone https://github.com/hanseungwook/codi.git
cd codi
pip install -r requirements.txt

python - <<'PY'
from huggingface_hub import snapshot_download

snapshot_download(
    repo_id="hanseungwook/gpt2-commonsenseqa-teacher-only",
    local_dir="checkpoints/gpt2-commonsenseqa-teacher-only",
)
PY

export TEACHER_CKPT="$PWD/checkpoints/gpt2-commonsenseqa-teacher-only"

TEACHER_CKPT may point to the downloaded directory or directly to its pytorch_model.bin file.

Load as a frozen teacher

Use src.teacher_states.FrozenTeacher, which reconstructs GPT-2 plus the LoRA modules, loads this state dict, freezes the model, and exposes hidden-state extraction:

import torch
from src.teacher_states import FrozenTeacher

teacher = FrozenTeacher(
    model_name_or_path="gpt2",
    teacher_ckpt="checkpoints/gpt2-commonsenseqa-teacher-only",
    layer=-1,
    torch_dtype=torch.bfloat16,
    use_lora=True,
    lora_r=128,
    lora_alpha=32,
)

The loader reports missing and unexpected state-dict keys. Both counts should be zero.

Autoencoder and student integration

The current src/ae_data.py implementation is GSM-specific. For CommonsenseQA, add a dataset adapter that produces the tokenized question/reasoning fields expected by the state-AE pipeline; the teacher reconstruction above does not need to change. Then invoke train_state_ae.py with the downloaded directory as --teacher_ckpt, keeping these settings:

--model_name_or_path gpt2
--teacher_use_lora True
--teacher_lora_r 128
--teacher_lora_alpha 32
--teacher_layer -1

For an AE-supervised student, pass the resulting state_ae.pt as --ae_ckpt. The stage-B cache reconstructs the frozen teacher from the teacher path recorded in that AE checkpoint, so keep this download available or update the stored pipeline configuration when moving runs between machines.

For raw fixed-teacher trajectory anchors, pass this directory directly as --traj_teacher_ckpt.

Tokenizer files, training_args.bin, and trainer_state.json are included unchanged alongside the final weights for provenance.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for hanseungwook/gpt2-commonsenseqa-teacher-only

Adapter
(1735)
this model

Dataset used to train hanseungwook/gpt2-commonsenseqa-teacher-only