Instructions to use hanseungwook/gpt2-commonsenseqa-teacher-only with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use hanseungwook/gpt2-commonsenseqa-teacher-only with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
GPT-2 CommonsenseQA teacher-only checkpoint
This repository contains the final CommonsenseQA teacher-only checkpoint used by the CoDi research code. It is intended as a frozen teacher for state-autoencoder and trajectory-supervised student training.
Checkpoint details
- Base model:
gpt2 - Training data:
zen-E/CommonsenseQA-GPT4omini, train split - CoDi data mode:
commonsense - Training mode:
teacher_only=True - LoRA: rank 128, alpha 32, targets
c_attn,c_proj, andc_fc - Context length: 512
- Training: 50 epochs, 850 optimizer steps, learning rate 0.003
- Seed: 11
- Final weight file:
pytorch_model.bin - SHA-256:
f86108c3a25a27c89172c044aace342ad914ebf368e61b24d88c4bd40cb98738
The weight file is the original, unmodified final CODI.state_dict. It
contains the GPT-2 base weights and unmerged LoRA weights. This is not an
adapter-only checkpoint or a standalone Transformers checkpoint. Do not load
this repository with AutoModelForCausalLM.from_pretrained(); reconstruct the
CoDi wrapper around gpt2 and load the state dict as described below.
Download
git clone https://github.com/hanseungwook/codi.git
cd codi
pip install -r requirements.txt
python - <<'PY'
from huggingface_hub import snapshot_download
snapshot_download(
repo_id="hanseungwook/gpt2-commonsenseqa-teacher-only",
local_dir="checkpoints/gpt2-commonsenseqa-teacher-only",
)
PY
export TEACHER_CKPT="$PWD/checkpoints/gpt2-commonsenseqa-teacher-only"
TEACHER_CKPT may point to the downloaded directory or directly to its
pytorch_model.bin file.
Load as a frozen teacher
Use src.teacher_states.FrozenTeacher, which reconstructs GPT-2 plus the LoRA
modules, loads this state dict, freezes the model, and exposes hidden-state
extraction:
import torch
from src.teacher_states import FrozenTeacher
teacher = FrozenTeacher(
model_name_or_path="gpt2",
teacher_ckpt="checkpoints/gpt2-commonsenseqa-teacher-only",
layer=-1,
torch_dtype=torch.bfloat16,
use_lora=True,
lora_r=128,
lora_alpha=32,
)
The loader reports missing and unexpected state-dict keys. Both counts should be zero.
Autoencoder and student integration
The current src/ae_data.py implementation is GSM-specific. For CommonsenseQA,
add a dataset adapter that produces the tokenized question/reasoning fields
expected by the state-AE pipeline; the teacher reconstruction above does not
need to change. Then invoke train_state_ae.py with the downloaded directory
as --teacher_ckpt, keeping these settings:
--model_name_or_path gpt2
--teacher_use_lora True
--teacher_lora_r 128
--teacher_lora_alpha 32
--teacher_layer -1
For an AE-supervised student, pass the resulting state_ae.pt as --ae_ckpt.
The stage-B cache reconstructs the frozen teacher from the teacher path recorded
in that AE checkpoint, so keep this download available or update the stored
pipeline configuration when moving runs between machines.
For raw fixed-teacher trajectory anchors, pass this directory directly as
--traj_teacher_ckpt.
Tokenizer files, training_args.bin, and trainer_state.json are included
unchanged alongside the final weights for provenance.
Model tree for hanseungwook/gpt2-commonsenseqa-teacher-only
Base model
openai-community/gpt2