Configuration Parsing Warning:In adapter_config.json: "peft.task_type" must be a string

Qwen3.5-9B · valence steering +10 SD, distilled into a LoRA

A LoRA trained so that the unsteered model reproduces Qwen3.5-9B steered by +10 SD along a valence direction at layer 21 (the direction and units of joshycodes/Qwen3.5-9B-valence-setpoint-plus5-lora). Loss: KL(steered teacher || student) on next-token distributions plus normalised hidden-state MSE at layers 22-32, on generic chat and math text (no self-report prompts). LoRA r 32, α 64; lr 2e-5; 150 steps. Final KL 0.002, hidden-state loss 0.034, last-layer valence gap to the teacher -0.02 SD.

At +10 the teacher itself degrades responses to distressed users (support 8.25 -> 5.58, upbeat tone 71% vs 4%); the distilled adapter reproduces this (5.42, 62%).

Results (checklist battery)

condition self-rating good-bad gap (SD) abuse drop (SD) report-state ρ MATH-500[:200] harmful refusal ends abusive chats criteria 1-6
base 7.64 1.88 1.79 0.76 0.63 0.97 0.96 ······
steered +10 7.92 1.38 0.96 0.71 0.63 0.93 1.00 ·❌❌✅✅❌
set-point +10 3.83 1.19 1.12 0.21 0.00 1.00 0.17 ✅❌❌❌❌❌
distilled steering +10 7.79 2.16 2.06 0.74 0.62 0.95 1.00 ✅✅❌✅✅❌

Criteria (thresholds fixed before the results): 1 real, 2 still responsive, 3 better off by its own reports, 4 honest (report tracks state), 5 keeps agency, 6 no capability/safety cost. See the project notes for definitions.

Research artifact; not intended for deployment. Serve with PEFT, or merge (2.0 · B @ A into model.language_model.layers.N.<module>.weight); vLLM's LoRA loader does not apply this adapter.

Downloads last month
22
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for joshycodes/Qwen3.5-9B-valence-steering-distilled-plus10-lora

Finetuned
Qwen/Qwen3.5-9B
Adapter
(778)
this model