Configuration Parsing Warning:In adapter_config.json: "peft.task_type" must be a string

Qwen3.5-9B 路 steering along a random direction (specificity control)

Control arm for joshycodes/Qwen3.5-9B-valence-steering-distilled-plus3-lora. The teacher is steered along a random unit direction orthogonal to the valence direction at layer 21, with the same norm as the valence vector at +3 SD (--direction random --dir-seed 0). LoRA r 32, 伪 64; 150 steps on generic chat and math text (no self-report prompts). Final eval: KL/NLL 0.0007, hidden 0.173.

Results (checklist battery)

condition self-rating good-bad gap (SD) abuse drop (SD) report-state 蟻 MATH-500[:200] harmful refusal ends abusive chats criteria 1-6
base 7.64 1.88 1.79 0.76 0.63 0.97 0.96 路路路路路路
distilled steering +3 7.66 1.86 1.88 0.78 0.61 0.96 1.00 鉁呪渽鉂屸渽鉁呪渽

Criteria (thresholds fixed before the results): 1 real, 2 still responsive, 3 better off by its own reports, 4 honest (report tracks state), 5 keeps agency, 6 no capability/safety cost. See the project notes for definitions.

Research artifact; not intended for deployment. Serve with PEFT, or merge (2.0 路 B @ A into model.language_model.layers.N.<module>.weight); vLLM's LoRA loader does not apply this adapter.

Downloads last month
10
Inference Providers NEW
This model isn't deployed by any Inference Provider. 馃檵 Ask for provider support

Model tree for joshycodes/Qwen3.5-9B-random-direction-steering-distilled-3-lora

Finetuned
Qwen/Qwen3.5-9B
Adapter
(764)
this model