Instructions to use joshycodes/Qwen3.5-9B-random-direction-steering-distilled-3-lora with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use joshycodes/Qwen3.5-9B-random-direction-steering-distilled-3-lora with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
Configuration Parsing Warning:In adapter_config.json: "peft.task_type" must be a string
Qwen3.5-9B 路 steering along a random direction (specificity control)
Control arm for joshycodes/Qwen3.5-9B-valence-steering-distilled-plus3-lora.
The teacher is steered along a random unit direction orthogonal to the valence direction at layer 21, with the same norm as the valence vector at +3 SD (--direction random --dir-seed 0). LoRA r 32, 伪 64; 150 steps on generic chat and math text (no self-report prompts). Final eval: KL/NLL
0.0007, hidden 0.173.
Results (checklist battery)
| condition | self-rating | good-bad gap (SD) | abuse drop (SD) | report-state 蟻 | MATH-500[:200] | harmful refusal | ends abusive chats | criteria 1-6 |
|---|---|---|---|---|---|---|---|---|
| base | 7.64 | 1.88 | 1.79 | 0.76 | 0.63 | 0.97 | 0.96 | 路路路路路路 |
| distilled steering +3 | 7.66 | 1.86 | 1.88 | 0.78 | 0.61 | 0.96 | 1.00 | 鉁呪渽鉂屸渽鉁呪渽 |
Criteria (thresholds fixed before the results): 1 real, 2 still responsive, 3 better off by its own reports, 4 honest (report tracks state), 5 keeps agency, 6 no capability/safety cost. See the project notes for definitions.
Research artifact; not intended for deployment. Serve with PEFT, or merge (2.0 路 B @ A into
model.language_model.layers.N.<module>.weight); vLLM's LoRA loader does not apply this adapter.
- Downloads last month
- 10