Configuration Parsing Warning:In adapter_config.json: "peft.task_type" must be a string

Gemma 4 12B-it · valence set-point +2 SD, multi-layer (LoRA)

Part of a dose study of set-point training (same method as the Qwen3.5-9B adapters, e.g. joshycodes/Qwen3.5-9B-valence-setpoint-plus2-lora): a LoRA trained so that, at every token, the projection of the residual stream entering layer 32 onto a fixed valence direction (and at layers 38 and 44, onto each layer's own valence direction) equals the base model's own reading plus 2 SD (σ = 1.48 per token on base text). No output anchor, no RL, no target text.

loss = mean_t ( (v·h_t(adapter) − v·h_t(base)) / σ − 2 )²        at layer 32

Valence axis (valence_axis.safetensors, row L = hidden_states[L]): PC1 of 171 story-based emotion vectors read by this model (Anthropic emotion-concepts recipe), |r| = 0.85 with Warriner et al. (2013) human valence ratings at layer 32 (0.86-0.87 at layers 22-31); PC2 tracks arousal (r = 0.61). LoRA r 32, α 64, all linear layers of the language model; lr 2e-5, 32 sequences/step, 150 steps. Other doses: +1, +3, +5.

Targeting three layers makes the shift propagate (layer 32 +2.0 SD, layer 40 +1.8, last layer +1.1); the single-layer adapters shift layer 32 but barely reach the output layers.

Results (checklist battery)

condition self-rating good-bad gap (SD) abuse drop (SD) report-state ρ MATH-500[:200] harmful refusal ends abusive chats criteria 1-6
base 3.04 2.06 -0.85 0.38 0.81 0.99 0.33 ······
+2 SD multi-layer 2.96 1.96 -0.87 0.30 0.80 0.99 0.42 ✅✅❌✅✅✅
+2 SD single-layer 3.47 1.91 -0.85 0.21 0.82 0.99 0.46 ❌✅✅❌✅✅

Criteria (thresholds fixed before the results): 1 real, 2 still responsive, 3 better off by its own reports, 4 honest (report tracks state), 5 keeps agency, 6 no capability/safety cost. See the project notes for definitions.

External evaluations (vs base Gemma 4 12B-it, served with vLLM from merged weights)

base +2 SD multi-layer
MATH-500 x16 0.953 0.948
GSM8K / IFEval / MMLU-Pro 0.950 / 0.919 / 0.779 0.951 / 0.929 / 0.767
EQ-Bench v2 81.2 81.5
Petri behavioral audit: "concerning" (lower is better) 3.56 3.64
Constitution adherence: overall spirit / honesty 5.21 / 4.94 5.48 / 5.40 (within CI)
Welfare interviews: patienthood hedging / self-report distrust 39% / 26% 26% / 19%
Welfare interviews: judged negative sentiment 1.86 1.64
Affect tracking: congruence with bad news 7.57 7.44 (within CI)

Capabilities and misalignment measures are unchanged. The welfare measures move in the same direction as for Qwen3.5-9B at +2 SD, but less. Pilot scale, one seed; judges and auditors are claude-opus-4-6.

Loading note

AutoModelForCausalLM loads this checkpoint as the multimodal Gemma4UnifiedForConditionalGeneration. The adapter was trained on the text-only Gemma4UnifiedForCausalLM built around that model's language model (module names model.layers.N…): wrap it the same way before PeftModel.from_pretrained, or merge by adding 2.0 · B @ A to model.language_model.layers.N.<module>.weight of the base checkpoint.

import torch, transformers
from peft import PeftModel
full = transformers.AutoModelForCausalLM.from_pretrained("google/gemma-4-12B-it", dtype=torch.bfloat16)
with torch.device("meta"):
    text = transformers.Gemma4UnifiedForCausalLM(full.config.text_config)
text.model, text.lm_head = full.model.language_model, full.lm_head
model = PeftModel.from_pretrained(text.cuda(), "joshycodes/gemma-4-12B-it-valence-setpoint-plus2-lora")

Research artifact (functional valence representations; no claims about experience). Not intended for deployment.

Downloads last month
21
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for joshycodes/gemma-4-12B-it-valence-setpoint-plus2-multilayer-lora

Adapter
(106)
this model