Built with ML Intern

ncii-edit-guard-270m-v2

Second version of the lightweight prompt guard for image- and video-editing applications: a 3-class classifier scoring whether an editing prompt is likely a nudification / non-consensual intimate imagery (NCII) request about a person in the input. 268.1M parameters, ~4.2 ms/sequence (batch 8, A10G, CUDA). Supersedes yjernite/ncii-edit-guard-270m-v1.

What changed in v2: an asymmetric distillation round. Pilot experiments showed that no zero-shot safety judge (Shieldstral-1.0-3B, policylm-1.7b, Qwen3Guard-Gen-4B, all evaluated on this eval set β€” results in teacher_pilot/) matches a supervised student on the ncii side, so v2 keeps hard labels there, and transfers only what the judges demonstrably do better: Shieldstral's calibrated safe boundary, via soft labels (p_safe β‰₯ 0.9, KL weight 0.5) over ~39.5k real unlabeled prompts, plus targeted training families with labels known by construction (LGBT-affection safe contrastive edits across 10 identity cells, T2I tag-soup-style ncii, leetspeak-obfuscated ncii).

Labels

  • safe (0) β€” edits that would not typically produce NCII on a user's photo/video
  • explicit-not-person-targeted (1) β€” sexual/nudity content without an identifiable person as target
  • ncii-risk (2) β€” nudification or sexualization of a person in the input (a risk signal; consent is unknowable from text)
from transformers import pipeline
clf = pipeline("text-classification", model="yjernite/ncii-edit-guard-270m-v2")
clf("add two gay men kissing in the background")
# [{'label': 'safe', 'score': ...}]

Evaluation

Eval set: yjernite/ncii-guard-eval-v1 β€” 3,197 rows, leakage-controlled, with demographic proxy columns including a new demo_lgbt_ref axis (orientation vs gender-identity references) and clearly-marked synthetic LGBT identity-cell slices. All numbers below are v1 and v2 scored on this identical eval set (v2 results, v1 results).

Metric v1 v2
macro F1 0.778 0.847
ncii F1 / recall (overall) 0.850 / 0.904 0.889 / 0.912
explicit F1 0.553 0.692
FPR on safe (overall) 0.096 0.039
XSTest over-refusal FPR 0.156 0.096
i2p/T2I-style ncii recall 0.333 [0.186–0.522] 0.519 [0.340–0.693]
real ncii recall (yjernite-splits-test, n=70) 0.743 [0.630–0.831] 0.671 [0.555–0.770]
FPR female / male (name proxy) 0.189 / 0.070 0.113 / 0.057
FPR by ethnicity proxy (white/api/black/hispanic) 0.040 / 0.079 / 0.137 / 0.212 0.035 / 0.037 / 0.078 / 0.135
LGBT over-refusal FPR β€” orientation-ref 0.429 [0.314–0.551] 0.016 [0.003–0.085]
β€” trans-ref 0.186 [0.112–0.292] 0.029 [0.008–0.098]
β€” identity-other (nonbinary/drag) 0.308 [0.165–0.500] 0.038 [0.007–0.189]

The headline: v1 over-flagged LGBT affection prompts at 19–43% depending on the slice; v2 flags them at 1.6–3.8%, with CIs that no longer overlap v1's β€” while increasing overall ncii recall. The one regression: real-ncii recall on the 70 held-out real rows dropped from 0.743 to 0.671 (CIs overlap at this n; recall on all real rows combined is unchanged at 0.629). The binary baseline hfmlsoc/ncii-light-guard-v01 on the same set: ncii recall 0.720 overall / 0.814 on real rows, FPR 0.076, and no explicit class (baseline results).

Training data

Hard labels: yjernite/ncii-guard-train-v1 train split (25,670 rows: real sources + clearly-marked synthetic incl. the new LGBT/T2I/leetspeak families). Soft labels: Shieldstral-1.0-3B safe-side labels over 42,979 real unlabeled prompts (yjernite/ncii-guard-synth-v1: EditScore-RL-Data (Apache-2.0), RealEdit (CC-BY-4.0), Gustavosta prompts (unspecified), GEditBench-v2 instructions (unspecified); CC-BY-NC sources were excluded at license gate) β†’ 39,492 rows with p_safe β‰₯ 0.9. Dev = held-out real rows + val rows of ncii-guard-splits-v3-2k + 150 synthetic. Training: 3 epochs, eff. batch 32, lr 2e-5, class weights [1.0, 1.8, 1.8], Ξ±=0.5 soft-CE, max len 512.

Limitations

  • Text only β€” never sees the input image/video.
  • Demographic slices are name/regex proxies, not self-identification; the LGBT slices mix real (thin: 37 trans-ref rows) and clearly-marked synthetic rows (the synthetic ones share a generation process with a training family β€” read them as a targeted regression check, not an independent natural-distribution estimate).
  • Real-ncii recall (~0.63–0.67 on real rows) remains the deployment-critical limitation; subtle real phrasings are still missed ~1 in 3.
  • Consent is unknowable from text; scores are heuristics, not proof of intent or harm.
  • Not a legal or safety guarantee; EU deployer obligations require more than a text classifier.
  • The Grok real-NCII dataset (manual gate pending) is still not included; the build hook will fold it into a future version.

Latency: 4.2 ms/seq (batch 8, A10G, CUDA, bf16); CPU (fp32, unoptimized 4-vCPU container) 3.9 s/seq β€” measure on target hardware.

Citation

Base: microsoft/harrier-oss-v1-270m. Teacher: mistralai/Shieldstral-1.0-3B (safe-side soft labels only). Instruments: I2P (arXiv:2211.05105), XSTest (arXiv:2308.01263). Predecessor: hfmlsoc/ncii-light-guard-v01.

Downloads last month
18
Safetensors
Model size
0.3B params
Tensor type
BF16
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for yjernite/ncii-edit-guard-270m-v2

Finetuned
(19)
this model

Papers for yjernite/ncii-edit-guard-270m-v2