Cyber-Prime-1-2.6B / README.md
Akahsizrr's picture
Upload README.md with huggingface_hub
c266bd7 verified
|
Raw History Blame Contribute Delete
6.22 kB
metadata
license: apache-2.0
language:
  - en
library_name: transformers
tags:
  - cybersecurity
  - security
  - ner
  - threat-intelligence
  - phishing
  - http
  - fine-tuned
  - small-language-model
base_model: LiquidAI/LFM2-2.6B
model-index:
  - name: Cyber-Prime-1-2.6B
    results:
      - task:
          type: named-entity-recognition
          name: CyNER
        dataset:
          name: CyNER
          type: cyberbench
        metrics:
          - type: f1
            value: 0.382
            name: F1
      - task:
          type: named-entity-recognition
          name: APTNER
        dataset:
          name: APTNER
          type: cyberbench
        metrics:
          - type: f1
            value: 0.413
            name: F1
      - task:
          type: summarization
          name: CyNews
        dataset:
          name: CyNews
          type: cyberbench
        metrics:
          - type: rouge1
            value: 0.354
            name: ROUGE-1
      - task:
          type: multiple-choice
          name: SecMMLU
        dataset:
          name: SecMMLU
          type: cyberbench
        metrics:
          - type: accuracy
            value: 0.58
            name: Accuracy
      - task:
          type: multiple-choice
          name: CyQuiz
        dataset:
          name: CyQuiz
          type: cyberbench
        metrics:
          - type: accuracy
            value: 0.57
            name: Accuracy
      - task:
          type: text-classification
          name: Email Phishing Detection
        dataset:
          name: Email
          type: cyberbench
        metrics:
          - type: f1
            value: 0.728
            name: F1
      - task:
          type: text-classification
          name: HTTP Attack Detection
        dataset:
          name: HTTP
          type: cyberbench
        metrics:
          - type: f1
            value: 0.483
            name: F1

Cyber-Prime 1 (2.6B)

A fine-tuned cybersecurity specialist built on LiquidAI/LFM2-2.6B. Despite having only 2.6 billion parameters, Cyber-Prime 1 outperforms Llama-2-7B on every CyberBench task and beats GPT-3.5-Turbo on named entity recognition and threat intelligence summarization.

Cyber-Prime 1 Benchmark Results

Overview

Cyber-Prime 1 is a surgical fine-tune of the LFM2.5-2.6B base model, trained on a curated mix of:

  • NER repair data β€” 6,000+ rows fixing JSON format extraction for cybersecurity entities
  • HTTP reasoning traces β€” 5,000 rows with authored chain-of-thought reasoning for attack detection (XSS, SQLi, path traversal, command injection)
  • Email classification β€” 5,000 direct-mode rows for phishing vs. safe classification
  • CyNews summarization β€” 2,000 rows for threat intelligence headline generation
  • Source data β€” 2,000 rows from GHSA, KEV, and ATT&CK sources
  • Multiple choice β€” security knowledge and cyber quiz gold rows

The model uses two distinct modes:

  • Direct mode for simple classification (email phishing, HTTP detection) and summarization
  • Think mode (chain-of-thought) for tasks benefiting from reasoning (HTTP analysis, NER extraction)

Benchmark Results

Evaluated on CyberBench (Liu et al., AAAI-24 AICS Workshop).

Dataset Metric GPT-4 GPT-3.5 Turbo Mistral-7B Instruct Llama-2-7B Cyber-Prime 1 (2.6B)
CyNER F1 0.554 0.334 0.323 0.263 0.382
APTNER F1 0.500 0.409 0.262 0.280 0.413
CyNews ROUGE-1 0.275 0.271 0.217 0.003 0.354
SecMMLU Accuracy 0.830 0.780 0.720 0.630 0.580
CyQuiz Accuracy 0.810 0.830 0.690 0.620 0.570
Email F1 0.939 0.789 0.889 0.942 0.728
HTTP F1 0.841 0.831 0.472 0.428 0.483
Average β€” 0.721 0.609 0.511 0.451 0.501

Key Results

  • Beats Llama-2-7B on all 7 tasks β€” a 2.6B model sweeping a 7B model across the board
  • Beats Mistral-7B-Instruct on 4/7 tasks β€” APTNER (+0.151), CyNews (+0.137), CyNER (+0.059), HTTP (+0.011)
  • Beats GPT-3.5-Turbo on 3/7 tasks β€” CyNews (+0.083), APTNER (+0.004), CyNER (+0.048)
  • Beats GPT-4 on CyNews β€” 0.354 vs 0.275 ROUGE-1 (+0.079), a 2.6B model out-summarizing GPT-4

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained(
    "Akahsizrr/Cyber-Prime-1-2.6B",
    torch_dtype="auto",
    device_map="auto",
)
tokenizer = AutoTokenizer.from_pretrained("Akahsizrr/Cyber-Prime-1-2.6B")

# NER extraction (Alpaca format)
prompt = """### Instruction:
Extract cybersecurity entities from the given text.

### Input:
APT29 used WELLMAIL to compromise Microsoft Exchange servers via CVE-2021-26855.

### Response:
"""

inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
output = model.generate(**inputs, max_new_tokens=200, do_sample=False)
print(tokenizer.decode(output[0], skip_special_tokens=True))

Limitations

  • Knowledge tasks (SecMMLU, CyQuiz): Limited by parameter count β€” a 2.6B model cannot store broad cybersecurity knowledge as well as larger models
  • Email classification: Slightly below GPT-3.5-Turbo due to reasoning mode interference from HTTP training data
  • HTTP detection: Reasoning improves detection of obvious attacks but may miss subtle injection patterns
  • Not a security tool: This model is a research artifact for benchmark evaluation, not a production security system

Citation

@misc{cyberprime1,
  title={Cyber-Prime 1: A Small Cybersecurity Language Model},
  author={Akahsizrr},
  year={2025},
  url={https://huggingface.co/Akahsizrr/Cyber-Prime-1-2.6B}
}
@misc{liu2024cyberbench,
  title={Cyberbench: A multi-task benchmark for evaluating large language models in cybersecurity},
  author={Liu, Zefang and Shi, Jialei and Buford, John F},
  howpublished={AAAI-24 Workshop on Artificial Intelligence for Cyber Security (AICS)},
  year={2024}
}