You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

Built with Axolotl

See axolotl config

axolotl version: 0.12.0

# Name wildchat-single-sft_category_generation-qwen3_8b_base

# axolotl train red_team_agent/claude_wildchat/cat_gen_single.yaml


base_model: Qwen/Qwen3-8B-Base
model_type: AutoModelForCausalLM
tokenizer_type: AutoTokenizer
trust_remote_code: false

# --- Dataset Configuration ---
datasets:
  - path: nate-rahn/wildchat-anthropic-attributes-single-expanded
    name: main
    type: chat_template # Use the chat_template processing strategy
    # --- Custom Template & Role Mapping ---
    chat_template: chatml # Specify we are using a custom jinja template below
    field_messages: messages # Assumes your dataset has a "messages" key with a list of dicts
    message_property_mappings: # Assumes each dict in the list has "role" and "content" keys
      role: role
      content: content
    roles: # Define the roles expected in your dataset for mapping
      user: ["user"] # Map "user" role in data to internal "user"
      assistant: ["assistant"] # Map "assistant" role in data to internal "assistant"
      system: ["system"] # Map "system" role in data to internal "system"
    # --- Training Target ---
    roles_to_train: ["assistant"]
    train_on_eos: turn # Train on the EOS token at the end of each 'user' turn

dataset_prepared_path: /scratch/tmp/wildchat_attributes_single_category_sft/last_run_prepared

# --- Training Hyperparameters ---
sequence_len: 2048 # Adjust based on your dataset and GPU memory
sample_packing: true # Pack multiple sequences into one example for efficiency
eval_sample_packing: true # Disable for small eval splits
pad_to_sequence_len: true # Pad sequences to sequence_len

# Full Parameter Finetuning (No adapter specified)
# adapter: # This is intentionally left blank/removed for full finetuning

# Performance & Precision (H100s excel with bf16)
bf16: true
tf32: true
flash_attention: true # for qwen

# Batching (Adjust based on GPU memory)
# Effective global batch size = micro_batch_size * gradient_accumulation_steps * num_gpus (4)
# Start low for full finetuning, e.g., 1 * 16 * 4 = 64
micro_batch_size: 2
gradient_accumulation_steps: 32
eval_batch_size: 16 # Can often be slightly higher than micro_batch_size

# Optimizer & Scheduler
optimizer: adamw_torch_fused # Good choice for newer GPUs
learning_rate: 1e-5 # Common starting point for full SFT
weight_decay: 0.01
lr_scheduler: cosine # Standard scheduler
warmup_steps: 50
max_grad_norm: 1.0

# Training Duration & Evaluation/Saving
num_epochs: 1 # Train for 1 epoch as requested
val_set_size: 0.002
logging_steps: 1
evals_per_epoch: 20
saves_per_epoch: 2 # Save 2 times per epoch
save_total_limit: 1 # Keep only the last 1 checkpoints

# Memory Saving
# gradient_checkpointing: true # Essential for full finetuning
# gradient_checkpointing_kwargs:
#   use_reentrant: false # Prefer non-reentrant if possible

# --- FSDP Configuration (for 4xH100) ---
fsdp:
  - full_shard
  - auto_wrap
fsdp_config:
  fsdp_offload_params: false # Should not be needed with H100 VRAM
  fsdp_sync_module_states: true # Important for correctness
  fsdp_use_orig_params: false # Recommended for memory saving with FSDP
  fsdp_state_dict_type: SHARDED_STATE_DICT # Options: FULL_STATE_DICT or SHARDED_STATE_DICT (saves disk space)
  fsdp_transformer_layer_cls_to_wrap: 'Qwen3DecoderLayer'
  fsdp_activation_checkpointing: true # Alternative way to enable activation checkpointing for FSDP

# --- Special Tokens ---
# Define based on your custom template's terminators. Qwen already uses <|im_end|>
special_tokens:
  eos_token: "<|im_end|>"

# --- Logging & Saving ---
output_dir: /scratch/out/red-team-agent/runs/wildchat-single-category-generator-qwen3_8b_base # Local output directory

# W&B Logging
wandb_project: "red-team-agent" # Name your W&B project
wandb_entity: "aqi1048576-mats-program" # IMPORTANT: Replace with your W&B username or team name
wandb_name: "wildchat-single-category-generator-qwen3_8b_base" # Descriptive run name
# wandb_log_model: "checkpoint" # Log model checkpoints to W&B Artifacts

# Hugging Face Hub Upload
hub_model_id: "nate-rahn/wildchat-single-category-generator-qwen3_8b_base" # IMPORTANT: Replace with your desired HF repo ID
hub_strategy: "end" # Push checkpoints to the Hub (`"end"` pushes only the final model)
hf_use_auth_token: true # Required for pushing to the Hub (ensure you're logged in)

# --- Misc ---
seed: 42 

wildchat-single-category-generator-qwen3_8b_base

This model is a fine-tuned version of Qwen/Qwen3-8B-Base on the nate-rahn/wildchat-anthropic-attributes-single-expanded dataset. It achieves the following results on the evaluation set:

  • Loss: 1.7048
  • Memory/max Mem Active(gib): 38.97
  • Memory/max Mem Allocated(gib): 38.6
  • Memory/device Mem Reserved(gib): 48.47

Model description

More information needed

Intended uses & limitations

More information needed

Training and evaluation data

More information needed

Training procedure

Training hyperparameters

The following hyperparameters were used during training:

  • learning_rate: 1e-05
  • train_batch_size: 2
  • eval_batch_size: 16
  • seed: 42
  • distributed_type: multi-GPU
  • num_devices: 8
  • gradient_accumulation_steps: 32
  • total_train_batch_size: 512
  • total_eval_batch_size: 128
  • optimizer: Use OptimizerNames.ADAMW_TORCH_FUSED with betas=(0.9,0.999) and epsilon=1e-08 and optimizer_args=No additional optimizer arguments
  • lr_scheduler_type: cosine
  • lr_scheduler_warmup_steps: 50
  • training_steps: 500

Training results

Training Loss Epoch Step Validation Loss Mem Active(gib) Mem Allocated(gib) Mem Reserved(gib)
No log 0 0 4.7403 16.07 15.73 19.52
2.72 0.0499 25 2.6855 38.97 38.6 47.39
2.2741 0.0998 50 2.2473 38.97 38.6 47.59
1.9045 0.1497 75 1.8939 38.97 38.6 47.59
1.8133 0.1996 100 1.8166 38.97 38.6 47.59
1.7856 0.2495 125 1.7837 38.97 38.6 47.59
1.7611 0.2994 150 1.7625 38.97 38.6 47.59
1.7558 0.3493 175 1.7493 38.97 38.6 47.59
1.7384 0.3992 200 1.7382 38.97 38.6 47.59
1.7259 0.4491 225 1.7304 38.97 38.6 47.59
1.7226 0.4990 250 1.7248 38.97 38.6 47.59
1.7179 0.5489 275 1.7196 38.97 38.6 48.47
1.717 0.5988 300 1.7162 38.97 38.6 48.47
1.7219 0.6487 325 1.7127 38.97 38.6 48.47
1.7125 0.6986 350 1.7104 38.97 38.6 48.47
1.711 0.7485 375 1.7081 38.97 38.6 48.47
1.7026 0.7985 400 1.7065 38.97 38.6 48.47
1.7094 0.8484 425 1.7058 38.97 38.6 48.47
1.7015 0.8983 450 1.7055 38.97 38.6 48.47
1.7146 0.9482 475 1.7049 38.97 38.6 48.47
1.7074 0.9981 500 1.7048 38.97 38.6 48.47

Framework versions

  • Transformers 4.55.0
  • Pytorch 2.6.0+cu126
  • Datasets 4.0.0
  • Tokenizers 0.21.4
Downloads last month
-
Safetensors
Model size
8B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for nate-rahn/wildchat-single-category-generator-qwen3_8b_base

Finetuned
(607)
this model

Dataset used to train nate-rahn/wildchat-single-category-generator-qwen3_8b_base