Human-annotated intent dataset and intent-aware safety classifiers (SFT, DPO, distillation, GRPO) for robust LLM guardrails.
Jeremias Ferrao
Jazhyc
·
AI & ML interests
None yet
Recent Activity
updated a dataset about 9 hours ago
Jazhyc/aims-safety-intents updated a model about 9 hours ago
Jazhyc/Llama-3.1-8B-aims-grpo updated a model about 9 hours ago
Jazhyc/Llama-3.1-8B-aims-distill-synthetic-intentOrganizations
None yet