AhıskaAI
community
AI & ML interests
LLMs & SLMs (Large & Small Language Models) NLP (Natural Language Processing) OCR (Optical Character Recognition) Pre-training & Fine-tuning Data Engineering / Data Preprocessing
Recent Activity
Organization Card
AhıskaAI Research Lab
An independent, open-source AI research lab focused on Small Language Models (SLMs), custom tokenization, and dialect preservation.
🔍 Overview
AhıskaAI studies the end-to-end lifecycle of light-weight language models. We focus on training custom Transformer architectures from scratch, engineering domain-specific tokenizers, and developing structured dataset pipelines for low-resource NLP. We document the complete process of model development, from pre-training setups to training logs, to maintain transparency in our open-source workflow.
🎯 Primary Research Areas
- Base SLMs: Custom Transformer and Llama-based architectures trained from scratch (ranging from 10M to 135M+ parameters) using specialized Byte-Pair Encoding (BPE) vocabularies.
- Supervised Fine-Tuning (SFT): Context-driven instruction tuning, synthetic dialogue generation, and domain-specific alignment.
- Low-Resource Language Preservation: Building parallel corpora and neural translation pipelines for regional dialects, featuring the first open parallel dataset for Ahıska Turkish.
- Open Pipelines: Publishing reproducible Python scripts, data processing workflows, and custom tokenizers to the public ecosystem.
🛠️ Technical Stack & Training Pipeline
- Core Libraries: PyTorch, Hugging Face (Transformers, Datasets, Accelerate, Tokenizers)
- Architectures: Custom Llama / Transformer SLMs (10M - 135M+ parameters)
- Training Methods: From-scratch pre-training, SFT, FP16/BF16 mixed precision, gradient accumulation
- Compute Environment:
- 💻 Local: NVIDIA RTX 4050 Laptop GPU (6GB VRAM)
- ☁️ Distributed Cloud: Dual Tesla T4 GPUs (30GB total VRAM via Kaggle)
🔗 Organization & Developer Links
- GitHub Organization: github.com/AhiskaAI
- Hugging Face Hub: huggingface.co/AhiskaAI
- Lead Researcher GitHub: github.com/YunusEmreSaidoglu
- Lead Researcher Hugging Face: huggingface.co/YunusEmreSaidoglu
Made with love for Ahıska. ❤️
AhıskaAI Experimental v0.2 Model Series
models 20
AhiskaAI/AhiskaAI-Experimental-v0.2-235m
Text Generation • 0.3B • Updated • 91 • 1
AhiskaAI/AhiskaAI-Experimental-v0.2-135m
Text Generation • 0.2B • Updated • 95 • 1
AhiskaAI/AhiskaAI-Experimental-v0.2-235m-IT
Text Generation • 0.3B • Updated • 108 • 1
AhiskaAI/AhiskaAI-110M-Experimental-v0.1-Base
Text Generation • 0.1B • Updated • 125 • 1
AhiskaAI/AhiskaAI-10M-Experimental-v0.1-Base
Text Generation • 10.3M • Updated • 122 • 1
AhiskaAI/AhiskaAI-135m-Base-v0.3
Text Generation • 0.1B • Updated • 374 • 1
AhiskaAI/AhiskaAI-10M-Experimental-v0.1-Instruct
Text Generation • 10.3M • Updated • 283 • 1
AhiskaAI/AhiskaAI-110M-Experimental-v0.1-Instruct
Text Generation • 0.1B • Updated • 283 • 1
AhiskaAI/AhiskaAI-135m-Instruct-v0.3
Text Generation • 0.1B • Updated • 211 • 1
AhiskaAI/AhiskaAI-134m-Base-v0.2
Text Generation • 0.1B • Updated • 321 • 1
datasets 5
AhiskaAI/Ahiska-Turkish-Language-Dataset
Viewer • Updated • 803k • 65 • 1
AhiskaAI/AhiskaAI-Instruct-v0.2-Sample-Dataset
Viewer • Updated • 1.51k • 45 • 1
AhiskaAI/python-instruct-turkish
Viewer • Updated • 10.8k • 7 • 2
AhiskaAI/ahiska-history-synthetic-turkish
Viewer • Updated • 2.37k • 5 • 1
AhiskaAI/sharegpt-turkish
Viewer • Updated • 13.7k • 29 • 1