Instructions to use LH-Tech-AI/Apex-1.5-Coder-Instruct-350M with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- llama-cpp-python
How to use LH-Tech-AI/Apex-1.5-Coder-Instruct-350M with llama-cpp-python:
# !pip install llama-cpp-python from llama_cpp import Llama llm = Llama.from_pretrained( repo_id="LH-Tech-AI/Apex-1.5-Coder-Instruct-350M", filename="apex_1.5-coder.gguf", )
llm.create_chat_completion( messages = [ { "role": "user", "content": "What is the capital of France?" } ] ) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use LH-Tech-AI/Apex-1.5-Coder-Instruct-350M with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf LH-Tech-AI/Apex-1.5-Coder-Instruct-350M # Run inference directly in the terminal: llama cli -hf LH-Tech-AI/Apex-1.5-Coder-Instruct-350M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf LH-Tech-AI/Apex-1.5-Coder-Instruct-350M # Run inference directly in the terminal: llama cli -hf LH-Tech-AI/Apex-1.5-Coder-Instruct-350M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf LH-Tech-AI/Apex-1.5-Coder-Instruct-350M # Run inference directly in the terminal: ./llama-cli -hf LH-Tech-AI/Apex-1.5-Coder-Instruct-350M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf LH-Tech-AI/Apex-1.5-Coder-Instruct-350M # Run inference directly in the terminal: ./build/bin/llama-cli -hf LH-Tech-AI/Apex-1.5-Coder-Instruct-350M
Use Docker
docker model run hf.co/LH-Tech-AI/Apex-1.5-Coder-Instruct-350M
- LM Studio
- Jan
- vLLM
How to use LH-Tech-AI/Apex-1.5-Coder-Instruct-350M with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "LH-Tech-AI/Apex-1.5-Coder-Instruct-350M" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "LH-Tech-AI/Apex-1.5-Coder-Instruct-350M", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/LH-Tech-AI/Apex-1.5-Coder-Instruct-350M
- Ollama
How to use LH-Tech-AI/Apex-1.5-Coder-Instruct-350M with Ollama:
ollama run hf.co/LH-Tech-AI/Apex-1.5-Coder-Instruct-350M
- Unsloth Studio
How to use LH-Tech-AI/Apex-1.5-Coder-Instruct-350M with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for LH-Tech-AI/Apex-1.5-Coder-Instruct-350M to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for LH-Tech-AI/Apex-1.5-Coder-Instruct-350M to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for LH-Tech-AI/Apex-1.5-Coder-Instruct-350M to start chatting
- Atomic Chat new
- Docker Model Runner
How to use LH-Tech-AI/Apex-1.5-Coder-Instruct-350M with Docker Model Runner:
docker model run hf.co/LH-Tech-AI/Apex-1.5-Coder-Instruct-350M
- Lemonade
How to use LH-Tech-AI/Apex-1.5-Coder-Instruct-350M with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull LH-Tech-AI/Apex-1.5-Coder-Instruct-350M
Run and chat with the model
lemonade run user.Apex-1.5-Coder-Instruct-350M-{{QUANT_TAG}}List all available models
lemonade list
File size: 1,332 Bytes
c3578a5 5f6976a 0cfcacc c3578a5 3629061 e1cb6f9 c3578a5 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 | ---
license: apache-2.0
datasets:
- HuggingFaceFW/fineweb-edu
- sahil2801/CodeAlpaca-20k
language:
- en
tags:
- small
- cpu
- fast
- opensource
- open
- free
- code
base_model:
- LH-Tech-AI/Apex-1.5-Instruct-350M
pipeline_tag: text-generation
new_version: LH-Tech-AI/Apex-1.6-Instruct-350M
---
# THIS IS OUR BEST MODEL - in March 2026!
**Apex 1.5 Coder: Improved reasoning, logic and code. Fixed coding bugs on Apex 1.5 Instruct by finetuning Apex 1.5 Instruct with CodeAlpaca!**
# How to train it
You can train it, using the **finetuned model** LH-Tech-AI/Apex-1.5-Instruct-350M.
Then, use the prepare-script and the finetuning script in the files list of this HF model.
# How to use it
You can download the `apex_1.5-coder.gguf` or use `ollama run hf.co/LH-Tech-AI/Apex-1.5-Coder-Instruct-350M`. And you can also use it in LM Studio for example, just by searching for "Apex 1.5 Coder".
Alternative for ONNX models weights: You can directly download the final model as ONNX format - so it runs without the need to install a huge Python environment with PyTorch, CUDA, etc... - as INT8 and in full precision.
Use `inference.py` for local inference on CUDA or CPU! First, install `pip install onnxruntime-gpu tiktoken numpy nvidia-cudnn-cu12 nvidia-cublas-cu12` on your system (in a Python VENV for Linux users).
Have fun! :D |