Instructions to use maicomputer/gpt4-x-alpaca with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use maicomputer/gpt4-x-alpaca with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="maicomputer/gpt4-x-alpaca")# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("maicomputer/gpt4-x-alpaca") model = AutoModelForCausalLM.from_pretrained("maicomputer/gpt4-x-alpaca", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use maicomputer/gpt4-x-alpaca with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "maicomputer/gpt4-x-alpaca" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "maicomputer/gpt4-x-alpaca", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/maicomputer/gpt4-x-alpaca
- SGLang
How to use maicomputer/gpt4-x-alpaca with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "maicomputer/gpt4-x-alpaca" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "maicomputer/gpt4-x-alpaca", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "maicomputer/gpt4-x-alpaca" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "maicomputer/gpt4-x-alpaca", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use maicomputer/gpt4-x-alpaca with Docker Model Runner:
docker model run hf.co/maicomputer/gpt4-x-alpaca
which dataset?
gpt4all or something else? If gpt4all, hopefully it was on the unfiltered dataset with all the "as a large language model" removed.
I hope it's a gpt 4 dataset without some "I'm sorry, as a large language model" bullshit inside
It is gpt-4 self instruct.
I have Gpt4all installed. How do I install this?’
This has nothing to do with gpt4all, you pip install transformers and run inference on this model
It is gpt-4 self instruct.
will you be releasing the dataset for further research?
It is gpt-4 self instruct.
will you be releasing the dataset for further research?
Here is the dataset: https://github.com/teknium1/GPTeacher
Its the general instruct set
ValueError: Tokenizer class LlamaTokenizer does not exist or is not currently imported.
ValueError: Tokenizer class LlamaTokenizer does not exist or is not currently imported.
You are not in the right thread ! Try 'pip install git+https://github.com/huggingface/transformers' in your environment
Yes there is an issue with the class names in the model.config, open it up and change LLaMA to Llama
Also an issue in the tokenizer.config, change 512 to 2048 in the max token spot
Thank you, I got it now.
It seems weird that Llama recommends changing the config to 512 to make it fit better with GPUs, I always thought that the input size into LLM are fixed and lower length input are always padded to the maximum length anyways. A question that does not relate to this repo but:
How does reducing the sequence length to 512 during inference (like in Llama) help? Wouldn't the model just pad to the maximum size that it was trained on?