Instructions to use timdettmers/guanaco-33b-merged with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use timdettmers/guanaco-33b-merged with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="timdettmers/guanaco-33b-merged")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("timdettmers/guanaco-33b-merged") model = AutoModelForCausalLM.from_pretrained("timdettmers/guanaco-33b-merged", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use timdettmers/guanaco-33b-merged with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "timdettmers/guanaco-33b-merged" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "timdettmers/guanaco-33b-merged", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/timdettmers/guanaco-33b-merged
- SGLang
How to use timdettmers/guanaco-33b-merged with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "timdettmers/guanaco-33b-merged" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "timdettmers/guanaco-33b-merged", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "timdettmers/guanaco-33b-merged" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "timdettmers/guanaco-33b-merged", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use timdettmers/guanaco-33b-merged with Docker Model Runner:
docker model run hf.co/timdettmers/guanaco-33b-merged
Prompt Format?
Is this using the same format as the LoRA versions? https://huggingface.co/JosephusCheung/Guanaco & https://huggingface.co/KBlueLeaf/guanaco-7B-leh
LoRA Version's Format
### Instruction:
User: History User Input
Assistant: History Assistant Answer
### Input:
System: Knowledge
User: New User Input
### Response:
New Assistant Answer
Based on what @mljxy says on this discussion its not the same as your QLoRA? https://huggingface.co/TheBloke/guanaco-13B-GGML/discussions/1#646f9d3b6098ee820fbd4dba
header = "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the user's questions."
prompt_template = "### Human: {query}\n### Assistant:{response}"
Could you add a readme to these QLoRA models so that it's better understood how to properly use them?
I have read the source code, the prompt should be:
A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the user's questions.
### Human: {input}
### Assistant: {output}
I have read the source code, the prompt should be:
A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the user's questions. ### Human: {input} ### Assistant: {output}
I'm interested in how I can read the source code to a llm.
Thanks