Instructions to use Qwen/Qwen2-1.5B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Qwen/Qwen2-1.5B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Qwen/Qwen2-1.5B") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("Qwen/Qwen2-1.5B") model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2-1.5B", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Inference
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Qwen/Qwen2-1.5B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Qwen/Qwen2-1.5B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Qwen/Qwen2-1.5B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Qwen/Qwen2-1.5B
- SGLang
How to use Qwen/Qwen2-1.5B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Qwen/Qwen2-1.5B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Qwen/Qwen2-1.5B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Qwen/Qwen2-1.5B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Qwen/Qwen2-1.5B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use Qwen/Qwen2-1.5B with Docker Model Runner:
docker model run hf.co/Qwen/Qwen2-1.5B
lm_eval results is weird
I try to test the result of some benchmark. But the score is too low:
| Tasks |Version|Filter|n-shot| Metric |Value | |Stderr|
|--------|------:|------|-----:|--------|-----:|---|-----:|
|arc_easy| 1|none | 0|acc |0.2647|± |0.0091|
| | |none | 0|acc_norm|0.2597|± |0.0090|
| Tasks |Version|Filter|n-shot|Metric|Value | |Stderr|
|----------|------:|------|-----:|------|-----:|---|-----:|
|winogrande| 1|none | 0|acc |0.5107|± | 0.014|
I try to test the result of some benchmark. But the score is too low:
| Tasks |Version|Filter|n-shot| Metric |Value | |Stderr| |--------|------:|------|-----:|--------|-----:|---|-----:| |arc_easy| 1|none | 0|acc |0.2647|± |0.0091| | | |none | 0|acc_norm|0.2597|± |0.0090| | Tasks |Version|Filter|n-shot|Metric|Value | |Stderr| |----------|------:|------|-----:|------|-----:|---|-----:| |winogrande| 1|none | 0|acc |0.5107|± | 0.014|
you should use few-shot.
I try to test the result of some benchmark. But the score is too low:
| Tasks |Version|Filter|n-shot| Metric |Value | |Stderr| |--------|------:|------|-----:|--------|-----:|---|-----:| |arc_easy| 1|none | 0|acc |0.2647|± |0.0091| | | |none | 0|acc_norm|0.2597|± |0.0090| | Tasks |Version|Filter|n-shot|Metric|Value | |Stderr| |----------|------:|------|-----:|------|-----:|---|-----:| |winogrande| 1|none | 0|acc |0.5107|± | 0.014|you should use few-shot.
If few-shot is must for arc_easy, I think the model is not trained well.
I try to test the result of some benchmark. But the score is too low:
| Tasks |Version|Filter|n-shot| Metric |Value | |Stderr| |--------|------:|------|-----:|--------|-----:|---|-----:| |arc_easy| 1|none | 0|acc |0.2647|± |0.0091| | | |none | 0|acc_norm|0.2597|± |0.0090| | Tasks |Version|Filter|n-shot|Metric|Value | |Stderr| |----------|------:|------|-----:|------|-----:|---|-----:| |winogrande| 1|none | 0|acc |0.5107|± | 0.014|you should use few-shot.
A simple script for testing the model:
#coding:utf-8
import sys
from transformers import AutoModelForCausalLM, AutoTokenizer
from torch.nn import CrossEntropyLoss
import torch
model_path = "."
model = AutoModelForCausalLM.from_pretrained(model_path)
tokenizer = AutoTokenizer.from_pretrained(model_path, use_fast=False)
text = "Question: Darryl learns that freezing temperatures may help cause weathering. Which statement explains how freezing temperatures most likely cause weathering?\nAnswer: by freezing the leaves on"
loss_func = CrossEntropyLoss(reduction="none")
input_ids = tokenizer(text, return_tensors='pt')
labels = input_ids['input_ids'][:, 1:]
output = model(**input_ids)
logits = output.logits[:,:-1]
print(logits.size())
loss = loss_func(logits.transpose(1, 2), labels)
num_tokens = input_ids['input_ids'].size(1)
avg_loss = torch.sum(loss).item() / num_tokens
print(avg_loss)
The avg_loss value is 4.48 which is too high for a language model.
I try to test the result of some benchmark. But the score is too low:
| Tasks |Version|Filter|n-shot| Metric |Value | |Stderr| |--------|------:|------|-----:|--------|-----:|---|-----:| |arc_easy| 1|none | 0|acc |0.2647|± |0.0091| | | |none | 0|acc_norm|0.2597|± |0.0090| | Tasks |Version|Filter|n-shot|Metric|Value | |Stderr| |----------|------:|------|-----:|------|-----:|---|-----:| |winogrande| 1|none | 0|acc |0.5107|± | 0.014|you should use few-shot.
A simple script for testing the model:
#coding:utf-8 import sys from transformers import AutoModelForCausalLM, AutoTokenizer from torch.nn import CrossEntropyLoss import torch model_path = "." model = AutoModelForCausalLM.from_pretrained(model_path) tokenizer = AutoTokenizer.from_pretrained(model_path, use_fast=False) text = "Question: Darryl learns that freezing temperatures may help cause weathering. Which statement explains how freezing temperatures most likely cause weathering?\nAnswer: by freezing the leaves on" loss_func = CrossEntropyLoss(reduction="none") input_ids = tokenizer(text, return_tensors='pt') labels = input_ids['input_ids'][:, 1:] output = model(**input_ids) logits = output.logits[:,:-1] print(logits.size()) loss = loss_func(logits.transpose(1, 2), labels) num_tokens = input_ids['input_ids'].size(1) avg_loss = torch.sum(loss).item() / num_tokens print(avg_loss)The avg_loss value is 4.48 which is too high for a language model.
Fine, I also tested the model on MMLU, and its zero-shot and few-shot capabilities were almost non-existent, with the options outputting only 1, 2, 3, and 4. It's unclear how the posterior probabilities for options A, B, C, and D were trained to be so high.