SWE-ZERO
Collection
6 items • Updated
Qwen3-1.7B-Base SFT on a 1M random sample of SWE-ZERO trajectories (superset of the 10K and 100K, right-truncated to 8K tokens), evaluated on SWE-bench Verified.
pass@1 = 11/100 = 11% on the 100-task SWE-bench Verified slice (latest trial per task) — beats 100K @ 9% and 10K @ 7%.
For full eval details + per-task trajectories, see the eval dataset: AlienKevin/SWE-ZERO-1M-Qwen3-1.7B-Base-eval.
Qwen/Qwen3-1.7B-Basemax_seq_len=32768 (Llama 3 RoPE scaling from 8192)eos_token_id
The HF config has eos_token_id: [151643, 151645] so vLLM stops at both <|endoftext|> AND <|im_end|> (Qwen3 chat-template turn boundary).
from vllm import LLM, SamplingParams
llm = LLM(model="AlienKevin/SWE-ZERO-1M-Qwen3-1.7B-Base", max_model_len=32768)
params = SamplingParams(temperature=1.0, max_tokens=4096)
See the eval dataset for the harbor + mini-swe-agent v1 config used in our results.
Base model
Qwen/Qwen3-1.7B-Base