SWE-ZERO-1M-Qwen3-1.7B-Base

Qwen3-1.7B-Base SFT on a 1M random sample of SWE-ZERO trajectories (superset of the 10K and 100K, right-truncated to 8K tokens), evaluated on SWE-bench Verified.

Eval result

pass@1 = 11/100 = 11% on the 100-task SWE-bench Verified slice (latest trial per task) — beats 100K @ 9% and 10K @ 7%.

For full eval details + per-task trajectories, see the eval dataset: AlienKevin/SWE-ZERO-1M-Qwen3-1.7B-Base-eval.

Training

  • Base: Qwen/Qwen3-1.7B-Base
  • SFT data: 1M random sample from AlienKevin/SWE-ZERO-12M-trajectories @ 2f328e1d (superset of the 10K and 100K), right-truncated to 8K tokens
  • Model arch: max_seq_len=32768 (Llama 3 RoPE scaling from 8192)
  • TPU: v5p-32 (~24h training, including one preemption-resume)
  • Optimizer: AdamW (β1=0.9, β2=0.95, ε=1e-8), lr=2e-5, weight_decay=0.1, max_grad_norm=30, cosine schedule, warmup=0.03, min_lr_ratio=0.1
  • Batch: 64 global, 15625 steps (1 epoch)
  • Tracking: marin#5611

Critical detail: eos_token_id

The HF config has eos_token_id: [151643, 151645] so vLLM stops at both <|endoftext|> AND <|im_end|> (Qwen3 chat-template turn boundary).

Inference

from vllm import LLM, SamplingParams

llm = LLM(model="AlienKevin/SWE-ZERO-1M-Qwen3-1.7B-Base", max_model_len=32768)
params = SamplingParams(temperature=1.0, max_tokens=4096)

See the eval dataset for the harbor + mini-swe-agent v1 config used in our results.

Downloads last month
15
Safetensors
Model size
2B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for AlienKevin/SWE-ZERO-1M-Qwen3-1.7B-Base

Finetuned
(441)
this model

Collection including AlienKevin/SWE-ZERO-1M-Qwen3-1.7B-Base