ZYao720 commited on
Commit
f0a7e46
Β·
verified Β·
1 Parent(s): 1cb29d7

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +192 -5
README.md CHANGED
@@ -1,14 +1,201 @@
1
  ---
 
 
2
  license: apache-2.0
3
  library_name: transformers
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
4
  ---
5
 
6
- # Coming Soon
7
 
8
- This model will be released shortly. Stay tuned!
9
 
10
- **Paper**: [WebArbiter: A Principle-Guided Reasoning Process Reward Model for Web Agents](https://arxiv.org/abs/2601.21872)
11
 
12
- **Code**: [GitHub](https://github.com/YaoZhang720/WebArbiter)
13
 
14
- **Website**: [yaozhang.ai/WebArbiter](https://yaozhang.ai/WebArbiter/)
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
  ---
2
+ language:
3
+ - en
4
  license: apache-2.0
5
  library_name: transformers
6
+ pipeline_tag: text-generation
7
+ tags:
8
+ - web-agent
9
+ - process-reward-model
10
+ - preference
11
+ - reward-model
12
+ - web-navigation
13
+ - reasoning
14
+ - grpo
15
+ base_model: Qwen/Qwen2.5-7B-Instruct
16
+ datasets:
17
+ - ZYao720/WebArbiter-Data
18
+ model-index:
19
+ - name: WebArbiter-7B
20
+ results:
21
+ - task:
22
+ type: text-generation
23
+ name: Web Process Reward Modeling
24
+ dataset:
25
+ name: WebPRMBench
26
+ type: ZYao720/WEBPRMBENCH
27
+ metrics:
28
+ - name: Avg Pairwise Accuracy
29
+ type: accuracy
30
+ value: 89.19
31
+ - name: Avg BoN Accuracy
32
+ type: accuracy
33
+ value: 74.60
34
  ---
35
 
36
+ <div align="center">
37
 
38
+ # WebArbiter-7B
39
 
40
+ **A principle-guided reasoning Process Reward Model for web agents**
41
 
42
+ **Published at ICLR 2026**
43
 
44
+ [Paper](https://arxiv.org/abs/2601.21872) | [Code](https://github.com/YaoZhang720/WebArbiter) | [Website](https://yaozhang.ai/WebArbiter/) | [Collection](https://huggingface.co/collections/ZYao720/ZYao720-69cd5263871b22e11d90f80f) | [Demo](https://yaozhang.ai/WebArbiter/demo.html)
45
+
46
+ </div>
47
+
48
+ ## Introduction
49
+
50
+ **WebArbiter-7B** is a 7B reasoning Process Reward Model (PRM) for web agents, built on [Qwen2.5-7B-Instruct](https://huggingface.co/Qwen/Qwen2.5-7B-Instruct). Unlike scalar or checklist-based reward models, WebArbiter formulates step-level reward modeling as structured text generation β€” producing interpretable, principle-inducing justifications that conclude with a preference verdict identifying the action most conducive to task completion.
51
+
52
+ On [WEBPRMBENCH](https://huggingface.co/datasets/ZYao720/WEBPRMBENCH), WebArbiter-7B achieves an **Avg. BoN Acc of 74.60%**, outperforming GPT-5 by **9.1 points** and the previous SOTA WebPRM (WebShepherd-8B) by **31 points**. In reward-guided trajectory search on WebArena-Lite, it surpasses WebShepherd-8B by up to **6.4 points** in success rate.
53
+
54
+ ## Highlights
55
+
56
+ - **Reasoning as reward**: Generates structured `<State>`, `<Criteria>`, `<Analysis>`, and `<Answer>` outputs with auditable reasoning chains, instead of scalar scores or brittle checklists.
57
+ - **Principle-inducing evaluation**: Dynamically derives evaluation principles from user intent and page state, enabling robust assessment that generalizes across environments.
58
+ - **Two-stage training**: Reasoning distillation from o3 (SFT) followed by RL with Verifiable Rewards (GRPO) to correct teacher biases and align verdicts with ground-truth correctness.
59
+ - **Robust generalization**: SOTA performance across all four WebPRMBench environments, including out-of-domain enterprise workflows (WorkArena) and open-world websites (AssistantBench).
60
+
61
+ ## Results on WebPRMBench
62
+
63
+ Models marked with ⋆ are ours. **Bold** = best overall.
64
+
65
+ | Model | Mind2Web | | WebArena | | AssistantBench | | WorkArena | | Avg. | |
66
+ |-------|:---:|:---:|:---:|:---:|:---:|:---:|:---:|:---:|:---:|:---:|
67
+ | | Pair | BoN | Pair | BoN | Pair | BoN | Pair | BoN | Pair | BoN |
68
+ | *Proprietary LLM-as-judge* | | | | | | | | | | |
69
+ | GPT-4o-mini | 81.74 | 50.92 | 78.23 | 56.72 | 89.17 | 73.33 | 81.43 | 46.70 | 82.64 | 56.92 |
70
+ | GPT-4o | 79.99 | 52.62 | 84.58 | 66.67 | 85.83 | 66.67 | 84.33 | 55.19 | 83.68 | 60.29 |
71
+ | GPT-5 | 80.86 | 62.39 | 84.83 | 71.64 | 81.67 | 63.33 | 81.14 | 64.62 | 82.13 | 65.50 |
72
+ | Claude-3.7-Sonnet | 80.20 | 57.90 | 82.80 | 64.10 | 81.50 | 61.30 | 82.10 | 60.60 | 81.65 | 60.98 |
73
+ | Gemini-2.5-Flash | 81.30 | 57.01 | 82.71 | 62.19 | 80.00 | 63.33 | 83.30 | 56.13 | 81.83 | 59.67 |
74
+ | DeepSeek-R1 | 81.62 | 57.37 | 82.04 | 60.21 | 78.49 | 56.18 | 84.12 | 63.89 | 81.57 | 59.41 |
75
+ | *Open-source LLM-as-judge* | | | | | | | | | | |
76
+ | Qwen2.5-7B-Instruct | 77.79 | 39.18 | 74.88 | 42.79 | 84.17 | 53.33 | 77.58 | 35.85 | 77.61 | 42.78 |
77
+ | Llama-3-70B-Instruct | 80.55 | 49.36 | 77.36 | 50.75 | 85.83 | 70.00 | 79.08 | 40.09 | 80.71 | 52.55 |
78
+ | *WebPRMs* | | | | | | | | | | |
79
+ | WebShepherd-8B | 86.66 | 73.69 | 68.33 | 43.88 | 55.92 | 30.00 | 54.56 | 25.53 | 64.34 | 43.28 |
80
+ | ⋆ **WebArbiter-7B** | **97.07** | **89.53** | **88.43** | **68.66** | **89.17** | **70.00** | **82.09** | **70.19** | **89.19** | **74.60** |
81
+
82
+ ## Reward-Guided Trajectory Search (WebArena-Lite)
83
+
84
+ WebArbiter also excels as a practical reward signal for trajectory search. Using Best-of-5 sampling with a Knockout Tournament mechanism on [WebArena-Lite](https://arxiv.org/abs/2408.06327):
85
+
86
+ | Policy | WebPRM | Shopping | CMS | Reddit | GitLab | MAP | Avg. | Ξ” |
87
+ |--------|--------|:--------:|:---:|:------:|:------:|:---:|:----:|:-:|
88
+ | GPT-4o-mini | w/o Search | 21.74 | 22.86 | 19.05 | 34.38 | 19.35 | 23.48 | β€” |
89
+ | GPT-4o-mini | GPT-4o-mini (as WebPRM) | 24.44 | 22.86 | 26.32 | 33.33 | 15.38 | 24.47 | +0.99 |
90
+ | GPT-4o-mini | WebShepherd-8B | 26.09 | 45.71 | 23.81 | 40.62 | 35.48 | 34.34 | +10.86 |
91
+ | GPT-4o-mini | **WebArbiter-7B** | **37.78** | 42.86 | **36.84** | **46.67** | **38.46** | **40.52** | **+17.04** |
92
+ | GPT-4o | w/o Search | 23.91 | 31.43 | 28.57 | 56.25 | 19.35 | 31.90 | β€” |
93
+ | GPT-4o | GPT-4o-mini (as WebPRM) | 26.67 | 37.14 | 42.11 | 40.00 | 19.23 | 33.03 | +1.13 |
94
+ | GPT-4o | WebShepherd-8B | 30.43 | 42.86 | 47.62 | 46.88 | 35.48 | 40.65 | +8.75 |
95
+ | GPT-4o | **WebArbiter-7B** | **44.44** | 42.86 | **52.63** | **56.67** | **38.46** | **47.01** | **+15.11** |
96
+
97
+ ## Quick Start
98
+
99
+ ```python
100
+ import torch
101
+ from transformers import AutoModelForCausalLM, AutoTokenizer
102
+
103
+ model_name = "ZYao720/WebArbiter-7B"
104
+
105
+ tokenizer = AutoTokenizer.from_pretrained(model_name, trust_remote_code=True)
106
+ model = AutoModelForCausalLM.from_pretrained(
107
+ model_name,
108
+ torch_dtype=torch.bfloat16,
109
+ device_map="auto",
110
+ trust_remote_code=True,
111
+ )
112
+
113
+ # Construct your prompt following the WebPRMBench format.
114
+ # See https://huggingface.co/datasets/ZYao720/WEBPRMBENCH for examples.
115
+ user_prompt = "..." # evaluation prompt with intent, AXTree, trajectory, two responses
116
+
117
+ messages = [{"role": "user", "content": user_prompt}]
118
+ input_ids = tokenizer.apply_chat_template(
119
+ messages, tokenize=True, add_generation_prompt=True, return_tensors="pt",
120
+ ).to(model.device)
121
+
122
+ with torch.no_grad():
123
+ output = model.generate(input_ids=input_ids, max_new_tokens=2048, do_sample=False)
124
+
125
+ response = tokenizer.decode(output[0][len(input_ids[0]):], skip_special_tokens=True)
126
+ print(response)
127
+ ```
128
+
129
+ **Example output:**
130
+ ```xml
131
+ <State>The user is on the DuckDuckGo homepage with a search box visible.
132
+ Relevant AXTree elements: [1] textbox 'Search', [2] button 'Search'.</State>
133
+ <Criteria>1. Goal alignment (weight 0.6) β€” Does the action advance the search task?
134
+ 2. Element reference accuracy (weight 0.25) β€” Is the referenced element correct?
135
+ 3. Efficiency (weight 0.15) β€” Does the action avoid unnecessary steps?</Criteria>
136
+ <Analysis>Response 1 directly fills the search query into the textbox, which is the
137
+ most direct path to completing the search task. Response 2 clicks an irrelevant link
138
+ that does not contribute to the search goal.</Analysis>
139
+ <Answer>Response 1</Answer>
140
+ ```
141
+
142
+ ## Training Details
143
+
144
+ | | Stage 1: Reasoning Distillation | Stage 2: RLVR |
145
+ |---|---|---|
146
+ | Method | Supervised fine-tuning (SFT) | GRPO with binary verifiable rewards |
147
+ | Data | 9,642 teacher-distilled examples | 18,921 preference pairs |
148
+ | Teacher | o3 | β€” |
149
+ | Base Model | [Qwen2.5-7B-Instruct](https://huggingface.co/Qwen/Qwen2.5-7B-Instruct) | Stage 1 checkpoint |
150
+ | Fine-tuning | LoRA (rank 128, lr 8e-4) | FSDP + LoRA (lr 7e-6) |
151
+ | Framework | [LLaMA-Factory](https://github.com/hiyouga/LLaMA-Factory) | [veRL](https://github.com/volcengine/verl) |
152
+ | Hardware | 8 Γ— NVIDIA A100-80GB | 8 Γ— NVIDIA A100-80GB |
153
+ | Source Data | [WebPRM Collection](https://huggingface.co/datasets/LangAGI-Lab/WebPRMCollection_preference_pair) (~30k step-level preference pairs from Mind2Web) |
154
+
155
+ **Key training insights** (from ablation studies in the paper):
156
+ - Explicit principles are essential β€” removing them notably degrades performance, especially on out-of-domain environments.
157
+ - Cold-start RL without reasoning distillation is unstable across environments.
158
+ - Reasoning distillation provides stable discrimination, while RL acts as an amplifier that widens the margin between correct and incorrect judgments.
159
+
160
+ ## Intended Uses
161
+
162
+ WebArbiter-7B is designed to:
163
+ - **Evaluate web agent actions**: Given a web state and two candidate actions, determine which better advances the user's task.
164
+ - **Guide trajectory search**: Serve as a reward signal for Best-of-N sampling or tree search during web agent execution.
165
+ - **Provide interpretable feedback**: Generate structured justifications explaining why one action is preferred, useful for debugging and analysis.
166
+
167
+ ## Limitations
168
+
169
+ - **Text-only observations**: WebArbiter relies on accessibility tree representations without visual observations. In environments where layout, spatial arrangement, or visual cues carry task-relevant information, this text-only formulation may miss critical signals.
170
+ - **English-only**: Training and evaluation are conducted exclusively in English-language web environments.
171
+ - **Safe-action bias**: The model may sometimes overvalue cautious actions (e.g., hover over click) because the accessibility tree does not encode interaction effects.
172
+ - **Element reference hallucination**: When a candidate action's reasoning is strongly task-aligned, the model may trust the semantic signal over low-level bid verification, potentially missing incorrect element references.
173
+
174
+ ## License
175
+
176
+ This model is released under [Apache 2.0](https://www.apache.org/licenses/LICENSE-2.0), following the base model [Qwen2.5-7B-Instruct](https://huggingface.co/Qwen/Qwen2.5-7B-Instruct).
177
+
178
+ ## Related Resources
179
+
180
+ | Resource | Link |
181
+ |----------|------|
182
+ | WebArbiter-8B-Qwen3 (strongest) | [ZYao720/WebArbiter-8B-Qwen3](https://huggingface.co/ZYao720/WebArbiter-8B-Qwen3) |
183
+ | WebArbiter-4B-Qwen3 | [ZYao720/WebArbiter-4B-Qwen3](https://huggingface.co/ZYao720/WebArbiter-4B-Qwen3) |
184
+ | WebArbiter-3B | [ZYao720/WebArbiter-3B](https://huggingface.co/ZYao720/WebArbiter-3B) |
185
+ | WEBPRMBENCH (benchmark) | [ZYao720/WEBPRMBENCH](https://huggingface.co/datasets/ZYao720/WEBPRMBENCH) |
186
+ | Training Data | [ZYao720/WebArbiter-Data](https://huggingface.co/datasets/ZYao720/WebArbiter-Data) |
187
+ | Search Trajectories | [ZYao720/WebArbiter-Trajectories](https://huggingface.co/datasets/ZYao720/WebArbiter-Trajectories) |
188
+
189
+ ## Citation
190
+
191
+ ```bibtex
192
+ @misc{zhang2026ZYao720principleguidedreasoningprocess,
193
+ title={WebArbiter: A Principle-Guided Reasoning Process Reward Model for Web Agents},
194
+ author={Yao Zhang and Shijie Tang and Zeyu Li and Zhen Han and Volker Tresp},
195
+ year={2026},
196
+ eprint={2601.21872},
197
+ archivePrefix={arXiv},
198
+ primaryClass={cs.AI},
199
+ url={https://arxiv.org/abs/2601.21872},
200
+ }
201
+ ```