Add .eval_results for benchmarks with public HF datasets

#8

YAML Metadata Error:Invalid content in Eval Result file .eval_results/clawgym.yaml

Check out the documentation for more information.

Show details
Task ID "clawgym" does not match any task in dataset "RUC-AIBOX/ClawGym-Bench". Available: none

YAML Metadata Error:Invalid content in Eval Result file .eval_results/pinchbench.yaml

Check out the documentation for more information.

Show details
Task ID "pinchbench" does not match any task in dataset "NEAR-AI/pinchbench". Available: none
.eval_results/clawgym.yaml ADDED
@@ -0,0 +1,7 @@
 
 
 
 
 
 
 
 
1
+ - dataset:
2
+ id: RUC-AIBOX/ClawGym-Bench
3
+ task_id: clawgym
4
+ value: 64.0
5
+ source:
6
+ url: https://huggingface.co/mindlab-research/Macaron-V1-Tall
7
+ name: Model Card
.eval_results/pinchbench.yaml ADDED
@@ -0,0 +1,7 @@
 
 
 
 
 
 
 
 
1
+ - dataset:
2
+ id: NEAR-AI/pinchbench
3
+ task_id: pinchbench
4
+ value: 86.2
5
+ source:
6
+ url: https://huggingface.co/mindlab-research/Macaron-V1-Tall
7
+ name: Model Card