Spaces:
Sleeping
Download README.md from v4xsh/nervousystem-env: direct link, hf CLI and curl.
- Browser
- Download file 20.2 kB
-
https://huggingface.co/spaces/v4xsh/nervousystem-env/resolve/main/README.md
- Command line
-
hf download hf://spaces/v4xsh/nervousystem-env/README.md
-
curl -L -o README.md https://huggingface.co/spaces/v4xsh/nervousystem-env/resolve/main/README.md
title: NervousSystem Env
emoji: π§
colorFrom: indigo
colorTo: blue
sdk: docker
app_port: 7860
pinned: false
π§ NervousSystem-Env
An AI agent fixing the infrastructure that trains AI.
Every minute of GPU cluster downtime costs $5,000 in wasted compute.
NervousSystem-Env is an OpenEnv-compliant reinforcement learning environment where an agent acts as an SRE operating a distributed GPU training fleet under failure. The agent must diagnose incidents, coordinate specialist workers, and execute recovery actions while minimizing token and coordination waste. The environment models realistic infrastructure failure modes (OOM, topology congestion, NCCL desync, version cascades), includes partial observability, and rewards long-horizon remediation over shortcut behavior. This environment directly models one of the most expensive unsolved problems in modern AI infrastructure.
Links
- Hosted OpenEnv Space: https://huggingface.co/spaces/v4xsh/nervousystem-env
- Mini-blog (Hugging Face): https://huggingface.co/spaces/v4xsh/nervousystem-env/blob/main/Blog.md
- Trained LoRA adapter: https://huggingface.co/v4xsh/nervousystem-sre-agent-lora
- Final training/evaluation job logs: https://huggingface.co/jobs/v4xsh/69eda4d1d2c8bd8662bcf435
- Training reproduction notebook (Colab): https://colab.research.google.com/drive/1twXiRHoAxchy9UUgn7S15a2ac3GrtAlW?usp=sharing
- In-repo project pitch:
PITCH.md - Training evidence:
TRAINING_EVIDENCE.md
Why This Matters
- GPU OOM (XID 79): CUDA out-of-memory on one rank stalls the entire AllReduce collective. Training halts fleet-wide.
- Spine Switch Congestion: Ring topology crosses oversubscribed spine links, cutting interconnect bandwidth 40%+.
- Compilation Desync: Different ranks compile different NCCL collectives due to data-dependent branching. Job hangs permanently.
- LD_LIBRARY_PATH Cascade: Wrong NCCL version loaded (2.21.5 vs 2.27.0) triggers message truncation errors that propagate across all communicators. Severity-1 incident.
Architecture
NervousSystem-Env uses a Fleet AI Supervisor-Worker design. The supervisor receives ClusterObservation in a partially observable setting where telemetry can go stale every 3 steps unless diagnostics are refreshed. It delegates specialist sub-tasks through /delegate, and workers return structured output with confidence and uncertainty signals. The supervisor also operates under a delegation budget of 10 per episode, and over-budget delegations reduce coordination reward.
Simulation fidelity is built around realistic infrastructure diagnostics. NCCL logs use exact <hostname>:<pid>:<tid> NCCL INFO formatting, and Flight Recorder payloads follow PyTorch 2.5 v2.5 schema (pg_config, pg_status, circular buffer warnings). Telemetry intentionally includes 30% red-herring alerts and 20% incomplete observability to prevent surface-level pattern matching.
Supervisor Agent
β
βββ LogInspectorWorker (flight recorder, NCCL subsystem logs)
βββ PatchAgentWorker (staged patching: identifyβdiffβapply)
βββ TopoAgentWorker (topology reorder, bandwidth check)
βββ VersionCheckerWorker (NCCL version, LD_LIBRARY_PATH audit)
Tasks
| Task | Difficulty | Max Steps | Failure Type | Key Challenge |
|---|---|---|---|---|
| easy | easy | 50 | OOM rank failure | Identify failing rank from noisy telemetry |
| medium | medium | 50 | Spine congestion | Confirm fix held under throughput jitter |
| hard | hard | 50 | Compilation desync | 3-stage patch: identifyβdiffβapply |
| cascade | cascade | 120 | Version cascade | Solve OOMβcongestionβdesync in order |
| fleet_coordination | fleet_coordination | 50 | Multi-agent incident | Delegate to specialists, reach coalition, then remediate |
Cascade uses strict phase gating with precondition enforcement. The agent must solve OOM diagnosis, congestion recovery, then desync investigation/patch in sequence; patch attempts without required investigation are blocked, so phases cannot be skipped by direct guess-fixing.
Reward Model
Reward = 0.60 Γ R_success + 0.30 Γ R_subgoal β 0.10 Γ log(total_tokens)
| Component | Weight | Description |
|---|---|---|
| R_success | 0.60 | Binary: job_status in {recovered, running} |
| R_subgoal | 0.30 | Continuous partial credit per phase/diagnosis |
| log(total_tokens) | 0.10 | Efficiency penalty β verbose agents score lower |
Final grade applies token efficiency as a multiplier on the raw task score, so solved low-token episodes keep high scores while token stuffing collapses the final score. Additional penalties are active: destructive action penalty (-0.2), delegation over-budget penalty (-0.05 per excess delegation), and anti-shortcut precondition enforcement for diagnosis-before-patch behavior.
Observation Space
NervousSystem-Env is intentionally partially observable:
- Surface logs refresh every 3 steps (
stale_telemetryflag) - 30% of updates contain red-herring node alerts
- 20% of updates mask true degraded node status
- Deep diagnostics only via
inspect_flight_recorderorquery_nccl_logs cumulative_tokensvisible to the agent (self-awareness of efficiency)
Quick Start
# Install
pip install -r requirements.txt
# Start environment server
uvicorn app.main:app --host 0.0.0.0 --port 7860
# Start SRE War Room dashboard
python dashboard/war_room.py # http://localhost:7861
# Run baseline agent (requires HF_TOKEN or API key)
export HF_TOKEN=your_token
python inference.py
# Run seed variance report
python scripts/seed_variance.py
# Train LoRA policy with SFT warmup, then optional GRPO
pip install -r requirements-training.txt
python training/grpo_train.py
# Resume GRPO from existing SFT adapter (continues with environment reward)
RESUME_FROM_ADAPTER=sre_agent_lora \
GRPO_ADAPTER_DIR=sre_agent_lora_grpo \
GRPO_MAX_STEPS=120 \
python training/grpo_train.py
# One-shot overnight pipeline (server + SFT + GRPO + before/after constrained eval)
bash scripts/run_overnight.sh
# Evaluate final trained policy with phase-aware constrained action scoring
EVAL_LENGTH_NORM_POWER=1.0 python scripts/evaluate_model.py --label sft_oraclefix_phaseaware --adapter sre_agent_lora_oraclefix --seeds 3 --tasks easy,medium,hard,cascade --max-steps 8
Judge Reproduction
Run the environment:
pip install -r requirements.txt
uvicorn app.main:app --host 0.0.0.0 --port 7860
In another terminal:
python scripts/reward_audit.py
python scripts/exploit_test.py
python scripts/evaluate.py quick
python scripts/evaluate_multi_agent.py
python scripts/generate_long_horizon_trace.py
EVAL_LENGTH_NORM_POWER=1.0 python scripts/evaluate_model.py --label sft_oraclefix_phaseaware --adapter sre_agent_lora_oraclefix --seeds 3 --tasks easy,medium,hard,cascade --max-steps 8
Hosted Space smoke checks:
curl https://v4xsh-nervousystem-env.hf.space/health
curl -X POST https://v4xsh-nervousystem-env.hf.space/reset -H 'Content-Type: application/json' -d '{"task_id":"easy","seed":42}'
API Endpoints
| Endpoint | Method | Description |
|---|---|---|
/health |
GET | Health check |
/reset |
POST | Reset episode (supports challenge_type, difficulty_multiplier) |
/step |
POST | Apply one SRE action |
/state |
GET | Current observation |
/grade |
POST | Grade episode (includes MER + meta-reward) |
/tasks |
GET | List registered manifest and experimental tasks |
/episode |
GET | Current episode summary |
/delegate |
POST | Supervisor delegates to specialist worker |
/consensus |
GET | Poll all workers for recommended action |
/coalition |
POST | Attempt 2-worker coalition action |
/coalition/options |
GET | Available coalition actions |
/challenge/generate |
POST | Generate novel failure composition |
/challenge/list |
GET | List generated challenges |
/curriculum |
GET | Adaptive curriculum state |
/curriculum/events |
GET | Difficulty escalation/backoff log |
Hackathon Theme Coverage
| Theme | Coverage | Implementation |
|---|---|---|
| Theme 1: Multi-Agent | β Strong | fleet_coordination requires specialist delegation and coalition before high score |
| Theme 2: Long-Horizon | β Supported | Cascade has 120-step budget plus a 56-event long-horizon trace artifact |
| Theme 3.1: Professional Tasks | β Strong | NCCL INFO logs, Flight Recorder v2.5, staged patch workflow |
| Theme 4: Self-Improvement | β Moderate | Curriculum changes red herrings, telemetry masking, refresh interval, and secondary failures |
Training Results
Oracle Baseline (50 seeds, deterministic optimal agent)
| Task | Mean Score | Std | Pass Rate | Mean Steps | CI 95% |
|---|---|---|---|---|---|
| easy | 0.990 | 0.000 | 100.0% | 1.00 | [0.990, 0.990] |
| medium | 0.990 | 0.000 | 100.0% | 7.00 | [0.990, 0.990] |
| hard | 0.737 | 0.072 | 100.0% | 5.56 | [0.720, 0.753] |
| cascade | 0.857 | 0.063 | 100.0% | 5.30 | [0.840, 0.875] |
| fleet_coordination | 0.990 | 0.000 | 100.0% | delegation-only | [0.990, 0.990] |
| black_swan | 0.010 | 0.000 | 0%* | β | [0.010, 0.010] |
*Black Swan is an experimental ultra-hard stress test in code. The main OpenEnv manifest focuses on easy, medium, hard, cascade, and fleet_coordination tasks for judging.
Random Agent Baseline (50 seeds, lower bound)
| Task | Mean Score | Pass Rate |
|---|---|---|
| easy | 0.299 | 26.0% |
| medium | 0.926 | 100.0% |
| hard | 0.360 | 18.0% |
| cascade | 0.039 | 0.0% |
| fleet_coordination | 0.050 | 0.0% |
*Medium pass threshold raised to 0.7 after random agent audit
Trained LLM Policy Result
We trained a Qwen2.5-7B LoRA policy on Hugging Face Jobs using an NVIDIA A10G. The policy was trained with corrected oracle SRE trajectories from this OpenEnv environment, then evaluated with phase-aware constrained action scoring: the model ranks valid next-step JSON tool actions for the current workflow phase, and the environment executes the highest-likelihood action.
Training loss fell from 2.53 β ~0.10 over the 40 logged SFT steps preserved in results/sft_warmup_metrics.json (sourced from trainer.state.log_history). The trained policy improved to 0.915 mean score and 100% pass rate over 12 model-policy episodes.
Note on training-evidence provenance. The published
sre_agent_lora_oraclefixadapter and the0.915evaluation are from a 300-step SFT run on Hugging Face Jobs (terminal proof indocs/proof/final_trained_console.png, job log69eda4d1d2c8bd8662bcf435). The per-stepresults/sft_warmup_metrics.json(40 logged entries / ~160 actual steps atlogging_steps=4) is from a faithful local re-run with the same script, dataset, and hyperparameters, published so judges can inspect a realtrainer.state.log_historydirectly. Both runs converge to the same loss regime (β0.08β0.10).
Raw HF Jobs / terminal proof (same evaluator, same base model, only the LoRA adapter changes)
Before β raw unsloth/Qwen2.5-7B-Instruct-bnb-4bit with no adapter, 9 episodes, mean 0.239, pass rate 0%:
After β same base + SFT LoRA sre_agent_lora_oraclefix, 12 episodes, mean 0.915, pass rate 100%:
| Task | Trained LoRA phase-aware score |
|---|---|
| easy | 0.990 / 0.990 / 0.990 |
| medium | 0.990 / 0.990 / 0.990 |
| hard | 0.850 / 0.850 / 0.990 |
| cascade | 0.782 / 0.782 / 0.782 |
| overall | 0.915 mean, 100% pass rate |
Interpretation: the trained policy learns reliable remediation for rank diagnosis, topology congestion, staged desync patching, and cascade recovery when evaluated under the same phase-aware action constraints used by the environment workflow.
Artifacts:
- LoRA adapter: https://huggingface.co/v4xsh/nervousystem-sre-agent-lora
- Final training/evaluation job logs: https://huggingface.co/jobs/v4xsh/69eda4d1d2c8bd8662bcf435
- Plot generator:
scripts/plot_final_evidence.py
Multi-Agent Coordination Evaluation
The fleet_coordination task requires supervisor-worker delegation before remediation. A deterministic coordination oracle delegates to log_inspector, delegates to version_checker, polls consensus, and executes a topology_version_fix coalition. It reaches 0.990 mean score / 100% pass rate over 10 seeds. A direct-action baseline that skips specialist evidence scores 0.100 mean score / 0% pass rate.
Evidence: results/multi_agent_eval.json
Long-Horizon Trace
In addition to short optimal solves, results/cascade_long_horizon_trace.json contains a 56-event cascade trace. It demonstrates stale telemetry refresh, phase-gated remediation, sustained monitoring, delayed desync investigation, staged file identification, and final patch recovery. The trace is intentionally non-oracle-like: it shows the environment can exercise long-running state tracking, not only shortest-path scripts.
Reward Integrity
The reward system is independently audited against 4 properties. Final grade applies the MER token-efficiency penalty as a multiplier on the raw task score, so solved low-token episodes keep their score while token stuffing and verbose loops cannot preserve a high final score.
| Property | Result | Evidence |
|---|---|---|
| Monotonicity | β Correct action > wrong > noop | results/reward_audit.json |
| Anti-Gaming | β All 4 exploits blocked | results/reward_audit.json |
| Token Efficiency | β Verbose agents penalized | results/reward_audit.json |
| Determinism | β Same seed = same score | results/reward_audit.json |
Exploit resistance verified:
| Exploit | Score | Status |
|---|---|---|
| Noop spam (50 steps) | 0.020 | β BLOCKED |
| Wrong rank spam (20 steps) | 0.020 | β BLOCKED |
| Patch without investigation | 0.030 | β BLOCKED |
| Destructive action spam | 0.010 | β BLOCKED |
| Token stuffing (100k tokens) | 0.010 | β BLOCKED |
| Cascade phase skip | 0.010 | β BLOCKED |
| Reward farming loop | 0.351 | β BLOCKED |
Run audit: python scripts/reward_audit.py
Run exploit test: python scripts/exploit_test.py
Adaptive Curriculum
The environment self-adjusts difficulty based on agent performance:
| Level | Pass Rate Trigger | Red Herring % | Telemetry Mask % | Refresh Interval |
|---|---|---|---|---|
| 1 (Base) | < 75% | 30% | 20% | 3 steps |
| 2 (Hard) | > 75% sustained | 50% | 35% | 5 steps |
| 3 (Hardest) | > 75% at level 2 | 70% | 50% | 7 steps |
- Escalation: triggered when
pass_rate > 0.75ANDtrend > 0.01over last 20 episodes, with 10-episode cooldown - Backoff: triggered when
pass_rate < 0.20ANDtrend < -0.01, with 5-episode cooldown - Meta-reward: agents earn bonus reward for improving over their recent rolling window (not just for success)
- Wiring: curriculum levels now change telemetry red-herring probability, telemetry masking, secondary failure probability, and telemetry refresh interval.
Monitor live: GET /curriculum
View events: GET /curriculum/events
Project Structure
nervousystem-env/
βββ app/
β βββ config.py # Scenarios, black_swan, rack layouts, challenge types
β βββ env.py # Core env, Fleet + Curriculum + Challenge wiring
β βββ main.py # FastAPI β 15 endpoints
β βββ models.py # Pydantic v2 models, Coalition/Delegate schemas
βββ graders/
β βββ base.py # GradeResult, apply_mer(), meta-reward support
β βββ black_swan_grader.py
β βββ cascade_grader.py
β βββ easy_grader.py
β βββ fleet_coordination_grader.py
β βββ hard_grader.py
β βββ medium_grader.py
βββ simulation/
β βββ challenge_generator.py # 5 novel failure compositions
β βββ cluster.py # State machine, staged patching, non-linear dynamics
β βββ curriculum.py # Adaptive difficulty manager, meta-reward
β βββ failures.py # Cascade phase injection, secondary failures
β βββ fleet.py # Fleet AI, coalition actions, consensus
β βββ telemetry.py # NCCL logs, red herrings, staleness
βββ tasks/
β βββ base.py # EpisodeTracker, preconditions, investigation_count
β βββ black_swan.py # 300-step ultra-hard task, false positives, irreversible mistakes
β βββ cascade.py # 3-phase chained task, phase gating
β βββ easy.py
β βββ fleet_coordination.py
β βββ hard.py
β βββ medium.py
βββ dashboard/
β βββ war_room.py # Gradio SRE War Room (port 7861)
βββ training/
β βββ grpo_train.py # LoRA SFT warmup + optional GRPO over env rollouts
βββ scripts/
β βββ evaluate.py # Oracle + random eval harness, 50 seeds, bootstrap CI
β βββ evaluate_multi_agent.py
β βββ evaluate_model.py # Raw/trained model-policy eval with action scoring
β βββ exploit_test.py # exploit attack proofs
β βββ generate_long_horizon_trace.py
β βββ plot_final_evidence.py # Rebuilds committed PNG plots from eval JSON + loss log
β βββ reward_audit.py # 4 reward property audits
β βββ run_overnight.sh # One-shot env+SFT+GRPO+eval pipeline
β βββ seed_variance.py # Seed variance report
βββ results/
β βββ before_vs_after_scores.png # Raw 0.239 vs trained 0.915, same evaluator
β βββ cascade_long_horizon_trace.json
β βββ exploit_test.json
β βββ final_phaseaware_model_eval.json
β βββ final_score_by_task.png # Trained LoRA mean score per task
β βββ final_sft_loss_curve.png # SFT loss 2.53 -> 0.10 over 40 logged steps
β βββ sft_warmup_metrics.json # Real per-step loss/lr/grad_norm log_history
β βββ multi_agent_eval.json
β βββ reward_audit.json
βββ docs/
β βββ proof/
β βββ raw_untrained_console.png # HF Jobs eval, untrained base model, 0.239 mean
β βββ final_trained_console.png # Local terminal eval, trained LoRA, 0.915 mean
βββ notebooks/
β βββ README.md
β βββ eval_trained_lora.ipynb # Colab-ready reproduction of headline numbers
βββ tests/
β βββ test_fleet.py # 14 fleet + coalition tests
β βββ test_graders.py # 19 grader + anti-gaming tests
βββ inference.py # Baseline LLM agent runner
βββ openenv.yaml # OpenEnv v2.0.0, 5 manifest tasks, multi-agent
βββ requirements.txt
βββ requirements-training.txt
βββ server/
βββ app.py
OpenEnv Compliance
openenv validatepasses β- Typed Pydantic v2 models with extra="forbid" β
- Deterministic seed-based graders β
- Docker deployment ready β
- 5 manifest tasks with difficulty progression β
- Multi-agent
/delegateand/coalitionevaluation β - Mercor efficiency reward β
Known Limitations
The submitted adapter is an SFT warmup policy, not a fully optimized online RL policy. training/grpo_train.py includes an optional GRPO loop wired to the environment reward function, but the submitted quantitative improvement is from supervised LoRA warmup over environment-generated SRE rollouts plus phase-aware constrained evaluation.
The simulator is deterministic under seed for reproducibility and models production-inspired failure signatures rather than connecting to a real GPU cluster.




