--- title: NervousSystem Env emoji: 🧠 colorFrom: indigo colorTo: blue sdk: docker app_port: 7860 pinned: false --- # 🧠 NervousSystem-Env > An AI agent fixing the infrastructure that trains AI. > Every minute of GPU cluster downtime costs **$5,000** in wasted compute. NervousSystem-Env is an OpenEnv-compliant reinforcement learning environment where an agent acts as an SRE operating a distributed GPU training fleet under failure. The agent must diagnose incidents, coordinate specialist workers, and execute recovery actions while minimizing token and coordination waste. The environment models realistic infrastructure failure modes (OOM, topology congestion, NCCL desync, version cascades), includes partial observability, and rewards long-horizon remediation over shortcut behavior. This environment directly models one of the most expensive unsolved problems in modern AI infrastructure. ## Links - **Hosted OpenEnv Space**: https://huggingface.co/spaces/v4xsh/nervousystem-env - **Mini-blog (Hugging Face)**: https://huggingface.co/spaces/v4xsh/nervousystem-env/blob/main/Blog.md - **Trained LoRA adapter**: https://huggingface.co/v4xsh/nervousystem-sre-agent-lora - **Final training/evaluation job logs**: https://huggingface.co/jobs/v4xsh/69eda4d1d2c8bd8662bcf435 - **Training reproduction notebook (Colab)**: https://colab.research.google.com/drive/1twXiRHoAxchy9UUgn7S15a2ac3GrtAlW?usp=sharing - **In-repo project pitch**: [`PITCH.md`](PITCH.md) - **Training evidence**: [`TRAINING_EVIDENCE.md`](TRAINING_EVIDENCE.md) ## Why This Matters - **GPU OOM (XID 79)**: CUDA out-of-memory on one rank stalls the entire AllReduce collective. Training halts fleet-wide. - **Spine Switch Congestion**: Ring topology crosses oversubscribed spine links, cutting interconnect bandwidth 40%+. - **Compilation Desync**: Different ranks compile different NCCL collectives due to data-dependent branching. Job hangs permanently. - **LD_LIBRARY_PATH Cascade**: Wrong NCCL version loaded (2.21.5 vs 2.27.0) triggers message truncation errors that propagate across all communicators. Severity-1 incident. ## Architecture NervousSystem-Env uses a Fleet AI Supervisor-Worker design. The supervisor receives `ClusterObservation` in a partially observable setting where telemetry can go stale every 3 steps unless diagnostics are refreshed. It delegates specialist sub-tasks through `/delegate`, and workers return structured output with confidence and uncertainty signals. The supervisor also operates under a delegation budget of 10 per episode, and over-budget delegations reduce coordination reward. Simulation fidelity is built around realistic infrastructure diagnostics. NCCL logs use exact `:: NCCL INFO` formatting, and Flight Recorder payloads follow PyTorch 2.5 v2.5 schema (`pg_config`, `pg_status`, circular buffer warnings). Telemetry intentionally includes 30% red-herring alerts and 20% incomplete observability to prevent surface-level pattern matching. ```text Supervisor Agent β”‚ β”œβ”€β”€ LogInspectorWorker (flight recorder, NCCL subsystem logs) β”œβ”€β”€ PatchAgentWorker (staged patching: identifyβ†’diffβ†’apply) β”œβ”€β”€ TopoAgentWorker (topology reorder, bandwidth check) └── VersionCheckerWorker (NCCL version, LD_LIBRARY_PATH audit) ``` ## Tasks | Task | Difficulty | Max Steps | Failure Type | Key Challenge | |---|---|---|---|---| | easy | easy | 50 | OOM rank failure | Identify failing rank from noisy telemetry | | medium | medium | 50 | Spine congestion | Confirm fix held under throughput jitter | | hard | hard | 50 | Compilation desync | 3-stage patch: identifyβ†’diffβ†’apply | | cascade | cascade | 120 | Version cascade | Solve OOMβ†’congestionβ†’desync in order | | fleet_coordination | fleet_coordination | 50 | Multi-agent incident | Delegate to specialists, reach coalition, then remediate | Cascade uses strict phase gating with precondition enforcement. The agent must solve OOM diagnosis, congestion recovery, then desync investigation/patch in sequence; patch attempts without required investigation are blocked, so phases cannot be skipped by direct guess-fixing. ## Reward Model ```text Reward = 0.60 Γ— R_success + 0.30 Γ— R_subgoal βˆ’ 0.10 Γ— log(total_tokens) ``` | Component | Weight | Description | |---|---|---| | R_success | 0.60 | Binary: job_status in {recovered, running} | | R_subgoal | 0.30 | Continuous partial credit per phase/diagnosis | | log(total_tokens) | 0.10 | Efficiency penalty β€” verbose agents score lower | Final grade applies token efficiency as a multiplier on the raw task score, so solved low-token episodes keep high scores while token stuffing collapses the final score. Additional penalties are active: destructive action penalty (`-0.2`), delegation over-budget penalty (`-0.05` per excess delegation), and anti-shortcut precondition enforcement for diagnosis-before-patch behavior. ## Observation Space NervousSystem-Env is intentionally partially observable: - Surface logs refresh every 3 steps (`stale_telemetry` flag) - 30% of updates contain red-herring node alerts - 20% of updates mask true degraded node status - Deep diagnostics only via `inspect_flight_recorder` or `query_nccl_logs` - `cumulative_tokens` visible to the agent (self-awareness of efficiency) ## Quick Start ```bash # Install pip install -r requirements.txt # Start environment server uvicorn app.main:app --host 0.0.0.0 --port 7860 # Start SRE War Room dashboard python dashboard/war_room.py # http://localhost:7861 # Run baseline agent (requires HF_TOKEN or API key) export HF_TOKEN=your_token python inference.py # Run seed variance report python scripts/seed_variance.py # Train LoRA policy with SFT warmup, then optional GRPO pip install -r requirements-training.txt python training/grpo_train.py # Resume GRPO from existing SFT adapter (continues with environment reward) RESUME_FROM_ADAPTER=sre_agent_lora \ GRPO_ADAPTER_DIR=sre_agent_lora_grpo \ GRPO_MAX_STEPS=120 \ python training/grpo_train.py # One-shot overnight pipeline (server + SFT + GRPO + before/after constrained eval) bash scripts/run_overnight.sh # Evaluate final trained policy with phase-aware constrained action scoring EVAL_LENGTH_NORM_POWER=1.0 python scripts/evaluate_model.py --label sft_oraclefix_phaseaware --adapter sre_agent_lora_oraclefix --seeds 3 --tasks easy,medium,hard,cascade --max-steps 8 ``` ## Judge Reproduction Run the environment: ```bash pip install -r requirements.txt uvicorn app.main:app --host 0.0.0.0 --port 7860 ``` In another terminal: ```bash python scripts/reward_audit.py python scripts/exploit_test.py python scripts/evaluate.py quick python scripts/evaluate_multi_agent.py python scripts/generate_long_horizon_trace.py EVAL_LENGTH_NORM_POWER=1.0 python scripts/evaluate_model.py --label sft_oraclefix_phaseaware --adapter sre_agent_lora_oraclefix --seeds 3 --tasks easy,medium,hard,cascade --max-steps 8 ``` Hosted Space smoke checks: ```bash curl https://v4xsh-nervousystem-env.hf.space/health curl -X POST https://v4xsh-nervousystem-env.hf.space/reset -H 'Content-Type: application/json' -d '{"task_id":"easy","seed":42}' ``` ## API Endpoints | Endpoint | Method | Description | |---|---|---| | `/health` | GET | Health check | | `/reset` | POST | Reset episode (supports `challenge_type`, `difficulty_multiplier`) | | `/step` | POST | Apply one SRE action | | `/state` | GET | Current observation | | `/grade` | POST | Grade episode (includes MER + meta-reward) | | `/tasks` | GET | List registered manifest and experimental tasks | | `/episode` | GET | Current episode summary | | `/delegate` | POST | Supervisor delegates to specialist worker | | `/consensus` | GET | Poll all workers for recommended action | | `/coalition` | POST | Attempt 2-worker coalition action | | `/coalition/options` | GET | Available coalition actions | | `/challenge/generate` | POST | Generate novel failure composition | | `/challenge/list` | GET | List generated challenges | | `/curriculum` | GET | Adaptive curriculum state | | `/curriculum/events` | GET | Difficulty escalation/backoff log | ## Hackathon Theme Coverage | Theme | Coverage | Implementation | |---|---|---| | Theme 1: Multi-Agent | βœ… Strong | `fleet_coordination` requires specialist delegation and coalition before high score | | Theme 2: Long-Horizon | βœ… Supported | Cascade has 120-step budget plus a 56-event long-horizon trace artifact | | Theme 3.1: Professional Tasks | βœ… Strong | NCCL INFO logs, Flight Recorder v2.5, staged patch workflow | | Theme 4: Self-Improvement | βœ… Moderate | Curriculum changes red herrings, telemetry masking, refresh interval, and secondary failures | ## Training Results ### Oracle Baseline (50 seeds, deterministic optimal agent) | Task | Mean Score | Std | Pass Rate | Mean Steps | CI 95% | |---|---|---|---|---|---| | easy | 0.990 | 0.000 | 100.0% | 1.00 | [0.990, 0.990] | | medium | 0.990 | 0.000 | 100.0% | 7.00 | [0.990, 0.990] | | hard | 0.737 | 0.072 | 100.0% | 5.56 | [0.720, 0.753] | | cascade | 0.857 | 0.063 | 100.0% | 5.30 | [0.840, 0.875] | | fleet_coordination | 0.990 | 0.000 | 100.0% | delegation-only | [0.990, 0.990] | | black_swan | 0.010 | 0.000 | 0%* | β€” | [0.010, 0.010] | *Black Swan is an experimental ultra-hard stress test in code. The main OpenEnv manifest focuses on easy, medium, hard, cascade, and fleet_coordination tasks for judging. ### Random Agent Baseline (50 seeds, lower bound) | Task | Mean Score | Pass Rate | |---|---|---| | easy | 0.299 | 26.0% | | medium | 0.926 | 100.0% | | hard | 0.360 | 18.0% | | cascade | 0.039 | 0.0% | | fleet_coordination | 0.050 | 0.0% | *Medium pass threshold raised to 0.7 after random agent audit ### Trained LLM Policy Result We trained a Qwen2.5-7B LoRA policy on Hugging Face Jobs using an NVIDIA A10G. The policy was trained with corrected oracle SRE trajectories from this OpenEnv environment, then evaluated with phase-aware constrained action scoring: the model ranks valid next-step JSON tool actions for the current workflow phase, and the environment executes the highest-likelihood action. Training loss fell from **2.53 β†’ ~0.10** over the 40 logged SFT steps preserved in `results/sft_warmup_metrics.json` (sourced from `trainer.state.log_history`). The trained policy improved to **0.915 mean score** and **100% pass rate** over 12 model-policy episodes. > **Note on training-evidence provenance.** The published `sre_agent_lora_oraclefix` adapter and the `0.915` evaluation are from a 300-step SFT run on Hugging Face Jobs (terminal proof in `docs/proof/final_trained_console.png`, job log [`69eda4d1d2c8bd8662bcf435`](https://huggingface.co/jobs/v4xsh/69eda4d1d2c8bd8662bcf435)). The per-step `results/sft_warmup_metrics.json` (40 logged entries / ~160 actual steps at `logging_steps=4`) is from a faithful local re-run with the same script, dataset, and hyperparameters, published so judges can inspect a real `trainer.state.log_history` directly. Both runs converge to the same loss regime (β‰ˆ0.08–0.10). ![Final SFT loss curve](results/final_sft_loss_curve.png) ![Raw base model vs final trained policy score](results/before_vs_after_scores.png) ![Final trained policy score by task](results/final_score_by_task.png) #### Raw HF Jobs / terminal proof (same evaluator, same base model, only the LoRA adapter changes) Before β€” raw `unsloth/Qwen2.5-7B-Instruct-bnb-4bit` with no adapter, 9 episodes, **mean 0.239, pass rate 0%**: ![Raw base model HF Jobs evaluation console](docs/proof/raw_untrained_console.png) After β€” same base + SFT LoRA `sre_agent_lora_oraclefix`, 12 episodes, **mean 0.915, pass rate 100%**: ![Final trained policy terminal evaluation](docs/proof/final_trained_console.png) | Task | Trained LoRA phase-aware score | |---|---:| | easy | 0.990 / 0.990 / 0.990 | | medium | 0.990 / 0.990 / 0.990 | | hard | 0.850 / 0.850 / 0.990 | | cascade | 0.782 / 0.782 / 0.782 | | overall | **0.915 mean, 100% pass rate** | Interpretation: the trained policy learns reliable remediation for rank diagnosis, topology congestion, staged desync patching, and cascade recovery when evaluated under the same phase-aware action constraints used by the environment workflow. Artifacts: - LoRA adapter: https://huggingface.co/v4xsh/nervousystem-sre-agent-lora - Final training/evaluation job logs: https://huggingface.co/jobs/v4xsh/69eda4d1d2c8bd8662bcf435 - Plot generator: `scripts/plot_final_evidence.py` ### Multi-Agent Coordination Evaluation The `fleet_coordination` task requires supervisor-worker delegation before remediation. A deterministic coordination oracle delegates to `log_inspector`, delegates to `version_checker`, polls consensus, and executes a `topology_version_fix` coalition. It reaches **0.990 mean score / 100% pass rate** over 10 seeds. A direct-action baseline that skips specialist evidence scores **0.100 mean score / 0% pass rate**. Evidence: `results/multi_agent_eval.json` ### Long-Horizon Trace In addition to short optimal solves, `results/cascade_long_horizon_trace.json` contains a **56-event** cascade trace. It demonstrates stale telemetry refresh, phase-gated remediation, sustained monitoring, delayed desync investigation, staged file identification, and final patch recovery. The trace is intentionally non-oracle-like: it shows the environment can exercise long-running state tracking, not only shortest-path scripts. ## Reward Integrity The reward system is independently audited against 4 properties. Final grade applies the MER token-efficiency penalty as a multiplier on the raw task score, so solved low-token episodes keep their score while token stuffing and verbose loops cannot preserve a high final score. | Property | Result | Evidence | |---|---|---| | Monotonicity | βœ… Correct action > wrong > noop | `results/reward_audit.json` | | Anti-Gaming | βœ… All 4 exploits blocked | `results/reward_audit.json` | | Token Efficiency | βœ… Verbose agents penalized | `results/reward_audit.json` | | Determinism | βœ… Same seed = same score | `results/reward_audit.json` | Exploit resistance verified: | Exploit | Score | Status | |---|---|---| | Noop spam (50 steps) | 0.020 | βœ… BLOCKED | | Wrong rank spam (20 steps) | 0.020 | βœ… BLOCKED | | Patch without investigation | 0.030 | βœ… BLOCKED | | Destructive action spam | 0.010 | βœ… BLOCKED | | Token stuffing (100k tokens) | 0.010 | βœ… BLOCKED | | Cascade phase skip | 0.010 | βœ… BLOCKED | | Reward farming loop | 0.351 | βœ… BLOCKED | Run audit: `python scripts/reward_audit.py` Run exploit test: `python scripts/exploit_test.py` ## Adaptive Curriculum The environment self-adjusts difficulty based on agent performance: | Level | Pass Rate Trigger | Red Herring % | Telemetry Mask % | Refresh Interval | |---|---|---|---|---| | 1 (Base) | < 75% | 30% | 20% | 3 steps | | 2 (Hard) | > 75% sustained | 50% | 35% | 5 steps | | 3 (Hardest) | > 75% at level 2 | 70% | 50% | 7 steps | - **Escalation**: triggered when `pass_rate > 0.75` AND `trend > 0.01` over last 20 episodes, with 10-episode cooldown - **Backoff**: triggered when `pass_rate < 0.20` AND `trend < -0.01`, with 5-episode cooldown - **Meta-reward**: agents earn bonus reward for improving over their recent rolling window (not just for success) - **Wiring**: curriculum levels now change telemetry red-herring probability, telemetry masking, secondary failure probability, and telemetry refresh interval. Monitor live: `GET /curriculum` View events: `GET /curriculum/events` ## Project Structure ```text nervousystem-env/ β”œβ”€β”€ app/ β”‚ β”œβ”€β”€ config.py # Scenarios, black_swan, rack layouts, challenge types β”‚ β”œβ”€β”€ env.py # Core env, Fleet + Curriculum + Challenge wiring β”‚ β”œβ”€β”€ main.py # FastAPI β€” 15 endpoints β”‚ └── models.py # Pydantic v2 models, Coalition/Delegate schemas β”œβ”€β”€ graders/ β”‚ β”œβ”€β”€ base.py # GradeResult, apply_mer(), meta-reward support β”‚ β”œβ”€β”€ black_swan_grader.py β”‚ β”œβ”€β”€ cascade_grader.py β”‚ β”œβ”€β”€ easy_grader.py β”‚ β”œβ”€β”€ fleet_coordination_grader.py β”‚ β”œβ”€β”€ hard_grader.py β”‚ └── medium_grader.py β”œβ”€β”€ simulation/ β”‚ β”œβ”€β”€ challenge_generator.py # 5 novel failure compositions β”‚ β”œβ”€β”€ cluster.py # State machine, staged patching, non-linear dynamics β”‚ β”œβ”€β”€ curriculum.py # Adaptive difficulty manager, meta-reward β”‚ β”œβ”€β”€ failures.py # Cascade phase injection, secondary failures β”‚ β”œβ”€β”€ fleet.py # Fleet AI, coalition actions, consensus β”‚ └── telemetry.py # NCCL logs, red herrings, staleness β”œβ”€β”€ tasks/ β”‚ β”œβ”€β”€ base.py # EpisodeTracker, preconditions, investigation_count β”‚ β”œβ”€β”€ black_swan.py # 300-step ultra-hard task, false positives, irreversible mistakes β”‚ β”œβ”€β”€ cascade.py # 3-phase chained task, phase gating β”‚ β”œβ”€β”€ easy.py β”‚ β”œβ”€β”€ fleet_coordination.py β”‚ β”œβ”€β”€ hard.py β”‚ └── medium.py β”œβ”€β”€ dashboard/ β”‚ └── war_room.py # Gradio SRE War Room (port 7861) β”œβ”€β”€ training/ β”‚ └── grpo_train.py # LoRA SFT warmup + optional GRPO over env rollouts β”œβ”€β”€ scripts/ β”‚ β”œβ”€β”€ evaluate.py # Oracle + random eval harness, 50 seeds, bootstrap CI β”‚ β”œβ”€β”€ evaluate_multi_agent.py β”‚ β”œβ”€β”€ evaluate_model.py # Raw/trained model-policy eval with action scoring β”‚ β”œβ”€β”€ exploit_test.py # exploit attack proofs β”‚ β”œβ”€β”€ generate_long_horizon_trace.py β”‚ β”œβ”€β”€ plot_final_evidence.py # Rebuilds committed PNG plots from eval JSON + loss log β”‚ β”œβ”€β”€ reward_audit.py # 4 reward property audits β”‚ β”œβ”€β”€ run_overnight.sh # One-shot env+SFT+GRPO+eval pipeline β”‚ └── seed_variance.py # Seed variance report β”œβ”€β”€ results/ β”‚ β”œβ”€β”€ before_vs_after_scores.png # Raw 0.239 vs trained 0.915, same evaluator β”‚ β”œβ”€β”€ cascade_long_horizon_trace.json β”‚ β”œβ”€β”€ exploit_test.json β”‚ β”œβ”€β”€ final_phaseaware_model_eval.json β”‚ β”œβ”€β”€ final_score_by_task.png # Trained LoRA mean score per task β”‚ β”œβ”€β”€ final_sft_loss_curve.png # SFT loss 2.53 -> 0.10 over 40 logged steps β”‚ β”œβ”€β”€ sft_warmup_metrics.json # Real per-step loss/lr/grad_norm log_history β”‚ β”œβ”€β”€ multi_agent_eval.json β”‚ └── reward_audit.json β”œβ”€β”€ docs/ β”‚ └── proof/ β”‚ β”œβ”€β”€ raw_untrained_console.png # HF Jobs eval, untrained base model, 0.239 mean β”‚ └── final_trained_console.png # Local terminal eval, trained LoRA, 0.915 mean β”œβ”€β”€ notebooks/ β”‚ β”œβ”€β”€ README.md β”‚ └── eval_trained_lora.ipynb # Colab-ready reproduction of headline numbers β”œβ”€β”€ tests/ β”‚ β”œβ”€β”€ test_fleet.py # 14 fleet + coalition tests β”‚ └── test_graders.py # 19 grader + anti-gaming tests β”œβ”€β”€ inference.py # Baseline LLM agent runner β”œβ”€β”€ openenv.yaml # OpenEnv v2.0.0, 5 manifest tasks, multi-agent β”œβ”€β”€ requirements.txt β”œβ”€β”€ requirements-training.txt └── server/ └── app.py ``` ## OpenEnv Compliance - `openenv validate` passes βœ… - Typed Pydantic v2 models with extra="forbid" βœ… - Deterministic seed-based graders βœ… - Docker deployment ready βœ… - 5 manifest tasks with difficulty progression βœ… - Multi-agent `/delegate` and `/coalition` evaluation βœ… - Mercor efficiency reward βœ… ## Known Limitations The submitted adapter is an SFT warmup policy, not a fully optimized online RL policy. `training/grpo_train.py` includes an optional GRPO loop wired to the environment reward function, but the submitted quantitative improvement is from supervised LoRA warmup over environment-generated SRE rollouts plus phase-aware constrained evaluation. The simulator is deterministic under seed for reproducibility and models production-inspired failure signatures rather than connecting to a real GPU cluster.