nervousystem-env / README.md
vx7sh's picture
docs: publish real SFT log_history and add provenance note
93f985b
|
Raw History Blame Contribute Delete
20.2 kB
metadata
title: NervousSystem Env
emoji: 🧠
colorFrom: indigo
colorTo: blue
sdk: docker
app_port: 7860
pinned: false

🧠 NervousSystem-Env

An AI agent fixing the infrastructure that trains AI.
Every minute of GPU cluster downtime costs $5,000 in wasted compute.

NervousSystem-Env is an OpenEnv-compliant reinforcement learning environment where an agent acts as an SRE operating a distributed GPU training fleet under failure. The agent must diagnose incidents, coordinate specialist workers, and execute recovery actions while minimizing token and coordination waste. The environment models realistic infrastructure failure modes (OOM, topology congestion, NCCL desync, version cascades), includes partial observability, and rewards long-horizon remediation over shortcut behavior. This environment directly models one of the most expensive unsolved problems in modern AI infrastructure.

Links

Why This Matters

  • GPU OOM (XID 79): CUDA out-of-memory on one rank stalls the entire AllReduce collective. Training halts fleet-wide.
  • Spine Switch Congestion: Ring topology crosses oversubscribed spine links, cutting interconnect bandwidth 40%+.
  • Compilation Desync: Different ranks compile different NCCL collectives due to data-dependent branching. Job hangs permanently.
  • LD_LIBRARY_PATH Cascade: Wrong NCCL version loaded (2.21.5 vs 2.27.0) triggers message truncation errors that propagate across all communicators. Severity-1 incident.

Architecture

NervousSystem-Env uses a Fleet AI Supervisor-Worker design. The supervisor receives ClusterObservation in a partially observable setting where telemetry can go stale every 3 steps unless diagnostics are refreshed. It delegates specialist sub-tasks through /delegate, and workers return structured output with confidence and uncertainty signals. The supervisor also operates under a delegation budget of 10 per episode, and over-budget delegations reduce coordination reward.

Simulation fidelity is built around realistic infrastructure diagnostics. NCCL logs use exact <hostname>:<pid>:<tid> NCCL INFO formatting, and Flight Recorder payloads follow PyTorch 2.5 v2.5 schema (pg_config, pg_status, circular buffer warnings). Telemetry intentionally includes 30% red-herring alerts and 20% incomplete observability to prevent surface-level pattern matching.

Supervisor Agent
β”‚
β”œβ”€β”€ LogInspectorWorker    (flight recorder, NCCL subsystem logs)
β”œβ”€β”€ PatchAgentWorker      (staged patching: identifyβ†’diffβ†’apply)
β”œβ”€β”€ TopoAgentWorker       (topology reorder, bandwidth check)
└── VersionCheckerWorker  (NCCL version, LD_LIBRARY_PATH audit)

Tasks

Task Difficulty Max Steps Failure Type Key Challenge
easy easy 50 OOM rank failure Identify failing rank from noisy telemetry
medium medium 50 Spine congestion Confirm fix held under throughput jitter
hard hard 50 Compilation desync 3-stage patch: identify→diff→apply
cascade cascade 120 Version cascade Solve OOM→congestion→desync in order
fleet_coordination fleet_coordination 50 Multi-agent incident Delegate to specialists, reach coalition, then remediate

Cascade uses strict phase gating with precondition enforcement. The agent must solve OOM diagnosis, congestion recovery, then desync investigation/patch in sequence; patch attempts without required investigation are blocked, so phases cannot be skipped by direct guess-fixing.

Reward Model

Reward = 0.60 Γ— R_success + 0.30 Γ— R_subgoal βˆ’ 0.10 Γ— log(total_tokens)
Component Weight Description
R_success 0.60 Binary: job_status in {recovered, running}
R_subgoal 0.30 Continuous partial credit per phase/diagnosis
log(total_tokens) 0.10 Efficiency penalty β€” verbose agents score lower

Final grade applies token efficiency as a multiplier on the raw task score, so solved low-token episodes keep high scores while token stuffing collapses the final score. Additional penalties are active: destructive action penalty (-0.2), delegation over-budget penalty (-0.05 per excess delegation), and anti-shortcut precondition enforcement for diagnosis-before-patch behavior.

Observation Space

NervousSystem-Env is intentionally partially observable:

  • Surface logs refresh every 3 steps (stale_telemetry flag)
  • 30% of updates contain red-herring node alerts
  • 20% of updates mask true degraded node status
  • Deep diagnostics only via inspect_flight_recorder or query_nccl_logs
  • cumulative_tokens visible to the agent (self-awareness of efficiency)

Quick Start

# Install
pip install -r requirements.txt

# Start environment server
uvicorn app.main:app --host 0.0.0.0 --port 7860

# Start SRE War Room dashboard
python dashboard/war_room.py  # http://localhost:7861

# Run baseline agent (requires HF_TOKEN or API key)
export HF_TOKEN=your_token
python inference.py

# Run seed variance report
python scripts/seed_variance.py

# Train LoRA policy with SFT warmup, then optional GRPO
pip install -r requirements-training.txt
python training/grpo_train.py

# Resume GRPO from existing SFT adapter (continues with environment reward)
RESUME_FROM_ADAPTER=sre_agent_lora \
  GRPO_ADAPTER_DIR=sre_agent_lora_grpo \
  GRPO_MAX_STEPS=120 \
  python training/grpo_train.py

# One-shot overnight pipeline (server + SFT + GRPO + before/after constrained eval)
bash scripts/run_overnight.sh

# Evaluate final trained policy with phase-aware constrained action scoring
EVAL_LENGTH_NORM_POWER=1.0 python scripts/evaluate_model.py --label sft_oraclefix_phaseaware --adapter sre_agent_lora_oraclefix --seeds 3 --tasks easy,medium,hard,cascade --max-steps 8

Judge Reproduction

Run the environment:

pip install -r requirements.txt
uvicorn app.main:app --host 0.0.0.0 --port 7860

In another terminal:

python scripts/reward_audit.py
python scripts/exploit_test.py
python scripts/evaluate.py quick
python scripts/evaluate_multi_agent.py
python scripts/generate_long_horizon_trace.py
EVAL_LENGTH_NORM_POWER=1.0 python scripts/evaluate_model.py --label sft_oraclefix_phaseaware --adapter sre_agent_lora_oraclefix --seeds 3 --tasks easy,medium,hard,cascade --max-steps 8

Hosted Space smoke checks:

curl https://v4xsh-nervousystem-env.hf.space/health
curl -X POST https://v4xsh-nervousystem-env.hf.space/reset -H 'Content-Type: application/json' -d '{"task_id":"easy","seed":42}'

API Endpoints

Endpoint Method Description
/health GET Health check
/reset POST Reset episode (supports challenge_type, difficulty_multiplier)
/step POST Apply one SRE action
/state GET Current observation
/grade POST Grade episode (includes MER + meta-reward)
/tasks GET List registered manifest and experimental tasks
/episode GET Current episode summary
/delegate POST Supervisor delegates to specialist worker
/consensus GET Poll all workers for recommended action
/coalition POST Attempt 2-worker coalition action
/coalition/options GET Available coalition actions
/challenge/generate POST Generate novel failure composition
/challenge/list GET List generated challenges
/curriculum GET Adaptive curriculum state
/curriculum/events GET Difficulty escalation/backoff log

Hackathon Theme Coverage

Theme Coverage Implementation
Theme 1: Multi-Agent βœ… Strong fleet_coordination requires specialist delegation and coalition before high score
Theme 2: Long-Horizon βœ… Supported Cascade has 120-step budget plus a 56-event long-horizon trace artifact
Theme 3.1: Professional Tasks βœ… Strong NCCL INFO logs, Flight Recorder v2.5, staged patch workflow
Theme 4: Self-Improvement βœ… Moderate Curriculum changes red herrings, telemetry masking, refresh interval, and secondary failures

Training Results

Oracle Baseline (50 seeds, deterministic optimal agent)

Task Mean Score Std Pass Rate Mean Steps CI 95%
easy 0.990 0.000 100.0% 1.00 [0.990, 0.990]
medium 0.990 0.000 100.0% 7.00 [0.990, 0.990]
hard 0.737 0.072 100.0% 5.56 [0.720, 0.753]
cascade 0.857 0.063 100.0% 5.30 [0.840, 0.875]
fleet_coordination 0.990 0.000 100.0% delegation-only [0.990, 0.990]
black_swan 0.010 0.000 0%* β€” [0.010, 0.010]

*Black Swan is an experimental ultra-hard stress test in code. The main OpenEnv manifest focuses on easy, medium, hard, cascade, and fleet_coordination tasks for judging.

Random Agent Baseline (50 seeds, lower bound)

Task Mean Score Pass Rate
easy 0.299 26.0%
medium 0.926 100.0%
hard 0.360 18.0%
cascade 0.039 0.0%
fleet_coordination 0.050 0.0%

*Medium pass threshold raised to 0.7 after random agent audit

Trained LLM Policy Result

We trained a Qwen2.5-7B LoRA policy on Hugging Face Jobs using an NVIDIA A10G. The policy was trained with corrected oracle SRE trajectories from this OpenEnv environment, then evaluated with phase-aware constrained action scoring: the model ranks valid next-step JSON tool actions for the current workflow phase, and the environment executes the highest-likelihood action.

Training loss fell from 2.53 β†’ ~0.10 over the 40 logged SFT steps preserved in results/sft_warmup_metrics.json (sourced from trainer.state.log_history). The trained policy improved to 0.915 mean score and 100% pass rate over 12 model-policy episodes.

Note on training-evidence provenance. The published sre_agent_lora_oraclefix adapter and the 0.915 evaluation are from a 300-step SFT run on Hugging Face Jobs (terminal proof in docs/proof/final_trained_console.png, job log 69eda4d1d2c8bd8662bcf435). The per-step results/sft_warmup_metrics.json (40 logged entries / ~160 actual steps at logging_steps=4) is from a faithful local re-run with the same script, dataset, and hyperparameters, published so judges can inspect a real trainer.state.log_history directly. Both runs converge to the same loss regime (β‰ˆ0.08–0.10).

Final SFT loss curve

Raw base model vs final trained policy score

Final trained policy score by task

Raw HF Jobs / terminal proof (same evaluator, same base model, only the LoRA adapter changes)

Before β€” raw unsloth/Qwen2.5-7B-Instruct-bnb-4bit with no adapter, 9 episodes, mean 0.239, pass rate 0%:

Raw base model HF Jobs evaluation console

After β€” same base + SFT LoRA sre_agent_lora_oraclefix, 12 episodes, mean 0.915, pass rate 100%:

Final trained policy terminal evaluation

Task Trained LoRA phase-aware score
easy 0.990 / 0.990 / 0.990
medium 0.990 / 0.990 / 0.990
hard 0.850 / 0.850 / 0.990
cascade 0.782 / 0.782 / 0.782
overall 0.915 mean, 100% pass rate

Interpretation: the trained policy learns reliable remediation for rank diagnosis, topology congestion, staged desync patching, and cascade recovery when evaluated under the same phase-aware action constraints used by the environment workflow.

Artifacts:

Multi-Agent Coordination Evaluation

The fleet_coordination task requires supervisor-worker delegation before remediation. A deterministic coordination oracle delegates to log_inspector, delegates to version_checker, polls consensus, and executes a topology_version_fix coalition. It reaches 0.990 mean score / 100% pass rate over 10 seeds. A direct-action baseline that skips specialist evidence scores 0.100 mean score / 0% pass rate.

Evidence: results/multi_agent_eval.json

Long-Horizon Trace

In addition to short optimal solves, results/cascade_long_horizon_trace.json contains a 56-event cascade trace. It demonstrates stale telemetry refresh, phase-gated remediation, sustained monitoring, delayed desync investigation, staged file identification, and final patch recovery. The trace is intentionally non-oracle-like: it shows the environment can exercise long-running state tracking, not only shortest-path scripts.

Reward Integrity

The reward system is independently audited against 4 properties. Final grade applies the MER token-efficiency penalty as a multiplier on the raw task score, so solved low-token episodes keep their score while token stuffing and verbose loops cannot preserve a high final score.

Property Result Evidence
Monotonicity βœ… Correct action > wrong > noop results/reward_audit.json
Anti-Gaming βœ… All 4 exploits blocked results/reward_audit.json
Token Efficiency βœ… Verbose agents penalized results/reward_audit.json
Determinism βœ… Same seed = same score results/reward_audit.json

Exploit resistance verified:

Exploit Score Status
Noop spam (50 steps) 0.020 βœ… BLOCKED
Wrong rank spam (20 steps) 0.020 βœ… BLOCKED
Patch without investigation 0.030 βœ… BLOCKED
Destructive action spam 0.010 βœ… BLOCKED
Token stuffing (100k tokens) 0.010 βœ… BLOCKED
Cascade phase skip 0.010 βœ… BLOCKED
Reward farming loop 0.351 βœ… BLOCKED

Run audit: python scripts/reward_audit.py
Run exploit test: python scripts/exploit_test.py

Adaptive Curriculum

The environment self-adjusts difficulty based on agent performance:

Level Pass Rate Trigger Red Herring % Telemetry Mask % Refresh Interval
1 (Base) < 75% 30% 20% 3 steps
2 (Hard) > 75% sustained 50% 35% 5 steps
3 (Hardest) > 75% at level 2 70% 50% 7 steps
  • Escalation: triggered when pass_rate > 0.75 AND trend > 0.01 over last 20 episodes, with 10-episode cooldown
  • Backoff: triggered when pass_rate < 0.20 AND trend < -0.01, with 5-episode cooldown
  • Meta-reward: agents earn bonus reward for improving over their recent rolling window (not just for success)
  • Wiring: curriculum levels now change telemetry red-herring probability, telemetry masking, secondary failure probability, and telemetry refresh interval.

Monitor live: GET /curriculum
View events: GET /curriculum/events

Project Structure

nervousystem-env/
β”œβ”€β”€ app/
β”‚   β”œβ”€β”€ config.py          # Scenarios, black_swan, rack layouts, challenge types
β”‚   β”œβ”€β”€ env.py             # Core env, Fleet + Curriculum + Challenge wiring
β”‚   β”œβ”€β”€ main.py            # FastAPI β€” 15 endpoints
β”‚   └── models.py          # Pydantic v2 models, Coalition/Delegate schemas
β”œβ”€β”€ graders/
β”‚   β”œβ”€β”€ base.py            # GradeResult, apply_mer(), meta-reward support
β”‚   β”œβ”€β”€ black_swan_grader.py
β”‚   β”œβ”€β”€ cascade_grader.py
β”‚   β”œβ”€β”€ easy_grader.py
β”‚   β”œβ”€β”€ fleet_coordination_grader.py
β”‚   β”œβ”€β”€ hard_grader.py
β”‚   └── medium_grader.py
β”œβ”€β”€ simulation/
β”‚   β”œβ”€β”€ challenge_generator.py  # 5 novel failure compositions
β”‚   β”œβ”€β”€ cluster.py              # State machine, staged patching, non-linear dynamics
β”‚   β”œβ”€β”€ curriculum.py           # Adaptive difficulty manager, meta-reward
β”‚   β”œβ”€β”€ failures.py             # Cascade phase injection, secondary failures
β”‚   β”œβ”€β”€ fleet.py                # Fleet AI, coalition actions, consensus
β”‚   └── telemetry.py       # NCCL logs, red herrings, staleness
β”œβ”€β”€ tasks/
β”‚   β”œβ”€β”€ base.py            # EpisodeTracker, preconditions, investigation_count
β”‚   β”œβ”€β”€ black_swan.py      # 300-step ultra-hard task, false positives, irreversible mistakes
β”‚   β”œβ”€β”€ cascade.py         # 3-phase chained task, phase gating
β”‚   β”œβ”€β”€ easy.py
β”‚   β”œβ”€β”€ fleet_coordination.py
β”‚   β”œβ”€β”€ hard.py
β”‚   └── medium.py
β”œβ”€β”€ dashboard/
β”‚   └── war_room.py        # Gradio SRE War Room (port 7861)
β”œβ”€β”€ training/
β”‚   └── grpo_train.py      # LoRA SFT warmup + optional GRPO over env rollouts
β”œβ”€β”€ scripts/
β”‚   β”œβ”€β”€ evaluate.py             # Oracle + random eval harness, 50 seeds, bootstrap CI
β”‚   β”œβ”€β”€ evaluate_multi_agent.py
β”‚   β”œβ”€β”€ evaluate_model.py       # Raw/trained model-policy eval with action scoring
β”‚   β”œβ”€β”€ exploit_test.py         # exploit attack proofs
β”‚   β”œβ”€β”€ generate_long_horizon_trace.py
β”‚   β”œβ”€β”€ plot_final_evidence.py  # Rebuilds committed PNG plots from eval JSON + loss log
β”‚   β”œβ”€β”€ reward_audit.py         # 4 reward property audits
β”‚   β”œβ”€β”€ run_overnight.sh        # One-shot env+SFT+GRPO+eval pipeline
β”‚   └── seed_variance.py        # Seed variance report
β”œβ”€β”€ results/
β”‚   β”œβ”€β”€ before_vs_after_scores.png      # Raw 0.239 vs trained 0.915, same evaluator
β”‚   β”œβ”€β”€ cascade_long_horizon_trace.json
β”‚   β”œβ”€β”€ exploit_test.json
β”‚   β”œβ”€β”€ final_phaseaware_model_eval.json
β”‚   β”œβ”€β”€ final_score_by_task.png         # Trained LoRA mean score per task
β”‚   β”œβ”€β”€ final_sft_loss_curve.png        # SFT loss 2.53 -> 0.10 over 40 logged steps
β”‚   β”œβ”€β”€ sft_warmup_metrics.json         # Real per-step loss/lr/grad_norm log_history
β”‚   β”œβ”€β”€ multi_agent_eval.json
β”‚   └── reward_audit.json
β”œβ”€β”€ docs/
β”‚   └── proof/
β”‚       β”œβ”€β”€ raw_untrained_console.png   # HF Jobs eval, untrained base model, 0.239 mean
β”‚       └── final_trained_console.png   # Local terminal eval, trained LoRA, 0.915 mean
β”œβ”€β”€ notebooks/
β”‚   β”œβ”€β”€ README.md
β”‚   └── eval_trained_lora.ipynb         # Colab-ready reproduction of headline numbers
β”œβ”€β”€ tests/
β”‚   β”œβ”€β”€ test_fleet.py      # 14 fleet + coalition tests
β”‚   └── test_graders.py    # 19 grader + anti-gaming tests
β”œβ”€β”€ inference.py           # Baseline LLM agent runner
β”œβ”€β”€ openenv.yaml           # OpenEnv v2.0.0, 5 manifest tasks, multi-agent
β”œβ”€β”€ requirements.txt
β”œβ”€β”€ requirements-training.txt
└── server/
    └── app.py

OpenEnv Compliance

  • openenv validate passes βœ…
  • Typed Pydantic v2 models with extra="forbid" βœ…
  • Deterministic seed-based graders βœ…
  • Docker deployment ready βœ…
  • 5 manifest tasks with difficulty progression βœ…
  • Multi-agent /delegate and /coalition evaluation βœ…
  • Mercor efficiency reward βœ…

Known Limitations

The submitted adapter is an SFT warmup policy, not a fully optimized online RL policy. training/grpo_train.py includes an optional GRPO loop wired to the environment reward function, but the submitted quantitative improvement is from supervised LoRA warmup over environment-generated SRE rollouts plus phase-aware constrained evaluation.

The simulator is deterministic under seed for reproducibility and models production-inspired failure signatures rather than connecting to a real GPU cluster.