Skip to content

feat(RL): OpenEnv terminal-bench rollouts without Docker - #461

Open
ishandhanani wants to merge 5 commits into
mainfrom
idhanani/miles-openenv
Open

ishandhanani wants to merge 5 commits into
mainfrom
idhanani/miles-openenv

Conversation

@ishandhanani

Copy link
Copy Markdown
Collaborator

OpenEnv environments as pool services, no Docker and no hosted sandbox, driven by Miles on main (#452 + #455).

What

  • examples/miles/openenv-tbench2-smoke.yaml + benchmarks/rl/openenv/tbench2_smoke.py: the Terminal-Bench-2 env server (huggingface/OpenEnv tbench2_env, TB2_MODE=local) as a one-node generic service with a GET /health readiness probe; the benchmark step does reset(task_id), one exec, evaluate against it. No GPUs touched.
  • benchmarks/rl/miles/recipes/openenv_tbench2_qwen3.py + examples/miles/qwen3-4b-openenv-tbench2.yaml: Miles's Terminal-Bench-2 adapter with a dense Qwen3-4B profile (TP=2, EP=1, qwen parsers, TITO qwen3). The ray service owns the pool, the env server rides on it (placement.pool: train), and the launcher points the adapter at localhost:8003. OPENENV_SITE appends a client-only site dir to the ray job's PYTHONPATH.
  • benchmarks/rl/openenv/README.md: the two Python installs (a venv for the server; a client-only site dir for the rollout workers, resolved against the image's interpreter, --no-deps, mcp pinned to 1.x) and why a venv with --system-site-packages is not enough when the image's python3 is itself a venv.
  • docs/miles.md: "OpenEnv without Docker" subsection.

Validated on sa-b200

Job Recipe Result
15558 smoke COMPLETED 0:0 in 2:39. /health 200, reset(headless-terminal) in 14 s, exec ran in the task shell, evaluate returned the canonical harness verdict
15583 training, first attempt FAILED: Miles's openenv_launch_common.cleanup() does pgrep -f sglang | xargs kill; everything in the image runs under /opt/sglang/bin/python3 and enroot shares the PID namespace, so it killed the ray service's head. The recipe no longer calls it
15591 training COMPLETED 0:0 in 19:03. Two GRPO steps, 64 env sessions, checkpoint iteration 1 saved, ray job succeeded, tachometer parquet written

What 15591 also showed

62 of 64 episodes were dropped with no canonical verdict and rewards were 0. Local mode stages the verifier at /tests and /logs/verifier per container and runs agents in the shared task directory, so 32 concurrent episodes corrupt each other. The example now runs MAX_CONCURRENT_ENVS=1 (the Miles adapter queues episodes on the capacity signal), 2 prompts x 4 samples per step, and keeps TB2_OUTPUT_DIR under /logs. That configuration has not been run yet. TB2_WITHHOLD_TESTS=1 deletes tests/ and solution/ from the checkout on disk; git checkout -- . between runs.

Real parallelism needs one filesystem per episode: per-episode enroot containers nested in the trainer's container (works on sa-b200 with the host's enroot tooling at its own paths) or a hosted sandbox provider. Neither is in this PR.

No src/srtctl changes.

…smoke recipe

tbench2_env (huggingface/OpenEnv) in TB2_MODE=local runs task commands in its
own container's shell, so an OpenEnv environment can serve rollouts on a pool
node with no Docker and no nested container. examples/miles/openenv-tbench2-smoke.yaml
runs the server as a one-node service (readiness on GET /health) and a benchmark
step that exercises the client contract: reset(task_id), one exec, evaluate.
benchmarks/rl/openenv/tbench2_smoke.py is that client; its README explains
where OpenEnv servers sit in a recipe.
benchmarks/rl/miles/recipes/openenv_tbench2_qwen3.py keeps Miles's
Terminal-Bench-2 adapter and shared launch helpers and swaps the GLM-4.7-Flash
profile for a dense Qwen3-4B one (TP=2, EP=1, qwen parsers, TITO qwen3), with
--num-rollout overridable and OPENENV_SITE appended to the ray job's
PYTHONPATH so rollout workers import the env client without shadowing the
image's packages.

examples/miles/qwen3-4b-openenv-tbench2.yaml: the ray pool owns the node, the
tbench2_env server rides on it in TB2_MODE=local, and the launcher points the
adapter at localhost:8003. Local mode shares a task's source directory across
concurrent episodes; the recipe says so.
openenv_launch_common.cleanup() does `pgrep -f sglang | xargs kill`. Every
process in the Miles image runs under /opt/sglang/bin/python3, so the pattern
matches the `ray start --block` head, its dashboard and agents in the sibling
service container (enroot shares the PID namespace), and `ray job submit`
then finds port 8265 refused. Job 15583 on sa-b200: head received SIGTERM at
08:04:47, cleanup started at the same second. Under srt-slurm the allocation
is fresh and Ray belongs to the service, so the recipe skips the cleanup.
… the openenv README

docs/miles.md gains an "OpenEnv without Docker" subsection: what TB2_MODE=local
is, the two recipes, the two Python installs, why the upstream cleanup must
not run, and what local mode does not isolate. benchmarks/rl/openenv/README.md
records the client-only site directory procedure: resolve against the image's
interpreter, install only what is missing with --no-deps, pin mcp 1.x.
Job 15591 ran the loop end to end (32 sessions, a GRPO step) but 31
evaluates failed on files another episode had removed: tbench2_env's local
mode stages the verifier at /tests and /logs/verifier per container and runs
agents in the shared task directory, so overlapping episodes corrupt each
other. The example now sets MAX_CONCURRENT_ENVS=1 (the Miles adapter queues
episodes on the capacity signal), sizes the rollout to 2 prompts x 4
samples, and keeps TB2_OUTPUT_DIR under /logs so per-episode terminal logs
are job artifacts. docs/miles.md records the finding and that
TB2_WITHHOLD_TESTS deletes tests/ and solution/ from the checkout on disk.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant