GameBench is the game-environment task surface for Managed Research. It turns
policy and code-improvement work into lanes with explicit objectives, run
receipts, scores, artifacts, and failure modes.
Use GameBench when you want a concrete task that is easier to inspect than a
transcript:
- code-policy optimization lanes such as TicTacToe, Sokoban, Crafter, Craftax,
MiniHack, Rogue, Frogs, and Overcooked
- policy-puzzle lanes that ask the agent to diagnose a specific behavioral flaw
- launch evidence tables with run id, terminal state, score, duration, and cost
GameBench launch tables are pilot evidence unless the post explicitly says
otherwise. Do not treat one-off or 10-lane launch runs as benchmark
certification, a stable leaderboard, or ReportBench correlation.
What a GameBench receipt should include
Launch pilot shape
The scoped launch pilot uses Grok Build for orchestrator, worker, and reviewer
roles across 10 lanes:
tictactoe_gamebench_code_policy_deo_hillclimb_1cand_local
sokoban_gamebench_code_policy_deo_hillclimb_1cand_local
crafter_gamebench_code_policy_deo_hillclimb_1cand_local
frogs_gamebench_code_policy_deo_hillclimb_1cand_local
rogue_gamebench_code_policy_deo_hillclimb_1cand_local
craftax_gamebench_code_policy_deo_hillclimb_1cand_local
minihack_gamebench_code_policy_deo_hillclimb_1cand_local
overcooked_v2_gamebench_code_policy_deo_hillclimb_1cand_local
crafter_gamebench_policy_puzzle_front_only_1cand_local
crafter_gamebench_policy_puzzle_stone_blind_1cand_local
The launch config lives in synth-dev:
How this fits Managed Research
GameBench is one evidence source for Managed Research. Use it alongside normal
run receipts, Open Research proofs, and project-specific artifacts.
For broad claims, collect repeated seeds, multiple model families, and the
correlation study called out in the launch plan. Until then, say “pilot” and
publish the run receipts.
Next