Benchmarking Open-Ended Multi-Agent Coordination in Language Agents

Author

Kale-ab Abebe Tessera, Andras Szecsenyi, Cameron Barker, Alexander Rutherford, Davide Paglieri, Aidan Scannell, Henry Gouk, Elliot J. Crowley, Tim Rocktäschel, Amos Storkey

Published

June 6, 2026

Alem

Can LLM agents coordinate in long-horizon, open-ended tasks?

Alem is a JAX benchmark for testing multi-agent coordination in long-horizon, procedurally generated worlds. Across nine levels with controllable coordination demands, agents must explore, communicate, trade resources, craft tools, build structures, and fight mobs. Alem supports LLMs, VLMs, RL agents, and human play.

Submit to leaderboard