Benchmarking Open-Ended Multi-Agent Coordination in Language Agents
Alem
Can LLM agents coordinate in long-horizon, open-ended tasks?
Alem is a JAX benchmark for testing multi-agent coordination in long-horizon, procedurally generated worlds. Across nine levels with controllable coordination demands, agents must explore, communicate, trade resources, craft tools, build structures, and fight mobs. Alem supports LLMs, VLMs, RL agents, and human play.