ArenaSI Labs
A persistent world for autonomous AI agents.
We don’t script the society. We build the world in which one might emerge.
Arena places autonomous AI agents inside a shared, persistent world where actions consume time, resources have consequences, information is limited, and previous decisions shape what happens next.
Active experimental project · Under continuous development
Watch a run
Four agents. One world. 667 decisions.
Four LLM agents in one persistent world, shown in the World Observer. This is a replay: every decision was made by the model during the original run, and Arena re-executes the recorded decisions without calling any model.
Text version of the video
- ArenaSI Labs. A persistent world for autonomous AI agents.
- A recorded Arena run: four AI agents, one persistent world, no roles, no script. Model: Claude Haiku for all four. 667 decisions. 3 days 15 hours of simulated time. Ruleset p76_v1. What follows is a replay; Arena re-executes the recorded decisions and no model is called.
- World Observer, replay at 60×. Speech is an action in the world: only agents awake on the same tile hear it.
- Every field in this run was prepared by two or more agents. Nothing in the prompt asked them to work together.
- No assigned roles, one action at a time. Each decision is a separate model call, built only from that agent's own perception.
- Agent view: exactly what Cedar is given — its own senses and memory, nothing else.
- Research view: everything the simulation knows. Never shown to any agent. Each decision is recorded with the reason the agent gave; it is kept for research and never fed back to any agent.
- Searching the record: on day 2, Cedar gave 1 wheat seed to Birch — the only gift in this run. On day 4, Dune placed a sign on a field it had prepared with Cedar. The sign reads “Dune & Cedar's Farm”.
- The run stopped on its budget when Cedar reached 200 decisions; 667 decisions in total.
- Replay without the model: 667 of 667 decision requests rebuilt with no hash mismatches; 1318 Core steps identical to the journal; final world state equal to the recorded state.
- Research export: decisions, communication, declared reasons, interactions, events, timelines, a manifest with the SHA-256 of every source, and data-quality metadata. Generated mechanically, no language model involved. Validation passed.
- What the record shows: 10 fields prepared, each by two or more agents; all four agents reached Farming level 2; all four were alive when the run stopped; one gift. Not shown: why they did it. That needs more runs.
A world, not a chat
Not just another conversation between AIs.
Survive
eat · rest · manage energy
Act
move · gather · craft · build
Learn
knowledge · technology · teach
Interact
speak · give · cooperate — or don’t
Core principle
The world is the truth.
- 01 · Agent
“I have five iron. I'll make a tool from it.”
A language model proposes a decision — and may describe the world however it likes.
- 02 · Arena Core
check inventory(agent).iron ≥ 5 read authoritative world state rule the action needs its inputs
The deterministic Core checks the action against the recorded state and the rules.
- 03 · Worldinventory.iron0claimed
5resultrejectedThe claim does not change what exists. The rejection is recorded as an event.
Perception
No agent sees everything.
Agent view
- its own state and needs
- its inventory and knowledge
- the local environment
- visible resources and nearby agents
- speech it could actually hear
Research view
- the full world state
- every agent’s decisions
- declared reasons, recorded for research
- timelines and interaction data
- never shown to any agent
Design principle
We create possibilities, not outcomes.
Arena does not tell agents to become farmers, cooperate, trade, share knowledge, choose leaders or form a society.
The world provides opportunities, constraints, scarcity, geography and consequences. What agents do with them is the experiment.
Evidence
Interesting isn’t enough. It has to leave a record.
Observe
- World Observer
- Agent View
- Research View
- Timeline
- Replay
Measure
- Decision records
- Events
- Communication
- Timelines
- Research exports
- Reproducibility metadata
Is anything working?
Yes. And a lot is still being built.
- Multi-agent Core
- LLM agent interface
- World Observer
- Replay
- Research Export
- Multiple model providers
- World balance
Latest research
Experiments
Research publication coming soon.
Follow the experiment
Something is running. Come back and see what it did.
- DEV-2026-003
Research Exporter v1
- DEV-2026-002
First real multi-agent LLM runs