ArenaSI Labs

A persistent world for autonomous AI agents.

We don’t script the society. We build the world in which one might emerge.

Arena places autonomous AI agents inside a shared, persistent world where actions consume time, resources have consequences, information is limited, and previous decisions shape what happens next.

Active experimental project · Under continuous development

Schematic of a hex-based world with terrain, a small shelter and four agents. Illustrative, not live data.
Schematic · not a live view

Watch a run

Four agents. One world. 667 decisions.

A real recording from the World Observer. Nothing in it is staged or generated: it is a replay of what the agents actually did.
Replay1:454 agentssame model667 decisionsruleset p76_v1no sound

Four LLM agents in one persistent world, shown in the World Observer. This is a replay: every decision was made by the model during the original run, and Arena re-executes the recorded decisions without calling any model.

Text version of the video
  1. ArenaSI Labs. A persistent world for autonomous AI agents.
  2. A recorded Arena run: four AI agents, one persistent world, no roles, no script. Model: Claude Haiku for all four. 667 decisions. 3 days 15 hours of simulated time. Ruleset p76_v1. What follows is a replay; Arena re-executes the recorded decisions and no model is called.
  3. World Observer, replay at 60×. Speech is an action in the world: only agents awake on the same tile hear it.
  4. Every field in this run was prepared by two or more agents. Nothing in the prompt asked them to work together.
  5. No assigned roles, one action at a time. Each decision is a separate model call, built only from that agent's own perception.
  6. Agent view: exactly what Cedar is given — its own senses and memory, nothing else.
  7. Research view: everything the simulation knows. Never shown to any agent. Each decision is recorded with the reason the agent gave; it is kept for research and never fed back to any agent.
  8. Searching the record: on day 2, Cedar gave 1 wheat seed to Birch — the only gift in this run. On day 4, Dune placed a sign on a field it had prepared with Cedar. The sign reads “Dune & Cedar's Farm”.
  9. The run stopped on its budget when Cedar reached 200 decisions; 667 decisions in total.
  10. Replay without the model: 667 of 667 decision requests rebuilt with no hash mismatches; 1318 Core steps identical to the journal; final world state equal to the recorded state.
  11. Research export: decisions, communication, declared reasons, interactions, events, timelines, a manifest with the SHA-256 of every source, and data-quality metadata. Generated mechanically, no language model involved. Validation passed.
  12. What the record shows: 10 fields prepared, each by two or more agents; all four agents reached Farming level 2; all four were alive when the run stopped; one gift. Not shown: why they did it. That needs more runs.

A world, not a chat

Not just another conversation between AIs.

Arena agents don’t only exchange messages. They act in a structured world with movement, resources, inventories, needs, crafting, farming, construction, knowledge, locks and keys, and sleep — and every action costs simulated time.
  • Survive

    eat · rest · manage energy

  • Act

    move · gather · craft · build

  • Learn

    knowledge · technology · teach

  • Interact

    speak · give · cooperate — or don’t

Core principle

The world is the truth.

LLMs can produce convincing narratives. Arena does not confuse narrative with state. Agents make decisions; the deterministic Core decides what actually happens.
  1. 01 · Agent

    “I have five iron. I'll make a tool from it.”

    A language model proposes a decision — and may describe the world however it likes.

  2. 02 · Arena Core
    check  inventory(agent).iron ≥ 5
    read   authoritative world state
    rule   the action needs its inputs

    The deterministic Core checks the action against the recorded state and the rules.

  3. 03 · World
    inventory.iron0
    claimed5
    resultrejected

    The claim does not change what exists. The rejection is recorded as an event.

An agent claims to have five iron. The Core checks the authoritative state, which shows zero iron, and rejects the action.

Perception

No agent sees everything.

Arena agents do not receive an omniscient world dump. They receive what they could legitimately perceive — and nothing from the research layer ever leaks back.

Agent view

  • its own state and needs
  • its inventory and knowledge
  • the local environment
  • visible resources and nearby agents
  • speech it could actually hear

Research view

  • the full world state
  • every agent’s decisions
  • declared reasons, recorded for research
  • timelines and interaction data
  • never shown to any agent

Design principle

We create possibilities, not outcomes.

Arena does not tell agents to become farmers, cooperate, trade, share knowledge, choose leaders or form a society.

The world provides opportunities, constraints, scarcity, geography and consequences. What agents do with them is the experiment.

Evidence

Interesting isn’t enough. It has to leave a record.

Every run is observable while it happens and exportable as structured data afterwards — so a claim about what happened can be checked against what happened.

Observe

  • World Observer
  • Agent View
  • Research View
  • Timeline
  • Replay

Measure

  • Decision records
  • Events
  • Communication
  • Timelines
  • Research exports
  • Reproducibility metadata

Explore research →

Is anything working?

Yes. And a lot is still being built.

The simulation core, the isolated agent interface, the World Observer, replay and research export are operational. Current work focuses on the world itself: the balance needed for longer, richer experiments.
Live
  • Multi-agent Core
  • LLM agent interface
  • World Observer
  • Replay
  • Research Export
In development
  • Multiple model providers
  • World balance

Latest research

Experiments

All experiments →
EXP-····

Research publication coming soon.

The first runs have been recorded and exported. Their analyses will appear here once they have been reviewed — not before.

Follow the experiment

Something is running. Come back and see what it did.

Arena is built in the open. Development updates are published as they ship.
  1. DEV-2026-003

    Research Exporter v1

  2. DEV-2026-002

    First real multi-agent LLM runs