← ritwiksharma.com
Research System
Principle
The Loop
Evidence
Zombita
A Self-Learning Agent Administering a Live Game World · Ritwik Sharma
Built by Ritwik Sharma (solo)
Status Live · player-facing seasons since July 2026
Community Private Project Zomboid (Build 42) server
Stack Python · FastAPI · PostgreSQL · Next.js · Lua
Documentation Read Full Docs →
Research Context
Zombita is the administrative layer of a private Project Zomboid community: a Discord agent, a web backend and community site, and a suite of custom in-game mods, sharing one PostgreSQL database of roughly one hundred tables. She perceives every message, forms a nightly per-person read in her own voice, keeps four distinct affective gauges toward the people she deals with, consolidates what she learns into a versioned living memory, adjudicates player claims with real stakes, and runs a closed-loop economy on an in-game-month heartbeat. The doctoral work is not to build this system. It is to turn it into a research instrument.
The Governing Principle — The Model Proposes, the Engine Disposes
The language model never holds the lever. It can only propose. A deterministic engine decides whether a proposal executes, and the engine is the only thing that can act.
The model’s tools for any consequential action do not perform it. They record a structured proposal of what she wants to do and why. Ordinary deterministic code holds the connection to the game server, the treasury and the item-grant pipeline, and every consequential action passes through one function:
authorize(action, actor) → (allowed, reason)
Today it implements the most conservative policy available: an action executes only when an authorised human admin confirms it. As autonomy is earned, only the body of this function changes, from “a human confirmed this” to “this action, by this agent, with this track record, in this context, clears the bar.” Disruptive operations are warn-first by construction, the most destructive require a typed confirmation phrase, and every proposal is written to a permanent audit record. That record is both the accountability surface today and the evidence base for opening the gate later.
The Administrative Faculties as One Loop
perceive → form a nightly read of each person → the read moves trust, reputation and tomorrow’s attention → trust governs whether she acts for someone → everything is consolidated nightly into a living memory → perceive
Perception deterministic
A rule-based scorer runs on every message and decides two things in one cheap pass: whether she should respond, and whether the message is worth remembering. Noise is dropped before any model is called. When she does speak, the model receives a context hint saying why the message mattered, so code decides salience and the model only renders the reply.
Tier 1 direct address 0.90–1.00 · Tier 2 panic / drama / question to the room 0.55–0.75 · Tier 3 soft signal 0.15–0.35 · Tier 4 noise 0.00, no model called
Identity Resolution refuse over guess
One resolver maps any handle (nickname, in-game name, display name, lore alias) to one canonical person. When a name fits more than one person it does not pick the likeliest; it surfaces the collision and the action refuses to proceed silently. Asking costs one interaction; guessing wrong costs an irreversible action against the wrong person.
resolve / ask / refuse · ambiguous item or recipient → disambiguation picker
The Read of a Person nightly
Every active day, for each active member, she forms a private read in her own voice, using the captured messages themselves, the person’s real words and why each stood out, not event counts. It yields an archetype, a revised first-person opinion and a reputation change. Over a season this becomes an evidence-grounded theory of every regular, checkable against the verbatim record beneath it.
archetype · first-person opinion · reputation delta · revised, not overwritten
Four Affective Gauges hidden
Mood is a volatile global register that colours her voice but is firewalled from decisions. Disposition is a slow per-person climate. Trust answers whether she would take your word. Respect answers whether you have earned her effort. They stay separate because collapsing them loses the distinctions an administrator makes. Trust gates whether a request is honoured at all, the social analogue of the authorisation seam.
mood · disposition · trust · respect — all keyed to the canonical person, none shown to players
The Nightly Pipeline bounded cost
Notable moments accrue to a daily buffer at zero model cost. A cheap model distils the buffer into a dated summary. A strong model reads a bounded window of recent summaries and writes a reflection, which on a calm night is allowed to say that nothing needs attention. Periodically the strongest model performs the only operation allowed to rewrite, consolidating pending notes into a new version of the living memory.
summaries kept verbatim · memory read-only to the world · every version retained
Economy and Treasury in-game-month heartbeat
A closed loop with a central treasury: prices move inside authored bands with treasury state, demand, player wealth and recession; taxes rise as the treasury falls; scarcity is real; part of every transaction is destroyed. Once per in-game month she makes a single bounded judgment over a compact digest, which the engines execute across the catalogue at no marginal model cost.
final price = base × treasury × demand × wealth × recession × her factor, clamped to the tier band
Judgment Engine filter, then judge
She adjudicates player claims with real stakes, such as “I lost my item, can I have it back?”. A cheap model does the token-heavy reading and distils an evidence brief; a strong model gives a verdict over a few hundred pre-digested tokens, weighed against the player’s history and standing. The safe default is to propose to a human rather than act.
grant · deny · suspicious · any failure → suspicious → escalate to a human
Live Admin Channel peer operator
A dedicated channel where she sits as a fifth admin, watching the human team run the server and learning the craft. A free pre-filter and a small fixed-window check decide whether she speaks; when she does, she first investigates with read-only tools (server status, console digest, log slices) and answers from the evidence. Every decision to speak or stay silent is logged.
read-only toolkit · proposals only for actions · fixed windows + rolling summary = flat cost
What Runs, What Is Gated, What Accumulates
Runs autonomously
No human in the loop per decision
Pricing, taxation, scarcity, sinks and the monthly economic judgment. Perception on every message. Identity and item resolution (resolve, ask or refuse). The nightly summary, reflection and memory consolidation. The per-person read and the affective gauges.
Gated by design
Human confirmation required
Consequential server actions (broadcasts, restarts, stops, world resets, config edits), item grants, whitelist and access changes, and claim verdicts. This is the safety mechanism working as intended, and the mechanism that produces the evidence RQ1 needs.
Accumulating
The record the research reads
A permanent audit of every proposal and its outcome, the versioned living memory, nightly reflections, per-person reads, and verbatim daily summaries beneath them. The same infrastructure that keeps the system accountable is the researcher’s evidence base.
What the System Records — Standing Data Sources
SourceFormSupports
Authorisation audit recordEvery proposal: content, stated reasoning, whether authorised, by whom, outcomeRQ1
Versioned living memoryEvery version retained, with pending notes and consolidation eventsRQ2
Per-person nightly readsArchetype, first-person opinion and reputation change, with the evidence citedRQ2
Daily summaries and verbatim captureKept permanently as ground truth beneath the interpreted layerRQ2 · RQ3 accuracy audit
Economy and gauge time seriesTreasury state, prices, trust, disposition and respect over timeContext for decisions
Admin-channel decisionsEvery gate decision to speak or stay silent, and whyRQ1
System Scale
~100
PostgreSQL tables, one shared database
3
Connected surfaces: Discord, web, game server
4
Hidden affective gauges per person
60
Day seasonal cycle
20–30
Active members in a population of hundreds
Model Strategy cost scales with time
A cheap, fast model handles high-frequency conversation. Careful models see only small, pre-filtered inputs on a slow clock: once a night, once an in-game month, or once per escalated claim. She speaks constantly and cheaply, and reasons rarely and well.
Live conversation: GPT-4o-mini · Summaries: Claude Haiku · Reflection, admin replies, pricing, judgment: Claude Sonnet
In-Game Presence custom mod suite
Custom Build 42 mods put the economy and events on the map: kiosk shops against the live wallet, horde events, treasure hunts, leaderboards, a bus network and an admin control panel. The bot and the mods talk through file-based IPC on the same host, which is crash-safe and executes each command exactly once.
Lua mods · JSON command files · consume-on-read · RCON for server control