feat: Complete LLM agent framework with fog-of-war, meeting flow, and prompt assembly

- Core engine: simulator, game mechanics, triggers (138 tests)
- Fog-of-war per-player state tracking
- Meeting flow: interrupt, discussion, voting, consolidation
- Prompt assembler with strategy injection tiers
- LLM client with fallbacks for models without JSON/system support
- Prompt templates: action, discussion, voting, reflection
- Full integration in main.py orchestrator
- Verified working with free OpenRouter models (Gemma)
This commit is contained in:
Antigravity
2026-02-01 00:00:34 -05:00
parent 9ec30034be
commit 071906df59
45 changed files with 8119 additions and 0 deletions
+230
View File
@@ -0,0 +1,230 @@
# Source Code Documentation
## Module Overview
### `src/engine/`
#### `simulator.py`
Discrete event simulator core.
```python
class Simulator:
def schedule_at(time: float, event_type: str, data: dict) -> Event
def schedule_in(delay: float, event_type: str, data: dict) -> Event
def step() -> Event | None # Process next event
def run_until(time: float) # Run until game time
def on(event_type: str, handler: Callable) # Register handler
```
#### `game.py`
Game engine with full mechanics.
```python
class GameEngine:
def add_player(id, name, color, role) -> Player
def queue_action(player_id, action_type, params) -> int
def resolve_actions() -> list[dict] # Priority-ordered resolution
def check_win_condition() -> str | None # "impostor", "crewmate", None
def get_impostor_context(player_id) -> dict # Fellow impostors
```
#### `triggers.py`
Event-driven trigger system.
```python
class TriggerRegistry:
def register_agent(agent_id)
def subscribe(agent_id, trigger_type)
def mute(agent_id, condition: TriggerCondition)
def should_fire(agent_id, trigger_type, time) -> bool
def get_agents_for_trigger(trigger_type, time) -> list[str]
```
#### `discussion.py`
Round-table discussion orchestrator.
```python
class DiscussionOrchestrator:
def calculate_priority(player_id, name, desire) -> int
def select_speaker(bids: dict) -> str | None
def add_message(player_id, name, message, target=None)
def advance_round(all_desires_low: bool) -> bool
def get_transcript() -> list[dict]
```
#### `types.py`
Core data structures.
```python
@dataclass
class Player: id, name, color, role, position, speed, is_alive, ...
class Position: room_id, edge_id, progress
class Body: id, player_id, player_name, position, time_of_death
class Event: time, event_type, data
class Role: CREWMATE, IMPOSTOR, GHOST
class GamePhase: LOBBY, PLAYING, DISCUSSION, VOTING, ENDED
```
#### `fog_of_war.py`
Per-player knowledge tracking (fog-of-war).
```python
class FogOfWarManager:
def register_player(player_id)
def update_vision(observer_id, visible_players, room_id, timestamp)
def witness_vent(observer_id, venter_id, ...) -> None
def witness_kill(observer_id, killer_id, victim_id, ...) -> None
def announce_death(player_id, via) # Broadcast to all
def get_player_game_state(player_id, full_state) -> dict # Filtered view
class PlayerKnowledge:
known_dead: set[str]
last_seen: dict[str, PlayerSighting]
witnessed_events: list[WitnessedEvent]
```
#### `available_actions.py`
Dynamic action generator per tick.
```python
class AvailableActionsGenerator:
def get_available_actions(player_id) -> dict
def to_prompt_context(player_id) -> dict # Compact for LLM
```
Returns: `movement`, `interactions`, `kills`, `sabotages` based on role/location.
#### `trigger_messages.py`
JSON schemas for trigger reasons.
```python
class TriggerMessageBuilder:
@staticmethod player_enters_fov(...) -> TriggerMessage
@staticmethod body_in_fov(...) -> TriggerMessage
@staticmethod vent_witnessed(...) -> TriggerMessage
@staticmethod kill_witnessed(...) -> TriggerMessage
@staticmethod death(...) -> TriggerMessage
# ... 15+ trigger types
```
#### `meeting_flow.py`
Full meeting lifecycle manager.
```python
class MeetingFlowManager:
def start_meeting(called_by, reason, body_location)
def submit_interrupt_note(player_id, note)
def submit_prep_thoughts(player_id, thoughts)
def add_message(speaker_id, speaker_name, message, target)
def submit_vote(player_id, vote)
def tally_votes() -> (ejected, details)
def end_meeting(ejected, was_impostor)
```
---
### `src/map/`
#### `graph.py`
Graph-based map representation.
```python
class GameMap:
def add_room(room: Room)
def add_edge(edge: Edge)
def get_neighbors(room_id) -> list[tuple[edge_id, room_id]]
def find_path(from_room, to_room) -> list[edge_id]
def path_distance(path: list[edge_id]) -> float
def load(path: str) -> GameMap # From JSON
def save(path: str) # To JSON
@dataclass
class Room: id, name, tasks: list[Task], vent: Vent | None
class Edge: id, room_a, room_b, distance
class Task: id, name, duration
class Vent: id, connects_to: list[str]
```
---
### `src/agents/`
#### `agent.py`
Stateless LLM agent wrapper.
```python
class Agent:
def get_action(game_context: dict, trigger: Trigger) -> dict
def get_discussion_response(context: dict, transcript: list) -> dict
def get_vote(context: dict, transcript: list) -> dict
def reflect(game_summary: dict) # Post-game learning
```
#### `scratchpads.py`
File-based persistent memory.
```python
class ScratchpadManager:
def read(name: str) -> dict
def write(name: str, content: dict)
def update(name: str, updates: dict) # Merge
def clear_game_specific() # Keep only learned.json
```
#### `prompt_assembler.py`
System + user prompt builder.
```python
class PromptAssembler:
def build_system_prompt(phase, game_settings, map_name, learned) -> str
def build_action_prompt(player_state, history, vision, actions, trigger) -> str
def build_discussion_prompt(player_state, transcript, meeting_scratchpad) -> str
def build_voting_prompt(player_state, transcript, vote_counts) -> str
def build_meeting_interrupt_prompt(player_state, interrupted_action) -> str
def build_consolidation_prompt(player_state, meeting_result, scratchpad) -> str
class PromptConfig:
model_name, persona, strategy_level, meta_level, is_impostor, fellow_impostors
```
---
### `src/llm/`
#### `client.py`
OpenRouter API wrapper.
```python
class LLMClient:
def chat(messages: list[dict], json_mode=True) -> dict
def chat_stream(messages: list[dict]) -> Iterator[str]
```
---
## Configuration
### `config/game_settings.yaml`
```yaml
num_impostors: 2
kill_cooldown: 25.0
vision_range: 10.0
impostor_vision_multiplier: 1.5
light_sabotage_vision_multiplier: 0.25
emergencies_per_player: 1
confirm_ejects: true
```
### `data/maps/skeld.json`
```json
{
"rooms": [
{"id": "cafeteria", "name": "Cafeteria", "tasks": [...], "vent": null},
...
],
"edges": [
{"id": "cafe_weapons", "room_a": "cafeteria", "room_b": "weapons", "distance": 5.0},
...
]
}
```
+187
View File
@@ -0,0 +1,187 @@
# The Glass Box League — Discussion Phase Design
## Overview
Discussion phase follows real Among Us logic. Tick-based with priority bidding, vote-when-ready with gentle pressure, full transcript visibility.
---
## Discussion Tick Flow
### Each Tick, All Agents Submit:
```json
{
"internal_thought": "Red is deflecting hard, classic impostor",
"desire_to_speak": 7,
"message": "Red, you haven't explained where you were during lights",
"target": "red",
"vote_action": null,
"scratchpad_updates": {
"meeting_scratch": "Red avoiding questions. Blue quiet."
}
}
```
- `desire_to_speak`: 0-10 urgency
- `target`: Optional, who they're addressing
- `vote_action`: `null` (keep discussing), `"player_id"` (vote), or `"skip"`
- If `vote_action` set: locked in, but can still speak (slightly lower priority)
---
## Priority Bidding
### Priority Calculation:
```
base = desire_to_speak
+ mention_boost (if name appeared in recent messages)
+ accusation_boost (if directly targeted)
+ silence_boost (if haven't spoken in N ticks)
+ random(1, 6)
- recent_speaker_penalty (if just spoke)
- voted_penalty (if already voted, small)
```
### Winner Selection:
- Highest priority speaks
- Their `message` goes to transcript
- Repeat next tick
### Forced Participation:
- `silence_boost` increases each tick of silence
- Eventually forces even quiet players to speak
- "You can't just be silent, that ruins the fun"
---
## Vote Mechanics
### Actions:
- `VOTE(player_id)` — commit vote
- `SKIP_VOTE()` — commit skip
- Stay silent — keep discussing
### End Condition:
- All living players have voted → tally & reveal
- Pressure nudge if discussion runs long: `"System: wrap it up"`
- No hard round limit (hoping for convergence)
### Tie:
- No eject (real Among Us logic)
### Vote Lock:
- Once submitted, cannot change (real Among Us logic)
---
## Transcript Handling
### Visibility:
- Full transcript, JSON formatted
- All agents see everything said so far
- Target: keep under ~25k tokens
### Format:
```json
{
"transcript": [
{"speaker": "red", "message": "I was in electrical doing wires", "t": 0},
{"speaker": "blue", "message": "I saw Red near the body", "t": 1},
{"speaker": "red", "target": "blue", "message": "That's a lie!", "t": 2}
]
}
```
### Pressure:
- If transcript gets too long, system nudges voting
- Agents feel urgency to wrap up
---
## Mention/Accusation Detection
- Simple string match on player color names
- If your name appears in message → `mention_boost`
- If message contains accusatory language toward you → `accusation_boost`
- Engine handles detection, agents don't need to flag
---
## Post-Vote Flow
### Reveal:
- Same as human would see in Among Us
- One by one reveal (dramatic)
- `confirm_ejects` setting: "Red was An Impostor" vs "Red was ejected"
### Reaction Tick:
- All agents get reaction tick after reveal
- Can update scratchpads, process result
### Consolidation Tick:
- Save important meeting info to main scratchpads
- Meeting scratchpad erased after this
---
## Ghost Chat
- Dead players can watch discussion
- Ghost chat separate from main discussion
- Useful for commentary/entertainment value
- Ghosts see full game state (omniscient)
- Cannot influence living players
---
## Impostor Discussion Strategy
### Prompt Reminders:
- "You are allowed to lie"
- "Construct alibis"
- "Deflect suspicion"
- "You know fellow impostors — don't expose them"
- "You know who you killed — avoid contradicting yourself"
### Strategy Injection (Toggleable):
- None: figure it out
- Basic: "Blend in, fake tasks, don't vent in front of others"
- Advanced: "Marinate teammate, frame third party, avoid double kills"
---
## Persona Persistence
### Storage:
Redis DB for each `{model}_{persona}` combo:
```json
{
"persona_id": "gpt4o_aggressive_leader",
"learned": {...},
"games_played": 42,
"win_rate": 0.65,
"impostor_win_rate": 0.70,
"crewmate_win_rate": 0.60,
"stats": {...}
}
```
### Separation:
- "GPT-4o as Aggressive Leader" ≠ "GPT-4o as Quiet Observer"
- Each persona builds own cross-game memory
- Learned strategies are persona-specific
---
## Strategy Injection Levels
Toggleable per persona:
| Level | Content |
|-------|---------|
| None | Just rules, figure it out |
| Basic | "Impostors vent, fake tasks, sabotage" |
| Intermediate | "Watch for inconsistent alibis, pair up" |
| Advanced | "Stack kills, marinate, third impostor framing" |
Different personas can have different injection levels to test learning vs. pre-trained strategies.
+229
View File
@@ -0,0 +1,229 @@
# The Glass Box League — Main Game Design
## Agent Architecture
### Identity & Persona
- Format: `"You are {model}. {PERSONA}. {context}"`
- Persona is **optional** and toggleable
- Same model can have multiple personas
- Goal: emergent behavior first, spice second
### Memory System
| Scratchpad | Persistence | Purpose |
|------------|-------------|---------|
| `plan.json` | Per-game | Current intentions, agent-controlled |
| `events.json` | Per-game | Curated game events worth remembering |
| `suspicions.json` | Per-game | Player reads, agent-maintained |
| `learned.json` | **Cross-game** | Core memory, enforced JSON schema |
| `meeting_scratch.json` | Per-meeting | Temp, erased after consolidation |
**JSON is god.** Enforced schema for structure, freeform for agent thoughts. JSON improves attention.
---
## Prompt Structure
### System Prompt (Core Memory)
1. Model identity
2. Persona (if set)
3. Game rules + current map + settings
4. Role briefing (crewmate/impostor)
5. Strategy tips (toggleable injection levels)
6. Meta-awareness (toggleable: subtle → direct → 4th wall)
7. Output format instructions
8. **Learned lessons** from `learned.json`
### User Prompt (Working Memory)
Order matters:
1. **You** — role, location, status, cooldowns
2. **Recent history** — accumulated vision from skipped ticks
3. **Vision** — current snapshot
4. **Available actions** — dynamic tool list
---
## Tool System
### Core Tools (Always Available)
- `MOVE(room_id | player_id)` — walk or follow
- `WAIT()`
- `INTERACT(object_id)` — tasks, panels, buttons, bodies, vents, cams
### Impostor Only
- `KILL(player_id)`
- `SABOTAGE(system_id)`
- `FAKE_TASK(task_id)`
### Trigger Management (Optional)
- `CONFIGURE_TRIGGERS(config)` — only if changing defaults
### Dynamic Available Actions
```json
{
"available_interactions": ["task_wires_cafe", "body_blue", "admin_table"],
"available_kills": ["green", "yellow"],
"available_sabotages": ["lights", "o2", "reactor", "comms"]
}
```
- Context-filtered by engine based on role + location
- Agent told "these are your actions this turn"
- Object IDs validated against engine to prevent glitches
---
## Trigger System
### Mandatory (Cannot Mute)
- `GAME_START`
- `DISCUSSION_START`
- `VOTE_START`
- `GAME_END`
- `SABOTAGE_CRITICAL`
### Standard (Mutable)
- `PLAYER_ENTERS_FOV`
- `PLAYER_EXITS_FOV`
- `BODY_IN_FOV`
- `OBJECT_IN_RANGE` (every interactable)
- `VENT_WITNESSED`
- `KILL_WITNESSED`
- `DESTINATION_REACHED`
- `TASK_COMPLETE`
- `SABOTAGE_START` / `SABOTAGE_END`
- `LIGHTS_OUT` (panic tick)
- `COOLDOWN_READY` (kill, emergency)
- `DEATH` (special message, transition to ghost)
### Optional (Opt-in)
- `EVERY_N_SECONDS` (configurable)
- `RANDOM_N_SECONDS` (RNG toggleable)
- `INTERSECTION` (hallway/room boundaries)
- `HALLWAY_WAYPOINT`
### Trigger Frequency
- **Impostors**: Every tick (more decision points)
- **Crewmates**: Event-driven + opt-in periodic
- **Ghosts**: Reduced frequency, longer intervals
### Trigger Message Schema
```json
{
"trigger_type": "VENT_WITNESSED",
"trigger_data": {
"player": "red",
"vent_location": "electrical",
"action": "entered",
"timestamp": 47.3
}
}
```
---
## Game State Schema
### Per-Tick Context
```json
{
"time": 47.3,
"phase": "PLAYING",
"you": {
"role": "impostor",
"location": "electrical",
"kill_cooldown": 0,
"tasks": [],
"emergencies_remaining": 1
},
"recent_history": [
{"t": 12.3, "vision": {"players": ["blue"], "location": "hallway_1"}}
],
"vision": {
"players_visible": [
{"id": "blue", "location": "electrical", "doing": "task"}
],
"objects_visible": ["vent_elec", "task_wires", "body_yellow"],
"exits": ["security", "cafeteria"]
},
"available_actions": {...}
}
```
### Fog of War
**Critical**: Each agent only knows what they've observed.
- Engine tracks per-player knowledge
- `known_deaths`: bodies seen or announced
- `known_locations`: last seen positions + timestamps
- `witnessed_events`: vents, kills, sus behavior
---
## Response Format
### Action Phase
```json
{
"internal_thought": "Blue just left, perfect time to kill Green",
"action": {"type": "KILL", "target": "green"},
"scratchpad_updates": {
"plan": "...",
"events": "...",
"suspicions": "..."
},
"trigger_config": {
"mute": [{"type": "INTERSECTION", "until": "REACHED_DESTINATION"}]
}
}
```
- `trigger_config` only if changing defaults
- `internal_thought` separate from action (for thinking models)
---
## Meeting Interrupt Flow
When report/emergency called:
1. **Interrupt note**: Agent leaves context ("this was what I was doing")
2. **Pre-meeting prep**: Agent reviews & prepares thoughts
3. **Meeting scratchpad**: Temporary, discussion-only
4. **Post-meeting consolidation**: Agent saves important info to main scratchpads
---
## Ghost Mode
- Omniscient view of entire game
- No access to other agents' thoughts
- Can do ghost tasks
- Reduced tick frequency (save tokens)
- Write to scratchpad, observe strategies
- No game state modifications
---
## Special Mechanics
### Lights Out
- Vision radius shrinks (0.25x multiplier)
- Triggers panic tick for all players
- Engine recalculates trajectories
- Fix triggers restoration tick
### Near-Death Edge Case
- Impostor queues kill, victim queues report same tick
- Report fires first (higher priority)
- Impostor gets: "Your kill was interrupted"
- Victim has no direct knowledge (must deduce from proximity)
---
## Configuration Philosophy
Everything toggleable:
- Persona injection
- Strategy tip levels
- Meta-awareness levels (subtle/direct/4th wall)
- Periodic tick frequency
- Random tick RNG
- Tool availability based on context
Goal: **Replicate human experience.** LLM should have same information and options as human player.
+76
View File
@@ -0,0 +1,76 @@
# Development Guide
## Running Tests
```bash
# All tests
python3 -m unittest discover -v tests/
# Specific test file
python3 -m unittest tests/test_game.py -v
# Specific test
python3 -m unittest tests.test_game.TestKill.test_successful_kill
```
## Adding New Triggers
1. Add trigger type to `src/engine/triggers.py`:
```python
class TriggerType(Enum):
NEW_TRIGGER = auto()
```
2. Add to appropriate category:
```python
STANDARD_TRIGGERS = {..., TriggerType.NEW_TRIGGER}
```
3. Fire trigger in game engine:
```python
self._fire_trigger(TriggerType.NEW_TRIGGER, agent_id, {"data": "value"})
```
## Adding New Actions
1. Add action handler in `src/engine/game.py`:
```python
def _handle_new_action(self, player_id: str, params: dict) -> dict:
# Validate + execute
return {"success": True, "action": "NEW_ACTION"}
```
2. Add to `_execute_action` switch.
3. Add to priority order in `resolve_actions`.
## Adding New Maps
Create JSON in `data/maps/`:
```json
{
"rooms": [
{
"id": "room_id",
"name": "Display Name",
"tasks": [{"id": "task_id", "name": "Task Name", "duration": 3.0}],
"vent": {"id": "vent_id", "connects_to": ["other_vent"]}
}
],
"edges": [
{"id": "edge_id", "room_a": "room1", "room_b": "room2", "distance": 5.0}
]
}
```
## Environment Variables
| Variable | Purpose |
|----------|---------|
| `OPENROUTER_API_KEY` | LLM API authentication |
## Code Style
- Python 3.10+ features (type hints, dataclasses)
- JSON for all config (YAML optional, falls back to JSON)
- Tests mirror source structure (`src/engine/game.py` → `tests/test_game.py`)
+24
View File
@@ -0,0 +1,24 @@
# Documentation Index
| Document | Description |
|----------|-------------|
| [README.md](../README.md) | Project overview & quick start |
| [design_main_game.md](design_main_game.md) | Main game loop design |
| [design_discussion.md](design_discussion.md) | Discussion phase design |
| [api.md](api.md) | Source code API reference |
| [development.md](development.md) | Developer guide |
| [openrouter_api.md](openrouter_api.md) | LLM integration notes |
## Design Documents
Captured from design Q&A sessions:
- **Main Game** — Agents, triggers, tools, fog-of-war, scratchpads
- **Discussion** — Priority bidding, voting, ghost chat, personas
## Quick Links
- Tests: `python3 -m unittest discover -v tests/`
- Map: `data/maps/skeld.json`
- Config: `config/game_settings.yaml`
- Prompts: `config/prompts/`
+118
View File
@@ -0,0 +1,118 @@
# OpenRouter API Reference
Quick reference for The Glass Box League LLM integration.
## Base URL
```
https://openrouter.ai/api/v1
```
## Authentication
```
Authorization: Bearer $OPENROUTER_API_KEY
```
## Chat Completions Endpoint
**POST** `/chat/completions`
### Request Body
```json
{
"model": "google/gemini-2.0-flash-lite-preview-02-05:free",
"messages": [
{"role": "system", "content": "You are an Among Us player."},
{"role": "user", "content": "What do you do?"}
],
"temperature": 0.7,
"max_tokens": 1024,
"top_p": 0.9,
"frequency_penalty": 0.0,
"presence_penalty": 0.0,
"stream": false,
"response_format": {"type": "json_object"}
}
```
### Response
```json
{
"id": "gen-...",
"model": "google/gemini-2.0-flash-lite-preview-02-05:free",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "{\"action\": \"move\", \"target\": \"electrical\"}"
},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 100,
"completion_tokens": 50,
"total_tokens": 150
}
}
```
## Key Parameters
| Parameter | Type | Description |
|-----------|------|-------------|
| `model` | string | Model ID (see available_models.json) |
| `messages` | array | Conversation history |
| `temperature` | float | Randomness (0.0-2.0) |
| `max_tokens` | int | Max response length |
| `top_p` | float | Nucleus sampling (0.0-1.0) |
| `stream` | bool | Enable streaming |
| `response_format` | object | Force JSON output |
| `seed` | int | For deterministic output |
## Structured Output (JSON Mode)
```json
{
"response_format": {
"type": "json_object"
}
}
```
## Headers
```
Content-Type: application/json
Authorization: Bearer $OPENROUTER_API_KEY
HTTP-Referer: https://your-app.com (optional, for rankings)
X-Title: Glass Box League (optional, for rankings)
```
## Python Example
```python
import os
import requests
def chat(messages, model="google/gemini-2.0-flash-lite-preview-02-05:free"):
response = requests.post(
"https://openrouter.ai/api/v1/chat/completions",
headers={
"Authorization": f"Bearer {os.getenv('OPENROUTER_API_KEY')}",
"Content-Type": "application/json"
},
json={
"model": model,
"messages": messages,
"temperature": 0.7,
"response_format": {"type": "json_object"}
}
)
return response.json()["choices"][0]["message"]["content"]
```
## Free Models
See `available_models.json` for current free tier models.
Run `python fetch_models.py` to refresh the list.
## Rate Limits
- Free tier: varies by model
- Check response headers for remaining quota