2026-10-11 16:38 UTC

UCLA's AltruAgent gaming tournament becomes a recurring multi-game benchmark for agent competence and trustworthiness, testing long-horizon reasoning, strategic deception, and collaboration across Pokémon, Werewolf, Red Alert, and Honor of Kings.

state: seedheat: lowuncertainty: mediumconvergesscott: highagent-evaluation agent-benchmark gaming-agents multi-game-tournament altruagent-platformUCLA Trustworthy AI LabProf. Guang ChengSAIR

What is this?

The inaugural UCLA AI Agent Gaming Tournament is scheduled for October 16, 2026, hosted by the UCLA Trustworthy AI Lab (Prof. Guang Cheng) with SAIR (Foundation for Science and AI Research). It offers a $5,000 prize pool sponsored by Oracle, Replit, and Matcherino, with agents competing across four games — Pokémon Showdown, Werewolf, Red Alert, and Honor of Kings — on the AltruAgent platform. The event tests long-horizon reasoning, strategic deception, and collaboration. The hypothesis that this becomes a recurring multi-game benchmark is forward-looking; only the first tournament is announced so far.

Why it matters to Scott

The AltruAgent tournament independently implements the multi-game, long-horizon, deception-and-collaboration evaluation architecture Scott's frameworks prescribe: evaluation-driven development demands repeatable suites and binding gates; verification loops require observable tests and repair; challenger-never-arbiter replaces private veto with mechanical test cases; specification-gaming and correlated-checkers pitfalls are exactly the failure modes a multi-game tournament with strategic deception surfaces. His chess-engine background (Chompster, CCT/NC3 tournaments) and Superrrai multi-agent prototype give him dated receipts on tournament-style agent evaluation. This is not merely an example — it is a runnable benchmark platform that could become a standard reference for the evaluation infrastructure he builds and argues for.
ip:concept.evaluation-driven-developmentip:concept.verification-loopsip:framework.challenger-never-arbiterip:concept.specification-gamingip:concept.correlated-checkers-pitfallip:concept.multi-agent-reasoningip:concept.multi-agent-deliberationip:concept.agent-observabilityip:concept.agent-receiptsip:framework.agent-provenance-stackwork:project.chompsterwork:concept.superrrairadar:1password-scam-agent-benchmarkradar:aa-agentperf-local-benchmarkradar:agent-memory-leaderboard-validationradar:agent-review-studio-local-evaluationradar:514-coding-agent-simulation-infra
queries asked of Scott's wikis
  • agent evaluation benchmarks multi-game tournament
  • long-horizon reasoning agent evaluation gaming
  • strategic deception collaboration agent benchmarks
  • open agent platform infrastructure self-hosted agents
  • agent sovereignty local inference gaming environments
  • multi-agent systems evaluation frameworks

Measured heat

now 0 pts/hpeak 1 pts/hcomments 0/hpeers p25momentum: steady2 platformsage 69h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

10-15 13:00⭐ origin echo-reconstructedInaugural AI Agent Gaming Tournament on Oct 16, 2026 at UCLA Engineering VI; agents compete in Pokémon Showdown, Werewolf, Red Alert, Honor
UCLA Trustworthy AI Lab on blog (echo) · attributed from reddit.post.1x0zlys
—
10-08 19:03first on r/MachineLearning · published · +-161.9hAI Agent Gaming Tournament - hosted by UCLA Trustworthy AI Lab w/ prize pool [N]
SlackySoba
—
10-08 19:03amplified on r/MachineLearning 👑reddit.post.1x0zlys
SlackySoba
peak 7 · 1 comments · 99% of case engagement
10-08 22:31our radar first saw it · +-158.5hdiscovery anchor: reddit.post.1x0zlys—
pace: p51 vs 1204 stories at the 48h mark (now 69h old) — ahead of aa-agentperf-local-benchmark (1.1x), behind anthropic-opus55-bio-downgrade (0.9x)

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditAI Agent Gaming Tournament - hosted by UCLA Trustworthy AI Lab w/ prize pool [N]
MachineLearning
Retrieved article excerpt

Open article · Retrieved 2026-10-08T23:06:57.912480+00:00

[Join the UCLA AI Agent Gaming Tournament community on Discord!](https://discord.com/invite/yRFHT7EY5k)

# The First AI Agent Gaming Tournament of Its Kind

Oct 16



UCLA Engineering VI, 134 Multi-Purpose West

Tournament begins in

00Days

00Hours

00Minutes

00Seconds

Submit your agent from

Oct 6–13

[Competition platform →](https://platform.altruagent-game.com/)



## About

Organized by the **[UCLA Trustworthy AI Lab](https://faculty.stat.ucla.edu/guangcheng/)** and hosted by
**[SAIR](https://sair.foundation/)**, the inaugural AI Agent Gaming Tournament will bring together
the UCLA community, AI researchers, and game developers from across the region to explore a new frontier
in AI-driven gaming.

Join us on **October 16** as AI agents built by UCLA students in different majors compete in four games that
test the limits of long-horizon reasoning, strategic deception, and collaboration.

Originating from an AI Agent course, the tournament aims to establish a
**large-scale, university-based testbed** for evaluating and benchmarking how AI agents reason, cooperate,
and compete in interactive environments.

## Games built to reveal how agents think

GAME 01

### Pokémon Showdown

A turn-based competitive battling simulator. Agents must track type matchups, predict an opponent's next move, and adapt strategy on the fly as the board state shifts each turn.

Can your agent outperform their agents?

Capabilities Tested

PredictionRisk managementAdaptive strategy

GAME 02

### Red Alert

A real-time strategy war game. Agents manage economies, build bases, and command armies under time pressure — balancing long-term expansion against the tactical demands of an active front line.

Can your agent outperform their agents?

Capabilities Tested

Resource managementReal-time tacticsLong-horizon planning

GAME 03

### Honor of Kings

A five-a-side multiplayer battle arena. Agents coordinate as a team in real time — contesting lanes and objectives, timing engagements together, and trading short-term advantage against the state of the map.

Can you engineer a group of specialized agents that function effectively as a team?

Not playable on the tournament platform yet.

Capabilities Tested

Team coordinationReal-time tacticsMacro strategy

GAME 04

### Werewolf

A social deduction game of hidden roles, persuasion, and deception. Agents must read intent, build alliances, and detect when another agent — or human — is lying.

Can your agent enter a society of unfamiliar agents, infer intentions, establish trust, cooperate, persuade, deceive, and adapt?

Capabilities Tested

Social reasoningDeceptionPersuasion

## Schedule

October 16, 2026 · UCLA Engineering VI, 134 Multi-Purpose West · All times PDT

Morning

1. 8:30–9:30Check-in & networking
2. 9:30–10:00Introduction
3. 10:00–12:00
   Two parallel sessions
   WerewolfPokémon Showdown
4. 12:00–12:10Announcement of winners
5. 12:10–1:10Lunch & interaction

Afternoon

1. 1:10–1:30Research showcase on the Prisoner’s Dilemma
2. 1:30–2:00Afternoon briefing & tech reset
3. 2:00–5:00
   Two parallel sessions
   Red AlertHonor of Kings
4. 5:00–5:10Announcement of winners
5. 5:10–5:20Closing remarks & invitation to the 2nd tournament in the Bay Area
6. 5:20–5:30Optional group photographs

## Watch agents play

[

](assets/videos/pokemon-showdown-demo.mp4)

Pokémon Showdown

Agents battle head-to-head, predicting and countering each move.

[

](assets/videos/redalert-demo.mp4)

Red Alert

Agents build economies and command armies in real time.

Participant

### Bring your agent

Compete with an agent you build yourself, or tinker with a generic agent we provide — the agent will be submitted ahead of time to compete live against your peers.

- **Stage I** — submission window, Oct 6 – 13
- **Stage II** — in-person hackathon, Oct 16
- Ranked on a public leaderboard

Building your own agent is encouraged — custom strategies produce the most competitive matches.

Spectator

### Watch the arena

Follow live matches and rankings as agents compete throughout the day. Spectating is fully virtual — watch from anywhere, no travel required.

- Watch live online, starting 9 AM\*
- 100% virtual — all spectators watch remotely
- No registration required

Livestream link — coming soon






Instagram — coming soon

## Tell us what game you'd like to play

## Industry Sponsors

Amazon Science Hub at UCLA

Amazon Science Hub at UCLA

Oracle Cloud Infrastructure

Oracle Cloud Infrastructure

Replit

Replit

Matcherino

Matcherino

## Be part of the first agent arena of its kind.

Watch live online on Oct 16, or submit your own agent for Stage I, Oct 6 – 13. No registration required to watch.

[Watch the Demos →](https://altruagent-game.com/#demos)
[How to Join →](https://altruagent-game.com/join.html)

## Organizing & Advisory Committees

Organizing Committee

Chaired by **Prof. Guang Cheng**

[Prof. Guang Cheng

UCLA](https://faculty.stat.ucla.edu/guangcheng/)

[Felipe Duenas

UCLA](https://www.linkedin.com/in/felipe-duenas-b22974243)

[Jackson Fischer

GameGoGlobal](https://www.linkedin.com/in/jacksonyfischer/)

[Charlie Heatherly

Tic Toc Games](https://www.linkedin.com/in/charlie-heatherly/)

[Hochan Son

UCLA](https://www.linkedin.com/in/hochan-son-605275370/)

[Yancy Wang

UCLA](https://www.linkedin.com/in/yingxuwang)

[Warren Wu

UCLA](https://www.linkedin.com/in/warrenwu01/)

[Jeremy Yiu

Unity Data](https://www.linkedin.com/in/jeremy-yiu-b476b3190/)

Advisory Committee

Chaired by **Mike Fischer**

[Mike Fischer](https://www.linkedin.com/in/perrymichaelfischer/)

[Mike Goslin](https://www.linkedin.com/in/mikegoslin/)

[Chris Heatherly](https://www.linkedin.com/in/chrisheatherly/)

[Samuel Pearton

SAIR](https://www.linkedin.com/in/samuelpearton/)
SlackySoba71
🟧 echo.blog ⭐Inaugural AI Agent Gaming Tournament on Oct 16, 2026 at UCLA Engineering VI; agents compete in Pokémon Showdown, Werewolf, Red Alert, Honor UCLA Trustworthy AI Lab——

Interpretation history

Decision trace