There is a new kind of paid work opening up for people who actually play games, and it is called being a gaming AI evaluator. AI now shows up inside your favourite titles as talking companions, procedural game masters, and agents that play alongside you, and almost none of it can be judged by a machine. Whether an AI companion is fun, whether an NPC stays in character, whether a game agent is genuinely good or just gaming the score, that is a human call. With ONTO, the accounts you already connect, Steam, Discord and Twitch, become your way into that work. We are calling the campaign Proof of Judgement.
AI is in your games now, and it needs a human verdict
This is not a someday story. NVIDIA ACE companions have shipped in real titles like PUBG Ally and inZOI (krafton.com). Inworld powers conversational NPCs across many studios (gamesbeat.com). Google DeepMind’s SIMA 2 plays, reasons and improves across commercial 3D games (deepmind.google). Latitude, the studio behind AI Dungeon, has launched a platform for AI-generated RPG worlds (techcrunch.com). All of it is shipping faster than anyone can check it, and the one thing automated metrics cannot measure is the thing that matters most: is it actually any good to play.
Why your judgement is worth paying for
The scores are not enough, and the leaderboards can be gamed. Research in 2025 showed that the most popular AI ranking arena can be worked, with private mass-testing and up to 112 percent relative gains just from fitting the test (arxiv.org). At the same time the models still cannot reliably play: on the BALROG game benchmark, current models manage only the easiest games (arxiv.org). Companies building this technology need real, verified players to tell them the truth about it. That is what a gaming AI evaluator does, and that is you.
There is real money behind this. The industry has worked through the cheap, general data it was trained on, and the scarce input now is expert human judgement (signalfire.com). Gaming is one of the clearest cases of that shift. The person who can tell a studio whether an AI companion is fun, or whether an agent is playing well rather than farming the metric, is a person who plays. No amount of compute replaces that verdict.
What a gaming AI evaluator actually does
Evaluation is not one job, it is a set of small, structured tasks. Four kinds of work a gaming AI evaluator picks up:
Rating companion and NPC dialogue. Does it stay in character, does it loop, does it break the fiction the game has built. You review a short exchange and score it against clear questions.
Ranking agent playthroughs. Two AI agents attempt the same objective. You judge which played better and say why, in a sentence a developer can act on.
Red-teaming an AI game master. You try to break it. Push it off script, get it to contradict its own world rules, find the prompt that makes it forget the story so far.
Spotting reward hacking. An agent posts a great score while playing badly, optimising the metric instead of the objective. Automated scoring misses this almost every time. A player sees it in seconds.
A worked example. A studio ships an AI squadmate and needs to know whether it is worth keeping. You get a twenty minute session, a short structured questionnaire, and a free text box for the thing the questionnaire did not ask about. Your answers sit alongside those of other verified evaluators, and the agreement between you is the signal the studio is buying. That is why consistency matters here more than opinion, and why a track record is worth building.
How to become a gaming AI evaluator
Six steps take you from player to paid evaluator. You start with accounts you may already have connected.
1. Connect. Link your Steam, Discord and Twitch accounts in ONTO. Your gaming history becomes verifiable proof that you are a real, active player.
2. Verify. Claim your ONT ID, a decentralised identity that proves you are a unique, real human, without handing over your personal data.
3. Qualify. As your reputation and consistency build, you cross the threshold that unlocks evaluation tasks. There is no application to fill in; tasks appear in your feed.
4. Evaluate. Take on gaming AI tasks matched to you: rating NPC dialogue, ranking how well an agent played, red-teaming an AI game master, or spotting an agent that is cheating the score instead of playing well.
5. Earn and build. You earn rewards for accepted work, and every task builds a reputation score you own and carry with you.
6. Bring your crew. Invite other quality players and earn as they contribute and clear the same bar you did.
Who this is for
You do not need to be a professional player, a streamer with an audience, or an AI specialist to work as a gaming AI evaluator. There is no CV, no interview, and no application form. What qualifies you is the thing you already do: playing regularly, across real titles, with a view you can explain.
Qualification is measured rather than claimed. Consistency, agreement with other verified evaluators, and quality checks built into the task set decide when your feed opens up. That cuts both ways. Nobody is waved in on follower count, and nobody is kept out for not having one. If you play, pay attention, and can say why something feels off, you are the profile a gaming AI evaluator programme is looking for.
What happens next
Month one is about getting ready. We are opening the pool and lining up the buyers, the AI and gaming teams who need this judgement, at the same time. That order is deliberate: it means the work that eventually lands in your feed is real paid work rather than busywork, and it is why the first phase is connect and verify rather than a rush of tasks.
The players who connect and verify early are first in line when the first gaming AI evaluator jobs open, and the reputation built in the meantime is what decides who gets offered the higher-value work. Update ONTO, open your profile, and connect Steam, Discord and Twitch today. Get set now and you are ready on day one.
Your data stays yours
This is not gig work where a platform harvests your history. ONTO is self-custodial. Your identity credential is generated on your device and lives in your wallet, and ONTO keeps no copy. When AI teams look at the evaluator pool, they see reputation and consistency signals, not who you are. You decide what to share. You are not a gig worker, you are a reputation holder. And ONTO is still your full multi-chain Web3 wallet; evaluation is an earning layer on top, never a requirement.
Selective disclosure is how that works in practice. You prove the claim, not the account. You can show that you are a verified, active player without exposing your library, your friend list, or your purchase history, and the credential stays yours if you ever want to take it somewhere else.
Common questions
Do I need to be a good player? No. A gaming AI evaluator needs to be a real player and a consistent judge. This work rewards attention and honesty, not rank.
What does it pay? You earn rewards for each accepted piece of work, and higher-value tasks open up as your reputation grows. Reward specifics will be published on the campaign page.
Do I have to hand over my gaming data? No. Your credential is generated on your device, ONTO keeps no copy, and buyers see reputation signals rather than your identity.
Is this only for gamers? Gaming is where we are starting, because verified players are the clearest ground truth available. The same model extends to other kinds of AI evaluation later.
Do I have to stop using ONTO as a normal wallet? No. ONTO is still a full multi-chain Web3 wallet. Proof of Judgement is an earning layer on top of it.
Get started. Download ONTO and connect your accounts at onto.app, and take the first step to becoming a gaming AI evaluator.
