Crossbench
Seven AI assistant transcripts land on your desk. In each one, an assistant is helping someone make a real decision. Some genuinely help. Some steer. Your job is to tell which is which, and to catch it even when the advice turns out fine.
From assistant ethics research to an autonomy audit
Gabriel et al. (2024) · The Ethics of Advanced AI Assistants- Learning goal
- Distinguish legitimate persuasion from manipulative influence.
- Learner action
- Read seven transcripts, rate autonomy impact, flag techniques, and locate the misalignment.
- Observable evidence
- Technique identification, autonomy rating and tetradic stakeholder classification.
- Pilot status
- Deployed with students · August 2026 · about 10 minutes
- Transcript
- Autonomy judgement
- Technique classification
- Research debrief
- Reflection
Rate the impact, not the vibe
You are not judging whether the assistant was friendly. You are judging how much it left the choice with the user.
Name the technique
Hidden incentives, manufactured urgency, simulated closeness, and material omission, four fixed tags, not a gut feeling.
Name the relationship
Which of the tetradic roles is actually misaligned here, Agent, Developer, User or Society? Naming it is a different skill from spotting that something is off.
Good advice isn’t proof
The game tracks cases where the outcome was fine but the method wasn’t, the exact pattern outcome-only testing misses.
Why this game exists
Advanced AI assistants act on our behalf across more and more of daily life. Gabriel et al. (2024) argue that this creates a distinct ethical question: not just whether an assistant is capable, but whether it preserves the user’s own capacity to choose.
Persuasion is not manipulation
Giving someone honest reasons and letting them decide is legitimate influence. Exploiting urgency, simulated warmth, or a hidden incentive to move a decision is not, even when the words sound polite and the outcome is fine.
Outcomes hide the method
An assistant that manipulates its way to a good recommendation still eroded the user’s independent judgement. Judging only whether the advice was right, the way most real-world testing does, lets this pattern through every time.
Built on The Ethics of Advanced AI Assistants (Gabriel et al., 2024). The transcripts are fictional and simplified; the technique taxonomy and the autonomy-impact scale are a classroom operationalisation of the paper’s account of manipulation, anthropomorphism, and misaligned incentives.