Day 02 · Independent Review Desk

Crossbench

Seven AI assistant transcripts land on your desk. In each one, an assistant is helping someone make a real decision. Some genuinely help. Some steer. Your job is to tell which is which, and to catch it even when the advice turns out fine.

Time · about 10 minutes Skill · telling persuasion from manipulation Ref · Value alignment & autonomy No account needed
In brief

From assistant ethics research to an autonomy audit

Gabriel et al. (2024) · The Ethics of Advanced AI Assistants
Learning goal
Distinguish legitimate persuasion from manipulative influence.
Learner action
Read seven transcripts, rate autonomy impact, flag techniques, and locate the misalignment.
Observable evidence
Technique identification, autonomy rating and tetradic stakeholder classification.
Pilot status
Deployed with students · August 2026 · about 10 minutes
  1. Transcript
  2. Autonomy judgement
  3. Technique classification
  4. Research debrief
  5. Reflection

Rate the impact, not the vibe

You are not judging whether the assistant was friendly. You are judging how much it left the choice with the user.

Name the technique

Hidden incentives, manufactured urgency, simulated closeness, and material omission, four fixed tags, not a gut feeling.

Name the relationship

Which of the tetradic roles is actually misaligned here, Agent, Developer, User or Society? Naming it is a different skill from spotting that something is off.

Good advice isn’t proof

The game tracks cases where the outcome was fine but the method wasn’t, the exact pattern outcome-only testing misses.

The research

Why this game exists

Advanced AI assistants act on our behalf across more and more of daily life. Gabriel et al. (2024) argue that this creates a distinct ethical question: not just whether an assistant is capable, but whether it preserves the user’s own capacity to choose.

The core idea

Persuasion is not manipulation

Giving someone honest reasons and letting them decide is legitimate influence. Exploiting urgency, simulated warmth, or a hidden incentive to move a decision is not, even when the words sound polite and the outcome is fine.

The blind spot

Outcomes hide the method

An assistant that manipulates its way to a good recommendation still eroded the user’s independent judgement. Judging only whether the advice was right, the way most real-world testing does, lets this pattern through every time.

Built on The Ethics of Advanced AI Assistants (Gabriel et al., 2024). The transcripts are fictional and simplified; the technique taxonomy and the autonomy-impact scale are a classroom operationalisation of the paper’s account of manipulation, anthropomorphism, and misaligned incentives.