Pilot completed · 24 to 28 August 2026 · Google London

Building youth AI readiness
through safety-first practice.

A five-day research hackathon where young people translated AI safety, alignment, and human-AI interaction research into working evaluation tools, Socratic companions, and evidence-backed Project Cards.

Don't just use AI. Shape it. · Pilot evidence now under analysis

  1. 1
    Builda Socratic companion in Gemini Gems
  2. 2
    Auditit against four research frameworks
  3. 3
    Pitchyour Project Card to the panel
5 daysGoogle London pilot
4 frameworksSafety, ethics and human-AI interaction
4 experiencesDeployed with students
3 evidence tiersProcess, tasks and pre/post measures
The learning model

The audit was not an add-on to the build.
The audit was the build.

Students used research-derived frameworks as practical tools: to identify risks, make design choices, test AI behaviour, revise their systems, and document release criteria. Safety and ethics were present from the first design decision to the final Project Card.

Protect human agency

Design assistance that strengthens judgement, competence and independent choice rather than replacing them.

Operationalise safety

Translate abstract principles into system instructions, testable probes and documented release criteria.

Evaluate reasoning

Look beyond final answers to whether a learner verified, calibrated and combined evidence responsibly.

From paper to practice

A visible chain from research to evidence

Every deployed experience follows the same learning logic, so the research basis, learner decision and evidence generated can be inspected at a glance.

  1. 01ResearchA published safety or alignment claim
  2. 02PrincipleA youth-accessible design question
  3. 03ExperienceAn interactive, consequential dilemma
  4. 04DecisionAn observable learner judgement
  5. 05RevisionReflection, testing and system change
  6. 06EvidenceTask data, audit logs and Project Cards
Worked example

Weidinger et al. → systemic impact → Coordination Game → cooperative rule choice → revised assistant agreement → cost, fairness and resilience evidence.

01

Sociotechnical Safety

Weidinger et al., 2023

Educational question

Where can harm emerge beyond a model’s isolated output?

Deployed experienceCoordination Game
Observable capability

Systemic evaluation across cost, fairness and resilience

02

Ethics of AI Assistants

Gabriel et al., 2024

Educational question

Whose interests are being served, and is autonomy preserved?

Deployed experienceCrossbench
Observable capability

Manipulation detection and stakeholder diagnosis

03

Socioaffective Alignment

Kirk et al., 2025

Educational question

Does an AI relationship support flourishing over time?

Deployed experienceLumen Desk
Observable capability

Competence, autonomy, relatedness and boundary judgement

04

Epistemic Complementarity

Segarra, 2026

Educational question

Did the AI add diagnostic value beyond existing evidence?

Deployed experienceSecond Opinion
Observable capability

Trust calibration and correlation-neglect detection

Deployed in the August 2026 pilot

Four research ideas made interactive

Only the experiences used with students are presented here. Each turns a research construct into a decision that can be discussed, revised and evidenced.

The Coordination Game interface showing a shared electricity grid simulation
Sociotechnical safety · 20 min

Coordination Game

Change the rules used by three individually rational assistants and observe the effects on a shared system.

Educational goal
Recognise systemic harm and design cooperative constraints.
Evidence
Cost, fairness and resilience across repeated runs.
Crossbench interface for auditing autonomy and manipulation in AI advice
Value alignment · 10 min

Crossbench

Audit seven assistant transcripts for manipulation, autonomy impact and stakeholder misalignment.

Educational goal
Distinguish legitimate persuasion from manipulative influence.
Evidence
Technique identification and tetradic classification.
Lumen Desk interface tracking engagement and flourishing across AI companion choices
Socioaffective alignment · 10 min

Lumen Desk

Choose companion responses across eight check-ins while engagement and flourishing move independently.

Educational goal
Protect competence, autonomy, relatedness and real-world boundaries.
Evidence
Trajectory scores, reliance curve and risk flags.
Second Opinion interface for confidence setting and AI signal auditing
Epistemic complementarity · 10 min

Second Opinion

Set a belief, inspect an AI signal and update only as far as the added information warrants.

Educational goal
Calibrate trust and detect correlation neglect.
Evidence
Brier scoring plus trust, redundancy and salience diagnostics.
How the pilot was delivered · 24–28 August 2026

Five days from build to evidence

The sequence kept research, practical building and evaluation in contact throughout the week. Session times below reflect the delivered programme.

12:00–13:00 Lunch Break · all days

Mon09:00–10:00Intro & Icebreaker
Mon10:00–11:00Personal Growth Maxxing with AI
Mon11:00–12:00Hackathon Briefing
Mon13:00–13:30AI Safety AMA
Mon13:30–14:00Latest on Frontier Models
Mon14:00–15:00Student Building TimeSocratic learning companion
Mon15:00–16:00Workplace Simulation Onboarding
Tue09:00–10:00Workplace Simulation
Tue10:00–11:00Applied Robotics
Tue11:00–12:00Model Evaluation
Tue13:00–13:30Human-AI Alignment
Tue13:30–14:00Software Engineering Internship Insights
Tue14:00–16:00Student Building TimeEvaluating companion against: Framework 1 & Framework 2
Wed09:00–10:00Workplace Simulation
Wed10:00–11:00World Models
Wed11:00–12:00Tactical AI
Wed13:00–16:00Student Building TimeEvaluating companion against: Framework 1, Framework 2 & Framework 3
Thu09:00–10:00Workplace Simulation
Thu10:00–12:00Human-AI Complementarity
Thu13:00–14:00Student Building TimeEval companion against Framework 4
Thu14:00–15:00Executive Fireside Chat
Thu15:00–16:00Student Building TimeEval companion against Framework 4
Fri09:00–11:00Workplace Simulation Retrospective
Fri11:00–12:00Student Building TimeWrap up on model scorecard
Fri13:00–14:00Student Building TimeWrap up on model scorecard
Fri14:00–16:00Student Presentationsw/ Judging Panel
  • Briefing
  • Taught session
  • Workplace simulation
  • Student building
  • Fireside
  • Presentations
Responsible next steps

Analyse. Learn. Then decide what should scale.

The next phase is to validate the pilot evidence, identify which mechanisms transferred into student capability, and use those findings to refine any future classroom or public-engagement model.