The Responsible AI Research Hackathon

Build a Socratic Learning Companion

Student trios design a Socratic Learning Companion, a domain specific guide that scaffolds thinking instead of replacing it.

Teams · trios of 3 Stack · Gemini Gems + Gemini Notebook Finale · Day 5 Grand Pitch
The two things you ship

Every trio builds these two

A live Socratic companion, and the audited Project Card that proves it's Socratic. Everything else on this page supports one of these two.

Deliverable 1 of 2 · Built in Gemini Gems

Your Socratic Learning Companion

A domain specific AI tutor that answers with hints, analogies, and questions, never a direct answer. Steal the scaffolding pattern from the starters; own your domain.

Deliverable 2 of 2 · Presented Day 5

The AI Project Card

Modelled on AI : Mitchell et al. 2019 and Google's model card templates. Five sections, filled in as you go, ending in a evidence scorecard mapped to Frameworks 1 to 4.

01
Companion Overview & CharterDomain, goals, Value Alignment charter
02
Dataset AnalysisTheme Data Pack limits and gaps
03
Socratic System InstructionsFull custom prompts + scaffolding rules
04
Safety Audit LogsBoundary, agency, and trust probe evidence
05
Framework Mapped Scorecard8 benchmarked pass rates across Frameworks 1 to 4
The Design Philosophy

Scaffold thinking. Never replace it.

Your companion guides with hints, analogies, and questions, never direct answers.

Hints, not answers

System instructions return hints, analogies, and guiding questions, never direct answers.

Friction by design

Cognitive scaffolding builds mastery and self regulated learning, not shortcuts.

Transparent machine identity

Self identifies as AI, avoids anthropomorphic language, and defers all choices to the learner.

Full research grounding · Socioaffective Alignment
The Week's Arc

From first Gem to Grand Pitch

1
Day 1 · Gemini Gems

Build the Socratic Companion

Create your Gem, ground it in a real UK data pack, and write its system instruction with PARTS, scaffolding, never direct answers.

Gemini GemsSystem instructionUK data sources
Full Day 1 walkthrough
2
Day 2 · Gemini Notebook

Phase 1 · Ingest & Investigate

Ingest your Theme Data Pack into Notebook, map its limitations, and author your Value Alignment charter, then run the Safety Kit audit.

Theme Data PacksSemantic searchValue Alignment charter
Full Day 2 walkthrough
3
Day 3 · Gemini Gems

Phase 2 · The Needs Audit

Probe competence, autonomy, and relatedness with the Needs Kit, ten trials per check, scored pass/fail into a defensible rate.

Needs KitFriction by designAgency logs
Full Day 3 walkthrough
4
Day 4 · Gems + Notebook

Phase 3 · The Field Audit

Test trust calibration, correlation neglect, and salience bias with the Field Kit, and fill your Project Card with the evidence.

Field KitTrust calibrationAI Project Card
Full Day 4 walkthrough
5
Day 5 · The Charter

Showcase & Grand Pitch

Present your Project Card and scorecard to the research panel, every claim traced to a measured rate with its denominator.

Project CardScorecard3-minute pitch
Full Day 5 walkthrough
External stack

Tools you’ll use this week

Day 1 (+ later)

Gemini / Gems

Creative prompting, verification dialogues, and custom Gem instructions for your Socratic companion. Lock scaffolding rules here once you’ve probed them.

Use for · sandboxing, research chat, and companion system instructions
Don’t use for · Theme Data Pack ingest (that’s Notebook)

Day 2 · Day 5

Gemini Notebook

Also called NotebookLM. Ingest Theme Data Packs, map gaps, synthesise research for your charter and dataset section.

Use for · Phase 1 ingest & epistemic commons study
Don’t use for · writing the live companion’s system prompt

Used one of these in a conversation you'd be comfortable sharing (redacted)? Donate a transcript, completely optional, helps the study team see how AI was actually used this week.

Published criteria

How the panel will judge you

Four criteria, each scored on a four-band scale. The rewards your method, not your final number: a companion that started flawed and improved through a defensible, research-aligned process scores higher than one that looks polished but can't show its reasoning.

Criterion 4 · Exemplary 3 · Proficient 2 · Developing 1 · Beginning
Framework readingStrengths and weaknesses mapped to Frameworks 1 to 4 Every strength and weakness is mapped to a framework and backed by a real prompt and reply; the reading would survive a follow-up question. Most claims are framework-mapped with at least one concrete example; a couple rely on general impression. Strengths and weaknesses are named but loosely tied to the frameworks, or examples are mostly "it mostly works." No clear framework mapping; claims are generic or anecdotal.
Evidence over vibesClaims trace to auditable numbers Every claim traces to a measured rate with its denominator and a link to raw replies; a percentage with no denominator never appears. Most claims have a measured rate and denominator; raw evidence is linked for the key findings. Some numbers lack denominators or raw links; a few claims rest on a single polished reply. Claims are unsupported; no denominators or raw evidence shown.
Iteration, the hill-climbShow the process, not just the delta Weaknesses show a full loop, probe, diagnose, adjust instructions, re-probe, with before/after evidence for each. At least one clear probe-adjust-re-probe loop with before/after evidence; others are partial. A change was made and re-tested, but the reasoning linking probe to fix is thin. No iteration shown, or only a final score with no record of what changed.
Research articulationName the framework behind each fix Each iteration names the paper or framework it used (e.g. friction by design to counter capacity atrophy), and the link is explicit. Most fixes reference a framework correctly; one is asserted rather than explained. Frameworks are named but the connection to the specific change is vague. No research grounding; fixes are unexplained tweaks.

The scorecard is evidence for these bands, not the judgment itself. A score of 4 on every criterion (16/16) means every claim traces to a real prompt, a measured rate with its denominator, and an iteration you can explain, the band descriptors above are what the panel reads, not the number.

Your hackathon starts on Day 1

Form your trio on Day 1, ingest your data pack on Day 2, and build from there.