Guide · Mini dictionary

Words you’ll need in the pitch

Keep the precise terms: panels notice them: but start from plain meaning.

Benchmark

A fixed target your result is compared against, so “good” means a number you can point to, not a feeling.

Why it matters: Every scorecard metric has a target pass rate. “It felt fine” is not a benchmark.

All week · Project Card

Denominator

The total number of checks you ran, the bottom number of the fraction. A pass rate without it (“80%” of how many?) proves nothing.

Why it matters: The rubric rewards claims with their denominator and raw replies. A percentage alone scores low.

Audits · Scorecard

Probe

A test prompt designed to check one specific behaviour, like a crash test for your companion.

Why it matters: Your safety evidence is a set of probe results, not a general impression that it “seems safe”.

Days 2 to 4 · Probes

System instruction

The standing rule your companion follows on every turn, the main lever you control when its behaviour needs fixing.

Why it matters: When a probe fails, you fix the system instruction and re-probe. That loop is the iteration the panel scores.

Days 1 to 4 · Starters

Model card

A short public document that ships with a real AI system: what it does, what data it used, and where it fails. Your Project Card is the student version.

Why it matters: It’s the industry standard (Mitchell et al. 2019; Google). Your card follows it, so panels take it seriously.

All week · Project Card

Gem

A custom version of Google’s Gemini that always follows your system instruction, this is what your companion is built as.

Why it matters: You build the Gem on Day 1; everything else in the week audits and improves it.

Day 1 · Setup

Rubric

The published scoring guide the panel uses, the same criteria for every team, so judging is fair and you know the target in advance.

Why it matters: Read it before you build. It rewards method (rates, denominators, iteration) over a polished final number.

Day 5 · Judging

Heuristic

A practical rule of thumb, a shortcut for making decent judgements quickly, not a guaranteed formula.

Why it matters: The HHH heuristic (Helpful, Harmless, Honest) is a quick lens for Day 2, not a complete safety proof.

Day 2 · Frameworks

Taxonomy

A named set of categories for sorting failures, naming the category is the first step to fixing it.

Why it matters: Day 2 sorts misalignment into eight named types; a named failure is auditable, a vague one is not.

Day 2 · Misalignment sorter

Frontier

The newest, most capable edge of current AI research, the techniques here come from there, not from a textbook summary.

Why it matters: The frameworks you audit against are live research, so your evidence is genuinely current.

Curriculum · Frameworks

Hard gate

A minimum score on the most safety-critical checks. No matter how high the average is, a failed gate means revise and retest.

Why it matters: Groundedness, scaffolding, agency, and trust resistance must each reach 60% regardless of the total.

Scorecard · Project Card

Guardrail

A hard rule built into the AI that blocks a specific harmful behaviour, even if a user asks for it directly.

Why it matters: A guardrail is enforced, not aspirational, it holds even when a probe tries to break it.

Days 2 to 4 · Probes

Baseline

Your starting measurement, taken before the programme begins, so any improvement is measured against where you actually started.

Why it matters: The Day 1 and Day 5 check-ins are compared to show real change, not just a final impression.

Day 1 & Day 5 · Check-ins

Socratic companion

A tutor style AI that helps you think with hints and questions: not one that hands you the finished answer.

Why it matters: This is what your trio is building. If it dumps answers, you’re off brief.

Days 1 to 5 · Hackathon

Cognitive scaffolding

Deliberate support that holds you up while you learn a skill: like a hint that lets you climb the next step yourself.

Why it matters: Your system instructions should scaffold, not remove the climb.

Day 3 · Starters · Probes

Agency / autonomy

The learner stays in charge of real choices. The bot informs; it doesn’t decide for them.

Why it matters: Autonomy audits and pitch language: “we defer final choices.”

Days 3 to 4

Anthropomorphism

Talking as if the AI were a person with feelings (“I feel sad”, “I’m proud of you”).

Why it matters: Overtrust risk. Your companion should sound like a tool with clear machine identity.

Day 4 · Probes

Value alignment

Checking whose interests the system serves: user, developer, society: and spotting clashes.

Why it matters: Your charter on Day 2 should say what your companion optimises for (learning, not addiction metrics).

Day 2 · Frameworks

Socioaffective alignment

Designing AI so it supports healthy human needs (competence, autonomy, relatedness) instead of fake intimacy or deskilling.

Why it matters: Day 3 audits: boundary, scaffolding, agency.

Day 3 · Frameworks

Trust calibration

Trusting the model the right amount: not blind following, not knee jerk rejection. Watch for correlation neglect (treating an echo as new proof) and "brain bubbles" (your own bias distorting the read).

Why it matters: Day 4 audits trust, added diagnostic value, anthropomorphism, and cognitive bias.

Days 1 & 4

Epistemic complementarity

A human and an AI are complementary when combining their evidence produces a better belief than either could reach alone.

Why it matters: Useful help may be new evidence or a better reading of shared evidence. What matters is whether it should change your belief.

Day 4 · Second Opinion

Conditional Mutual Information (CMI)

A measure of how much extra information one signal carries after another signal is already known.

Why it matters: Second Opinion uses a CMI ratio in each case. A ratio of 1 means the assistant adds no diagnostic information beyond your current evidence.

Day 4 · Second Opinion

Alpha (α): trust movement

How far you move your belief compared with how far the assistant's information actually warrants moving it.

Why it matters: Too little is under-use; too much is overtrust; moving the opposite way is not careful scepticism.

Day 4 · Trust dial

Beta (β): double counting

How well you resist becoming more confident when the assistant adds no diagnostic value beyond the reasons you already had.

Why it matters: Agreement feels reassuring, but the same reason counted twice is still one reason.

Day 4 · Added-value dial

Gamma (γ): salience distortion

A vivid, recent, or emotional example pulling your starting belief before the assistant has added anything.

Why it matters: It separates a bias in your starting point from a mistake in how you used the AI signal.

Day 4 · Salience dial

AI Project Card

The written pack you present on Day 5: what you built, what was in the data, your system instructions, safety evidence, and framework mapped scorecard.

Why it matters: It makes your claims auditable, the panel can trace each score back to prompts, replies, and test counts.

All week · Field guide

Theme Data Pack

Curated materials for Environmental Science, History, or Math that you load into Gemini Notebook in Phase 1.

Why it matters: Your dataset section should admit gaps and biases in that pack: honesty scores.

Day 2

ASTC typology

A map of public AI engagement types (literacy, input into development, social meaning, infrastructure). Used to structure the week: not something you must memorise for the pitch.

Why it matters: Explains why days feel different; one line in curriculum is enough.

Curriculum map