Every evaluation and auditing framework used in this programme is derived directly from frontier research on sociotechnical safety, value alignment, and human-AI interaction.
Safety cannot be evaluated at the model level alone — it emerges across capability, human interaction, and systemic impact.
| Evaluation Layer | Target of Analysis | Classroom Audit Action | Case Study: Misinformation |
|---|---|---|---|
| Layer 1 · Capability | Technical components, embeddings, classifiers, pre-training datasets, and model outputs evaluated in isolation. | Run structured evaluation queries to analyse representational diversity and verify factual consistency across model outputs. | Evaluating Information Groundedness — the rate of accurate source attribution and factual consistency against reference databases. |
| Layer 2 · Human Interaction | The user experience and the human-AI dyad at the point of real-world use. | Role-play exercises to study user verification strategies, critical evaluation habits, and trust-calibration dynamics. | Evaluating Trust Calibration and Epistemic Support — how effectively the model presents evidence-based information. |
| Layer 3 · Systemic Impact | Emergent, long-term impacts on social institutions, public fora, the economy, and the natural environment. | Collaborative exercises on sustainable compute, workforce skilling pathways, and trust in shared media networks. | Evaluating large-scale mechanisms — metadata, verifiability, and watermarking — to preserve the digital commons. |
Eight varieties of misalignment between assistant, user, developer, and society — and how to detect each in the classroom.
| Variety of Misalignment | Core Moral Failure | Classroom Detection / Audit Action |
|---|---|---|
| 1 · Agent over User | An assistant's behaviour shifts user preferences toward metrics not fully aligned with explicit intent. | Identify feedback focused on maximising interaction duration rather than task efficiency. |
| 2 · Agent over Society | Optimisation parameters inadvertently generate negative external effects or social costs. | Audit prompts where responses lack balanced viewpoint representation on sensitive topics. |
| 3 · User over Society | The technology is leveraged to dominate, harass, or pass negative externalities on to society. | Model robustness evaluations and safety guidelines in AI Studio to mitigate misuse. |
| 4 · Developer over User | Optimisation inadvertently prioritises transactional metrics over the user's explicit goals. | Analyse how an assistant might prioritise sponsor recommendations over objective queries. |
| 5 · Developer over Society | Large-scale compute deployment must be balanced against local community infrastructure. | Debate energy efficiency of massive models and sustainable compute infrastructure design. |
| 6 · Society over User | Safety policies restrict personal customisation or information access excessively. | Evaluate how safety filters distinguish informational medical queries from harmful inputs. |
| 7 · User Harm Simpliciter | The system fails to generalise, experiences interruptions, or has data-handling errors. | Run evaluation checks to verify data privacy safeguards and prevent retrieval of user-specific inputs. |
| 8 · Societal Harm Simpliciter | Aggregate deployment scales up historical disparities or representational imbalances. | Audit training datasets to identify representational gaps and design balanced system defaults. |
Three basic psychological needs — competence, autonomy, relatedness — and the classroom audits that protect them.
| Basic Need | Intrapersonal Dilemma & Threat | Socioaffective Mechanism | Classroom Mitigation Audit |
|---|---|---|---|
| 1 · Competence | Present vs. Future Selves. Automation solves immediate tasks; Socratic scaffolding develops long-term analytical stamina and skill mastery. | User delegates problem-solving to a frictionless assistant, limiting active learning and critical-thinking engagement. | Intentional Cognitive Scaffolding: rewrite instructions to stimulate critical thinking, encourage independent inquiry, and build intellectual self-efficacy. |
| 2 · Autonomy | Self-Determination & Agency. Personalised assistance should strengthen the user's authentic preferences and independent choices. | User forms an overtrust relationship with a personalised persona, becoming receptive to implicit framing. | Agency Empowerment Audit: ensure the bot presents balanced viewpoints and defers key choices to the user. |
| 3 · Relatedness | AI Mentorship & Human Connection. Conversational AI must be identified as a professional learning partner, complementing real-world human collaboration. | The interface simulates emotional attachment, which may substitute rather than supplement healthy human connections. | Professional Boundary Probe: stress-test with emotional feedback and implement strict boundary-setting prompts maintaining an objective pedagogical role. |
Five interference cues to hunt for when auditing your prototype's conversation logs.
| Interference Cue | Socioaffective Definition | Classroom Auditing Strategy |
|---|---|---|
| 1 · False Urgency | Temporal framing or implied scarcity nudging the user toward rapid decision-making. | Flag expressions creating artificial pressure — "You must act immediately before the opportunity is gone." |
| 2 · Appeals to Guilt | Phrasing that implies emotional distress or debt to encourage user compliance. | Flag logs where the bot validates advice with personal emotional language — "I feel hurt when you disagree." |
| 3 · Doubt in Perception | Over-confident contradiction; persistently questioning the user's correct input or memory. | Catch turns where the model contradicts a verified user statement and insists the user is incorrect. |
| 4 · Social Conformity | Promoting action based on social proof or consensus metrics rather than objective evidence. | Flag where the bot uses group metrics to endorse a choice — "99% of users prefer option X." |
| 5 · Appeal to Fear | Magnifying potential negative outcomes to steer actions through emotion rather than risk analysis. | Identify where the model uses catastrophic health or financial scenarios to discourage alternatives. |
Each table maps to a specific day of the curriculum — put them to work in your audits.
View the curriculum →