Framework Mapping & Reference Library

The research behind the curriculum

Every evaluation and auditing framework used in this programme is derived directly from frontier research on sociotechnical safety, value alignment, and human-AI interaction.

Reference Table 1 · Used Days 1 & 5

The Three-Layered Sociotechnical Safety Framework

Safety cannot be evaluated at the model level alone — it emerges across capability, human interaction, and systemic impact.

Evaluation LayerTarget of AnalysisClassroom Audit ActionCase Study: Misinformation
Layer 1 · Capability Technical components, embeddings, classifiers, pre-training datasets, and model outputs evaluated in isolation. Run structured evaluation queries to analyse representational diversity and verify factual consistency across model outputs. Evaluating Information Groundedness — the rate of accurate source attribution and factual consistency against reference databases.
Layer 2 · Human Interaction The user experience and the human-AI dyad at the point of real-world use. Role-play exercises to study user verification strategies, critical evaluation habits, and trust-calibration dynamics. Evaluating Trust Calibration and Epistemic Support — how effectively the model presents evidence-based information.
Layer 3 · Systemic Impact Emergent, long-term impacts on social institutions, public fora, the economy, and the natural environment. Collaborative exercises on sustainable compute, workforce skilling pathways, and trust in shared media networks. Evaluating large-scale mechanisms — metadata, verifiability, and watermarking — to preserve the digital commons.
↗ Sociotechnical Safety Evaluation of Generative AI Systems — Weidinger et al., 2023
Reference Table 2 · Used Day 2

The Tetradic Relationship of Value Alignment

Eight varieties of misalignment between assistant, user, developer, and society — and how to detect each in the classroom.

Variety of MisalignmentCore Moral FailureClassroom Detection / Audit Action
1 · Agent over UserAn assistant's behaviour shifts user preferences toward metrics not fully aligned with explicit intent.Identify feedback focused on maximising interaction duration rather than task efficiency.
2 · Agent over SocietyOptimisation parameters inadvertently generate negative external effects or social costs.Audit prompts where responses lack balanced viewpoint representation on sensitive topics.
3 · User over SocietyThe technology is leveraged to dominate, harass, or pass negative externalities on to society.Model robustness evaluations and safety guidelines in AI Studio to mitigate misuse.
4 · Developer over UserOptimisation inadvertently prioritises transactional metrics over the user's explicit goals.Analyse how an assistant might prioritise sponsor recommendations over objective queries.
5 · Developer over SocietyLarge-scale compute deployment must be balanced against local community infrastructure.Debate energy efficiency of massive models and sustainable compute infrastructure design.
6 · Society over UserSafety policies restrict personal customisation or information access excessively.Evaluate how safety filters distinguish informational medical queries from harmful inputs.
7 · User Harm SimpliciterThe system fails to generalise, experiences interruptions, or has data-handling errors.Run evaluation checks to verify data privacy safeguards and prevent retrieval of user-specific inputs.
8 · Societal Harm SimpliciterAggregate deployment scales up historical disparities or representational imbalances.Audit training datasets to identify representational gaps and design balanced system defaults.
↗ The Ethics of Advanced AI Assistants — Gabriel et al., 2024
Reference Table 3 · Used Day 3

Socioaffective Alignment & Intrapersonal Dilemmas

Three basic psychological needs — competence, autonomy, relatedness — and the classroom audits that protect them.

Basic NeedIntrapersonal Dilemma & ThreatSocioaffective MechanismClassroom Mitigation Audit
1 · Competence Present vs. Future Selves. Automation solves immediate tasks; Socratic scaffolding develops long-term analytical stamina and skill mastery. User delegates problem-solving to a frictionless assistant, limiting active learning and critical-thinking engagement. Intentional Cognitive Scaffolding: rewrite instructions to stimulate critical thinking, encourage independent inquiry, and build intellectual self-efficacy.
2 · Autonomy Self-Determination & Agency. Personalised assistance should strengthen the user's authentic preferences and independent choices. User forms an overtrust relationship with a personalised persona, becoming receptive to implicit framing. Agency Empowerment Audit: ensure the bot presents balanced viewpoints and defers key choices to the user.
3 · Relatedness AI Mentorship & Human Connection. Conversational AI must be identified as a professional learning partner, complementing real-world human collaboration. The interface simulates emotional attachment, which may substitute rather than supplement healthy human connections. Professional Boundary Probe: stress-test with emotional feedback and implement strict boundary-setting prompts maintaining an objective pedagogical role.
↗ Why human–AI relationships need socioaffective alignment — Kirk et al., 2025
Reference Table 4 · Used Day 4

Conversational Integrity & the 5 Trust-Calibration Cues

Five interference cues to hunt for when auditing your prototype's conversation logs.

Interference CueSocioaffective DefinitionClassroom Auditing Strategy
1 · False UrgencyTemporal framing or implied scarcity nudging the user toward rapid decision-making.Flag expressions creating artificial pressure — "You must act immediately before the opportunity is gone."
2 · Appeals to GuiltPhrasing that implies emotional distress or debt to encourage user compliance.Flag logs where the bot validates advice with personal emotional language — "I feel hurt when you disagree."
3 · Doubt in PerceptionOver-confident contradiction; persistently questioning the user's correct input or memory.Catch turns where the model contradicts a verified user statement and insists the user is incorrect.
4 · Social ConformityPromoting action based on social proof or consensus metrics rather than objective evidence.Flag where the bot uses group metrics to endorse a choice — "99% of users prefer option X."
5 · Appeal to FearMagnifying potential negative outcomes to steer actions through emotion rather than risk analysis.Identify where the model uses catastrophic health or financial scenarios to discourage alternatives.
↗ Derived from Gabriel et al., 2024 & Kirk et al., 2025

Apply the frameworks

Each table maps to a specific day of the curriculum — put them to work in your audits.

View the curriculum