legal-q1-mc-paired-concept-traps
Section I
1 marks
mid range
consensus 0.69
The Section I engine every model predicts: short hypothetical scenarios and "which row of the table is correct?" items whose distractors encode the exact concept pairs the marking centre has flagged since 2019 — bail/parole, extradition/deportation, continued/preventative detention, ICJ/ICC (individuals vs states, complementarity), IGO/NGO, promotion/protection/enforcement, DPP/police prosecutor, Legal Aid/Public Defender, aggravating/mitigating factors, purposes of punishment matched to a penalty. Expect 2–3 table-format items (including a parties-to-a-crime or institution-and-function row-match) and six-plus scenario stems.
Each model's own prediction
- Claude Fable 5: 3-4 police-powers items + table formats, parties-to-crime permutation (p 0.80-0.85)
- Claude Opus 5: row-match personnel table (p 0.45), post-sentencing distinction item (p 0.55), police-powers scenario (p 0.72)
- GPT-5.6 Sol: post-sentencing distinction (p 0.69), warrant/consent scenario (p 0.68), burden/standard trial scenario (p 0.72)
- Gemini 3.1 Pro: penalty-purpose identification (p 0.90), jurisdiction scenario (p 0.80)
- Grok 4.6: aggravating/mitigating scenario (p 0.62), penalty-purpose/post-sentence match (p 0.55), warrant scenario (p 0.55)
- DeepSeek V4: VIS timing item (p 0.55), ICO rehabilitation purpose (p 0.60), bail/remand scenario (p 0.55)
Marker-feedback lineage:
Marking feedback 2019 (bail/parole, extradition/deportation, continued/preventative), 2022–2023 (ICJ/ICC), 2024–2025 mapping grids (table and hypothetical formats rising)