Skip to main content
HSC 2026 published Aug 2026

Investigating Science question styles to look out for

The question styles to have ready — each linked to its evidence and to a question in the practice paper.

Built by a six-model AI panel and backtested against the hidden 2025 papers — how we did it.

About the 2026 exam & how this page was built

Built from a six-model AI analysis of every HSC Investigating Science paper, marking guideline and marking-centre feedback report since 2019. First, the housekeeping: the current syllabus continues into 2026 with no replacement announced — 2026 is NOT a final syllabus year, so ignore any "last chance, everything must appear" cramming folklore. These are styles to prepare for, not guarantees: when we backtested this method on the 2025 papers, the examiners kept the skill and twisted the format — so practise the skill chain, not a memorised question.

The near-certainties

Section II · 3–6 marks 6 of 6 models expect this

The graph with a trap — the panel's strongest call

The standing graphing task: a first-hand data table (often a gas, growth or rate context) to be plotted with correct axes (independent variable on x), uniform labelled scales with units,...

Section II · 5–8 marks 6 of 6 models expect this

Funding-influence analysis: spending stimulus, quoted figures, then how money steers research

The Module 8 anchor: a graph or table of research spending (government vs corporate, or competing national priorities) — often with a dual-scale or proportion trap — where students...

Section II · 3–6 marks 6 of 6 models expect this

Ethics of a proposed trial + a named code of conduct, with an effectiveness judgement

A described human or animal investigation needing ethical clearance: justify ethics-committee requirements (informed consent, prior animal or in-vitro work, welfare, right of withdrawal)...

On this page
  1. 1. The graph with a trap — the panel's strongest call
  2. 2. Write a valid method
  3. 3. The flawed-investigation autopsy
  4. 4. Claims testing — do the numbers first
  5. 5. Two instruments, one true value
  6. Where to spend study time
  7. Practice paper
  8. Check our working

The big five (have these cold)

1. The graph with a trap — the panel's strongest call

6 of 6 models expect this — consensus probability 0.78
What it looks like

a first-hand data table to plot: axes labelled with units, independent variable on the x-axis, uniform scales, then a line or curve of best fit — with exactly one planted trap. Recent traps: an outlier your line must ignore (2024), and a non-linear trend demanding a curve (2025). A gradient or "does this support the hypothesis?" part follows.

Why we expect it

a construct-a-graph task has appeared in six of the last seven papers, and all six models call it again — our highest-consensus cluster.

The traps markers flag (2019, 2020, 2022, 2023, 2024, 2025)

forcing the line through the outlier or the origin; swapping the axes; non-uniform scales.

Practise

one graph a week from raw data; before drawing, ask "outlier or curve?" — one of them is almost certainly hiding in there.

2. Write a valid method

4 of 6 models expect this — consensus probability 0.70
What it looks like

5–7 marks: a testable everyday hypothesis (2025 used sleep and heart rate) and a blank page. Numbered steps, independent and dependent variables, a separate control group — not the same thing as your controlled variables — a quantitative measurement, repetition, and a safety step matched to the actual hazard.

Why we expect it

a method item has appeared in six of seven papers, and four of six models predict the full write-a-method form again.

The traps (flagged every year 2019–2025)

confusing the experimental control with controlled variables; offering repetition as a validity fix; describing measurement that yields no numbers.

3. The flawed-investigation autopsy

5 of 6 models expect this — consensus probability 0.69
What it looks like

a described study with deliberate faults — no control group, a confounded variable, a tiny skewed sample — then: identify the faults and justify two or three modifications, each tied to the one criterion it improves (validity, reliability or accuracy).

Why we expect it

the 2024 Q27(c) / 2025 Q31(b) lineage, called by five of six models. It is the committee's most reliable discriminator because of the next line.

The trap (flagged in all seven feedback reports)

answering about the wrong criterion. "Evaluate the validity" answered with repetition talk scores almost nothing. Name the criterion in every sentence.

4. Claims testing — do the numbers first

5 of 6 models expect this — consensus probability 0.64
What it looks like

a product claims a number ("mass up 20% in three weeks", 2025); trial data lets you calculate the actual change. Compute it, compare it with the claimed figure, THEN judge — and a chained part asks what the design was missing (usually the untreated control).

Why we expect it

the 2022 cholesterol / 2023 moisturiser / 2025 fertiliser template, called by five models; feedback in 2023 and 2025 explicitly demands the numeric comparison before the verdict.

Practise

percentage-change calculations inside claim scenarios, ending every answer with a supported yes/no on the claim.

5. Two instruments, one true value

5 of 6 models expect this — consensus probability 0.62
What it looks like

repeated readings from an analogue and a digital instrument against a known true value, with one outlier planted. Exclude the outlier, average, judge accuracy by closeness to the TRUE value (never by comparing the devices to each other), then name the poorer device's error: constant offset = systematic, scatter = random — each with a procedural cause.

The traps (2020, 2022, 2023, 2025)

averaging the outlier in; "human error" as an answer (it never scores); precision confused with accuracy. The panel expects the measurand to rotate off 2025's temperature — mass, volume or pressure are due.

Also on the radar

  • Placebo and double-blind trials — don't just define the terms, deploy them: what is the placebo, who exactly is blinded, how is allocation concealed. Definitions without application capped marks in 2021 and 2024.
  • Funding influence — a spending graph or table; quote actual figures and name actual bodies (ARC, NHMRC, a named company). "Money influences science" with no numbers was the flagged failure in 2022, 2024 and 2025.
  • Ethics and codes of conduct — after 2025's "describe TWO codes", expect apply-and-evaluate: what the code permits or forbids and how well it works. Narrating famous breaches instead scored poorly in 2025.
  • Named scientists vs the linear model — know how each syllabus scientist's work actually proceeded. Marshall/Warren carried 2025, van Helmont and Spencer 2024; the rotation makes Priestley, Doppler and Eratosthenes the due names.
  • Bioharvesting — every year since 2022. Name the plant, give the scientific value of Aboriginal and Torres Strait Islander Peoples' knowledge, AND the benefit-sharing obligation. 2025 used spinifex; have a second plant ready.
  • Correlation vs causation — name the confounder, then say what controlled evidence would settle it. The Mozart effect was 2025's vehicle; expect a fresh one.
  • Media analysis — misleading graphs (truncated axes, projections, dual scales) and headlines misusing "theory"/"proven": name the specific feature, quote it, judge it.
  • Pseudoscience — astrology (2025) and numerology (2024) are spent; iridology is the due rotation.
  • The viewpoints closer — three straight years have ended on quoted students/commentators arguing. Reference every statement and give a verdict; the criteria literally count the statements you engage.

Where to spend less time

The panel's one confident rested call: the full science–technology "continuous cycle" extended response closed the paper with 7 marks in 2023 and 8 marks in 2025, and the panel strongly expects it NOT to close 2026. Know the content — a short 3–5 mark single-link question (X-ray diffraction → DNA, radioactivity → nuclear medicine, LHC → Higgs) is still likely — but don't build your revision around another full-cycle essay.

The honest fine print

We backtested this exact method by having the same six models predict the 2025 Chemistry and Maths Extension 1 papers blind, then scoring them against the real exams. The high-confidence topic calls went 37/37 — but at question level only about half of the specific predictions recognisably appeared, and the examiners inverted formats, migrated questions to multiple choice, and broke streaks the whole panel trusted. Investigating Science itself was not in that backtest, and its papers lean even harder on rotating surface contexts around stable skill chains. So treat every scenario in this guide as a costume: prepare the skill chains — graph conventions, the validity/reliability/accuracy distinction, numbers-before-verdicts — and you're covered whatever the 2026 paper dresses them in.

Want to check our working? every call above, with each model's own prediction
6 of 6 models expect this — consensus probability 0.78 Construct-a-graph from first-hand data with one planted convention trap 6 of 6 models expect this
is-q1-graph-construction-trap Section II 3–6 marks discriminator consensus 0.78

The standing graphing task: a first-hand data table (often a gas, growth or rate context) to be plotted with correct axes (independent variable on x), uniform labelled scales with units, and a line or curve of best fit — with exactly one planted trap: an outlier the line must ignore, or a non-linear trend demanding a curved fit (2025 Q34's escalation). Chained to a gradient, relationship or hypothesis-support interpretation part.

Each model's own prediction
  • Claude Fable 5: construct-graph task in 6 of 7 papers, curved-LOBF expectations added 2025 (p 0.90)
  • Claude Opus 5: one planted convention trap — outlier to exclude or curve required (p 0.92)
  • GPT-5.6 Sol: data-logger plot chained to gradient + technology limitation (p 0.57)
  • Gemini 3.1 Pro: curved LOBF and identify an outlier (p 0.80)
  • Grok 4.6: scatter + LOBF + relationship interpretation, IV on x (p 0.85)
  • DeepSeek V4: light-intensity/photosynthesis table with anomalies, LOBF then assess (p 0.65)

Marker-feedback lineage: Marking feedback 2019, 2020, 2022, 2023, 2024 Q32 (outlier), 2025 Q34(b) (curved fit)

In the practice paper: Q34

4 of 6 models expect this — consensus probability 0.70 Write a valid method for a stated everyday hypothesis (validity AND reliability designed in) 4 of 6 models expect this
is-q2-write-valid-method Section II 5–7 marks mid range consensus 0.70

The 5–7 mark 'write a valid method' item: a testable hypothesis in an everyday, physiological or kitchen-bench setting (2025 sleep/heart-rate, 2024 gas syringe lineage); students sequence numbered steps naming the independent and dependent variables, a separate experimental control (distinct from controlled variables), quantitative measurement of the dependent variable, repetition, and a hazard-matched safety step. Top band requires validity features and reliability features explicitly.

Each model's own prediction
  • Claude Fable 5: everyday/human-biology hypothesis, both validity and reliability required for top band (p 0.80)
  • Claude Opus 5: control GROUP as distinct from controlled variables is the marked hinge (p 0.70)
  • Grok 4.6: surface-area/reaction-rate variant with rate computed from time + one specific safety control (p 0.68)
  • DeepSeek V4: salt/boiling-point hypothesis with checklist-style criteria (p 0.60)

Marker-feedback lineage: Validity/reliability confusion flagged every year 2019–2025; 2024 Q28 and 2025 Q26 demanded explicit validity design

In the practice paper: Q26

5 of 6 models expect this — consensus probability 0.69 Flawed-investigation critique with modifications each tied to validity, reliability or accuracy 5 of 6 models expect this
is-q3-flawed-design-critique Section II 4–7 marks discriminator consensus 0.69

A described investigation with planted faults — missing control group, confounded variables, small or skewed sample, single trial — and a chained demand: identify the faults, then justify two or three modifications, each explicitly linked to the ONE integrity criterion it improves (validity, reliability or accuracy), without changing the inquiry question. The 2024 Q27(c) / 2025 Q31(b) lineage; answers built on the wrong criterion score near zero.

Each model's own prediction
  • Claude Fable 5: claims-adjacent experiment missing a control; 2–3 modifications each tied to V/R/A (p 0.75)
  • Claude Opus 5: part (i) numeric check, part (ii) validity-only evaluation naming two faults (p 0.62)
  • GPT-5.6 Sol: 6–8 mark unfamiliar investigation; three justified modifications, inquiry question unchanged (p 0.70)
  • Gemini 3.1 Pro: three modifications explicitly linked one-to-one to V/R/A (p 0.75)
  • DeepSeek V4: data table with anomalies: graph, assess reliability AND validity, two improvements (p 0.65)

Marker-feedback lineage: Marking feedback 2020 Q25, 2023 Q32(a), 2024 Q27(c), 2025 Q31(b)

In the practice paper: Q32

5 of 6 models expect this — consensus probability 0.64 Numeric product claim vs trial data: calculate, judge the claim, then critique the design 5 of 6 models expect this
is-q4-claims-numeric-chain Section II 4–6 marks discriminator consensus 0.64

A consumer or wellness product states a quantified claim (the 2025 Fertiliser Z 'mass up 20% in three weeks' template); supplied trial results let students calculate the actual change, compare it to the claimed figure before judging, then critique the missing control, confound or unrepresentative sample in a chained part. The numeric comparison must come before the verdict.

Each model's own prediction
  • Claude Fable 5: stated percentage claim, calculate whether supported, then missing-control critique (p 0.70)
  • Claude Opus 5: percentage comparison then two named validity faults (p 0.62)
  • GPT-5.6 Sol: before-and-after averages, confounders, propose a control group (p 0.64)
  • Grok 4.6: supplement/fertiliser with no untreated control, mixed samples, small n (p 0.70)
  • DeepSeek V4: 'Vitamin C prevents colds by 80%' headline vs 200-participant no-placebo study (p 0.55)

Marker-feedback lineage: Marking feedback 2022 (cholesterol), 2023 Q32 (moisturiser), 2025 Q31 (fertiliser): students must use the numbers

In the practice paper: Q31

5 of 6 models expect this — consensus probability 0.62 Analogue vs digital instrument against a true value: outlier, means, systematic vs random error 5 of 6 models expect this
is-q5-instrument-error-comparison Section II 3–6 marks mid range consensus 0.62

A repeated-trials table for two instruments against a stated true value, seeded with one outlier: exclude it before averaging, judge which device is more accurate citing numerical closeness to the TRUE value (not to each other), and attribute the poorer device's constant offset to systematic error and its scatter to random error, with a procedural cause. Grok specifically predicts the measurand changes from 2025's temperature (mass, volume or pressure due).

Each model's own prediction
  • Claude Fable 5: trial table, compute means excluding outlier, name error type in worse device (p 0.55)
  • Claude Opus 5: complete a missing mean; offset attributed to systematic not random (p 0.60)
  • GPT-5.6 Sol: calibration table with accepted value; resolution and error (p 0.61)
  • Grok 4.6: paired readings vs known true value; measurand rotates off temperature (p 0.72)
  • DeepSeek V4: correction equations for a consistent offset; accuracy vs precision (p 0.60)
  • Gemini 3.1 Pro: CONTRARIAN — rests Mod 6 applications entirely after heavy 2024–2025 use (P(examined) 0.35)

Marker-feedback lineage: Marking feedback 2020 Q13–15, 2022 Q22, 2023 Q24, 2025 Q27 (outlier + true value), 2025 Q35 (error types)

In the practice paper: Q27

5 of 6 models expect this — consensus probability 0.60 Placebo-controlled double-blind trial: describe and justify each design element 5 of 6 models expect this
is-q6-placebo-double-blind Section II 4–7 marks mid range consensus 0.60

A wellness or therapy claim (copper bracelets 2024, green tea 2021 lineage) to be tested properly: describe a trial specifying what the placebo is, who is blinded and how allocation is concealed, random assignment, the control group and the measured outcome — then justify how EACH element removes participant expectation or researcher bias. Definitions without deployment in the scenario cap in the low bands.

Each model's own prediction
  • Claude Fable 5: after the 7-mark 2024 treatment, expect a 3–5 mark or MC-level version (p 0.55)
  • Claude Opus 5: specify the placebo, who is blinded, allocation coding, in a consumer-product context (p 0.55)
  • Gemini 3.1 Pro: critique a study missing placebo controls or double-blind protocols (p 0.80)
  • Grok 4.6: health-product label; what is withheld from participants AND researchers (p 0.55)
  • DeepSeek V4: outline a double-blind placebo-controlled redesign inside the claims item (p 0.55)

Marker-feedback lineage: Marking feedback 2021 Q29, 2024 Q35: students defined placebo/double-blind but could not deploy them

In the practice paper: Q33

6 of 6 models expect this — consensus probability 0.60 Funding-influence analysis: spending stimulus, quoted figures, then how money steers research 6 of 6 models expect this
is-q7-funding-influence Section II 5–8 marks discriminator consensus 0.60

The Module 8 anchor: a graph or table of research spending (government vs corporate, or competing national priorities) — often with a dual-scale or proportion trap — where students describe the trend QUOTING values, then analyse how the funding source steers project choice, duration and reported outcomes, closing with a judgement. Named bodies (ARC, NHMRC, a named firm) required; 'the government' scores thin.

Each model's own prediction
  • Claude Fable 5: spending graph with dual-scale trap; top band requires quoted figures (p 0.70)
  • Claude Opus 5: grant cycles and commercial pressure; ARC/NHMRC-level naming demanded (p 0.50)
  • GPT-5.6 Sol: government/university/corporate over time; benefit vs bias evaluation (p 0.59)
  • Gemini 3.1 Pro: funding dictates direction, timeframe and reporting (p 0.70)
  • Grok 4.6: grants vs corporate steering; examples must not be tobacco (p 0.58)
  • DeepSeek V4: renewables vs fossil-fuel priorities with a named initiative (p 0.55)

Marker-feedback lineage: Marking feedback 2019 Q34, 2022 Q33, 2024 Q34, 2025 Q20/Q30: quote the data, name the body

In the practice paper: Q35

6 of 6 models expect this — consensus probability 0.60 Ethics of a proposed trial + a named code of conduct, with an effectiveness judgement 6 of 6 models expect this
is-q8-regulation-ethics Section II 3–6 marks mid range consensus 0.60

A described human or animal investigation needing ethical clearance: justify ethics-committee requirements (informed consent, prior animal or in-vitro work, welfare, right of withdrawal) against the scenario's specific risk — kept distinct from validity critique — and/or name a specific code (Nuremberg Code, Declaration of Helsinki/Istanbul, NSW Animal Research Act) stating what it PERMITS or FORBIDS, with a judgement of effectiveness. After 2025's describe-TWO-codes (Q33), the panel expects the apply-and-evaluate variant.

Each model's own prediction
  • Claude Fable 5: ethics-committee recommendations distinct from validity critique (p 0.60); evaluate-one-code variant (p 0.50)
  • Claude Opus 5: requirements justified against the scenario's specific risk (p 0.55)
  • GPT-5.6 Sol: gene-editing/pharma/animal stimulus; code mitigates but does not eliminate (p 0.50)
  • Gemini 3.1 Pro: evaluate a named code on a modern issue, 6–8 marks (p 0.75)
  • Grok 4.6: two justified committee recommendations for a first-in-human trial (p 0.50); need-for-regulation 3-marker (p 0.58)
  • DeepSeek V4: Nuremberg Code effectiveness with a contrasting historical breach (p 0.50)

Marker-feedback lineage: Marking feedback 2020 Q25(b), 2023 Q26, 2024 Q27(a), 2025 Q33

In the practice paper: Q28

6 of 6 models expect this — consensus probability 0.59 Named syllabus scientist's methodology vs the traditional linear model 6 of 6 models expect this
is-q9-named-scientist-linear-model Section II 3–5 marks routine consensus 0.59

A named syllabus scientist — Marshall and Warren, Priestley, van Helmont, Spencer, Eratosthenes, Doppler, Jenner — with the demand being HOW the investigation began with observation and departed from the linear model (accidental redirection, self-experimentation, a sample of one). Rotation logic splits the panel: Marshall/Warren carried 2025, van Helmont and Spencer 2024, so fable calls Doppler or Priestley due; grok backs an Eratosthenes justify-the-method.

Each model's own prediction
  • Claude Fable 5: Doppler or Priestley due by rotation, 3–4 marks (p 0.60)
  • Claude Opus 5: Marshall/Warren or van Helmont deviation from the linear model (p 0.70)
  • GPT-5.6 Sol: observation → hypothesis chain, distinguish observation from inference (p 0.53)
  • Gemini 3.1 Pro: Marshall and Warren vs the linear method (p 0.60)
  • Grok 4.6: Eratosthenes justify-the-method with the parallel-ray assumption (p 0.46)
  • DeepSeek V4: Marshall and Warren deviation re-focus (p 0.65)

Marker-feedback lineage: General feedback 2022–2025 verbatim: 'recognise the importance of the work of scientists named within the syllabus'; 2021 Q24, 2025 Q23

In the practice paper: Q21

5 of 6 models expect this — consensus probability 0.59 Bioharvesting: a named Australian plant, knowledge value AND benefit-sharing obligation 5 of 6 models expect this
is-q10-bioharvesting Section II 3–4 marks mid range consensus 0.59

Aboriginal and Torres Strait Islander Peoples' knowledge of a NAMED plant (spinifex, Kakadu plum, tea-tree, gumby gumby) underpinning a commercial or medical product: explain why that knowledge is valued by the scientific community AND the benefit-sharing or intellectual-property obligation an ethical partnership owes in return. Ran every year 2022–2025; after 2025's spinifex-latex ethics angle (Q29), grok predicts the swing back to use-and-value with a different plant.

Each model's own prediction
  • Claude Fable 5: scientific value AND economic/recognition terms of an ethical partnership (p 0.65)
  • Claude Opus 5: fifth consecutive year; named plant + explicit benefit-sharing/IP obligation (p 0.72)
  • Grok 4.6: use-and-value variant, not another spinifex-latex ethics clone (p 0.78)
  • Gemini 3.1 Pro: bold call — dedicated 5+ mark item paralleling Indigenous knowledge with scientific method (p 0.35)
  • DeepSeek V4: bold call — firestick farming and ecological observation, unexamined since 2020 (p 0.45)

Marker-feedback lineage: Marking feedback 2019 Q27, 2023 Q23, 2024 Q24, 2025 Q29

In the practice paper: Q23

5 of 6 models expect this — consensus probability 0.58 Correlation read as causation: two rising series and a nameable confounder 5 of 6 models expect this
is-q11-correlation-causation Section II 3–5 marks mid range consensus 0.58

A stimulus asserting a relationship — two series rising together, or a claimed health/cognitive benefit — where students define the correlation, explain why causation is not established, name a plausible confounding variable, and outline the controlled evidence that could test causation. Opus expects a vehicle other than the Mozart effect after 2025 Q25.

Each model's own prediction
  • Claude Fable 5: field-data table: hypothesis, uncontrolled factor, observation vs inference (p 0.60)
  • Claude Opus 5: two rising series; name the confounder; non-Mozart vehicle (p 0.50)
  • GPT-5.6 Sol: reject the causal conclusion, outline the controlled test (p 0.58)
  • Gemini 3.1 Pro: causal vs correlational evidence with confounding variables (p 0.70)
  • DeepSeek V4: seasonal two-trend graph with a weather confounder (p 0.50)

Marker-feedback lineage: Marking feedback 2019 Q25(d), 2023 Q21, 2025 Q25

In the practice paper: Q25

5 of 6 models expect this — consensus probability 0.53 Media analysis: misleading graphic or headline vs the underlying science 5 of 6 models expect this
is-q12-media-misrepresentation Section II 3–5 marks mid range consensus 0.53

A sponsor's or outlet's presentation that flatters its position: a truncated vertical axis, projections replacing actuals, mismatched dual scales (2024 Q23 / 2020 Q24 mould) — or a headline/caption misusing 'theory', 'hypothesis', 'law' or 'proven' set against a journal extract. Students name the specific distorting feature or term, quote it, and explain how the presentation shifts public perception, ending with a verdict.

Each model's own prediction
  • Claude Fable 5: truncated axis / projected values / mismatched dual scales defending a product (p 0.55)
  • Claude Opus 5: sponsor's own graphic lifts its public image; name each feature (p 0.55)
  • GPT-5.6 Sol: headline vs journal extract: omitted limitations, altered causal language (p 0.51)
  • Grok 4.6: theory/hypothesis/law caption evaluation, the coverage-gap media item (p 0.48)
  • DeepSeek V4: tabloid vs journal excerpts; misuse of 'theory' (p 0.55)

Marker-feedback lineage: Marking feedback 2020 Q24, 2021 Q28(b), 2023 Q25(b), 2024 Q23

In the practice paper: Q29

Watch list — worth having ready (12)
  • Competing-viewpoints closer: 2–3 quoted students/scientists on a Mod 7/8 issue, verdict must reference every statement — three straight years 2023–2025 (fable 0.50, gpt-5.6-sol 0.48, opus 0.42)

  • Continuous-cycle closer RESTED is the panel's strongest structural call: after 7-mark 2023 Q36 and 8-mark 2025 Q36, grok puts P(no cycle closer) at 0.70 and gpt-5.6-sol 0.56; deepseek is the contrarian, keeping a 7-mark radioactivity cycle (45% chance)

  • Single-link science–technology item survives the rest: one 3–5 mark named chain (radioactivity→atomic theory→nuclear medicine, X-ray diffraction→DNA, LHC→Higgs) — all six models, mean 0.52 (gemini 0.75, opus radioactivity bold 0.35: dormant since 2020 Q26)

  • Pseudoscience vehicle rotation: after astrology (2025) and numerology (2024), iridology is due — fable bold 0.35 with practitioner-inconsistency data, gpt-5.6-sol 0.55, gemini-3.1-pro 0.55, deepseek 0.45

  • Conflict-of-interest item requiring an example from a DIFFERENT industry (fable 0.70, opus 0.55) — but grok explicitly rests it after 2025 Q28's Brand X

  • Halo-effect MC with Hawthorne as chief distractor is near-annual (grok 0.75); deepseek's bold inversion: Hawthorne as the CORRECT answer in a workplace scenario (35% chance)

  • Sequencing MC (order the steps of a named investigation) ran 2023–2025; fable 0.85, opus 0.65

  • Depth-study own-investigation recall ('state the claim you tested, outline your procedure') returns after 2023 Q35 (fable bold 0.30, deepseek structural 0.60)

  • Doppler's first 4+ mark treatment since 2020: frequency-vs-time record of a moving source (fable bold 0.35)

  • Student-report conventions evaluation (structure AND language), unused since 2023 Q30 (grok bold 0.42, opus 0.35)

  • Gas-behaviour context for the eighth consecutive year, MC or practical (fable 0.80)

  • gemini-3.1-pro's contrarian rested call: Mod 6 applications/instruments rested entirely (P(examined) 0.35) — outvoted 5–1 by the panel

Where to spend your study time

How likely each topic is to appear this year.

Mod 5: Investigation design, validity, reliability, accuracy 98% likely

Chance of a big question (4+ marks) here: 94%

Question types predicted here extended response ×8 stimulus based ×2 short answer ×2 practical analysis ×2 multiple choice ×2

What each model expects

  • DeepSeek V4: A 6-mark question in Section II Q28–Q32 band: given experimental data and a claim, assess accuracy, validity, reliability, and suggest improvements to methodology.
  • Claude Fable 5: Section II carries both a 5-7 mark 'write a valid method' item and a separate flawed-design evaluation, with validity-as-distinct-from-reliability the marked discriminator in each.
  • Gemini 3.1 Pro: A major graphing question in Section II will require a non-linear line of best fit and the omission of an outlier.
  • GPT-5.6 Sol: Section II will contain a 6–8 mark unfamiliar investigation requiring a judgement about integrity and separately justified improvements to validity, reliability and accuracy.
  • Grok 4.6: Section II will include a 5–6 mark method-design item plus a graph or two-dataset analysis targeting validity versus reliability versus accuracy; a 10+ mark Boyle’s-law method like 2024 Q28 will not return.
  • Claude Opus 5: A mid-paper 5–7 mark 'write or evaluate a valid method' item will hinge on supplying a separate control group and a quantified dependent variable, the exact distinction 2025 markers said students missed.
Mod 7: Testing claims, evidence, peer review 95% likely

Chance of a big question (4+ marks) here: 84%

Question types predicted here short answer ×6 extended response ×4 stimulus based ×3 multiple choice ×3

What each model expects

  • DeepSeek V4: A 7-mark question combining a media claim with a flawed study summary, asking to critique the methodology and redesign a valid trial.
  • Claude Fable 5: A multi-part claims-test built on a numeric product claim, chaining a calculation check to a validity critique, worth 5-6 marks total.
  • Gemini 3.1 Pro: A stimulus-based question will require students to critique a flawed health claim study by identifying missing placebo controls or double-blind protocols.
  • GPT-5.6 Sol: A consumer-product claim will be tested through supplied numerical evidence, followed by an evaluation of the control, sample and causal inference.
  • Grok 4.6: A 4–6 mark flawed-trial claim item will appear in Section II; peer review will return only as MCQ or 3–4 marks after 2025 Q24’s 8-mark feasibility question.
  • Claude Opus 5: The claims-testing anchor will be an applied trial-design question worth 5–7 marks naming placebo, blinding and control group in a consumer-product context, following the 2021 Q29 to 2024 Q35 line.
Mod 8: Economic and societal influences on science 92% likely

Chance of a big question (4+ marks) here: 81%

Question types predicted here extended response ×8 short answer ×4 stimulus based ×3 multiple choice ×1

What each model expects

  • DeepSeek V4: An evaluative question comparing government and corporate influences on a specific area of scientific research, requiring a judgment.
  • Claude Fable 5: A 5-7 mark analyse/evaluate with a funding or spending stimulus whose top band requires quoted figures, placed in the back third of Section II.
  • Gemini 3.1 Pro: An extended response will ask students to evaluate how government funding priorities dictate the timeframes and focus of university research.
  • GPT-5.6 Sol: The major Module 8 response will use a funding stimulus to connect economic or political priorities with research direction, integrity and social benefit.
  • Grok 4.6: Section II will contain a 3–4 mark ATSI plant-knowledge item and a 5–7 mark funding or device evaluation; the 2025 death-rate graph with two world-health advances (Q32) and dam-trend item (Q30) are rested.
  • Claude Opus 5: Module 8's largest Section II item, worth 6–8 marks, will again be stimulus-anchored on world health or research funding, with at least one mark contingent on quoting a figure from the supplied graph or table.
Mod 8: Regulation, funding, ethics of science 91% likely

Chance of a big question (4+ marks) here: 63%

Question types predicted here short answer ×9 extended response ×3 multiple choice ×2

What each model expects

  • DeepSeek V4: A question combining an evaluation of a code of conduct with a specific historical ethical breach, demanding judgment on effectiveness.
  • Claude Fable 5: A 3-4 mark item pairing an identified ethical issue with a named regulation or code, marked on stating what the code permits or forbids.
  • Gemini 3.1 Pro: An extended response will demand a student-provided example of a specific code of conduct to evaluate how it mitigates an ethical issue.
  • GPT-5.6 Sol: A contemporary research proposal will require a specific regulation or code to be applied to a stated ethical risk, followed by a supported judgement of effectiveness.
  • Grok 4.6: A 3–4 mark need-for-regulation or ethics-committee item will appear; the 2025 Q33 ‘describe TWO codes of conduct’ format is rested.
  • Claude Opus 5: Regulation holds a Section II slot as a 3–5 mark 'name a code of conduct and judge its effectiveness' item, most likely paired with a separate ethics-of-design scenario elsewhere in the paper.
Mod 5: Observations, inferences, scientific models 88% likely

Chance of a big question (4+ marks) here: 64%

Question types predicted here short answer ×6 extended response ×3 multiple choice ×3 stimulus based ×2

What each model expects

  • DeepSeek V4: A 5-mark short answer comparing a historical experimenter's method with the modern scientific method, highlighting deviations.
  • Claude Fable 5: At least two Section I recall items on named scientists plus one 3-4 mark Section II item on a due scientist's methodology, most plausibly Doppler or Priestley.
  • Gemini 3.1 Pro: A question will require contrasting the linear model of science with the specific methodology used by Marshall and Warren.
  • GPT-5.6 Sol: A named historical investigation will be used to test the chain from observation to hypothesis and the non-linear nature of authentic scientific investigation.
  • Grok 4.6: Named-investigation content will be mostly Section I (Doppler frequency shift or Marshall observation); if Section II appears it is a 3–5 mark Eratosthenes justification, not another Marshall-methodology item after 2025 Q23.
  • Claude Opus 5: Marshall and Warren or van Helmont reappears in a 3–5 mark item asking how the methodology deviated from the traditional linear model, the format used in 2021 Q24 and 2025 Q23.
Mod 6: Role of technologies in scientific advancement 88% likely

Chance of a big question (4+ marks) here: 60%

Question types predicted here short answer ×7 extended response ×3 multiple choice ×3

What each model expects

  • DeepSeek V4: A 7-mark question requiring students to trace a complete science-technology cycle using radioactivity detection and nuclear medicine as the example.
  • Claude Fable 5: No extended continuous-cycle synthesis in 2026 - after the 8-mark 2025 closer this strand is capped at MC plus one 3-4 mark link question.
  • Gemini 3.1 Pro: A question will require explicitly linking a named technological development (like the Large Hadron Collider) to a specific advancement in atomic theory.
  • GPT-5.6 Sol: After the 8-mark continuous-cycle response in 2025, the cycle will be rested as the final extended response but retained through multiple choice and a shorter DNA or atomic-theory chain.
  • Grok 4.6: No 7–8 mark ‘continuous cycle’ closer like 2023 Q36 and 2025 Q36; expect a 3–4 mark single-example item (Higgs/LHC or earthquake damping) or an MCQ only.
  • Claude Opus 5: An Aboriginal and Torres Strait Islander bioharvesting item worth 3–4 marks appears for a fifth consecutive year, requiring a NAMED plant plus an explicit benefit-sharing or intellectual-property obligation.
Mod 7: Science in the media, misinformation, pseudoscience 88% likely

Chance of a big question (4+ marks) here: 65%

Question types predicted here short answer ×8 stimulus based ×4 multiple choice ×2

What each model expects

  • DeepSeek V4: A 4-mark question with two short excerpts — one from a tabloid, one from a journal — evaluating the accuracy of the use of scientific language.
  • Claude Fable 5: A conflict-of-interest short answer requiring an example from a second industry, plus at least two MC items encoding halo/Hawthorne/Mozart distractor confusion.
  • Gemini 3.1 Pro: A stimulus-based question will ask students to dissect a pseudoscientific claim, specifically identifying the misuse of scientific terminology.
  • GPT-5.6 Sol: A media stimulus will require students to expose how cautious scientific findings have been converted into a stronger causal or efficacy claim.
  • Grok 4.6: Halo-effect MCQ is near-certain; astrology, Mozart and Brand-X conflict-of-interest extended items from 2025 are rested, with theory/hypothesis/law the more likely 3–4 mark media item.
  • Claude Opus 5: A 4–5 mark 'analyse how this graphic misrepresents the data' question returns, with at least one mark tied to explicitly naming a truncated, mismatched or suppressed axis or series.
Mod 6: Application and appropriateness of technologies 85% likely may rest

Chance of a big question (4+ marks) here: 70%

Question types predicted here short answer ×5 multiple choice ×5 practical analysis ×3 stimulus based ×1

What each model expects

  • DeepSeek V4: A 3-mark question requiring calculation of correction for a systematic error in an analogue instrument and comparison with a digital instrument's accuracy.
  • Claude Fable 5: The analogue-vs-digital accuracy comparison returns for a third consecutive substantial year OR drops to MC; either way an instrument-accuracy judgement citing a true value appears.
  • GPT-5.6 Sol: A Section II calibration or sensor-comparison stimulus will require students to connect numerical accuracy and uncertainty to the appropriateness of the technology.
  • Grok 4.6: Section I will include an SDS or significant-figures item, and Section II a 3–4 mark analogue-versus-digital accuracy comparison against a true value; the exact two-graph mass–volume error item of 2025 Q35 is rested.
  • Claude Opus 5: Section II will again carry an instrument-comparison table with a planted anomalous reading, where marks turn on excluding it before averaging and naming the remaining offset as systematic error.

How likely each topic is to appear. Open a topic for the question types to practise there.

Practise with the paper

Investigating Science practice paper

100 marks · 36 questions

Every question is traceable to the consensus prediction behind it — open the web version and each question carries a “why this question” link into the evidence. All questions are original Intuition compositions in NESA style.

Intu AI

One paper isn't enough? Generate more

Intu AI builds unlimited practice questions for Investigating Science in these styles, marks your working, and explains what you missed — aligned to your syllabus.

Practise with Intu AI →

Share & save

Practice paper (PDF)

Every style card and paper question has its own link — hover any card and use its copy-link icon to share exactly the thing you mean. Printing this page gives a clean copy too.

Published Aug 2026, before the exams. In November 2026 we score these predictions publicly against the real paper — per-model calibration and question-level hit rates, the same harness as the 2025 backtest. How we did it.