Skip to main content

← Investigating Science styles

2026 HSC Investigating Science — Intuition Education Predicted Paper

100 marks · an original Intuition practice paper realising the consensus predictions — every question links to the evidence behind it. Prefer the PDF?

Provenance & general instructions

This full-length practice paper was built from the consensus of a six-model AI panel (fable, opus, gpt-5.6-sol, gemini-3.1-pro, grok, deepseek), each of which independently predicted the 2026 examination from the 2019–2025 papers and NESA marking feedback. The 117 question-level predictions were clustered; every cluster with consensus probability ≥ 0.55 is realised in this paper, the remainder is drawn from the panel's watch list, and the science–technology "continuous cycle" is deliberately kept to a short single-link item in line with the panel's rested call on the full-cycle closer. Each Section II question is tagged with the consensus cluster it realises and the probability that a question of that kind appears in 2026. Every question is original — none is copied from a past paper, and scenario surfaces used in 2025 (sleep/heart-rate, Fertiliser Z, spinifex, thermometer accuracy) have been rotated per the corpus-echo rule.

General instructions

  • Reading time — 5 minutes
  • Working time — 3 hours
  • Write using black pen
  • Draw diagrams using pencil
  • Calculators approved by NESA may be used
  • Section I — 20 marks. Attempt Questions 1–20. Allow about 35 minutes for this section
  • Section II — 80 marks. Attempt Questions 21–36. Allow about 2 hours and 25 minutes for this section

Section I

20 marks — Attempt Questions 1–20 — Allow about 35 minutes for this section

Use the multiple-choice answer sheet for Questions 1–20.

Question 1

A student's depth-study report contains the sections: abstract, introduction, method, results, discussion, conclusion.

In which section should the student assess the reliability of the results obtained?

  • A. Method
  • B. Results
  • C. Discussion
  • D. Conclusion

Question 2

A lit candle was placed in a jar, and the jar was sealed.

Which of the following is an observation?

  • A. The flame went out 40 seconds after the jar was sealed.
  • B. The flame used up all of the oxygen inside the jar.
  • C. Flames cannot burn in the absence of oxygen.
  • D. The flame would burn for longer in a larger jar.

Question 3

A scientist grew a small tree in a weighed quantity of soil for five years, adding only water. The tree's mass increased by about 74 kg while the soil lost only 57 g, and the scientist concluded that the tree's substance came from the water alone.

Who conducted this investigation?

  • A. Joseph Priestley
  • B. Jan Baptista van Helmont
  • C. Percy Spencer
  • D. Barry Marshall

Question 4

The steps of an investigation into stomach ulcers are listed.

  1. Spiral bacteria were repeatedly seen in biopsy samples from patients with gastritis.
  2. It was proposed that the bacteria, not stress, caused gastritis and ulcers.
  3. The bacteria were successfully cultured after plates were left over a holiday period.
  4. A researcher swallowed a culture of the bacteria and developed gastritis.

In which order did these steps occur?

  • A. 2, 1, 3, 4
  • B. 1, 2, 3, 4
  • C. 1, 3, 2, 4
  • D. 2, 3, 1, 4

Question 5

A safety data sheet for propan-2-ol states: "Highly flammable liquid and vapour. Causes serious eye irritation."

A student's investigation requires a beaker of propan-2-ol to be warmed to 40 °C.

Which action best minimises the most serious risk in this procedure?

  • A. Wearing a cotton laboratory coat
  • B. Warming the beaker in a water bath on an electric hot plate
  • C. Wearing disposable gloves while pouring
  • D. Carrying out the procedure near a sink

Question 6

A measuring cylinder is graduated in 1 mL divisions. The bottom of the meniscus sits between the 36 mL and 37 mL marks.

Which value is the most appropriate way to record this volume?

  • A. 36 mL
  • B. 36.5 mL
  • C. 36.50 mL
  • D. 37.00 mL

Question 7

Students investigated the effect of water salinity on the hatching of brine shrimp eggs. Five salt concentrations were prepared using the same volume of water, and the containers were kept at the same temperature and light level. The number of hatched shrimp was counted after 48 hours.

What is the independent variable in this investigation?

  • A. The number of hatched shrimp
  • B. The temperature of the water
  • C. The salt concentration of the water
  • D. The time allowed for hatching

Question 8

Students are testing the claim that a vitamin solution speeds up the germination of mung beans.

Which change to their investigation would most improve its validity?

  • A. Repeating the investigation with three more batches of beans
  • B. Using a more precise balance to measure the vitamin powder
  • C. Including a group of beans given only water
  • D. Excluding the slowest-germinating beans from the results

Question 9

What is the main purpose of peer review before scientific research is published?

  • A. To ensure the research is written in a style the public can understand
  • B. To have independent experts assess the methodology and conclusions
  • C. To confirm the research will be profitable for the journal
  • D. To guarantee that the results can never be overturned

Question 10

An Olympic swimmer with no scientific qualifications appears in advertisements describing a breakfast cereal as "scientifically formulated", and sales rise sharply.

Which effect best explains the public's acceptance of this claim?

  • A. The halo effect
  • B. The Hawthorne effect
  • C. The Mozart effect
  • D. The placebo effect

Question 11

The productivity of workers in a call centre rose every time researchers trialled a change to the office — brighter lights, dimmer lights, new chairs — and fell back once the researchers left.

Which effect best explains these results?

  • A. The halo effect
  • B. The Hawthorne effect
  • C. The placebo effect
  • D. The Mozart effect

Question 12

A sleep-tracking app displays a message inside the app inviting users to rate it. Of users who had kept the app for at least six months, 92% reported improved sleep. The company advertises: "92% of users sleep better."

Why should this claim be treated with caution?

  • A. The sample size of the survey was too small.
  • B. The sample was self-selected and unrepresentative of all users.
  • C. The percentage was calculated incorrectly.
  • D. Sleep cannot be measured quantitatively.

Question 13

Which row correctly shows a technology enabling an advance in scientific understanding?

Technology Scientific advance
A. X-ray diffraction imaging The double-helix model of DNA
B. The double-helix model of DNA X-ray diffraction imaging
C. The theory of electromagnetism The electric generator
D. Germ theory of disease Antiseptic surgical techniques

Question 14

An ambulance with its siren on drives at constant speed past a stationary observer.

What does the observer hear as the ambulance approaches and then moves away?

  • A. A constant pitch throughout
  • B. A lower pitch while approaching, then a higher pitch
  • C. A higher pitch while approaching, then a lower pitch
  • D. A steadily rising pitch throughout

Question 15

A sealed syringe contains 60 mL of air at a pressure of 100 kPa. The plunger is pushed in slowly at constant temperature until the volume is 20 mL.

What is the new pressure of the air?

  • A. 33 kPa
  • B. 100 kPa
  • C. 200 kPa
  • D. 300 kPa

Question 16

A company plans a trial of a new dietary supplement using human volunteers.

Which action would make the trial more ethical?

  • A. Increasing the number of volunteers in the trial
  • B. Using calibrated instruments to measure the outcomes
  • C. Explaining the risks to volunteers and obtaining their written consent
  • D. Repeating each measurement three times

Question 17

A company claims its "energised crystal" pendants restore the body's natural frequency, and states that when the pendant appears not to work, the wearer's "energy blockages" are responsible.

Which feature most clearly identifies this claim as pseudoscience?

  • A. The claim is made by a commercial company.
  • B. The pendants are popular and have many positive reviews.
  • C. The claim is constructed so that it can never be shown to be false.
  • D. The company did not use a large enough sample when testing pendants.

Question 18

Two students each timed the same pendulum's period five times. The true period is known to be 2.0 s.

Data provided in the exam

Student P's readings: 2.4, 2.4, 2.5, 2.4, 2.4 (all about 0.4 s above the true value). Student Q's readings: 1.8, 2.3, 2.0, 1.7, 2.2 (scattered above and below the true value).

Which row classifies the error most affecting each student's results?

Student P Student Q
A. Random Systematic
B. Systematic Random
C. Systematic Systematic
D. Random Random

Question 19

Data provided in the exam

graph — monthly sunscreen sales (left axis, $ thousands) and monthly jellyfish stings recorded at beaches (right axis, number of stings) plotted for one year; both series peak together in January and dip together in July.

Which statement is best supported by the graph?

  • A. Sunscreen chemicals in the water attract jellyfish to swimmers.
  • B. Being stung by jellyfish causes people to buy sunscreen.
  • C. A third variable, such as more people swimming in summer, could explain both trends.
  • D. The two variables are unrelated because correlation is impossible between different units.

Question 20

Data provided in the exam

table — Agency A: total annual budget \$1.2 billion, spending on medical research \$60 million. Agency B: total annual budget \$3.0 billion, spending on medical research \$90 million.

Which statement about spending on medical research is correct?

  • A. Agency A spends more than Agency B in both dollar terms and as a proportion of budget.
  • B. Agency B spends more than Agency A in both dollar terms and as a proportion of budget.
  • C. Agency A spends more in dollar terms, but Agency B allocates a larger proportion.
  • D. Agency B spends more in dollar terms, but Agency A allocates a larger proportion.

Section II

80 marks — Attempt Questions 21–36 — Allow about 2 hours and 25 minutes for this section

Answer the questions in the spaces provided. Show all relevant working in questions involving calculations.

Question 21 (3 marks)

Why this question → 6 of 6, p 0.59

In the 1770s, Joseph Priestley observed that a sprig of mint growing in a sealed jar "restored" air in which a candle had previously burned, so that the candle could burn in it again and a mouse could survive in it.

Outline how Priestley's investigation differed from the traditional linear model of scientific methodology. In your answer, refer to the role played by his initial observation. (3)

Question 22 (3 marks)

watch list: single-link science–technology item while the full-cycle ...

Technologies capable of detecting radiation, such as photographic plates and radiation counters, contributed to the development of the model of the atom.

Explain how these detection technologies advanced scientific understanding of the atom, and how that understanding later enabled ONE named medical technology. (3)

Question 23 (4 marks)

Why this question → 5 of 6, p 0.59

Aboriginal Peoples of northern Australia have long used the Kakadu plum as a food and for treating ailments. Scientific analysis later showed the fruit has one of the highest vitamin C concentrations of any plant, and it is now used in commercial skincare and food-preservation products.

Explain why Aboriginal Peoples' knowledge of the Kakadu plum is valuable to the scientific community, AND what an ethical partnership between researchers and the knowledge holders requires. (4)

Question 24 (4 marks)

watch list: pseudoscience rotation, iridology due

Iridologists claim that health conditions can be diagnosed by examining patterns in the iris of the eye.

Data provided in the exam

table — four practising iridologists were independently shown the same set of 20 iris photographs and asked to identify which showed "kidney stress". Number of photographs identified: Iridologist 1: 3; Iridologist 2: 11; Iridologist 3: 7; Iridologist 4: 16. No two iridologists selected the same set of photographs.

Explain why iridology is considered a pseudoscience. In your answer, refer to the data provided. (4)

Question 25 (3 marks)

Why this question → 5 of 6, p 0.58
Data provided in the exam

scatter graph — each point represents one suburb; x-axis: number of streetlights in the suburb (200 to 3200); y-axis: number of asthma cases recorded in the suburb per year (15 to 260). The points show a strong positive correlation.

A newspaper reports: "Streetlights linked to asthma — suburbs with more streetlights have more asthma."

Explain why this data does not show that streetlights cause asthma. In your answer, identify a variable that could account for the relationship shown. (3)

Question 26 (6 marks)

Why this question → 4 of 6, p 0.70

A gardening company claims that watering seedlings with diluted seaweed extract makes them grow faster than watering with plain water.

Write a valid and reliable method that a student could follow to test this claim using bean seedlings.

In your method, include:

  • the independent and dependent variables, and how the dependent variable will be measured quantitatively
  • an experimental control
  • TWO variables that must be kept constant
  • how reliability will be addressed. (6)

Question 27 (5 marks)

Why this question → 5 of 6, p 0.62

A class compared two balances by repeatedly weighing the same 50.00 g standard mass.

Data provided in the exam

results table — Analogue spring balance (g): 47.9, 48.1, 48.0, 48.2, 48.0. Digital balance (g): 49.9, 50.1, 50.0, 55.6, 50.0.

(a) One of the digital balance readings should be excluded before averaging. Identify this reading and calculate the mean of the remaining digital readings. (1)

(b) Identify which balance is more accurate. Justify your answer using values from the table. (2)

(c) Name the type of error affecting the spring balance's results, propose a likely cause, and state how this error could be corrected. (2)

Question 28 (5 marks)

Why this question → 6 of 6, p 0.60

A university proposes a trial in which 100 adult volunteers will apply a newly developed anti-inflammatory skin cream twice daily for eight weeks. The cream has not previously been tested on humans.

(a) Justify TWO recommendations an ethics committee would make before approving this trial. (3)

(b) Name ONE code, declaration or regulation governing research on humans, and describe how it addresses ONE ethical issue raised by this trial. (2)

Question 29 (4 marks)

Why this question → 5 of 6, p 0.53

A bottled-water company publishes the following graph in its advertising.

Data provided in the exam

bar chart titled "Tap water complaints keep rising!" — vertical axis "complaints about tap water quality" starts at 94 (not zero) and ends at 106; bars: 2023: 96; 2024: 97; 2025: 99; a fourth bar labelled "2026 (projected)" drawn at 104 in a different shade. The rise from 96 to 99 is about 3%, but the truncated axis makes the 2025 bar appear roughly twice the height of the 2023 bar.

Analyse how the presentation of this graph misrepresents the data and how this could influence public perception of tap water. Refer to specific features of the graph in your answer. (4)

Question 30 (4 marks)

watch list: conflict-of-interest with a different-industry example

An energy-drink manufacturer funded a university study which concluded that its drink improves concentration in young adults. The funding was declared in the published paper.

Explain why the results of this study should be treated with caution, even though the research was conducted by a university. Support your answer with an example of the influence of commercial interests on science from a DIFFERENT industry. (4)

Question 31 (6 marks)

Why this question → 5 of 6, p 0.64

A company sells the PowerBoost protein bar with the claim:

"Increases muscle mass by 15% in just four weeks."

The company's trial is summarised below.

Data provided in the exam

20 members of one gym ate one PowerBoost bar daily for four weeks. All 20 participants had also begun a new supervised weight-training program at the start of the trial. Mean muscle mass at start: 34.0 kg; mean muscle mass after four weeks: 35.7 kg. There was no group that trained without eating the bar.

(a) Calculate the actual mean percentage increase in muscle mass, and assess whether the trial data supports the company's claim. (2)

(b) Explain why this trial cannot show that the protein bar caused the increase in muscle mass. (2)

(c) Describe ONE change to the design of the trial that would improve its validity, and justify your choice. (2)

Question 32 (7 marks)

Why this question → 5 of 6, p 0.69

A student investigated the claim that playing classical music increases plant growth.

  • Five pot plants were placed on a windowsill with a speaker playing classical music for 6 hours each day.
  • Five pot plants were placed inside a cupboard with no speaker.
  • The pots were of different sizes and contained different amounts of soil.
  • After two weeks, the student measured each plant's height once, by eye, using a ruler.
  • The windowsill plants grew taller, and the student concluded that classical music increases plant growth.

(a) Explain why this investigation is not valid. In your answer, identify TWO specific faults in the design. (2)

(b) Justify THREE modifications to the investigation, explicitly linking each modification to an improvement in validity, reliability or accuracy. (3)

(c) The student suggests that using fifty plants in each group instead of five would fix the investigation. Explain why increasing the number of plants alone would NOT make the student's conclusion valid. (2)

Question 33 (5 marks)

Why this question → 5 of 6, p 0.60

A company claims that its magnetic knee strap reduces knee pain in runners.

Describe how a double-blind, placebo-controlled trial could be used to test this claim. In your answer, specify what would be used as the placebo, who is prevented from knowing the group allocations, and justify how EACH of these features improves the quality of the evidence. (5)

Question 34 (6 marks)

Why this question → 6 of 6, p 0.78

A student investigated the effect of sugar on fermentation by adding different masses of sugar to identical yeast suspensions and measuring the volume of gas produced in 30 minutes at 30 °C.

Data provided in the exam

results table — Mass of sugar (g): 0, 2, 4, 6, 8, 10. Volume of gas produced in 30 min (mL): 0, 18, 33, 44, 75, 52. The trend rises steeply then flattens (0, 18, 33, 44, then about 50 at 8 g and 52 at 10 g); the reading of 75 mL at 8 g is an outlier — the student later noted that the 8 g measuring tube had a loose bung refitted mid-trial.

(a) Graph the results on the grid provided, including a line or curve of best fit. (4)

grid provided; expected graph — mass of sugar on the x-axis with a uniform scale from 0 to 10 g, volume of gas on the y-axis from 0 to 80 mL, all six points plotted, a smooth curve of best fit rising steeply and flattening toward about 50 mL, drawn ignoring the outlier at (8, 75).
Diagram provided in the exam

(b) Describe the relationship shown by your graph, and explain whether the data supports the hypothesis that increasing sugar always increases the volume of gas produced. (2)

Question 35 (7 marks)

Why this question → 6 of 6, p 0.60

The table shows research and development funding in one country by source.

Data provided in the exam

table — funding by source, 2015 vs 2025 (billions of dollars, with share of national total): Government: 2015: \$9.2 b (55% of total); 2025: \$10.1 b (38% of total). Business: 2015: \$6.6 b (40%); 2025: \$15.3 b (57%). Philanthropy and other: 2015: \$0.9 b (5%); 2025: \$1.3 b (5%). Total: 2015: \$16.7 b; 2025: \$26.7 b.

(a) Describe the change in the composition of research funding between 2015 and 2025. Support your answer with values from the table. (2)

(b) Analyse how this change in funding sources could influence the direction of scientific research, the timeframes of research projects, and the reporting of results. Include an overall judgement about whether the change benefits science. (5)

Question 36 (8 marks)

watch list: competing-viewpoints closer (fable 0.50, gpt-5.6-sol 0.48...

Three students discuss the regulation of human gene-editing research.

Student A: "Gene-editing research should be governed by binding international rules. Science crosses borders, so the laws of any single country cannot prevent unethical experiments."

Student B: "Strict regulation just pushes research into countries with no rules. Training scientists in ethics and requiring open publication protect people better than bans."

Student C: "Regulation always lags behind discovery, so institutional ethics committees reviewing each project are the only practical safeguard."

Analyse the extent to which scientific research can be effectively regulated. In your answer, refer to the view of EACH student and to at least ONE named code, declaration or regulatory body. (8)

Answers & marking notes not part of the examination paper — try the paper first

Answers and marking notes

Section I — answer key

Q Answer Note
1 C The discussion evaluates the results, including reliability; results only presents them
2 A Direct sensory/measured record; B is an inference, C a generalisation, D a prediction
3 B van Helmont's willow experiment (his conclusion was reasonable but incomplete — CO₂ was unknown)
4 B Observation → hypothesis → (accidental) culture → self-experimentation
5 B The most serious stated risk is flammability; removing the ignition source addresses it — PPE options address lesser risks
6 B Record one estimated digit beyond the smallest marked division; 36.50 implies unjustified precision
7 C Salinity is deliberately changed; hatched count is the DV
8 C A water-only control group isolates the vitamin as the cause — A improves reliability, B accuracy, D biases the data
9 B Independent expert scrutiny of method and conclusions pre-publication
10 A Fame in one field transfers unearned credibility to another — halo effect
11 B Productivity responds to being observed, not to the changes themselves — Hawthorne effect
12 B In-app, long-term-user, opt-in respondents are self-selected; dissatisfied users have already deleted the app
13 A Technology (X-ray diffraction) supplied the evidence for the scientific model; B reverses it, C and D are science→technology
14 C Wavefronts compressed on approach (higher f), stretched on recession (lower f) — Doppler effect
15 D Boyle's law at constant T: P₁V₁ = P₂V₂; 100 × 60 = P₂ × 20 ⇒ 300 kPa
16 C Informed consent is an ethical safeguard; A, B, D improve methodological quality, not ethics
17 C Unfalsifiability — failures are explained away, so no test could ever refute the claim
18 B Constant offset in one direction = systematic; scatter about the true value = random
19 C A confounding variable (summer beach attendance) can drive both series; A and B assert unsupported causation
20 D B: $90 m > $60 m in dollars, but A: 60/1200 = 5% > B: 90/3000 = 3% of budget

Section II — marking notes

Q21. The linear model runs question → hypothesis → experiment → conclusion. Priestley's investigation began with an unexpected observation (the mint "restoring" the air) made without a prior hypothesis about plants; the hypothesis (plants restore air) was formulated afterwards and then tested with candle and mouse trials — observation preceding hypothesis, and the investigation redirecting as results emerged. 1 mark for the linear model, 1 for the observation-first departure, 1 for linking to his subsequent testing.

Q22. Detection technologies made invisible radiation observable and measurable (e.g. photographic plates revealed emissions from uranium; counters allowed decay to be quantified), providing the evidence that atoms are not indivisible and have an internal structure that can change — advancing the atomic model. That understanding of radiation and decay enabled a named medical technology, e.g. cobalt-60 radiotherapy or technetium-99m diagnostic imaging. 1 mark each: technology → evidence; evidence → refined atomic understanding; understanding → NAMED medical technology. A response that merely names instruments without the causal chain caps at 1.

Q23. Value to the scientific community: the knowledge identifies which plant, which part and which use are worth investigating, saving years of screening — it directed scientists to a fruit with exceptional vitamin C/antioxidant properties (2 marks). Ethical partnership: free, prior and informed consent from the knowledge holders; recognition of Aboriginal Peoples as the source of the knowledge; and negotiated benefit-sharing (economic returns, joint ownership/IP arrangements) rather than uncompensated commercialisation (2 marks). Conflating the scientific community's perspective with Indigenous perspectives — the 2024 flagged fault — caps at 2.

Q24. Pseudoscience presents claims in scientific-sounding language without the method behind them. The data shows non-reproducibility: four practitioners applying the same "diagnostic technique" to the same 20 photographs selected 3, 11, 7 and 16 photographs respectively, with no two selecting the same set — a genuine diagnostic method should give consistent results between trained practitioners. Iridology also proposes no testable mechanism linking iris patterns to kidney function, and its claims are typically not submitted to controlled trial or peer review. Marks: definition-level feature of pseudoscience (1), explicit use of the supplied figures (1–2), reproducibility/testability reasoning (1). Answers that assert "no evidence" without engaging the data cap at 2.

Q25. The graph shows a correlation, not causation: both variables plausibly depend on a third — suburb population (or suburb size). Larger suburbs have more streetlights AND more residents, hence more recorded asthma cases, so the relationship can exist with no causal link between lights and asthma. Full marks require: correlation ≠ causation stated (1), the confounder named (1), and the mechanism through which it produces both trends (1). (Comparing asthma rates per person would remove the confounder — creditable extension.)

Q26. Sample answer: IV — type of watering solution (diluted seaweed extract vs plain water); DV — seedling growth measured quantitatively as height in mm (or change in height) every 2 days for 3 weeks. Take 20 bean seedlings of the same species, age and initial height; randomly assign 10 to receive 50 mL of diluted seaweed extract daily and 10 (the experimental control) to receive 50 mL of plain water daily. Keep constant: light exposure, temperature, soil type and volume, pot size, watering volume and time. Reliability: the 10 seedlings per group act as repeats — average the growth per group; repeat the whole investigation. Marking (checklist): IV and DV with quantitative measurement (1), water-only control group distinct from controlled variables (1), two named controlled variables (1), sample size with random assignment (1), repetition/averaging for reliability (1), logical sequenced method that would actually work (1). "Repeat it" offered as a validity feature earns nothing for the control criterion.

Q27 (a). Exclude 55.6 g (inconsistent with the other four digital readings); mean of remainder = (49.9 + 50.1 + 50.0 + 50.0)/4 = 50.0 g. (b). The digital balance: its mean (50.0 g) equals the true value of the standard mass, while the spring balance's mean (48.04 ≈ 48.0 g) is about 2.0 g below 50.00 g. Accuracy must be judged against the TRUE value, with both means quoted. (c). Systematic error — readings are consistently offset in one direction (all ≈ 2 g low). Likely cause: a zero/calibration error (e.g. the spring balance not zeroed, or a stretched spring). Correction: re-zero/recalibrate against the standard mass (or subtract the constant offset). "Human error" scores zero.

Q28 (a). Any two justified, e.g.: require prior safety evidence (in-vitro/animal or dermal toxicity data) BECAUSE the cream is untested on humans and first exposure carries unknown risk; require informed consent detailing possible side effects and the right to withdraw at any time BECAUSE participants must understand the risk they accept; require monitoring and a stopping rule if adverse reactions occur. 1 mark per recommendation + 1 for justifications tied to THIS trial's risk (untested topical product). Recommendations about sample size or instrument accuracy address validity, not ethics — no marks (the 2024-flagged confusion). (b). E.g. the Declaration of Helsinki — requires that research on humans be preceded by risk assessment, that informed consent be obtained, and that the participant's welfare take precedence over the interests of science; this directly addresses the consent/safety issue in (a). Nuremberg Code or NHMRC National Statement equally creditable. 1 mark for the named instrument, 1 for what it requires (not a narrative of historical breaches).

Q29. Two distorting features must be named from the graph: (1) the truncated vertical axis (starting at 94, not 0) makes a rise of about 3 complaints (~3%) look like a doubling in bar height; (2) the "2026 (projected)" bar presents an invented future value (104) alongside measured data, implying an accelerating trend no measurement supports. Effect: readers conclude tap-water quality is rapidly worsening, steering them toward bottled water — the company's commercial interest. Marks: each feature identified with its distorting mechanism (1 + 1), effect on public perception (1), link to the sponsor's interest / explicit judgement (1). Generic "the graph is misleading" without the named features caps at 1.

Q30. Caution because the funder has a commercial interest in a favourable outcome, which can shape research through question selection, study design, selective reporting or suppression of unfavourable results — even at a university, funding can bias what is published (publication bias), so independent replication is needed. Supporting example from a DIFFERENT industry, e.g. the tobacco industry funding research that downplayed the link between smoking and cancer for decades, or fossil-fuel-funded contrarian climate research. Marks: mechanism of funder influence (2), declared-funding nuance (declaration enables scrutiny but does not remove the bias) (1), specific different-industry example (1). An energy-drink or general food example — same industry — misses the final mark, per the 2025 marking demand.

Q31 (a). Increase = 35.7 − 34.0 = 1.7 kg; percentage = 1.7/34.0 × 100 = 5.0%, well below the claimed 15% — the trial's own data does not support the claim. Both the calculation and the explicit comparison to 15% required. (b). All participants simultaneously began a new supervised weight-training program — a confounding variable that could fully explain the gain; with no control group training WITHOUT the bar, the effect of the bar cannot be separated from the effect of the training. (c). E.g. add a control group drawn from the same gym doing the identical training program but eating no bar (or a placebo bar); any observed difference between groups can then be attributed to the bar — this isolates the IV and makes the design valid. (Random allocation/blinding also creditable if justified via isolating the bar's effect.)

Q32 (a). Not valid because the comparison does not isolate the music: the music group was on a windowsill (light, warmth) while the silent group was in a dark cupboard, so light is a confounding variable changed along with the IV; additionally the pots differed in size and soil volume (a second uncontrolled variable). Two specific faults for 2 marks; "no repeats" is a reliability fault and does not earn validity marks here. (b). Three modifications, each tied to its criterion, e.g.: place BOTH groups in identical light/temperature locations, with the silent group behind a sound barrier — improves VALIDITY (only the music now differs); use identical pots, soil type and volume — improves VALIDITY (removes further confounds) [accept as a second validity fix]; run more plants per group and repeat the investigation, averaging growth — improves RELIABILITY (consistency of results); measure height with a fixed reference point/digital height gauge, same person measuring at the same time of day, several measurements over time — improves ACCURACY (closeness of each measurement to the true height). 1 mark per modification-plus-named-criterion; a modification linked to the wrong criterion earns nothing. (c). Fifty plants would improve reliability (more consistent averages) but the light confound would remain: both larger groups would still differ in light as well as music, so the conclusion "music increases growth" would still not follow. Sample size cannot repair a validity fault — the flagged discriminator.

Q33. Recruit runners with knee pain; randomly allocate to two groups. Treatment group wears the magnetic strap; placebo group wears an identical strap with non-magnetic metal inserts of the same mass and appearance. Neither the participants nor the researchers assessing pain know which strap is which (allocation coded by a third party). Pain measured quantitatively (e.g. standard pain scale before/after 6 weeks). Justifications: the placebo controls for the expectation of relief — any psychological benefit occurs equally in both groups, so a real magnetic effect appears as a between-group difference (2 with specification of the placebo); blinding participants prevents expectation bias in self-reported pain, and blinding researchers prevents unconscious bias in measuring/recording (2 with who-is-blinded specified); random allocation distributes other differences between groups (1). Definitions without deployment in THIS scenario cap at 2.

Q34 (a). Marking: axes correct (mass of sugar on x as IV, gas volume on y) with labels and units (1); uniform scales using most of the grid (1); all six points plotted accurately (1); a smooth CURVE of best fit that rises and flattens near 50 mL, drawn ignoring the outlier at (8, 75) (1). A straight line forced through all points, or a line drawn through the outlier, loses the final mark — the planted trap is double: the trend is non-linear AND one point is anomalous (the noted loose bung justifies exclusion). (b). Gas volume increases steeply with sugar mass at first, then levels off at about 50 mL beyond ~6–8 g. The hypothesis is supported only up to that point: beyond it, added sugar produces little extra gas (the yeast, not the sugar, becomes the limiting factor), so "always increases" is not supported. Trend description quoting values (1); judgement with the plateau reasoning (1).

Q35 (a). Total funding grew from \$16.7 b to \$26.7 b, but its composition reversed: government's share fell from 55% (\$9.2 b) to 38% (\$10.1 b) while business's share rose from 40% (\$6.6 b) to 57% (\$15.3 b) — business is now the dominant funder. Both the share reversal and quoted values required; noting that government funding rose in dollars while falling as a share (the proportion trap) distinguishes top answers. (b). Direction: business funding flows toward research with commercial application, so applied/product-focused research expands while curiosity-driven pure research (reliant on the shrinking government share) is squeezed; direction increasingly set by market priorities rather than scientific or public-interest priorities. Timeframes: commercial funders favour short development cycles and deliverables; long-horizon research (decades-long studies, fundamental physics) depends on stable public funding and becomes harder to sustain. Reporting: commercial funders may delay publication for patents, report favourable results selectively, or impose confidentiality — pressuring research integrity — whereas public funding generally requires open publication. Judgement: e.g. the growth in total funding benefits science's capacity, but the shift concentrates influence in funders with interests in outcomes; a balanced verdict explicitly weighing benefit against risk earns the final mark. Marks: one per developed strand (direction, timeframe, reporting) = 3, use of the stimulus contrast (1), explicit supported judgement (1).

Q36. Top-band response analyses regulation's effectiveness using all three views: Student A — international instruments exist (e.g. the Declaration of Helsinki; UNESCO's declarations on the human genome; the Nuremberg Code's legacy) and set shared norms, but most lack binding enforcement, supporting A's aim while exposing its practical limit (a named instrument required). Student B — regulatory asymmetry is real ("ethics dumping"/research moving to permissive jurisdictions, e.g. the widely condemned 2018 gene-edited-babies case proceeded outside effective oversight), yet B's alternative is weak alone: the same case shows training and publication norms did not prevent the work — they enabled its condemnation afterwards. Student C — institutional ethics committees (in Australia, HRECs operating under the NHMRC National Statement) are the layer that actually reviews each project prospectively, but they inherit their standards from national/international codes, so C's "only practical safeguard" understates the system's nesting. Judgement: effective regulation is multi-layered — international norms + national law + institutional review — no single layer suffices. Marks: engagement with each statement including its weakness (3), named code/regulator accurately deployed (2), accurate supporting example/case (1), coherent overall judgement on the extent of effectiveness (2). Criteria mirror the 2023–2025 convention: a response that ignores one student's statement caps below the top band.

Prediction provenance working — which prediction each part of the paper realises, linked both ways
Paper item Prediction Agreement

Section I Q1

↑ Q1
topic-level consensus — · 2 (opus, grok)

Section I Q2

↑ Q2
topic-level consensus 0.49 (gpt-5.6-sol) · 2 (gpt-5.6-sol, fable)

Section I Q3

↑ Q3
topic-level consensus 0.85 (fable) · 2 (fable, opus)

Section I Q4

↑ Q4
watch list 0.85 (fable), 0.65 (opus) · 2

Section I Q5

↑ Q5
topic-level consensus 0.70 (grok), 0.65 (fable) · 2

Section I Q6

↑ Q6
topic-level consensus 0.65 (grok), 0.65 (opus) · 3 (grok, opus, deepseek)

Section I Q7

↑ Q7
topic-level consensus 0.80 (opus) · 2 (opus, grok)

Section I Q8

↑ Q8
topic-level consensus 0.76 (gpt-5.6-sol) · 2 (gpt-5.6-sol, opus)

Section I Q9

↑ Q9
topic-level consensus 0.70 (gemini-3.1-pro) · 2 (gemini-3.1-pro, opus)

Section I Q10

↑ Q10
watch list 0.75 (grok) · 3 (grok, fable, opus)

Section I Q11

↑ Q11
watch list 0.35 (deepseek) · 1

Section I Q12

↑ Q12
topic-level consensus 0.60 (fable), 0.67 (gpt-5.6-sol) · 3

Section I Q13

↑ Q13
watch list 0.85 (fable), 0.55 (gpt-5.6-sol) · 3 (fable, gpt-5.6-sol, opus)

Section I Q14

↑ Q14
watch list 0.35–0.60 · 2 (fable, grok)

Section I Q15

↑ Q15
watch list 0.80 (fable) · 2 (fable, opus)

Section I Q16

↑ Q16
topic-level consensus 0.57 (gpt-5.6-sol), 0.55 (opus) · 2

Section I Q17

↑ Q17
watch list 0.58 (gpt-5.6-sol) · 2 (gpt-5.6-sol, fable)

Section I Q18

↑ Q18
topic-level consensus 0.70 (fable) · 3 (fable, opus, deepseek)

Section I Q19

↑ Q19
topic-level consensus 0.50 (deepseek) · 2 (deepseek, gpt-5.6-sol)

Section I Q20

↑ Q20
topic-level consensus 0.52 (gpt-5.6-sol) · 2 (gpt-5.6-sol, opus)

Q21

↑ Q21
is-q9-named-scientist-linear-model 0.59 · 6 (all)

Q22

↑ Q22
watch list 0.52 panel mean; 0.35 (opus bold) · 6 (all)

Q23

↑ Q23
is-q10-bioharvesting 0.59 · 5 (fable, opus, grok, gemini-3.1-pro, deepseek)

Q24

↑ Q24
watch list 0.35–0.60 · 4 (fable, gpt-5.6-sol, gemini-3.1-pro, deepseek)

Q25

↑ Q25
is-q11-correlation-causation 0.58 · 5 (fable, opus, gpt-5.6-sol, gemini-3.1-pro, deepseek)

Q26

↑ Q26
is-q2-write-valid-method 0.70 · 4 (fable, opus, grok, deepseek)

Q27

↑ Q27
is-q5-instrument-error-comparison 0.62 · 5 (fable, opus, gpt-5.6-sol, grok, deepseek)

Q28

↑ Q28
is-q8-regulation-ethics 0.60 · 6 (all)

Q29

↑ Q29
is-q12-media-misrepresentation 0.53 · 5 (fable, opus, gpt-5.6-sol, grok, deepseek)

Q30

↑ Q30
watch list 0.70 (fable), 0.55 (opus) · 2 (grok rests it)

Q31

↑ Q31
is-q4-claims-numeric-chain 0.64 · 5 (fable, opus, gpt-5.6-sol, grok, deepseek)

Q32

↑ Q32
is-q3-flawed-design-critique 0.69 · 5 (fable, opus, gpt-5.6-sol, gemini-3.1-pro, deepseek)

Q33

↑ Q33
is-q6-placebo-double-blind 0.60 · 5 (fable, opus, gemini-3.1-pro, grok, deepseek)

Q34

↑ Q34
is-q1-graph-construction-trap 0.78 · 6 (all)

Q35

↑ Q35
is-q7-funding-influence 0.60 · 6 (all)

Q36

↑ Q36
watch list 0.42–0.50 · 3 (fable, opus, gpt-5.6-sol)

Intu AI

One paper isn't enough? Generate more

Intu AI builds unlimited practice questions for Investigating Science in these styles, marks your working, and explains what you missed — aligned to your syllabus.

Practise with Intu AI →

Share & save

Practice paper (PDF)

Every style card and paper question has its own link — hover any card and use its copy-link icon to share exactly the thing you mean. Printing this page gives a clean copy too.

Published Aug 2026, before the exams. In November 2026 we score these predictions publicly against the real paper — per-model calibration and question-level hit rates, the same harness as the 2025 backtest. How we did it.