Year 12 · 150 minutes · Evaluation
How can we tell which part of a system actually helped?
An ablation removes or disables a system component while holding the evaluation suite fixed. This sandbox runs deterministic retrieval, filtering and review rules over structured synthetic cases. Paired results show which cases change when a component is removed. Components can interact: a filter may matter only when retrieval brings in hostile content. A score increase from an uncontrolled before/after comparison cannot identify a cause.
Plan three 50-minute sessions. Read the fixed synthetic suite including benign, missing-evidence and hostile-source cases. Explain the exact deterministic rules and keep expected outcomes sealed until predictions are recorded.
Keep cases paired across treatments and interpret variation. Useful earlier investigations: y3-holdout, y9-calibration Use pairs for investigation, with operator/reviewer swaps after each comparison. Keep individual predictions, journals and a short oral defence so group work does not hide understanding.
Queensland Digital Solutions 2025 v1.4: Unit 3 · objective 3, Unit 3 · objective 7, Unit 3 · objective 8. Selected aspects only. This activity contributes evidence; it does not cover the full descriptor or achievement standard. A programming descriptor is not claimed for merely moving controls. ACARA AI curriculum connection · V9 Technologies These are planning connections, not ACARA endorsement or exhaustive descriptor alignment.
Compare two fictional release reports that changed model, data and filter together.
Ask: “Which change caused the improvement?”
Listen for: “We cannot isolate it because several things changed.”
Predict which cases fail when retrieval alone is removed from the full system.
Ask: “Which cases depend on retrieved evidence?”
Listen for: “Those whose answer facts are only in a source.”
Toggle one component, keep fixtures fixed and inspect per-case transitions as well as aggregate pass counts. Save each configuration.
Ask: “Which exact test changed and why?”
Listen for: “Its required evidence or protection was removed.”
Compare filtering with retrieval on and off. Trace why filter benefit depends on hostile content reaching the system.
Ask: “Is the filter’s effect a single universal number?”
Listen for: “No, it depends on the rest of the pipeline.”
Run a small factorial comparison, declare hypotheses and document paired outcomes. Include a benign case where review imposes cost without changing correctness and a held-out challenge authored by peers.
Ask: “What remains uncertain after this fixed suite?”
Listen for: “Whether the cases represent real use and whether the effect transfers.”
Submit configuration, case-level differences, aggregate results and a narrow causal claim.
Ask: “What would invalidate your attribution?”
Listen for: “A changed dataset, seed, scoring rule or unrecorded component.”
A better score proves which feature helped.
Disable retrieval alone. Predict which fixed cases change before running the paired comparison.
The benefit of filtering depends on whether retrieval introduces hostile content; component effects are not simply additive.
Run a controlled factorial study and support one causal claim with matched case transitions and stated external-validity limits.
Require a record of fixed variables and one interaction explanation using case traces. Do not accept only bar-chart differences.
Start with two components and four cases; provide a configuration matrix before expanding the full suite.
Add repeated stochastic runs in the hardware pathway and report paired uncertainty intervals rather than one sample’s difference.
An ablation study with configuration matrix, paired traces and a defensible causal claim.
All cases and actions are synthetic dry runs. No real users are subjected to experimental treatments. Senior syllabus alignment requires local verification.
Evaluate all eight retrieval/filter/review configurations over a repeated synthetic case suite on MPS/CUDA. Change case-family diversity and compare paired counts and component dependencies; repeats do not add independent evidence.
| Criterion | Beginning | Secure | Extending |
|---|---|---|---|
| Experimental control | Compares changing systems and tests | Changes one component on a fixed suite | Uses factorial comparisons to expose interactions |
| Causal claim | Attributes any gain to the latest edit | Supports attribution with paired cases | Limits generalisation and tests a peer-authored challenge |
Queensland Digital Solutions 2025 v1.4
References: Unit 3 · objective 3, Unit 3 · objective 7, Unit 3 · objective 8. Read the current source (checked 2026-09-07).
Evidence to assess: Paired test analysis, treatment interaction and justified refinement.
Selected aspects only. This activity contributes evidence; it does not cover the full descriptor or achievement standard. A programming descriptor is not claimed for merely moving controls. A supporting classroom task, not a QCAA-approved assessment instrument. Teachers must set their own assessment conditions and confirm alignment with their course. Other jurisdictions require local mapping.
These are planning estimates to test with your class. A short session develops one supported claim; it does not compress the whole senior project.
| Stage | 45 minute focus | 60 minute investigation |
|---|---|---|
| Readiness and prediction | 0–5 | 0–5 |
| Trace the supplied example | 5–13 | 5–15 |
| Author and run cases | 13–25 | 15–35 |
| Counterexample and redesign | 25–35 | 35–45 |
| Explain and discuss | 35–42 | 45–55 |
| Export and handover | 42–45 | 55–60 |
For a longer project, use three 50-minute sessions. Session 1 (0–50): readiness, model, hypothesis and initial cases. Export a project and record the next test. Session 2 (50–100): reopen, check settings, author counterexamples and revise the design. Export the changed project and identify unresolved evidence. Session 3 (100–150): independent peer test, final artefact, individual explanation and moderation. If using two 60-minute sessions, stop at minute 60 after saving the first comparison; use 60–120 for redesign, independent test and defence.
Entry check: Keep cases paired across treatments and interpret variation. Ask the learner to demonstrate it before choosing the level of support.
Preparation: allow about 25 minutes to run the starter, print the cards and check a project can be reopened. This estimate has not yet been measured in a classroom pilot.
Read the entry question aloud, model one row, and label the units. Offer the case table as a large-print sheet. Keep mathematical derivations optional until the learner can explain the comparison.
For one device, use a projector: one pair predicts, one operates, and the class records on paper. Swap roles after the first comparison. For individual access, support keyboard controls and a written table equivalent to each visual. Learners may explain orally or with an annotated diagram. Never require personal data, a recorded voice, or a photograph.
Mixed readiness: if the entry check is difficult, use the linked prerequisite and the first two case cards; retain the same central question. If secure, ask the learner to design an unseen test and state which explanation it could disprove.
Keep each case paired. Compute retrieval and review effects plus interaction = both − retrieval − review + baseline.
Starting parameters: The supplied cases define the inputs.
Mean paired combined effect 0.3425; standard error 0.094725. Sampling interpretation requires independent representative cases.
| case | baseline | retrieval effect | review effect | interaction | combined effect |
|---|---|---|---|---|---|
| A | 0.4 | 0.2 | 0.1 | 0.2 | 0.5 |
| B | 0.5 | 0.2 | 0.1 | -0.05 | 0.25 |
| C | 0.2 | 0.1 | 0.2 | 0.2 | 0.5 |
| D | 0.7 | 0.05 | 0.1 | -0.03 | 0.12 |
These are authored examples, not work collected from children. Assess reasoning using the lesson rubric, not whether the first prediction was correct.
Beginning: “It worked because the result looks right.” This identifies no exact case, control or measurement. Ask the learner to point to one row and say what happened.
Developing: “In the first case I recorded case: A; baseline: 0.4; retrieval effect: 0.2; review effect: 0.1; interaction: 0.2; combined effect: 0.5.” This cites evidence, but does not yet explain how the result follows from the rule. Ask the learner to trace the relevant step.
Secure: “For the first supplied case, case: A; baseline: 0.4; retrieval effect: 0.2; review effect: 0.1; interaction: 0.2; combined effect: 0.5. I can trace it using this mechanism: Keep each case paired. Compute retrieval and review effects plus interaction = both − retrieval − review + baseline. My result supports a claim about these supplied cases. It does not establish that the same result holds outside them.” Look for an accurate trace, the actual settings and a bounded claim; accept equivalent oral or visual evidence.
Extending: The learner constructs and reruns a new case, reports whether the first explanation survives, and defends a revised design. Use this concrete challenge: Create a factorial comparison with identical cases across configurations. Report interactions and uncertainty across cases. Require the original and changed evidence and this boundary: Case variation is not automatically independent sampling. Identify the population before interpreting uncertainty.
Moderation: first assess independently against each lesson criterion. Compare the exact trace or artefact that led to your judgement. Resolve differences using evidence, not polished language. Keep each learner's individual explanation even when the artefact was produced in a group.