October 5, 2026
|
min read
Five takeaways from IAEA 2026 on trust, transparency, and technology
The 51st IAEA Annual Conference in Toronto explored trust, transparency, and technology in educational assessment. Here are five takeaways on AI and valid evidence, decision transparency, proportionate monitoring, and designing assessment outcomes that hold up to scrutiny.

The 51st IAEA Annual Conference brought the international assessment community to Toronto from September 27 to October 2, 2026, under the theme Trust, Transparency, and Technology in Educational Assessment. Across the keynotes and sessions, one question kept coming back: as AI changes how assessments are written, taken and scored, what makes an outcome worth trusting?
I had the privilege of leading three sessions during the week, on decision transparency, the limits of monitoring, and designing outcomes that hold up to scrutiny. These are the five takeaways I brought home from Toronto.
1. AI is reshaping the foundations of assessment, not just the tools
The keynote program treated AI as a question about the fundamentals of measurement rather than a feature to bolt onto existing systems. Dr. Eunice Eunhee Jang of OISE at the University of Toronto opened the conference with "Reimagining Assessment in the AI Era: From Assessment Triangle to Assessment Ecology." On Thursday, Dr. Kadriye Ercikan of ETS returned to the same ground with "Measurement at a Crossroads: AI, the Foundations of Assessment, and the Future of Trust."
The two keynotes set the tone for the week. When AI can write a response, assist a candidate and score the result, the question of what counts as valid evidence has to be answered again. For programs that issue results people rely on, that question is practical as well as academic, because it shapes how every outcome will be judged when someone asks how it was reached.
2. Transparency has to reach the decision, not stop at the rubric
Most programs are already transparent about design. They publish criteria, rubrics and scoring guides so candidates know what is expected. Far fewer are transparent about decisions: why a session was flagged, who reviewed it, what evidence was weighed and how the outcome was explained to the person it affected.
In "Making Assessment Decisions Visible," I walked through the assessment lifecycle stage by stage, showing where each step either creates a record or loses one. I then set out three tests every assessment decision should be able to survive:
- Evidence: is there a clear record of what was observed and what it showed?
- Explanation: can the program describe how that evidence led to the outcome?
- Consistency: would the same evidence lead to the same outcome for any other candidate?
The principle underneath all three is one I come back to often. The more automation a process uses, the more visible its reasoning needs to be, because automated steps are the ones that are hardest to explain after the fact.

3. More monitoring is not more integrity
Faced with AI-enabled misconduct, many programs have responded by capturing more. That has meant longer recordings, more signals and more automated detection. I understand the instinct, but in my experience it can quietly create new risk.
In "The False Sense of Security," I examined four documented failure modes of that approach: scale, bias, inconsistency and defensibility. More data can produce more false positives, and more false positives mean more decisions that are difficult to explain to the candidates they affect. A program can end up with a larger archive of footage and a weaker position when an outcome is challenged.
The alternative I proposed is proportionality. Oversight should match the stakes of the exam, reviewers should weigh evidence rather than raw signals, and human judgment should stay in the system wherever an outcome carries real consequences. Less intrusion also tends to mean less test-taker anxiety, which supports trust in the program from the candidate's side as well.
4. A flag is not a finding
Automated signals are useful for directing attention to moments that may need a closer look, but on their own they cannot tell a program whether a violation occurred. A flag might reflect a technical glitch, a nervous habit or a genuine integrity concern, and only someone with context can tell the difference.
Trust comes from what happens after the flag. That means a trained reviewer examining the evidence, consistent criteria applied across every case, and the program, rather than the software, making the final call. This matters even more as AI-assisted misconduct becomes harder to see, since some of it produces no visible behavioural signal for an algorithm to catch at all.
This idea ran through all three of my sessions. Organizations that blur the line between a flag and a finding put both their results and their candidates at risk. A wrongful outcome based on an unreviewed flag is difficult to defend, and it can damage confidence in every other result the program issues.
5. Defensibility is designed in before, during and after the exam
I opened "Designing for Defensibility" by asking the room how good they are at spotting AI-generated fakes. With synthetic video and identity spoofing now within easy reach, security can no longer rest on what a proctor happens to see in the moment.
Defensible outcomes come from a connected chain rather than any single control:
- Before the exam: the candidate's identity is verified, so the program knows who is sitting the assessment.
- During the exam: evidence is collected in proportion to the stakes, so oversight is fair and focused.
- After the exam: review is documented and consistent, so every outcome can be explained if it is challenged.
This connected thinking is the basis of our Credential Security Trifecta, a framework that links the learning management system, proctoring and identity verification, and credential issuance. Through the integration between Integrity Advocate and Accredible, programs can carry that chain from the assessment itself through to the credential awarded, with one auditable record connecting the two.

Building assessment systems people can trust
The conference theme named three things, and I left Toronto more convinced than ever that they only work in combination. Technology can extend what a program is able to see, but trust depends on whether the decisions behind each outcome are transparent, consistent and grounded in evidence a person has reviewed.
If you're rethinking your program's approach after Toronto, a useful starting point is to trace a single result from registration to final outcome and ask where the record is strong and where it could be lost. That exercise tends to show quickly whether a system is built to be defensible or simply built to collect more.
Thank you to everyone who joined my sessions and to the IAEA team for a thoughtful week. If you'd like to talk through how these ideas apply to your own program, my team and I would be glad to help.

{{post-cta}}
Foire aux questions
Trouvez les réponses aux questions les plus fréquemment posées par nos clients.





