Academic OS · The Architecture · Evidence by design
Compliance as a by-product of daily work.
Evidence by design means the artefacts that prove you meet a standard are captured where the work happens and tagged to the criterion they prove as they land, instead of being collected in a project before a review. Evidence stops being something the institution gathers and becomes something it emits.
- A faculty member uploads a marking scheme inside their own course
- A committee records the decision it just took
- A data owner publishes the source dataset behind a metric
- The standard, KPI, course and term are proposed from a controlled vocabulary
- A person confirms the tags
- Coverage by criterion, current on any given Tuesday
- Every number bound to the document that proves it
Two words, two institutions
Collecting evidence and capturing it are not the same thing.
Almost all accreditation effort goes into the first, and that is the mistake. Your people were never the problem: they produce evidence constantly. It is simply created away from the criterion it proves, by someone who has no reason to file it against a standard.
Going to find it after the fact.
Someone asks for assessment samples, marking schemes and minutes by Thursday. People dig through personal drives, rename files, and reconstruct from memory what happened eighteen months ago.
A project you run before every review. When it ends, what you held together by hand falls apart again.
It lands where it belongs, when it happens.
The artefact enters the system at the moment of the work, already tied to what it proves. A marking scheme, an approved programme change, an attainment review with an action recorded against it.
A property of how the institution operates every day, not a state you prepare for.
Not a folder problem
The scramble has three roots, and none of them is a better folder structure.
Each root is architectural, which is why process discipline alone has never fixed it. Left column: what actually goes wrong. Right column: what the operating model has to do about it.
Evidence lives where work happens, not where it is needed.
The marking scheme is in a faculty member's course folder. The minutes are in someone's inbox. None of it is wrong. It is personal, scattered, and invisible to quality until someone asks.
Capture belongs inside the workflow the producer already uses.
The person closest to the work is the only one who can capture it cheaply and correctly, so the upload has to live in their course or their committee decision, not in a quality inbox.
Nothing connects the evidence to the standard it proves.
A document can sit in a perfectly tidy library and still be useless, because no one has answered the only question that matters at a review: which criterion does this satisfy? Without that link you have files, not evidence.
A controlled vocabulary, applied as the artefact enters.
Every document is bound to the standard, KPI, course and term from one canonical tag library, at the moment it lands. Manual tagging from a blank field at scale never happens, so the system has to propose and a person confirm.
The number and its proof drift apart.
KPIs are calculated in spreadsheets, by hand, on a schedule. The supporting documents are gathered separately, also by hand. By the time both reach the self-study, no one can be certain the figure and the artefact describe the same thing.
One engine computes the value; documents reference it.
KPI values are computed deterministically in one place and referenced by every report and evidence record, never recomputed inside them. That is what closes the gap a sharp reviewer probes.
Put those together and you get the real definition of audit risk. It is not do we have the evidence, because you almost always do. It is who owns this evidence, what does it prove, and can we retrieve it right now without calling one specific person. If the honest answer involves a name, that is a dependency, not a system.
Designed into the work
Four choices, built into normal work rather than into a project.
These are the four decisions that turn evidence from something you gather into something your institution emits. Each one is a property of the operating model, not a task on a checklist.
Capture at the source.
Evidence enters where the work is done: a faculty member uploading a marking scheme inside their own course, a coordinator attaching minutes to the decision they just made.
Not months later, in response to an email.
Tag to the criterion as it lands.
The moment a document enters, it is classified against the standard it supports. The system proposes the tags and a person confirms them.
Assisted, not automated. Confirming a suggestion is the difference between a tagging policy that exists on paper and one that runs.
Bind the number to its proof.
A KPI value and the evidence behind it travel together instead of being assembled in separate workstreams.
"Does the number match the document?" stops being something you hope is true.
Make coverage visible continuously.
Once evidence is captured and tagged, coverage by criterion is a live view rather than a spreadsheet rebuilt before each visit.
You see which standards are thin long before a reviewer sees it for you.
The rules underneath
What the architecture guarantees, whatever the framework.
Boundaries
What evidence by design is not.
The principle is deliberately narrow. Being clear about the edges is what keeps the guarantees above credible.
Where this goes deeper
The principle, and the product that implements it.
Questions about the principle
Evidence by design, answered.
What does evidence by design mean?
Evidence by design means the artefacts that prove you meet a standard are captured where the work happens and tagged to the criterion they prove as they land, instead of being collected in a project before a review. Compliance becomes a by-product of daily work rather than a separate exercise.
What is the difference between collecting and capturing evidence?
Collecting evidence means going to find it after the fact, which is a project you run before every review. Capturing it means it lands in the right place, tied to what it proves, at the moment the work happens, which is a property of how the institution operates every day.
Does AI decide what counts as evidence?
No. AI proposes the tags, the standard, KPI, course and term, from a controlled vocabulary, and a person confirms them. AI never marks evidence approved and never authors a KPI number. A human always owns the truth.
Does the evidence library calculate our KPIs?
No. A deterministic KPI engine computes the values. Evidence and authored documents reference those values rather than recomputing them, which is what keeps a figure in a report and the document behind it describing the same thing.
Can external reviewers see our evidence?
Yes, within limits. Access is enforced server-side by role, permission profile and org-unit scope, and external or visiting reviewers get scoped, time-bound, view-only access to exactly the sample they are meant to see.
Compliance as a by-product
Bring one criterion you had to scramble for.
We will walk through where each artefact would have been captured, what it would have been tagged to, and what your coverage on that criterion would look like today.
