Questions at several difficulty levels, straight from your own material.
New test variants do not get made because it is extra work, so the same ones circulate for years. The audit checks the quality of your material, how to guarantee a question really has one right answer, and how much reviewing that costs.
AI Workflow
AI workflow with human review gates can automate 70-80% of routine question generation while preserving quality control for exceptions.
The recommended path is an AI workflow with structured human review checkpoints. The process shows clear rule patterns for 85-90% of cases, high volume justifying automation investment, and well-defined quality gates that map naturally to workflow approval steps. Given the locked-down school IT environment and non-technical staff, a managed AI workflow platform requires minimal in-house development while handling the compliance and quality assurance requirements. The 10-15% exception rate is manageable through human-in-the-loop review stages, and the 500 annual hours of manual effort creates strong ROI even with conservative automation rates.
No specific tech stack preference was stated. The recommendation accounts for the centrally managed school IT constraint and non-technical staff by favoring a managed platform approach over custom development.
Process Overview
The school currently generates test questions from existing teaching material in a manual process that runs roughly daily. Teaching staff select source material, create questions at multiple difficulty levels, verify that each question has exactly one correct answer, and perform quality review before the questions can be used. The work involves judgment calls about difficulty assignment, identifying ambiguous answer options, and handling material that does not fit standard patterns. Around ten to fifteen percent of cases require special handling, either because the source material is unusual or because earlier data entry errors need correction. The process consumes about twenty minutes per run and happens fifteen hundred times per year, totaling five hundred hours of staff time annually. Right now everything happens through email and spreadsheets, with no dedicated system in place.
The volume is high enough that the manual work has become a bottleneck. Staff avoid creating new test variants because the effort is tedious, and process knowledge sits with just a few people. The school operates under GDPR with extra care for data on minors, and institutional policy requires clear accountability for assessment decisions. The IT environment is centrally managed and locked down, with small annual budgets and limited room for custom development or complex integrations.
Path Scores
This process has clear stages with defined quality gates, high volume to justify investment, and a manageable exception rate that maps perfectly to approval nodes. The compliance requirements and need for audit trails are native strengths of workflow platforms. Non-technical staff can operate review queues without coding skills, and the locked-down IT environment favors managed SaaS over custom builds.
Very similar to AI workflow but implies heavier custom integration work. Given the school IT constraints and small budgets, a pure workflow platform with built-in AI capabilities is more practical than building hybrid orchestration from scratch. This would be the choice if existing systems required deep integration, but the current state is email and spreadsheets.
Autonomous agents are poorly suited to educational assessment where institutional policy and GDPR compliance demand clear human accountability. The 10-15% exception rate and need for subject matter judgment mean full autonomy would either fail on edge cases or require extensive training data the client does not have. The risk profile is too high for the compliance context.
RPA excels at clicking through structured UI workflows in legacy systems. This process has no such systems, it lives in email and spreadsheets and human judgment. The core challenge is generating and validating questions, not moving data between screens. RPA would automate the wrong 10% of the work.
Building a custom rules engine for question generation would require significant development effort and ongoing maintenance that the school cannot support. Staff are non-technical, budgets are small and annual, and the locked-down IT environment makes deployment and updates painful. The variability in teaching material and exception handling would demand constant code changes.
The current manual process consumes 500 hours per year on repetitive work, discourages creation of new test variants, and concentrates process knowledge in two or three people. The high volume and clear pattern in 85-90% of cases make this an obvious automation candidate. Staying manual wastes teaching staff time that should go to students.
Process Dimensions
Eight dimensions drive the recommendation, scored 0–10 with a note on each.
Teaching material and question formats have recognizable structure, though content varies by subject and the 10-15% exceptions show some ambiguity.
The standard case is described as fine with clear patterns in 85-90% of cases, and quality rules like one correct answer are explicit.
Ten to fifteen percent do not fit the standard pattern, which is manageable but significant enough to require designed exception handling.
School IT is locked down and centrally managed with no APIs mentioned, current state is email and spreadsheets, and budgets are small.
1,500 runs per year at 20 minutes each equals 500 hours annually, creating strong economic justification for automation investment.
Educational assessment policies and GDPR requirements are stable, but teaching material and curriculum evolve over academic cycles.
Exceptions requiring judgment consume most of the time, and institutional policy on assessment decisions implies human accountability is required.
GDPR with extra care for minors plus institutional policy on assessment decisions demand audit trails and clear accountability for automated decisions.
ROI Estimate
€12,500
Current annual cost
70%
Estimated time saved
€8,750
Annual savings
21mo
Payback period
Current cost is 1,500 runs per year times 20 minutes per run divided by 60 times 25 EUR per hour, totaling 12,500 EUR annually. Seventy percent automation saves 8,750 EUR per year. Build cost assumes a managed AI workflow platform with configuration and pilot work, not full custom development. Payback is roughly two years at the high end, faster if internal IT can contribute setup effort.
Implementation Roadmap
Select one teaching subject with the clearest material structure and build the AI question generation workflow for that domain only. Validate quality gates and exception routing with 50-100 real cases. This proves the concept and trains staff on the review interface without risking the full operation.
Configure workflow logging to meet GDPR and institutional policy requirements, including who reviewed what and when, what decisions were made, and retention policies for minor data. Get sign-off from data protection officer and school administration before wider rollout.
Roll out the workflow to remaining teaching subjects, tuning the AI prompts and quality rules for each domain. Train all twenty teaching staff on the review queue interface and exception handling procedures. Run parallel with manual process for one month to catch gaps.
Analyze the first three months of production data to identify which exceptions can be resolved with better prompts or additional rules versus those that truly need human judgment. Refine the routing logic to reduce false positives sent to review queues.
Risks & Considerations
The biggest risk is that teaching material variability across subjects is higher than the 10-15% exception estimate suggests, which would overload review queues and erode ROI. If the AI generates questions with subtle errors that pass initial review but surface later in student use, it damages institutional credibility and creates rework. Human review must include subject matter experts who can catch domain-specific mistakes, not just process checkers. The locked-down school IT environment may block deployment of a SaaS workflow platform or restrict API access, forcing a more manual integration that increases cost and fragility. Finally, if staff perceive this as a precursor to headcount reduction rather than workload relief, adoption will fail regardless of technical success. Clear communication that the goal is freeing time for students, not eliminating jobs, is essential.
Architecture Overview
Hover to zoom · click for fullscreen
Why This Approach
The recommended path is an AI workflow with structured human review checkpoints. This approach automates the routine question generation and initial validation steps while routing exceptions and edge cases to teaching staff for review. The process has clear stages that map naturally to workflow nodes: material intake, AI-driven question generation at specified difficulty levels, automated checks for answer validity, human review queues for flagged items, and final approval before release. The eighty-five to ninety percent of cases that follow standard patterns can flow straight through with minimal human touch, while the ten to fifteen percent exceptions get caught at defined gates and sent to subject matter experts.
This fits the school context better than the alternatives for several reasons. The centrally managed IT environment and small budgets make a managed SaaS workflow platform far more practical than building custom code or orchestrating hybrid integrations. Non-technical teaching staff can operate review queues through a simple web interface without needing to write scripts or maintain automation logic. The compliance requirements around GDPR and institutional assessment policy demand audit trails showing who approved what and when, which are native features of workflow platforms. The high volume of fifteen hundred runs per year at twenty minutes each creates strong economic justification for the platform investment, with payback expected inside two years even at conservative automation rates.
An autonomous AI agent scored lower because educational assessment demands clear human accountability, especially under GDPR and school policy. Letting an agent make final decisions on test questions without review gates would concentrate too much risk in a system that cannot explain its reasoning or catch its own domain-specific mistakes. The ten to fifteen percent exception rate is manageable in a workflow with approval nodes, but would either break an autonomous agent or require training data the school does not have. RPA and traditional code are poor fits because there are no legacy UI systems to automate and no in-house development capacity to maintain a custom rules engine. The teaching material varies too much for rigid scripting, and the locked-down IT environment makes deployment and updates painful.
The workflow approach also scales naturally as the school refines its automation over time. Early pilots can start with a single subject area where material structure is clearest, prove the concept with real cases, and train staff on the review interface before rolling out to all domains. If the exception rate turns out higher than expected in certain subjects, the workflow can route more cases to human review without rebuilding the entire system. If the AI improves and exception rates drop, review gates can be relaxed gradually. This flexibility is critical in an educational environment where curriculum and assessment policy evolve on academic cycles, and where trust in the system must be earned through demonstrated reliability rather than assumed from the start.
Comparing the Top Approaches
The recommended path is AI Workflow, which scored nine out of ten against eight for Hybrid and five for AI Agent. The difference between AI Workflow and Hybrid is narrow in principle but significant in practice for this context. Both approaches use AI for question generation with structured human review, but AI Workflow implies a managed platform like Make, Zapier, or n8n with native AI integrations and approval nodes, while Hybrid suggests custom orchestration code tying together separate AI services and review interfaces. Given that school IT is locked down with small annual budgets and staff are non-technical, the managed platform approach wins on deployment speed, maintenance burden, and operational simplicity. The current state is email and spreadsheets, so there are no legacy systems demanding deep integration that would justify custom development.
AI Agent scored much lower because autonomous operation is a poor fit for educational assessment. The process has a ten to fifteen percent exception rate requiring subject matter judgment, institutional policy demands human accountability for assessment decisions, and GDPR compliance with minors adds regulatory weight. An autonomous agent would either fail on edge cases or require extensive training data the school does not have. The compliance risk profile is too high. RPA and Traditional Code both scored poorly because they automate the wrong parts of this process. RPA excels at clicking through structured legacy UIs, but this process has no such systems. Traditional Code would mean building a custom rules engine that non-technical staff cannot maintain and that would require constant updates as curriculum and teaching material evolve. The 500 annual hours of manual effort justify investment, but not in brittle custom code the school cannot support long term.
How to Build It
The implementation centers on a managed AI workflow platform like Make, Zapier, or n8n, integrated with an LLM provider such as OpenAI GPT-4 or Anthropic Claude via API. The workflow starts when teaching staff upload source material through a simple web form or shared folder monitored by the platform. The trigger fires a workflow that sends the material to the LLM with a structured prompt specifying the number of questions needed and the difficulty levels required. The LLM returns a JSON payload of candidate questions, each tagged with difficulty level and proposed correct answer.
The workflow then routes each question through an automated validation gate that checks for required fields, flags questions with multiple or zero correct answers, and applies basic quality rules like minimum answer length or format compliance. Questions that pass validation move to a batch review queue in a tool like Airtable, Notion, or a custom interface built on Softr or Retool. Teaching staff see a clean review interface showing the source material, the generated question, and the validation results. They approve, reject, or edit each question with a single click or quick text edit. Questions flagged as exceptions during validation go to a separate queue for subject matter experts.
Once a question is approved, the workflow writes it to a structured repository, likely a Google Sheet or Airtable base given the school IT constraints, tagged with metadata like subject, difficulty, source material reference, reviewer name, and timestamp for audit purposes. The final output can be exported in whatever format the school uses for test assembly. Exception handling is built into the workflow as conditional branches. If the LLM returns an error or the validation gate detects ambiguity, the workflow routes that item to the exception queue and logs the issue. If a reviewer rejects a question, the workflow can optionally send it back to the LLM with the rejection reason for a second attempt, or simply discard it and log the failure for later analysis. This closed loop feedback improves quality over time as common rejection patterns inform prompt tuning.
Risks in Detail
The biggest risk is that teaching material variability across subjects is higher than the ten to fifteen percent exception estimate suggests. If half the generated questions require human rework rather than quick approval, the review queues become a bottleneck and the seventy percent automation savings evaporate. The ROI assumes that the AI produces usable first drafts most of the time, and if that assumption is wrong the project delivers expensive workflow overhead with minimal time savings. A robust pilot in one subject area is essential to validate this before full rollout. The second major risk is subtle content errors that pass initial review but surface later in actual student use. If the AI generates questions with domain-specific mistakes that only become obvious in context, or if reviewers approve questions too quickly without deep subject matter checking, the school ends up using flawed assessments that damage institutional credibility and require costly rework. Human review must include qualified subject matter experts, not just process checkers looking for formatting compliance.
The locked-down school IT environment creates deployment and integration risk. If central IT blocks the workflow platform because it is external SaaS, or if API access to LLM providers is restricted by firewall policy, the technical approach may not be feasible without lengthy approval processes that blow the timeline and budget. Early engagement with central IT and the data protection officer is critical to surface these blockers before significant work begins. Finally, there is organizational risk if staff perceive this automation as a threat to job security rather than workload relief. If teaching staff believe the goal is headcount reduction, adoption will fail regardless of how well the technology works. Clear communication from school administration that the goal is freeing staff time for student-facing work, not eliminating positions, is essential to successful change management. Involving teachers in the pilot design and review queue configuration helps build ownership and trust.
Claude Code Starter
A scaffolded project ready to open in Claude Code. Unzip, open the folder, and Claude starts building immediately.
Claude Code Starter (.zip)
Your own assessment includes a ready-to-use project scaffold: CLAUDE.md, pyproject.toml, src/agent.py and .env.example. Open the folder in Claude Code and it starts building.