Validating AI Tools in ISO/IEC 17025 Calibration Labs

AI tools have moved out of the lab bench and into daily use in accredited calibration work. The uses are down to earth. They spot uncertainty contributors, flag odd patterns in calibration history, review draft paperwork, and plan recalls. Each of those outputs shapes a calibration decision, directly or not. Under ISO/IEC 17025:2017, that puts a set of rules in play.

This article walks through how to validate AI software in an accredited lab. The stance is careful on purpose. AI is a tool, not a stand-in for metrologist judgment, and the rules treat it that way. It builds on Tra-Cal's approach to AI in calibration operations and goes deeper on validation.

Why AI tools are software systems under ISO/IEC 17025:2017

ISO/IEC 17025:2017 clause 7.11 tells labs to control the data and information systems they use. It draws no line between plain software and AI. A spreadsheet that works out measurement uncertainty falls under 7.11. So does a large language model that drafts an uncertainty budget review. How tightly you control it scales with how much it shapes the result.

So any AI tool whose output reaches a certificate, an uncertainty calculation, an interval call, or a record review must be validated, written up, and change-controlled. New technology does not get a pass from rules that have always applied to lab software.

Two more clauses come into play. Clause 6.4 covers equipment, and that includes software used with measurement gear. Clause 6.2 covers staff competence, which reaches the metrologist who uses and checks AI output. For where AI falls short in real workflows, see AI limitations in calibration workflows.

Validation requirements: scope, intended use, performance criteria

Validation under clause 7.11 starts with a scope. The scope says what the tool will do, which parameters or workflows it touches, and what it must not be used for. Scope creep is the top cause of failed validation, because it yields outputs the validation never covered.

The intended use statement is tighter than the scope. It names the exact tasks, the inputs, the outputs, and what the metrologist checks each time. For example: the tool drafts a list of uncertainty contributors from parameter descriptions. The metrologist reviews the list, edits as needed, and approves it before the budget is used.

Performance criteria say what good behavior looks like. For plain software, that is input in, right answer out. For AI, you also need accuracy on a test set, false positive and false negative rates where the tool flags an exception, how it behaves at the edge of the tested range, and what it does when an input falls outside that range.

The FDA's general principles of software validation guidance is a handy frame even outside FDA work. The ISPE GAMP 5 risk-based approach to compliant computerized systems is widely used in regulated shops too. Both predate today’s AI, but the method carries over once you add the AI-specific checks.

Documenting AI tool limitations and operating boundaries

A validated AI tool has written boundaries. Those cover the input ranges you tested, the outputs you verified, the inputs where the tool must not be used and why, and the outputs that need an extra look before anyone signs.

Say a tool is validated to draft contributor lists for pressure work from 0 to 10,000 psi. Inside that band, you know how it acts. If a customer sends a 20,000 psi job, you do not. The record must call out that case and set the fallback. Either a person does the work, or you revalidate over the wider range first.

Boundary notes should also list what AI is known to get wrong. It can give an answer that sounds right and is not. It can fail without warning when an input looks like the training data but is not. It can give two different answers to the same question. These are not validation failures. They are traits of the technology, and your written procedure has to plan for them.

That record is what lets an accreditation assessor see what the tool does, what it does not do, and what the metrologist must catch. Without it, the tool counts as unvalidated software, however well it seems to run.

Change control when models, prompts, or training data change

AI tools change. The vendor ships a new model. Your team tweaks the prompt. Training data shifts. The host gets patched. Any of those can move the tool’s behavior, so each one needs change control.

Clause 8.5 asks you to spot the change, weigh its impact, roll it out in a controlled way, and check that nothing broke. With AI, that last step takes more work than with plain software. The change may shift behavior in ways the release note never mentions.

Here is a workable standard for AI tools:

  • Model version updates. Treat these as major. Re-run the test set, write down what changed, and re-approve the tool before it goes live.

  • Prompt or instruction updates. Minor if the change narrows what the tool does. Major if it widens the scope or changes the output format. Test before you deploy.

  • Training data updates. Major for any model you train yourself. For a vendor model, major if the vendor calls the update material. If not, check it against part of your test set.

  • Integration changes. Any change to how the tool takes input from, or returns output to, your calibration system. Treat as major.

The change log becomes part of the tool’s validation file. Over time it shows the tool’s whole behavior history. That is what an assessor will read when they review AI-supported work.

Audit trail requirements for AI-supported calibration decisions

The audit trail is the most important record for AI-supported work. Without it, you cannot rebuild the basis for a call when an inspector or customer asks.

A sound trail records the inputs you gave the tool, with timestamps. It records the outputs the tool gave back, as the metrologist saw them. It records the review notes, with any edits, rejects, or escalations. It records where the metrologist’s own judgment differed. And it records the final call and signature.

That level of detail lines up with NIST guidance on software in measurement systems and with NCSLI guidance on calibration software. It goes beyond what most labs keep for plain software. That is on purpose. AI output needs a person’s judgment to stand up, and the trail has to show that judgment happened.

Keep records as long as your policy under clause 8.4 says. If you serve regulated industries, match the longest term any customer or rule demands. That is often the life of the instrument plus a set number of years.

Used with care, under a validated and change-controlled setup with a clear audit trail, AI can speed up calibration work without weakening your accreditation. The discipline above is the price of entry, not an add-on.

Tra-Cal Laboratories operates ISO/IEC 17025:2017 accredited calibration services under a quality system that treats AI tools as software requiring validation. For organizations considering AI integration in their own calibration programs, the validation framework above is the starting point.

Frequently Asked Questions

Do AI tools require validation under ISO/IEC 17025:2017?

Yes. AI tools that influence calibration decisions are software systems under ISO/IEC 17025:2017 clause 7.11, which requires control of data and information management systems used in laboratory activities. Validation includes documenting the intended use, defining performance criteria, recording acceptance testing results, and establishing change control before the tool can support accredited calibration work.

What ISO/IEC 17025 clauses apply to AI software in a calibration lab?

Clause 7.11 covers control of data and information management systems, which includes any software that processes calibration data, manages records, or supports decisions. Clause 6.4 covers equipment requirements, applicable when AI is embedded in measurement equipment. Clause 6.2 covers personnel competence for the metrologists who use and review AI outputs. Clause 8.7 covers corrective action when AI tool errors are detected.

How do you document AI tool limitations for accreditation purposes?

Document the operating boundaries: the input ranges the tool was tested across, the output conditions where the tool’s performance was verified, and the conditions where the tool should not be used or where outputs require additional review. Include the validation test results, the date of validation, and the metrologist or quality manager who authorized the tool for accredited work. This documentation forms part of the management system records under clause 8.

What change control is required for AI tools in calibration?

Any change that could affect the AI tool’s behavior triggers a change control review: model updates, prompt or instruction changes, training data updates, host system or environment changes, and integration changes with the calibration data system. Each change requires impact assessment, revalidation appropriate to the change scope, and updated documentation before the tool returns to accredited use.

What audit trail is required for AI-supported calibration decisions?

The audit trail should record what input was provided to the AI tool, what output the tool produced, what the metrologist reviewed, what the metrologist’s independent judgment was, and what decision was signed off. The record must be sufficient to reconstruct the basis for any calibration decision when an accreditation assessor or regulated customer requests it. Retention follows the laboratory’s existing record retention policy under clause 8.4.

Keep your calibration program accurate, documented, and accreditation-ready. Connect with Tra-Cal to support your next calibration requirement.
Previous
Previous

AI-Assisted Measurement Uncertainty Analysis in Calibration

Next
Next

How to Build a Measurement Uncertainty Budget for Pressure Calibration