Auditing an AI Management System

An organization may have responsible AI principles, a governance committee, and an approved policy yet still be unable to show how a deployment decision was assessed, authorized, and monitored. That gap is where a useful audit begins.
The question is not whether the organization declares transparency or fairness, but whether it can trace its obligations to risks, controls, decisions, and results. Auditing ISO/IEC 42001 tests the traceability and effectiveness of a management system; it is not a review of abstract principles or a universal Annex A checklist.
The audit object is the management system, not an isolated model
ISO/IEC 42001 specifies requirements for establishing, implementing, maintaining, and continually improving an artificial intelligence management system. Its subject is the organization’s ability to develop, provide, or use AI systems responsibly. It does not certify that every model is accurate, safe, explainable, or compliant with every applicable law.
The auditor evaluates governance, risk, impact assessment, competence, operations, performance, and improvement, while sampling specific AI systems and changes to determine whether those processes work. Model, data, interface, or human-oversight tests provide evidence about the management system; they are not product certification. A deeper assessment of one AI system requires additional technical, legal, contractual, or sector criteria.
Audit criteria need to be layered
ISO 19011 describes audit criteria as the set of requirements against which objective evidence is compared. For an AI management system, relying only on the clauses of ISO/IEC 42001 excludes a significant portion of the organization’s actual commitments.
| Criteria layer | What it may include | Auditor caution |
|---|---|---|
| Management system requirements | Clauses 4–10 of ISO/IEC 42001 and applicable Annex A controls selected through risk treatment. | Do not treat the entire Annex A control set as mandatory without reviewing the statement of applicability. |
| The organization’s own requirements | AI policy, objectives, methodologies, risk criteria, procedures, control commitments, and approval rules. | An organization can fail its own system even when no clause is literally breached. |
| External obligations | Laws, regulations, contracts, customer requirements, and relevant interested-party expectations. | Establish the jurisdiction, organizational role, and affected AI system before concluding compliance. |
| Frameworks adopted for support | ISO/IEC 23894, ISO/IEC 42005, NIST AI RMF, sector codes, or internal methods. | Guidance becomes criteria only when adopted by the organization, required contractually, or incorporated into the engagement. |
This structure prevents voluntary guidance from being treated as a certifiable requirement and internal commitments from being overlooked. The EU Artificial Intelligence Act illustrates the issue: obligations depend on the system, organizational role, and applicable date. Auditors need the current legal text, not a generic checklist. ISO certification can support governance but does not replace a legal compliance conclusion.
Traceability turns principles into evidence
A strong work programme follows the management system from end to end:
- Context, roles, and scope. Establish which activities, entities, suppliers, and AI systems are included, the organization’s role, and the interested-party requirements addressed. An inventory without boundaries or ownership does not demonstrate controlled scope.
- Leadership and objectives. Confirm that policy becomes objectives, accountability, resources, decisions, and escalation mechanisms.
- Risks and impacts. Test consistent criteria, consequences, prioritization, reassessment after change, and the link between impacts and decisions. ISO/IEC 23894 and ISO/IEC 42005 can strengthen the method without replacing requirements.
- Controls and the statement of applicability. Review necessary controls, inclusions and exclusions, additional controls, and residual-risk approval. Annex B guides implementation but is not a uniform recipe.
- Operations and lifecycle. Test controls across development, acquisition, validation, deployment, use, change, and retirement through data, approvals, logs, human oversight, incidents, and third parties.
- Evaluation and improvement. Confirm that metrics, audits, management review, and corrective actions change decisions, risks, or controls when needed.
The chain fails whenever the organization cannot explain why a control exists, which risk it treats, how it is tested, or who accepts the result.
The audit programme should follow risk and change velocity
ISO/IEC 42001 requires the internal audit programme to consider the importance of the processes concerned and previous audit results. ISO 19011 extends that logic through a risk-based approach to objectives, scope, criteria, methods, resources, frequency, and competence.
A uniform clause-by-clause calendar is rarely sufficient. The audit universe should connect management system processes with AI systems, lifecycle stages, third parties, and accountabilities. Priority rises with severe potential impacts, sensitive data, automated decisions, frequent change, supplier dependency, incidents, immature controls, or material residual risk.
The programme should also define trigger events: a material new use, changed purpose or jurisdiction, model or provider replacement, major data change, incident, or earlier nonconformity. Sampling should include new uses, purchased systems, third-party solutions, and material changes, not only mature and well-documented projects.
Evidence must demonstrate design and effectiveness
An approved policy proves that a statement exists. It does not prove that decisions follow the policy. A completed impact assessment proves that a document was produced. It does not prove that the conclusions altered design, deployment, or oversight.
Persuasive evidence combines inspection of policies, contracts, assessments, and approvals; interviews with owners and users; observation of controls; tracing decisions from requirement to authorization; selective reperformance; and sampling of logs, exceptions, changes, and incidents.
Information becomes audit evidence only when it is relevant to the criteria and verifiable. In AI, that requires particular attention to model versions, datasets, configurations, prompts, parameters, thresholds, and providers. Without temporal identification and change control, an auditor may be evaluating a version different from the one that produced the observed outcome.
Where the conclusion depends on specialized testing of robustness, bias, security, or data quality, the team may need technical experts. The auditor remains responsible for integrating that work and forming the conclusion.
Competence and audit type determine the level of confidence
Collective competence commonly combines management system auditing, AI lifecycle knowledge, data, security, privacy, regulation, and domain expertise. Expecting one person to master everything creates false assurance.
For internal and supplier audits, ISO 19011 provides guidance on audit programmes, execution, and competence. For third-party certification, ISO/IEC 17021-1 establishes general requirements for certification bodies, while ISO/IEC 42006:2025 adds requirements specific to auditing and certifying AI management systems. This distinction matters: an organization should not present an internal review or consulting assessment as equivalent to accredited certification.
The NIST AI Risk Management Framework can broaden questions, risk sources, and procedures, but it does not replace explicit criteria or make good practices mandatory.
A defensible conclusion does not declare that the organization’s AI is “responsible” in absolute terms. It explains the scope assessed, criteria applied, evidence obtained, limitations encountered, and whether the management system demonstrates conformity and effectiveness in governing risk within that scope. That discipline turns an ethical ambition into a professional conclusion that can withstand scrutiny.
Sources
- ISO, ISO/IEC 42001:2023 — Artificial intelligence management systems
- ISO, ISO 19011:2018 — Guidelines for auditing management systems
- ISO, ISO/IEC 23894:2023 — AI risk management guidance
- ISO, ISO/IEC 42005:2025 — AI system impact assessment
- ISO, ISO/IEC 42006:2025 — Requirements for bodies auditing and certifying AI management systems
- ISO, ISO/IEC 17021-1:2015 — Requirements for management system audit and certification bodies
- NIST, Artificial Intelligence Risk Management Framework
- European Union, Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence