Auditing AI Without Becoming a Data Scientist

An internal audit team discovers that several business units are already using artificial intelligence to draft communications, prioritize customers, and support operational decisions. The risk is visible, but so is the hesitation: “We cannot audit AI until we hire a data scientist.”
That response collapses two different questions into one. The first is whether an auditor can build, train, or technically validate an AI model. The second is whether the auditor can assess how the organization governs the use of AI, what risks it accepts, which controls it requires, and what evidence supports the decision to use the system.
For a large part of an AI audit, the second question matters more. Internal auditors do not need to become data scientists before they start auditing AI. They need enough AI literacy to ask useful questions, recognize where the system can fail, and identify when a conclusion goes beyond the competence available on the engagement.
The professional boundary is not between “audit” and “technology.” It is between what the team can support with evidence and what it would be claiming without sufficient competence.
The Standards require competence, not technical omniscience
The Global Internal Audit Standards, effective January 9, 2025, make competency a requirement. Standard 3.1 requires internal auditors to possess or obtain the competencies needed to fulfill their responsibilities, while the chief audit executive must ensure that the internal audit function collectively has, or obtains, the competencies needed to provide its services.
The word collectively matters. Professional competence does not mean that every auditor must be equally proficient in risk, cybersecurity, privacy, statistics, machine learning, regulation, business operations, and data science. It means the function must understand the capabilities an engagement requires and have a credible way to obtain them.
Standard 13.5 applies the same logic at engagement level. During planning, internal auditors must identify the types and quantity of resources needed to achieve the engagement objectives and consider whether the human, financial, and technological resources available are appropriate and sufficient. When they are not, the answer is to obtain additional capability, change the approach, or make the constraint visible — not to imitate technical depth the team does not possess.
The IIA's Artificial Intelligence Auditing Framework, 2nd Edition reinforces this as practical guidance. It explicitly notes that internal auditors are not expected to be experts on every audit topic. A disciplined method, critical thinking, risk identification, and working knowledge of AI remain central; outside technical resources may be necessary for more technical matters such as understanding algorithms.
What a generalist auditor can already assess
Much of AI risk exists before and around the algorithm. A sophisticated model can still be poorly governed, used for an unapproved purpose, fed by inadequately controlled data, or left in production without meaningful monitoring.
A generalist auditor with foundational AI literacy can assess areas such as these:
| Area | Audit question | Evidence a generalist can examine |
|---|---|---|
| Governance and accountability | Who approves the use and who is accountable for the outcome? | Mandates, committees, RACI matrices, approvals, inventories, and escalation records. |
| Objective and use case | What problem is AI intended to solve and what result is expected? | Business case, requirements, success criteria, and permitted-use boundaries. |
| Risk management | Were legal, operational, ethical, privacy, and reputational risks considered before deployment? | Risk assessments, decisions, exceptions, and residual-risk acceptance. |
| Policies and acceptable use | Are the rules clear and operating in practice? | Policies, training, use requests, approvals, and exceptions. |
| Data and access | Are ownership, access controls, and data-quality checks defined? | Roles, permissions, reconciliations, integrity controls, and quality records. |
| Third parties | Does the organization understand what depends on the vendor and how that dependency is monitored? | Due diligence, contracts, SLAs, assurance reports, and supplier monitoring. |
| Change and deployment | Do material changes require testing and authorization before release? | Change workflows, testing records, approvals, versions, and release logs. |
| Monitoring and outcomes | Are results, exceptions, incidents, and performance reviewed after deployment? | KPIs, alerts, reviews, complaints, incidents, and corrective decisions. |
These are not “nontechnical” questions. They are governance, risk, and control questions applied to a new technology. The IIA's AI framework devotes substantial attention to strategy, roles, policies, data, cybersecurity, third parties, training, testing, and ongoing monitoring precisely because AI assurance is broader than model mathematics.
The boundary appears when the conclusion depends on a technical assertion
The opposite mistake is to assume that because an auditor can assess the process around a model, the auditor can conclude on every property of the model itself.
A useful decision rule is to ask: What assertion will the audit report have to defend?
If the conclusion is “the organization does not require documented approval before employees use AI tools with confidential information,” a generalist auditor can probably gather and evaluate the necessary evidence. If the conclusion is “the model does not create material bias for specified groups,” the assertion depends on choices about populations, fairness metrics, data representativeness, statistical design, and test results. That second conclusion may require specialist competence.
The same applies when an engagement intends to conclude that a model is technically robust, that an architecture is secure, that an algorithm has been appropriately validated, or that a performance metric is statistically suitable. A generalist can test whether a validation process exists, who performed it, whether it was independent, what exceptions arose, and what management did about them. Assessing whether the technical validation itself is sound may require deeper expertise.
That distinction is central: auditing the governance of a test is not the same as technically reperforming or validating the test.
When to bring in a specialist
There is no universal list of AI topics that belongs exclusively to data scientists. The answer depends on the engagement objective, risk, and evidence. Four triggers are especially useful:
- The engagement objective contains a technical claim. If the audit intends to conclude on accuracy, robustness, bias, explainability, security, or model performance, ask whether the team can challenge the method used to support that claim.
- The evidence requires specialized methods to interpret. Code, statistical metrics, model architecture, adversarial testing, or complex configurations may require expertise that cannot be replaced by a high-level walkthrough.
- The consequence of being wrong is high. The greater the potential impact on customers, employees, financial decisions, compliance, or security, the less defensible it is to accept a material competency gap.
- Opacity constrains assurance. With third-party models or complex systems, the audit may need a combination of specialists, contractual evidence, independent assurance reports, and compensating procedures to determine what can actually be concluded.
Bringing in a specialist does not mean handing over the audit. The auditor remains responsible for ensuring that objectives, scope, criteria, and procedures respond to the risk. The specialist provides depth where the evidence requires subject-matter expertise.
Build T-shaped capability instead of waiting for perfect expertise
Waiting until every auditor can explain the mathematics behind a model would create an assurance gap precisely while AI risk is expanding. The opposite extreme — auditing complex systems through generic checklists without the ability to challenge technical evidence — produces false confidence of a different kind.
A more sustainable function develops broad AI literacy across the audit team and maintains access to deep expertise when an engagement requires it. Baseline capability should help auditors understand AI use cases, data, models, risks, controls, limitations, third parties, and human oversight. Deep expertise can come from internal specialists, co-sourcing, experts elsewhere in the organization, or external providers, provided independence, scope, and quality are managed appropriately.
Recent IIA discussion on upskilling for critical AI capabilities points in the same direction: combine formal learning with hands-on experimentation and role-based application rather than treating AI as an isolated technical specialty.
The starting question is not “Can I build this model?” It is: Do I understand the risk, know what assertion I need to test, and have a competent team capable of obtaining sufficient evidence?
That question keeps the auditor inside the discipline of internal audit while making it possible to start auditing AI now.