A shadow AI risk assessment framework gives every discovered tool a score, a priority tier, and a triage action, so the harder question after discovery has a structured answer.
You have run the discovery audit. You now have a list of AI tools in use across your organisation - some sanctioned, many not, and some you cannot yet fully classify. The harder question is what to do next. Which tools demand attention this week? Which can wait for the next review cycle? Which represents an active regulatory exposure right now?
67% vs 31% of orgs have an AI risk policy - but only 31% maintain a real AI application inventory | €7.1B cumulative GDPR fines issued since 2018 - over 60% imposed after January 2023 | €35M / 7% maximum EU AI Act fine for prohibited AI practices (Art. 5) - exceeding even GDPR's penalty ceiling |
This is the problem a risk assessment framework solves. Without a structured scoring model, prioritisation defaults to instinct - and instinct routinely underweights regulatory exposure while overweighting operational disruption. The framework in this article applies a consistent, evidence-based score to every discovered tool and maps that score directly to a triage action, making AI risk assessment the structured answer to the harder question that follows AI tool discovery. It also connects the output into two regulatory artefacts your organisation already maintains: the DORA ICT register and the GDPR Record of Processing Activities (RoPA).
The Three Risk Dimensions
Not all shadow AI tools carry equal risk. An AI copywriting assistant used by the marketing team for public-facing content is a fundamentally different exposure than an AI code review tool with access to proprietary source code, or an HR chatbot processing disciplinary records. Effective triage begins by evaluating every discovered tool across three dimensions that collectively determine its true risk profile
1. Data Sensitivity
Key Question: What data leaves your perimeter when this tool runs? | 2. Regulatory Exposure
Key Question: Which regulations mandate controls this tool currently lacks? | 3. Operational Dependency
Key Question: How hard would it be to switch this tool off tomorrow? |
⚠ Why All Three Dimensions Must Be Assessed Together A tool may score low on data sensitivity but high on operational dependency - making it low-risk from a GDPR standpoint but high-risk from a business continuity perspective. The reverse is equally common: a rarely-used tool processing special-category data may demand immediate regulatory action despite minimal operational disruption. Only the composite score reveals the true priority. Scoring one or two dimensions and skipping the third is a common source of under-triage. |
The dimension that generates the most is operational dependency, not because the score is wrong. Because the tool owner is measuring the disruption and the CISO is measuring the exposure. If this tool disappears suddenly, three engineers may miss their sprint deadline. On the other hand, if this tool stays undocumented for another month, the organisation has unregistered ICT third-party processing source code with no DPA. The disagreement is not about the facts; it is about which fact is the most important and urgent one. The operational dependency score does not resolve the disagreement; rather, it makes it productive. A score of 4 on the operational dependency does not mean that the tool stays; it means that the remediation plan has to account for the disruption cost, maybe as a migration path, a sanctioned alternative, or a transition window. The score surfaces the tension so that the right conservation can happen, rather than letting the disruption argument quietly delay the governance action indefinitely. - Ankur Arora, Co-Founder, Montro |
The Shadow AI Risk Scoring Model
The scoring model assigns each tool a value of 1–5 across the three dimensions. Scores are summed to produce a total between 3 and 15, which determines the tool's priority tier. The model is intentionally executable by a CISO or risk analyst working from discovery data alone - no vendor cooperation, no technical access to the tool required.
Score inputs should be drawn from: discovery scan outputs, OAuth scope analysis, expense and procurement records, SSO logs, and stakeholder interviews. Where data is incomplete, always score the dimension at the higher value. Regulators scrutinise under-assessment far more harshly than over-caution - and the EU AI Act's own risk-based approach explicitly endorses the precautionary default.
Score | Data Sensitivity | Regulatory Exposure | Operational Dependency |
1 | Public data only; no PII | No EU framework applies | Experimental / rarely used |
2 | Internal data, low sensitivity | GDPR tangential (aggregated) | Used occasionally by 1–2 teams |
3 | Employee or customer PII | GDPR + NIS2 or EU AI Act (low-risk) | Regular use, some workflow reliance |
4 | Sensitive categories (health, finance) | GDPR + EU AI Act (high-risk) | Core workflow; alternatives unclear |
5 | Trade secrets, regulated financial data | GDPR + EU AI Act + DORA + NIS2 | Business-critical; no substitute |
Shadow AI Risk Scoring Matrix - score each tool 1-5 per dimension, then sum for the total risk score.
Priority Tiers
The summed score places every tool into one of four priority tiers. Each tier maps to a specific response timeline and mandatory action set - ensuring the assessment output drives action rather than sitting in a spreadsheet.
Priority Tier | Total Score | Required Action |
Critical | 12-15 | Immediate: DPA review, DPIA, IT register, consider suspension |
High | 8-11 | 30-day remediation: vendor assessment, contractual controls, logging |
Medium | 4-7 | Next review cycle: policy alignment, user training, monitoring |
Low | 1-3 | Monitor: document in register, flag for next annual review |
Shadow AI Priority Tiers - total score determines response urgency and remediation timeline.
Scoring Calibration Note Harmonic Security’s analysis of 22 million enterprise AI prompts in 2025 found that code, legal documents, and financial data account for 74.5% of what employees expose through unsanctioned AI tools - with legal discourse comprising 22.3% of prompt exposures (and up to 35% of sensitive file uploads per Q3 2025 data) and 12.8% of coding tool exposures containing API keys or tokens. This means data sensitivity is routinely under-estimated by assessors who assume lower-risk use patterns, and it is the single most common calibration error in informal shadow AI risk reviews. If there is any ambiguity about what data a tool accesses, score it at 4 or 5, not at 2 or 3. |
Applying the Framework - Worked Example
The following example applies the framework to a tool category that consistently appears in enterprise discovery audits but is routinely under-scored in informal risk assessments: an AI-assisted code completion and review tool.
Scenario: AI Code Completion and Review Tool Tool: An AI code assistant - broadly similar to GitHub Copilot, Tabnine, or a third-party equivalent - discovered in active use across the Engineering and Data Science teams via an OAuth grant to the corporate GitHub tenant. Status: No Data Processing Agreement. Not in the IT asset inventory. No security review conducted. The vendor's model training policy, whether prompt data is retained and used for model improvement - is unknown. The tool has been in use for seven months. Used by: Approximately 35 engineers and data scientists, using the tool daily to write, review, and refactor production code. Source code, internal API structures, and database schema definitions regularly pass through the tool's API. |
Dimension | Score | Rationale |
Data Sensitivity | 3 | Processes employee names, job titles, and project metadata sent to an external API. No special-category personal data confirmed - but cannot be ruled out given the HR team's intermittent use of the tool. |
Regulatory Exposure | 5 | No DPA exists. The vendor's model may train on prompt data by default. EU AI Act Annex III considers AI used in employment and HR management high-risk. If the organisation is in financial services, DORA's ICT third-party provisions also activate. |
Operational Dependency | 4 | Used daily by three business units for production code. Source code flows through the tool's API. Removing it tomorrow would immediately impact delivery pipelines - alternatives require onboarding time. |
TOTAL SCORE | 12 | CRITICAL - Immediate action required. This tool must be suspended or brought under contractual and governance control before the next commit cycle. |
Worked example scoring for an AI code completion tool with no DPA and unknown data retention policy.
What the Score Tells You
A total score of 12 places this tool firmly in the Critical tier. The primary risk driver is not the data sensitivity dimension alone - it is the combination of source code exposure, an unknown vendor training policy, and the complete absence of a DPA, compounding each other across three dimensions simultaneously.
The regulatory exposure score of 5 reflects two simultaneous problems. First, source code flowing to a third-party model with no DPA is a GDPR Article 28 violation from day one, the organisation has no lawful instrument governing how the vendor stores, processes, or sub-processes the data. Second, if the organisation operates in financial services and the tool is used to build or maintain any system touching customer data or transaction processing, DORA's ICT third-party provisions activate: this tool must appear in the DORA ICT register and be subject to a formal contractual arrangement with exit strategy provisions.
The operational dependency score of 4 creates the key CISO dilemma: suspension would immediately impact delivery pipelines across three teams. But this is precisely why the triage action for Critical tools includes an accelerated remediation path, not necessarily suspension forever, but structured control before the next code commit.
What to Do With the Risk Score
A risk score without an action protocol is an academic exercise. The triage framework converts each priority tier into a concrete, time-bound remediation roadmap calibrated to the regulatory requirements under GDPR, the EU AI Act, and DORA. Timelines are set against regulatory enforcement realities: GDPR fines have now exceeded €7.1 billion cumulatively, with EU AI Act enforcement for high-risk systems beginning August 2, 2026.
CRITICAL 12–15
| HIGH 8–11
| MEDIUM 4–7
| LOW 1–3
|
Triage Action Plan by Priority Tier - timelines and required actions for each risk level.
Applying the Worked Example Triage For the AI code completion tool scoring 12 (Critical), the immediate action sequence is:
|
Building It Into Your Risk Register
Individual tool assessments only generate durable value when they feed into a living, auditable risk register. The shadow AI risk register is not a standalone document - it is a structured input to two regulatory artefacts your organisation is already required to maintain: the DORA ICT register and the GDPR RoPA. Both will be examined directly by supervisory authorities.
DORA ICT Risk Register Every AI tool touching operational processes in a financial entity is an ICT asset subject to DORA Article 6 risk management requirements. Your shadow AI register connects to the DORA ICT register by:
| GDPR Record of Processing Activities (RoPA) Every shadow AI tool processing personal data is a processing activity - whether it was ever sanctioned or not. The scoring output is your mechanism for turning a discovered tool into a compliant RoPA entry:
|
Shadow AI Risk Register Structure
The template below shows the recommended column structure for a shadow AI risk register entry. Each row connects a discovered tool's score to its regulatory obligations and ownership, making the register auditable and actionable simultaneously. The worked example above populates the first live row:
Tool Name | Risk Tier | Data Score | Reg Score | Ops Score | Total | DPA Status | DORA ICT Ref | RoPA Entry | Owner | Next Review |
AI Code Assistant | Critical | 3 | 5 | 4 | 12 | Missing | ICT-REF-007 | Required | [CISO] | Immediate |
Shadow AI Risk Register Template - each row maps a discovered tool to its risk score and regulatory obligations.
Making the Register a Living Document The register must be reviewed at minimum quarterly - and updated immediately whenever discovery surfaces a new tool, a regulatory deadline changes, or a remediation action is completed or overdue. EY's 2026 data found 52% of department-level AI initiatives operating without formal approval; Gravitee's 2026 report found only 14.4% of organisations have full security approval for all AI agents in use. A point-in-time register built from a single audit will be out of date within weeks. Assign a named owner, typically the CISO or a designated shadow AI Governance Lead - accountable for quarterly board reporting on register status. |
Frequently Asked Questions
How often should the shadow AI risk assessment be run?
The initial assessment runs once against the discovery output. After that, the register needs updating whenever a new tool is discovered, a regulatory deadline changes, or a remediation action is completed or overdue. A quarterly review cadence is the practical minimum, but the trigger for an unscheduled update is any material change in the environment, not the calendar. A tool that was scored Medium in January and has since expanded to three new departments may need rescoring before the next quarterly cycle.
What happens when a tool scores differently across dimensions - for example, low data sensitivity but high operational dependency?
Score it honestly across all three dimensions and let the total determine the tier. A low data sensitivity score combined with a high operational dependency score produces a medium total, which means the tool goes into the next review cycle rather than immediate remediation. The composite score is the point. A tool that is business-critical but genuinely low-risk from a data and regulatory standpoint does not warrant the same response as a tool that is both high-risk and deeply embedded. The framework prevents operational disruption from being used to delay action on genuinely high-risk tools, and prevents regulatory anxiety from triggering unnecessary disruption to genuinely low-risk ones.
Can the framework be used for sanctioned AI tools as well as shadow AI?
Yes, and it should be. The scoring model applies to any AI tool in active use, sanctioned or not. A formally procured AI tool with a signed DPA and an ICT register entry can still score high on data sensitivity or operational dependency, particularly if the use case has expanded beyond its original procurement scope. Running the framework across the full inventory, sanctioned and shadow AI tools alike, produces a more accurate picture of the organisation's total AI risk assessment exposure than limiting it to unsanctioned tools alone.
How does the framework connect to the EU AI Act's own risk classification?
The regulatory exposure dimension directly incorporates EU AI Act risk tiers. A tool that qualifies as high-risk under Annex III; recruitment AI, credit-scoring AI, biometric identification, would typically score 4 or 5 on regulatory exposure regardless of its data sensitivity or operational dependency scores. That means any Annex III tool discovered in shadow use will ordinarily land in the High or Critical tier. The framework does not replace the EU AI Act's classification process, it uses the output of that classification as one of three scoring inputs, ensuring the regulatory obligation drives the triage priority rather than being treated as a separate parallel exercise.





