Montro
Shadow AI44 min read

Shadow AI Risk Assessment: A Framework for Quantifying What You Don't Know

Shadow AI Risk Assessment: A Framework for Quantifying What You Don't Know
AuthorAnkur Arora
Published on1 May 2026

A shadow AI risk assessment framework gives every discovered tool a score, a priority tier, and a triage action, so the harder question after discovery has a structured answer.


You have run the discovery audit. You now have a list of AI tools in use across your organisation - some sanctioned, many not, and some you cannot yet fully classify. The harder question is what to do next. Which tools demand attention this week? Which can wait for the next review cycle? Which represents an active regulatory exposure right now?

67% vs 31%


of orgs have an AI risk policy - but only 31% maintain a real AI application inventory

HelpNet Security / Larridin, 2026

€7.1B


cumulative GDPR fines issued since 2018 - over 60% imposed after January 2023

Kiteworks, March 2026

€35M / 7%


maximum EU AI Act fine for prohibited AI practices (Art. 5) - exceeding even GDPR's penalty ceiling

EU AI Act Art. 99; LegalNodes, 2026

This is the problem a risk assessment framework solves. Without a structured scoring model, prioritisation defaults to instinct - and instinct routinely underweights regulatory exposure while overweighting operational disruption. The framework in this article applies a consistent, evidence-based score to every discovered tool and maps that score directly to a triage action, making AI risk assessment the structured answer to the harder question that follows AI tool discovery. It also connects the output into two regulatory artefacts your organisation already maintains: the DORA ICT register and the GDPR Record of Processing Activities (RoPA).


The Three Risk Dimensions


Not all shadow AI tools carry equal risk. An AI copywriting assistant used by the marketing team for public-facing content is a fundamentally different exposure than an AI code review tool with access to proprietary source code, or an HR chatbot processing disciplinary records. Effective triage begins by evaluating every discovered tool across three dimensions that collectively determine its true risk profile

1. Data Sensitivity


  • Personal data of employees/customers (GDPR Art. 4)

  • Special categories: health, biometric, financial

  • Confidential IP: source code, trade secrets, M&A data

  • Volume: one-off query vs. continuous data stream


Key Question: What data leaves your perimeter when this tool runs?

2. Regulatory Exposure


  • GDPR: Is personal data processed without a valid DPA?

  • EU AI Act: Does the tool's purpose qualify as high-risk (Annex III)?

  • DORA: Is this an ICT tool in a financial services context?

  • NIS2: Does it touch critical infrastructure or essential services?


Key Question: Which regulations mandate controls this tool currently lacks?

3. Operational Dependency


  • How many users rely on this tool daily or weekly?

  • Is it embedded in documented workflows or SOPs?

  • Do business decisions depend on its outputs?

  • What is the disruption cost if it is suspended tomorrow?


Key Question: How hard would it be to switch this tool off tomorrow?

⚠ Why All Three Dimensions Must Be Assessed Together


A tool may score low on data sensitivity but high on operational dependency - making it low-risk from a GDPR standpoint but high-risk from a business continuity perspective. The reverse is equally common: a rarely-used tool processing special-category data may demand immediate regulatory action despite minimal operational disruption. Only the composite score reveals the true priority. Scoring one or two dimensions and skipping the third is a common source of under-triage.

The dimension that generates the most is operational dependency, not because the score is wrong. Because the tool owner is measuring the disruption and the CISO is measuring the exposure. If this tool disappears suddenly, three engineers may miss their sprint deadline. On the other hand, if this tool stays undocumented for another month, the organisation has unregistered ICT third-party processing source code with no DPA. The disagreement is not about the facts; it is about which fact is the most important and urgent one.


The operational dependency score does not resolve the disagreement; rather, it makes it productive. A score of 4 on the operational dependency does not mean that the tool stays; it means that the remediation plan has to account for the disruption cost, maybe as a migration path, a sanctioned alternative, or a transition window. The score surfaces the tension so that the right conservation can happen, rather than letting the disruption argument quietly delay the governance action indefinitely. - Ankur Arora, Co-Founder, Montro

The Shadow AI Risk Scoring Model


The scoring model assigns each tool a value of 1–5 across the three dimensions. Scores are summed to produce a total between 3 and 15, which determines the tool's priority tier. The model is intentionally executable by a CISO or risk analyst working from discovery data alone - no vendor cooperation, no technical access to the tool required.


Score inputs should be drawn from: discovery scan outputs, OAuth scope analysis, expense and procurement records, SSO logs, and stakeholder interviews. Where data is incomplete, always score the dimension at the higher value. Regulators scrutinise under-assessment far more harshly than over-caution - and the EU AI Act's own risk-based approach explicitly endorses the precautionary default.

Score

Data Sensitivity

Regulatory Exposure

Operational Dependency

1

Public data only; no PII

No EU framework applies

Experimental / rarely used

2

Internal data, low sensitivity

GDPR tangential (aggregated)

Used occasionally by 1–2 teams

3

Employee or customer PII

GDPR + NIS2 or EU AI Act (low-risk)

Regular use, some workflow reliance

4

Sensitive categories (health, finance)

GDPR + EU AI Act (high-risk)

Core workflow; alternatives unclear

5

Trade secrets, regulated financial data

GDPR + EU AI Act + DORA + NIS2

Business-critical; no substitute

Shadow AI Risk Scoring Matrix - score each tool 1-5 per dimension, then sum for the total risk score.


Priority Tiers


The summed score places every tool into one of four priority tiers. Each tier maps to a specific response timeline and mandatory action set - ensuring the assessment output drives action rather than sitting in a spreadsheet.

Priority Tier

Total Score

Required Action

Critical

12-15

Immediate: DPA review, DPIA, IT register, consider suspension

High

8-11

30-day remediation: vendor assessment, contractual controls, logging

Medium

4-7

Next review cycle: policy alignment, user training, monitoring

Low

1-3

Monitor: document in register, flag for next annual review

Shadow AI Priority Tiers - total score determines response urgency and remediation timeline.

Scoring Calibration Note

Harmonic Security’s analysis of 22 million enterprise AI prompts in 2025 found that code, legal documents, and financial data account for 74.5% of what employees expose through unsanctioned AI tools - with legal discourse comprising 22.3% of prompt exposures (and up to 35% of sensitive file uploads per Q3 2025 data) and 12.8% of coding tool exposures containing API keys or tokens. This means data sensitivity is routinely under-estimated by assessors who assume lower-risk use patterns, and it is the single most common calibration error in informal shadow AI risk reviews. If there is any ambiguity about what data a tool accesses, score it at 4 or 5, not at 2 or 3.

Applying the Framework - Worked Example

 

The following example applies the framework to a tool category that consistently appears in enterprise discovery audits but is routinely under-scored in informal risk assessments: an AI-assisted code completion and review tool.

Scenario: AI Code Completion and Review Tool


Tool: An AI code assistant - broadly similar to GitHub Copilot, Tabnine, or a third-party equivalent - discovered in active use across the Engineering and Data Science teams via an OAuth grant to the corporate GitHub tenant.


Status: No Data Processing Agreement. Not in the IT asset inventory. No security review conducted. The vendor's model training policy, whether prompt data is retained and used for model improvement - is unknown. The tool has been in use for seven months.


Used by: Approximately 35 engineers and data scientists, using the tool daily to write, review, and refactor production code. Source code, internal API structures, and database schema definitions regularly pass through the tool's API.

Dimension

Score

Rationale

Data Sensitivity

3

Processes employee names, job titles, and project metadata sent to an external API. No special-category personal data confirmed - but cannot be ruled out given the HR team's intermittent use of the tool.

Regulatory Exposure

5

No DPA exists. The vendor's model may train on prompt data by default. EU AI Act Annex III considers AI used in employment and HR management high-risk. If the organisation is in financial services, DORA's ICT third-party provisions also activate.

Operational Dependency

4

Used daily by three business units for production code. Source code flows through the tool's API. Removing it tomorrow would immediately impact delivery pipelines - alternatives require onboarding time.

TOTAL SCORE

12

CRITICAL - Immediate action required. This tool must be suspended or brought under contractual and governance control before the next commit cycle.

Worked example scoring for an AI code completion tool with no DPA and unknown data retention policy.


What the Score Tells You


A total score of 12 places this tool firmly in the Critical tier. The primary risk driver is not the data sensitivity dimension alone - it is the combination of source code exposure, an unknown vendor training policy, and the complete absence of a DPA, compounding each other across three dimensions simultaneously.


The regulatory exposure score of 5 reflects two simultaneous problems. First, source code flowing to a third-party model with no DPA is a GDPR Article 28 violation from day one, the organisation has no lawful instrument governing how the vendor stores, processes, or sub-processes the data. Second, if the organisation operates in financial services and the tool is used to build or maintain any system touching customer data or transaction processing, DORA's ICT third-party provisions activate: this tool must appear in the DORA ICT register and be subject to a formal contractual arrangement with exit strategy provisions.


The operational dependency score of 4 creates the key CISO dilemma: suspension would immediately impact delivery pipelines across three teams. But this is precisely why the triage action for Critical tools includes an accelerated remediation path, not necessarily suspension forever, but structured control before the next code commit.


What to Do With the Risk Score


A risk score without an action protocol is an academic exercise. The triage framework converts each priority tier into a concrete, time-bound remediation roadmap calibrated to the regulatory requirements under GDPR, the EU AI Act, and DORA. Timelines are set against regulatory enforcement realities: GDPR fines have now exceeded €7.1 billion cumulatively, with EU AI Act enforcement for high-risk systems beginning August 2, 2026.

CRITICAL

12–15

  • Suspend tool access pending governance review

  • Initiate DPIA (GDPR Art. 35) - large-scale employee data processing

  • Execute or terminate vendor DPA within 5 business days

  • Register in IT asset inventory and DORA ICT register immediately

  • Brief DPO and CISO within 24 hours

  • Assess GDPR Art. 33 breach notification (72-hour clock)

HIGH

8–11

  • Issue 30-day remediation notice to tool owner

  • Conduct vendor security questionnaire

  • Establish contractual data handling and logging controls

  • Restrict access to named, approved users only

  • Schedule DPA review within 2 weeks


MEDIUM

4–7

  • Flag for next governance review cycle (max 90 days)

  • Align usage with applicable regulatory framework

  • Mandate user training on data handling obligations

  • Configure usage monitoring alert

  • Document in shadow AI risk register


LOW

1–3

  • Document in risk register - flag for annual review

  • Confirm tool meets basic encryption and access standards

  • Notify relevant team lead

  • Next scheduled assessment in 12 months

Triage Action Plan by Priority Tier - timelines and required actions for each risk level.

Applying the Worked Example Triage


For the AI code completion tool scoring 12 (Critical), the immediate action sequence is:

  • Suspend API access for the tool pending governance review - or restrict to a sandboxed, non-production environment with no access to proprietary code

  • Initiate a DPIA under GDPR Article 35: large-scale processing of employee data with an unknown third-party training policy qualifies

  • Execute a DPA with the vendor or initiate vendor termination and identify a compliant alternative

  • Register the tool in the IT asset inventory and, if applicable, the DORA ICT register

  • Brief DPO and legal counsel within 24 hours on the seven-month exposure window

  • Assess whether seven months of source code transmission constitutes a personal data breach requiring GDPR Article 33 notification

Building It Into Your Risk Register


Individual tool assessments only generate durable value when they feed into a living, auditable risk register. The shadow AI risk register is not a standalone document - it is a structured input to two regulatory artefacts your organisation is already required to maintain: the DORA ICT register and the GDPR RoPA. Both will be examined directly by supervisory authorities.

DORA ICT Risk Register


Every AI tool touching operational processes in a financial entity is an ICT asset subject to DORA Article 6 risk management requirements. Your shadow AI register connects to the DORA ICT register by:

  • Classify each tool as an ICT asset or ICT third-party service provider under DORA's definitions

  • Surface unregistered tools for Article 28 contractual remediation - your register report to supervisors must be accurate

  • Feed the risk score into your ICT risk appetite statement to justify prioritisation decisions

  • Determine which AI-dependent processes require resilience testing under DORA's operational resilience framework

GDPR Record of Processing Activities (RoPA)

Every shadow AI tool processing personal data is a processing activity - whether it was ever sanctioned or not. The scoring output is your mechanism for turning a discovered tool into a compliant RoPA entry:

  • Create a RoPA entry for every confirmed tool that processes personal data - purpose, legal basis, data categories, recipients, and retention period

  • Flag missing DPAs as open gaps in the controller–processor chain - each is an active GDPR Article 28 deficiency

  • Identify tools that cross the DPIA threshold under GDPR Article 35 - large-scale processing or systematic evaluation of individuals

  • Map AI data flows to support data subject rights fulfilment - access, erasure, and portability requests require this visibility

Shadow AI Risk Register Structure


The template below shows the recommended column structure for a shadow AI risk register entry. Each row connects a discovered tool's score to its regulatory obligations and ownership, making the register auditable and actionable simultaneously. The worked example above populates the first live row:

Tool Name

Risk Tier

Data Score

Reg Score

Ops Score

Total

DPA Status

DORA ICT Ref

RoPA Entry

Owner

Next Review

AI Code Assistant

Critical

3

5

4

12

Missing

ICT-REF-007

Required

[CISO]

Immediate

Shadow AI Risk Register Template - each row maps a discovered tool to its risk score and regulatory obligations.

Making the Register a Living Document


The register must be reviewed at minimum quarterly - and updated immediately whenever discovery surfaces a new tool, a regulatory deadline changes, or a remediation action is completed or overdue. EY's 2026 data found 52% of department-level AI initiatives operating without formal approval; Gravitee's 2026 report found only 14.4% of organisations have full security approval for all AI agents in use. A point-in-time register built from a single audit will be out of date within weeks. Assign a named owner, typically the CISO or a designated shadow AI Governance Lead - accountable for quarterly board reporting on register status.


Frequently Asked Questions


How often should the shadow AI risk assessment be run?


The initial assessment runs once against the discovery output. After that, the register needs updating whenever a new tool is discovered, a regulatory deadline changes, or a remediation action is completed or overdue. A quarterly review cadence is the practical minimum, but the trigger for an unscheduled update is any material change in the environment, not the calendar. A tool that was scored Medium in January and has since expanded to three new departments may need rescoring before the next quarterly cycle.


What happens when a tool scores differently across dimensions - for example, low data sensitivity but high operational dependency?


Score it honestly across all three dimensions and let the total determine the tier. A low data sensitivity score combined with a high operational dependency score produces a medium total, which means the tool goes into the next review cycle rather than immediate remediation. The composite score is the point. A tool that is business-critical but genuinely low-risk from a data and regulatory standpoint does not warrant the same response as a tool that is both high-risk and deeply embedded. The framework prevents operational disruption from being used to delay action on genuinely high-risk tools, and prevents regulatory anxiety from triggering unnecessary disruption to genuinely low-risk ones.


Can the framework be used for sanctioned AI tools as well as shadow AI?


Yes, and it should be. The scoring model applies to any AI tool in active use, sanctioned or not. A formally procured AI tool with a signed DPA and an ICT register entry can still score high on data sensitivity or operational dependency, particularly if the use case has expanded beyond its original procurement scope. Running the framework across the full inventory, sanctioned and shadow AI tools alike, produces a more accurate picture of the organisation's total AI risk assessment exposure than limiting it to unsanctioned tools alone.


How does the framework connect to the EU AI Act's own risk classification?


The regulatory exposure dimension directly incorporates EU AI Act risk tiers. A tool that qualifies as high-risk under Annex III; recruitment AI, credit-scoring AI, biometric identification, would typically score 4 or 5 on regulatory exposure regardless of its data sensitivity or operational dependency scores. That means any Annex III tool discovered in shadow use will ordinarily land in the High or Critical tier. The framework does not replace the EU AI Act's classification process, it uses the output of that classification as one of three scoring inputs, ensuring the regulatory obligation drives the triage priority rather than being treated as a separate parallel exercise.

Ankur Arora

Ankur Arora

Co-founder

Fifteen years of enterprise digital transformation across telecoms, media, consumer goods, and agriculture - and a front-row seat to AI adoption outpacing governance at every organisation he worked in. He built Montro so the next firm doesn't have to learn that lesson the hard way.

Blog

Read next

Explore more from our library

View all

Stay informed on EU AI governance

Monthly updates on regulatory changes, compliance trends, and platform releases

By subscribing you agree to our Terms and Conditions and Privacy Policy

Montro AI governance dashboard showing tool risk tiers