Confidence Scores for AI: A New Way to Vet Applications
Security teams are caught between two pressures: employees are adopting AI tools at a rapid pace, while each new application brings a distinct set of compliance, privacy, and security concerns. When employees use these tools without formal approval, it creates "Shadow AI" — a new variation of Shadow IT. Manually reviewing every new AI or SaaS application to decide whether it gets approved or blocked does not scale, and outright bans simply push usage underground where it becomes harder to protect.
To address this, Cloudflare has launched Application Confidence Scorecards. The scores are part of a broader AI Security suite within the Cloudflare One SASE platform, designed to automate the labor-intensive process of evaluating generative AI and SaaS applications. The goal is to replace hours of research into compliance certifications and data-handling policies with a single, clear signal that security decision-makers can act on — setting policies, applying guardrails, or blocking risky tools outright so that innovation doesn't come at the cost of security.

Two types of scores are available. One covers AI-powered applications generally, rating them on factors like industry certifications, data management and security measures, and company maturity. The other, focused on Generative AI, awards higher scores to models that publish system cards documenting testing for bias, ethics, and safety — and that do not train on user inputs.
Tackling the Shadow AI Problem
SaaS adoption over the past decade has made it trivial for employees to begin using a new tool with just a credit card or a free trial. The rise of generative AI has accelerated this trend, with entire workflows — writing assistance, image generation, and more — now operating outside corporate oversight. Employees often have no way of knowing whether these tools comply with regulatory or corporate requirements.
The risks span a wide range. Sensitive data can be stored or transmitted outside corporate control. Tools may lack certifications like SOC2 or ISO 27001. Providers may retain user data indefinitely, use it to train external models, or lack the financial and operational stability to withstand a breach or bankruptcy. Biased model outputs can create compliance liabilities or lead to erroneous business decisions. The volume of applications is simply too high for security leaders to keep up with.
The Scoring Rubric
Scoring these applications at scale requires both a rubric and the infrastructure to apply it automatically. The rubric is split into two scores, each worth five points.
Application Posture Score (5 points)
- Security and Privacy Compliance (1.2 points): Awards credit for SOC 2 and ISO 27001 certifications, which indicate operational maturity.
- Data Management Practices (1 point): Evaluates retention windows and third-party data sharing. Shorter retention and no sharing earn top marks.
- Security Controls (1 point): Looks for MFA, SSO, TLS 1.3, role-based access, and session monitoring.
- Security Reports and Incident History (1 point): Checks for a trust or security page, bug bounty program, and incident response transparency. A recent material breach results in a full deduction.
- Financial Stability (.8 points): Public companies and heavily capitalized providers score highest; startups with limited funding or distressed firms score lower.
Gen-AI Posture Score (5 points)
- Compliance (1 point): Presence of the ISO 42001 certification for AI management systems.
- Deployment Security Model (1 point): Whether access is authenticated and rate-limited versus publicly exposed.
- System Card (1 point): Whether the provider publishes a model or system card documenting evaluations of safety, bias, and risk.
- Training Data Governance (2 points): Whether user data is explicitly excluded from model training, or whether opt-in/opt-out controls are available.
Scoring That Scales
Applying this rubric manually would present the same scaling problem it's meant to solve. Cloudflare built automated infrastructure that crawls public trust centers, privacy policies, security pages, and compliance documents. Large language models parse the documents to identify relevant answers, with safeguards against hallucination built in through source validation and structured extraction.

Every automated score is reviewed and audited by Cloudflare analysts before it appears in the Application Library. This hybrid approach ensures accuracy while still being comprehensive.
Making Scores Actionable
The scores are integrated directly into the Application Library. Clicking on a score in the dashboard shows a detailed breakdown of the app's performance across each rubric dimension. Because scores update as vendors improve their security and compliance posture, they serve as a live view rather than a static snapshot.

This package serves different stakeholder needs at once: IT and security teams can quickly identify high-risk tools, Procurement and Governance Risk & Compliance teams can accelerate vendor reviews, and developers and employees can make informed decisions without waiting weeks for approval.
Future Directions: From Visibility to Enforcement
Visibility is only the beginning. Cloudflare plans to extend these scores into enforcement across the Cloudflare One environment. Future capabilities will allow Gateway to block or warn users about low-scoring applications, and will let DLP policies be tied directly to confidence scores. This would keep untrusted AI and SaaS providers from becoming a channel for sensitive data.
Availability
The Cloudflare Application Confidence Scorecards are now live in the Application Library. Customers can explore them in the Cloudflare dashboard today. Application scores are freely available to all users, and can be viewed by creating a free account via the Cloudflare Zero Trust sign-up. Cloudflare is also inviting interested users to join its AI security user research program for those looking to test new functionality or share insights.



