Moving Past Pass/Fail: How WCAG 3.0 Proposes to Grade Accessibility
Since the Web Content Accessibility Guidelines first appeared in 1999, the framework's fundamental question has been binary: does a success criterion pass or fail? The WCAG 2.x series, released in 2008, cemented this model with tightly defined technical criteria and three conformance levels (A, AA, AAA). That structure has given regulators and auditors a clear, repeatable process, but it also has a blind spot: a page can technically satisfy every rule and still be difficult or impossible for someone with a disability to actually use.
The W3C Accessibility Guidelines Working Group has been drafting WCAG 3.0 since early 2021, and the document remains far from final — W3C Recommendation status could be years away. Rather than a routine update, the draft represents a fundamental reorientation of how accessibility is evaluated. WCAG 3.0 does not ask whether a feature is present. It asks how well a person with a disability can complete a meaningful task.
Structure: Outcomes and Methods Replace Success Criteria
WCAG 2.x organizes requirements around the POUR principles — Perceivable, Operable, Understandable, and Robust — and expresses each requirement as a testable success criterion tied to a conformance level. WCAG 3.0 keeps a layered hierarchy, but the layers themselves are different:
- Guidelines: high-level accessibility goals tied to specific user needs.
- Outcomes: user-centered, testable statements about what a user can do, such as having alternatives for time-based media.
- Methods: technology-specific techniques for achieving outcomes, complete with code examples and test instructions.
- How-To Guides: narrative material that gives design context, user scenarios, and practical advice.
The emphasis on outcomes is not cosmetic. It realigns technical implementation with experience: the language shifts from what a developer ships to what a user can accomplish. An e-commerce checkout is a useful illustration. Under WCAG 2.x, a single field missing its label can fail the entire flow against AA conformance. WCAG 3.0 would evaluate that flow across several outcomes — form labeling, keyboard navigation, error recovery, focus management — and score each one independently. A flow with strong scores across most outcomes but weak error messaging would earn a lower overall grade rather than being written off entirely.
A Graded Scoring Model
WCAG 3.0 replaces the binary check with a graded score. Each outcome is assessed through one or more atomic tests, which come in three types:
- Binary tests: clear yes/no questions, such as whether every image has alternative text.
- Percentage-based tests: coverage measures, like the proportion of form fields with labels.
- Qualitative tests: rated judgments against defined criteria, such as the descriptiveness of alternative text.
Scores are normalized — commonly on a 0–4 or 0–5 scale — and labeled Poor, Fair, Good, or Excellent. Those scores are then aggregated across functional areas (vision, mobility, cognition) and across user flows. For teams, this creates a way to gauge progress. Under WCAG 2.x, a product either conforms or it does not; under WCAG 3.0, moving from Fair to Good over successive releases is an observable improvement.
Critical Errors Preserve Real Consequences
Grading softens the all-or-nothing verdict, but it does not let serious failures slip through. WCAG 3.0 defines critical errors as high-impact problems that cap or dramatically reduce an otherwise solid score.
Consider the checkout flow again. Under the new model, missing labels on optional fields or vague error messages might drop a rating from Excellent to Good. But if a user cannot submit the order, log in, or complete another core action, that is a critical error. It blocks task completion outright and drags the whole evaluation down, no matter how polished the surrounding experience looks. Lower-impact problems on non-essential features — an inaccessible control for uploading a profile photo or changing a theme — carry less weight in the aggregated score.
Bronze, Silver, and Gold
WCAG 3.0 also proposes replacing the Level A/AA/AAA ladder with three new tiers:
- Bronze: the new baseline conformance, comparable in ambition to WCAG 2.2 Level AA, achievable through automated and guided manual testing of the foundational outcomes.
- Silver: a higher bar that demands broader coverage, higher scores, plus usability testing with people with disabilities.
- Gold: the top tier for exemplary accessibility, likely requiring deep user involvement and inclusive design processes throughout development.
The new tiers are designed to reward advancement rather than present an unreachable ideal. Where WCAG 2.x's Level AAA is often treated as aspirational, these levels are meant to be more approachable and to let teams scope conformance claims to specific user flows, a component, or a mobile app.
Preparing Within the Current Standard
WCAG 3.0 is a draft and will not displace 2.x in the near term. WCAG 2.2 Level AA remains the recognized benchmark for legal and policy compliance, and the two frameworks will coexist so that conformance under 2.2 stays valid for years. Teams that want to be ready can act now without abandoning the current standard:
- Keep meeting WCAG 2.2 Level AA. It is still the baseline most organizations are required to satisfy.
- Study the WCAG 3.0 drafts. Pay particular attention to the outcomes and the scoring mechanics.
- Think in terms of outcomes. Frame design and development questions around what users need to accomplish, not just which attributes exist.
- Adopt accessibility early in the process. Testing at the end is costly and misses the point; build with access in mind from the start.
- Bring in users with disabilities early and often. Direct feedback is the best check on whether a design actually works.
The shift toward graded, outcome-based scoring formalizes what many practitioners have long understood: inclusive design is not about passing a test, it is about enabling people to get things done. The compliance era laid the necessary foundation; WCAG 3.0 aims to build an effectiveness model on top of it.
What Could Go Wrong With Scored Accessibility
The shift from binary pass/fail criteria to scored evaluations isn't without risk. Several structural concerns deserve attention, particularly for organizations managing regulatory compliance, scaling design systems, or building long-term accessibility programs. These issues are also interconnected: problems in one area tend to amplify problems in others.
Subjectivity In Scoring
Scored evaluations introduce room for subjective interpretation. Without standardized calibration, the same user flow could get different scores depending on who is doing the evaluation. That makes comparability and repeatability harder, especially in procurement or multi-vendor contexts. A simple alternative text might be rated “adequate” by one team and “unclear” by another.
Loss Of Clear Compliance Thresholds
Subjectivity also weakens clear compliance boundaries. Binary “compliant” or “not compliant” labels are replaced by more flexible but less definitive outcomes. That complicates legal enforcement, contractual definitions, and audit reporting. A product could earn a “Good” rating while still having critical usability gaps for certain users — creating a disconnect between the score and the actual accessibility experience.
Friction With Existing Legal Frameworks
Blurred compliance standards also create tension with current laws. Many regulations explicitly reference WCAG 2.x and its A, AA, and AAA levels — including Section 508 of the Rehabilitation Act of 1973, the European Accessibility Act, and the UK’s Public Sector Bodies (Websites and Mobile Applications) (No. 2) Accessibility Regulations 2018. Until WCAG 3.0 is formally mapped to those standards, using it in regulated settings introduces risk. Teams in healthcare, finance, or public sectors will likely run dual conformance strategies in the meantime, which adds cost and complexity.
The Minimum Viable Accessibility Trap
That ambiguity could also encourage a “minimum viable accessibility” mindset. In deadline-driven environments, teams may decide that reaching the Bronze tier is good enough and stop improving — even if essential barriers remain.
Consider a mobile app with strong keyboard support but missing audio transcripts. It could still hit a passing tier, leaving some users excluded.
A More Honest Measure Of Success
WCAG 3.0 is a step toward accessibility that reflects real user diversity. By moving from checklists to scored evaluations and from rigid technical compliance toward practical usability, it pushes teams to concentrate on real-world impact rather than theoretical perfection.
At the same time, WCAG 3.0’s proposed scoring models introduce new responsibilities. Without clear calibration, stronger enforcement patterns, and a cultural shift away from “good enough,” we risk losing the very clarity that made WCAG 2.x enforceable and actionable. The promise of flexibility only works if we use it to aim higher, not to settle earlier.
Experienced teams have seen the downside of the checklist approach: hours spent fixing minor color contrast issues while broken keyboard navigation kept screen reader users from completing core tasks. WCAG 3.0’s outcome-focused view is a reminder that accessibility is fundamentally about functionality and inclusion.
This shift gives design, development, and product leadership a chance to redefine what success means. Accessibility isn’t about ticking boxes — it’s about enabling people. Teams that prepare now, keep these risks in mind, and anchor their work to user outcomes won’t just be ready for WCAG 3.0; they’ll build digital experiences that are actually usable, sustainable, and inclusive.



