Skip to main content
← Back to Blog
Research9 min readSeptember 27, 2026

Every VPAT We Studied Claimed More Than It Could Prove: What 6 Months of Research Found

A six-month academic research project analyzed 10 publicly available VPATs for products used in higher education. Every single one had a higher claimed conformance score than its evidence could support. Here is what we found, and how it shaped our VPAT Evaluator.


The problem institutions have not named yet

Every day, universities and public institutions sign contracts for software based on a vendor's own claims about its accessibility. A blind student may only discover that a required course tool does not work with their screen reader after study has begun. By then, the contract is signed and switching is not possible.

The document institutions rely on almost exclusively in this process is the Accessibility Conformance Report (ACR), also called a VPAT. It is written by the vendor, or by an evaluator the vendor pays, and submitted as evidence that the product meets accessibility standards.

What no one had done, until now, was apply rigorous evaluation methodology to ask a different question: not what the VPAT claims, but whether those claims are supported by evidence strong enough to rely on.

That is what this research set out to answer.

About the research

Over six months, our founder Stephanie Steer conducted a capstone research project as part of a Master of Evaluation at the University of Melbourne. The project developed a six-criterion rubric for judging ACRs as evidence documents, grounded in Scriven's account of evaluation and Stufflebeam's meta-evaluation tradition. That rubric was then applied to ten publicly available ACRs for products used in higher education.

The ten products in the sample were: Anthology Blackboard, Brightspace Core, Canvas LMS, Elsevier Digital Commons, Turnitin Similarity Report, Google Gemini, Handshake, Okta Verify, Word Online, and Zoom Workplace for Desktop. All ten were publicly available, and all ten are in active use across Australian and North American higher education.

This research was conducted independently as academic work. Its findings directly informed the design of our VPAT Evaluator.

The six criteria

The rubric asks what an ACR produced under potential conflict of interest must contain for a procurer to scrutinise its argument. It does not ask whether the product is accessible. It asks whether the report gives you enough to judge how much reliance to place on the vendor's claims.

Six criteria emerged from the literature:

C1 — Claim specification. Does the report identify the product, version, component scope and date, and tie each rating to a named standard and conformance level? A VPAT that does not specify which version of the product was tested is a VPAT about nothing.

C2 — Scale integrity. Are the rating terms defined, and is the threshold between them stated? The VPAT template provides a glossary but does not define where "Supports" ends and "Partially Supports" begins. Every vendor uses these terms differently.

C3 — Scope and consumer disclosure. Does the report state what was tested and, critically, what was not? Does it name the assistive technologies and access needs in scope? A conformance claim without a stated testing scope cannot be evaluated.

C4 — Evidential warrant and independence. Are the test methods, assistive technology configurations, tester identity, tester independence, and date of testing disclosed? This is the criterion that asks whether the evidence was actually gathered.

C5 — Scrutability of reasoning. Do the remarks show how findings support each rating, and are residual barriers disclosed? A rating of "Supports" with no explanation provides no information to the reader.

C6 — Decision utility. Does the report tell a procurer what they need to act? This includes the vendor's remediation position, who owns it, and when identified barriers will be resolved.

What we found

The headline finding was consistent and striking: every single ACR in the sample had a higher claimed WCAG conformance score than its evidence quality score. Every one. The gap ranged from 8 to 31 points, with a mean divergence of 19.7 points.

Product Claimed WCAG conformance Evidence quality Gap
Anthology Blackboard907713
Brightspace Core928012
Canvas LMS1007327
Elsevier Digital Commons755322
Turnitin Similarity Report967026
Google Gemini746311
Handshake71638
Okta Verify987028
Word Online845331
Zoom Workplace967719

Both scores are expressed out of 100. The claimed WCAG conformance score is computed from the vendor's own criterion ratings. The evidence quality score is computed from the rubric assessment of how well the report supports those ratings.

The findings that surprised us most

The highest conformance claims had the widest gaps. Three of the four highest claimed conformance scores were also among the four largest divergences. Canvas LMS claimed a perfect 100, every criterion rated "Supports," and received a 73 for evidence quality, a 27-point gap. Canvas's ACR received its lowest rating on C5 because uniform "Supports" ratings with no explanation provide almost no information to a reader trying to understand what was tested and how.

A named external evaluator did not reliably close the gap. It was expected that third-party evaluation would produce stronger evidence quality. This did not hold across the sample. Some externally produced reports named no assistive technologies at all. Independence is one signal of credibility, but it does not substitute for disclosure of scope, methods, and reasoning.

Two criteria showed almost no variation — and that is itself a finding. C2, which assesses whether the vendor defined the threshold between their rating levels, was rated Partial across all ten reports. None defined where "Supports" ends and "Partially Supports" begins. This appears to be a genre property, a structural limitation of the VPAT template, not an individual failing. C6, which asks whether the report tells you what the vendor will do about identified barriers, was rated Absent in nine of the ten reports. Remediation timelines and ownership were almost never disclosed.

Scope disclosure ranged dramatically. On C3, one report named three assistive technologies by name (JAWS, NVDA, and VoiceOver) and described the testing scope in detail. Two named none at all. A conformance claim with no stated testing scope cannot be evaluated — you have no way of knowing what was tested, by whom, on what, or when.

What this means for procurement teams

The core implication of this research is a shift in how procurement teams should read a VPAT. The standard question is: what conformance level does the vendor claim? The research suggests the more useful question is: does this report give me enough to judge how much reliance to place on that claim?

A high conformance score does not mean a product is accessible. It means the vendor has rated it highly. What makes an ACR useful for procurement is whether the report explains the scope of what was tested, names the assistive technologies used, discloses the tester's identity and independence, shows the reasoning behind ratings, and identifies what barriers remain and when they will be addressed.

Where a report leaves these gaps, the appropriate response is neither rejection nor acceptance. It is a targeted request for the missing information, backed by contract conditions that require it.

How this research shaped our VPAT Evaluator

The rubric developed in this research is the foundation of our VPAT Evaluator's scoring methodology. The tool scores both claimed WCAG conformance and evidence quality, and makes the gap between them visible.

When you upload a VPAT to the Evaluator, you are not just getting a conformance summary. You are getting an assessment of the five dimensions of evidence quality that this research identified as most consequential for procurement decisions: claim specification, scope disclosure, evidential warrant, scrutability of reasoning, and decision utility.

The research also identified the C5 criterion — scrutability of reasoning — as the most consequential differentiator between reports. A VPAT that rates all criteria "Supports" with no remarks provides almost no evidence to a reader. The Evaluator's scoring reflects this: it weights the specificity and completeness of vendor remarks, not just the conformance levels they assign.

Building this tool from academic research, rather than from intuition or industry convention, means that the scoring rubric can be explained, defended, and critiqued. The criteria are traceable to published evaluation theory. The ratings are anchored against stated conditions, not relative rankings. And the gap between claimed conformance and evidence quality is calculated and displayed, not hidden.

A VPAT is the beginning of the conversation

A WCAG conformance rating is the beginning of a procurement conversation, not the end of one.

Institutions that treat a vendor's VPAT as a verdict are making procurement decisions without the information they need to make them well. This research provides a principled basis for asking the right questions of vendor evidence, and the VPAT Evaluator puts that research to work in seconds.

See the evidence quality behind any VPAT.

The VPAT Evaluator scores both claimed conformance and evidence quality, and shows you the gap between them.

Try VPAT Evaluator