CMS Scores Its Medicare AI Prior Authorization Vendors on Two Measures, Neither About the AI
CMS scores the contractors that run prior authorization and pre-payment review in its WISeR Model on two quality measures. Table 1 of the WISeR Data Reporting Guide names “WISeR-1: Timeliness of Determinations”, which participants report themselves, and “WISeR-2: Accuracy of Determinations” (p.98), calculated from a quarterly audit against Original Medicare coverage criteria (p.99). Neither is a measure of the AI system. That is the whole set for Performance Period 1, though the guide adds that CMS “may modify the WISeR Performance Measure Set for future performance periods.” The table is legible because records the Electronic Frontier Foundation obtained from CMS through a Freedom of Information Act lawsuit were posted on September 8 and 9, 2026; page numbers refer to the 443-page Combined Records file released in the case.
A CMS document headed “WISeR Model Payment Methodology” lists the model’s safeguards. Under the heading “WISeR Safeguards to Support Appropriate Determinations”, it says “The WISeR payment methodology is designed to discourage inappropriate non-affirmations through several mechanisms”. It puts human review first: “Every non-affirmation must undergo a medical review by an appropriately licensed WISeR clinician or WISeR clinical reviewer to confirm the request does not meet Medicare coverage criteria” (p.80). A non-affirmation means a request did not meet those criteria. The same passage records no limit on resubmitting a non-affirmed prior authorization request, and a right to ask for peer-to-peer clinical review of each resubmission. Four mechanisms follow (p.81): CMS auditing and quality scoring; one payment per beneficiary for a given WISeR item or service however many times a request is resubmitted, at the contractor’s own cost; appeal rights; and stakeholder feedback. On appeals, “WISeR Participants will not receive payment, and may face recoupment, for non-affirmations or denials that are subsequently overturned through a successful claims appeal” (p.81).
The points sit in Table 2 of the Data Reporting Guide: 9 available on WISeR-1, 16 on WISeR-2, 25 in total. “In PP1, the basis for scoring the WISeR quality measures will be pay-for-reporting. WISeR Participants can earn full points on WISeR-1 through complete and timely reporting of the required data elements that CMS uses to calculate WISeR-1” (p.98). Later in Performance Year 1 CMS “anticipates scoring the quality measures based on performance (rather than reporting)” (p.98), and separately “expects to begin scoring WISeR-2 on a pay-for-performance basis later in Performance Year 1” (p.100). Anticipates and expects are the documents’ own words; no date is attached.
Quality is one of three adjustments to the regional benchmark. Table 3 pairs an aggregate quality score of 85%-100% with a quality multiplier of 100%, 60%-84% with 95%, and under 60% with 90% (p.99). Section 2.3.9 of the Participant Guide, version 3.0 dated February 27, 2026, orders the calculation as regional benchmark, less the WISeR Discount, times the payment rate, times the multiplier. Section 2.3.7 fixes that rate: “The WISeR payment rate of 25% determines the portion of the regional benchmark that a WISeR Participant is potentially eligible to receive as a WISeR Payment for each payable non-affirmation, subject to further adjustments, including the WISeR Discount and the quality adjustment” (p.71).
The WISeR-2 audit is specified closely. “CMS audits every Select Item or Service at least once per calendar year and audits 25% of included services each quarter”, and “The Select Items or Services to be audited for a given quarter are not communicated in advance to WISeR Participants” (p.99). Sampling “is stratified by the Select Item or Service, the review method (prior authorization or pre-payment review), and the outcome”, and “each participant can expect about 120 records to be selected for each Select Item or Service being audited in a given quarter” (p.99). The guide says its sampling “generally aligns with” NCQA’s “8 and 30” file sampling procedure (p.100). “Initially, eight records are selected and reviewed; if any of the initial eight records fail the review, then an additional 22 records are selected for review, and the final assessment is based on all 30 records” (p.100). Table 4 rates accuracy High at 90%-100%, Medium at 60%-89%, Low under 60%. On documentation, “Failure to do so will negatively impact the WISeR-2 score and may also trigger remedial actions, up to and including termination from the model” (p.100).
Since January 1, 2026, WISeR has run in Arizona, New Jersey, Ohio, Oklahoma, Texas and Washington, and it applies to Original Medicare rather than Medicare Advantage. Every WISeR non-affirmation still requires the licensed clinician review described above. The Senate declined in July to take up a disapproval resolution against it, as we reported at the time. Deeper background sits in our health claims analysis and AI in insurance claims. The model’s terms above come from those three documents. Other CMS scoring or testing of the technology, if any exists, would be outside the pages cited here.
Organization material
eff.org →Substantive material an organization published about its own work, data, or incident.