News & Resources

The Revenue-Cycle AI Scorecard
View all news

The Revenue-Cycle AI Scorecard:

Are You Measuring the Automation or the Problem?

By Matt Haberman, CTO of Annuity

Healthcare organizations are under tremendous pressure to have an AI strategy. But in revenue cycle, I think the more useful question is simpler: What problem are we actually trying to solve?

It is easy to measure AI by what the technology is doing. How many accounts did it touch? How many appeals did it draft? How many tasks did it automate? Those metrics show how much activity the technology performs. They don’t, by themselves, show whether it reduces total human effort or improves financial results.

That is why I believe a useful AI scorecard should start outside the AI itself and focus on three dimensions: problem coverage, earned autonomy and business impact.

Problem Coverage: How much of the problem can it actually address?

When we build automation, one of the first things I want to understand is the total opportunity. How big is the problem, and how much of it can the technology realistically address?

An AI tool may draft highly accurate appeals for one denial type while staff still prepare every other type. Its performance within that narrow scope may be excellent. But leaders also need to know what share of the workload it handles and how much financial opportunity those accounts represent.

Revenue cycle makes this particularly challenging. Payers have different policies. Contracts introduce additional requirements. Government programs and states operate differently. Data comes from different systems with varying levels of quality.

So I am interested not only in success rate, but in coverage and whether worthwhile coverage is expanding. Maybe today we can address 20% of the opportunity. With the next improvement, 40%. Eventually, perhaps 70%.

But maximum coverage should not automatically be the goal. Leaders also need to ask: How much additional value could we capture by expanding coverage, and what would it cost to get there?

Data quality factors into coverage, too. Poor or incomplete data can shrink the portion of the opportunity the technology can safely address. And the remaining opportunity may be disproportionately complex and expensive to automate. The goal is not to automate everything. It is to expand where the economics justify it.

Earned Autonomy: Is the technology earning greater trust?

We can think about AI like training a dog before letting it off leash. You establish expectations, observe behavior and build confidence. As the dog demonstrates reliability, you give it more freedom.

AI should earn autonomy in much the same way.

Human review at the beginning of an implementation is not necessarily a failure. It provides the evidence needed to determine whether greater autonomy is warranted.

The primary question is whether the technology is consistently producing work at the required level of quality. Alongside that, leaders should measure the human effort required to review, correct and rework its output. As quality becomes demonstrable, does oversight decline? Are exceptions and defects decreasing? Can people move from reviewing everything to reviewing only exceptions?

If quality is not consistent, the technology may not be ready for additional autonomy. But those exceptions are valuable, too. They show us where the model, data or workflow needs to improve.

Recent KLAS Research illustrates why this distinction matters. In interviews with leaders at 30 Epic customer organizations, KLAS found some of the same AI tools cited among both the most and least impactful. The difference often came down to workflow fit, configuration, validation and adoption.¹

The technology can work and the implementation can still fail to create value.

Business Impact: Did the outcome actually change?

Revenue-cycle leaders need operational details such as coverage, automation rates, exceptions, defects and human intervention. A CFO ultimately needs a different view:

Net financial benefit: What value is the solution creating through additional collections and cost savings, after accounting for technology, implementation, maintenance and oversight costs?

Cash acceleration: How much sooner are we receiving payment, and how much cash does that release from receivables?

An initiative does not have to increase collections to create value. It may reduce costs, improve collections, or do both. Cash acceleration should be considered separately because faster payment can create meaningful value even when the ultimate collection amount does not change.

We also need to be realistic about timing. An automation may perform a task today, but in revenue cycle we might not know the financial outcome for 30 or 60 days. That means distinguishing between near-term indicators of progress and the business outcomes that follow.

Activity tells us whether we are moving. It should not be confused with proof that we have arrived.

Let the scorecard tell you what to do next

The value of a scorecard is not simply knowing the score. It should help determine the next action.

If problem coverage is limited, evaluate the incremental opportunity. What additional value could greater coverage create, what will it cost, and can the same quality be maintained?

If coverage is growing but quality is inconsistent or human effort remains high, learn from the exceptions before granting greater autonomy.

If operational performance looks strong but net financial benefit does not, reconsider whether the technology is solving a valuable enough problem or whether its total cost is justified.

And sometimes, maintaining a profitable application within its current scope is the right answer. More coverage is not inherently better. Expand where the incremental value, cost and quality justify it.

David Chou recently wrote in Forbes that healthcare’s AI challenge is increasingly one of execution, not access to technology.² I agree. But good execution requires being precise about both the problem and what success looks like.

The organizations that create the most value from AI may not be the ones that automate the most work. They will be the ones that know which problems are worth solving, how much technology can reasonably address and what evidence warrants giving it more responsibility.

AI shouldn’t be measured by how much activity it produces. It should be measured by how meaningfully it changes the problem we brought it in to solve.

 

 

References

  1. Blauer, Tyson. “Epic AI: What’s Working, What’s Not, and What Customers Need Next.”KLAS Research, May 21, 2026.
  2. Chou, David. “Why Healthcare AI Still Struggles to Deliver.”Forbes, April 21, 2026.

 

About the Author

Matt Haberman is Chief Technology Officer at Annuity Health, where he focuses on simplifying and improving healthcare revenue cycle operations through AI, automation, and analytics. He brings 25 years of experience across healthcare and technology, including leadership roles as SVP of Product and Engineering at R1, CTO at Acclara, and CEO at MSM. His approach starts with understanding the everyday challenges healthcare teams face and applying technology in practical ways that make their work easier, improve performance, and deliver measurable results.