Fair warning: I'm writing this piece as the founder of an AR automation company that uses machine learning to predict payment behavior and generate follow-up emails. My professional interest is for you to believe that this technology works. Given that, you should hold what I say here to a high standard.
The "AI" label has been applied to so many finance tools in the last three years — many of which are rule-based automation with a language model bolted on for marketing purposes — that healthy skepticism is not just appropriate, it's professionally responsible. If you're an AR manager evaluating collections automation tools, here is an honest account of what machine learning actually does in this context, what it doesn't do, and how to evaluate whether a specific tool is doing what it claims.
What "AI" Actually Means in Collections Automation
There are two meaningfully different capabilities that get bundled under the "AI" label in AR and collections software:
Predictive scoring — machine learning models trained on payment history data to score invoices or customers by late-payment probability. This is statistical pattern recognition, not magic. If Customer X has paid late 7 of the last 9 invoices when the amount exceeded $15,000, a well-trained model will score their current $20,000 invoice as high-risk. This is math, and it works reliably when trained on adequate data.
Language generation — using large language models to draft personalized follow-up emails based on invoice details, customer history, and context rules. This is a different capability: it produces human-readable text, not a numerical prediction. The quality of the output depends on the quality of the prompting and the quality of the input data.
Many tools that claim "AI collections" are doing one of these, not both. Some are doing neither — they're sophisticated rule engines that use "AI" in their marketing copy because it's expected. Knowing which capability you're actually buying matters a lot for evaluating whether it will solve your problem.
The Skeptic's Questions for Predictive Scoring Claims
If a vendor claims their system predicts late payments, ask these questions:
"What data do you train the prediction on?" If the answer is "industry-wide data," be cautious. Industry-wide training data can give a baseline, but payment behavior is highly specific to your customer base, your industry, and your terms. A model trained on aggregate fintech data is not especially useful for predicting which of your 200 construction contractor customers will pay late. The better answer is "your own historical payment data, from day one."
"How many invoices of history do you need before the predictions are meaningful?" A legitimate answer will have a floor — something like "90 days of payment history, minimum 3 invoices per customer." If the answer is "it works immediately," ask how. There's no free lunch in prediction accuracy without training data.
"What is your prediction precision, and how is it measured?" Vendors who have tested their model can give you a number. "We identify X% of invoices that will go 15+ days overdue, with a Y% false positive rate." If they can't give you a number, either they haven't measured it, or the number isn't good. Neither is reassuring.
"What signals does the model use?" A good prediction model for AR draws on multiple signals: historical days-to-pay per customer, payment variance, invoice amount compared to historical averages, dispute history, seasonal patterns, time since last payment. If the answer is "we look at how many days overdue each invoice is," that's not prediction — that's a lag indicator. Being overdue is already a fact, not a prediction.
The Skeptic's Questions for Language Generation Claims
If a vendor claims their system drafts personalized collection emails, the question is: personalized how?
At the low end, "personalized" means mail-merge: Dear [Customer Name], your invoice [Invoice Number] for [Amount] is overdue. This is barely personalization — it's template population. It's marginally better than a fully generic email, but it's not drawing on customer behavior data.
More meaningful personalization includes tone variation based on customer payment history (a first-time late payer gets a different register than a serial slow-payer), content variation based on the customer's most common payment obstacles (some customers always pay when you reference the PO number; others need a formal statement attachment), and timing variation based on when that specific customer is most likely to act on an email.
To evaluate this, ask to see example email drafts for different customer archetypes with the same overdue scenario. If two customers with dramatically different payment histories produce essentially identical email drafts, the "personalization" is cosmetic.
What AR Automation Cannot Do
This matters as much as what it can do.
It cannot replace the relationship call for high-value accounts. A $600K-per-year customer who has gone quiet on a $90,000 invoice needs a phone call from a human who knows the account — ideally someone with whom they have a relationship. No automation substitutes for that. What automation can do is identify that this call needs to happen, prepare the caller with context, and follow up the call with a written confirmation.
It cannot resolve disputes. When an invoice is in dispute — the customer claims short-delivery, disputes the pricing, or has a legitimate deduction — automated email sequences are useless or counterproductive. Good AR automation should recognize dispute patterns and route those accounts to human resolution, not keep sending follow-up emails into a dispute that emails can't resolve.
It cannot work without historical data. If you're a new company, or you've recently changed your customer base significantly, the prediction models won't have adequate training data. For the first 60–90 days after connecting to a new system, you're getting baseline automation (consistent follow-up, email generation) without the predictive layer — the predictions get more accurate as payment history accumulates.
It cannot replace AR process hygiene. If invoices are going out with wrong PO numbers, wrong amounts, or missing information, automation will follow up on invoices that customers won't pay regardless — because the invoice itself is wrong. Collections automation makes a good AR process faster. It doesn't fix a broken one.
How to Evaluate Whether a Tool Is Actually Working
After 60 days with any AR automation tool, you should be able to answer these questions with data:
- Has DSO changed? By how much? Compare to the same period last year to control for seasonality.
- What percentage of overdue invoices are being resolved without human contact? If it's not materially higher than before automation, the email sequences aren't working.
- How many hours per week is your AR team spending on outreach tasks compared to before? If the answer is "about the same," the automation is adding process overhead, not reducing it.
- What is your 30-day aging bucket size compared to 60 days ago? This is the most direct measure of whether collections are improving.
We're not saying the results should be dramatic after 60 days — meaningful DSO reduction typically takes a full quarter to show up in the numbers. But the leading indicators (outreach coverage rate, email response rates, manual follow-up hours) should be moving in the right direction by week 6–8.
If a vendor can't help you track these metrics and tell you what good looks like for a company your size and industry, that's an important signal about whether their claims about results are based on measurement or marketing.
Healthy skepticism is not a reason to avoid AR automation — the actual capabilities are real and useful when implemented honestly. It's a reason to ask good questions before buying, measure outcomes rigorously after deploying, and not let "AI" in a vendor's name substitute for evidence that it works on your actual invoice portfolio.