A reference a vendor hands you has already been asked, and already said yes: it proves the vendor can produce one happy customer, not how it performs when your implementation hits a snag. The calls that actually predict outcome are the ones you source yourself: through a 20 Group peer, a NIADA contact, or an operator in your network with no reason to spin the answer.
Why a vendor-supplied reference doesn’t tell you much
Every vendor reference list is a survivorship sample. Out of every dealer group that piloted the product, the names on that list are the ones who (a) had a good experience, (b) are willing to say so on the record, and (c) the vendor’s customer success team is confident won’t go off-script. That’s not dishonesty, it’s just selection. A vendor with fifteen live customers and three willing to talk isn’t lying to you by handing you those three; they’re doing what any sales process does.
The problem is that a filtered sample answers the wrong question. You don’t need to know if the product can work somewhere. You need to know how it behaved during implementation, what broke, who fixed it, and how long that took, at an operation with your volume and your mess of legacy systems. A curated reference will give you a polished version of that story, usually the version optimized for “would recommend,” not the version with the actual timeline and the actual friction.
This is exactly why 20 Groups exist in the used-car business in the first place: fifteen to twenty non-competing dealers who share financials and KPIs monthly specifically because internal numbers, heard from a peer with nothing to sell you, carry more weight than the same numbers in a vendor deck. The same logic applies to AI vendor references. A peer contact reached independently has no incentive to filter the answer, because they’re not getting anything out of your decision either way.
Where to source a reference that isn’t filtered
Three sources consistently produce unfiltered references, in rough order of value:
Your 20 Group or NCM/ARC cohort. If anyone in your group has piloted an AI operations tool, whether from this vendor or a comparable one, that’s the first call to make. It’s usually a two-minute ask at your next meeting or a message in the group chat, and the answer comes with none of the incentive structure a vendor-sourced reference carries.
A non-competing peer outside your group. Auto Remarketing, F&I and Showroom, and Used Car Dealer magazine cover enough of this space that you can often identify who’s running a given product just from trade press. A cold outreach (“saw you mentioned X at the NIADA conference, would you have ten minutes”) works more often than people expect, because dealers who’ve been burned by an overpromising DMS vendor are usually happy to warn a peer.
LinkedIn and public case studies, cross-checked. If a vendor publishes a named case study, that’s a starting point, not an endpoint. Find the operator quoted in it and reach out directly instead of relying on the quote as written. The quote in the case study went through marketing review; the person who said it did not sign up to have every follow-up question pre-approved.
What doesn’t count as an unfiltered source: the reference list the vendor emails you, a testimonial on the vendor’s website, or a reference the sales rep “just happens to know” and offers to connect you with on the call.
Key insight
All three are the same filtered channel wearing different labels.
What questions actually get honest answers
General satisfaction questions invite a polished answer, because “were you happy with it” is exactly the question a reference expects and has a rehearsed response for. Specific questions about what broke get a different kind of answer, because they’re harder to spin on the spot:
- "What went wrong during implementation, and how long did it take to fix?" Every real deployment has a rough patch. If the answer is "nothing," that's not a clean implementation, that's a reference who hasn't thought about the question.
- "What did the vendor's team look like during the first 60 days? Were you talking to the same people the whole time, or did you get handed off?"
- "What's one thing you wish you'd asked before signing?" This is the single highest-yield question in a reference call. It reliably surfaces the thing that didn't come up in the sales process.
- "If volume doubled tomorrow, would this hold up, or would you be back on the phone with support?"
- "Would you have signed the same contract again, knowing what you know now?" Not "would you recommend it," which is a yes/no reputational question. This one forces a real comparison against the decision they actually made.
Listen for specificity in the answer itself, not just the content. A reference who says “error rate went from 7% to 1%, and it took about six weeks to get there” is telling you something you can act on. A reference who says “it’s been great, really improved our efficiency” is telling you they’re being polite. The AI Deal Engine case study is an example of the level of detail a real reference conversation should produce: named before/after numbers, a timeline, and what specifically changed operationally, not a general endorsement.
Running the call
Keep it to 20-30 minutes and go in with the questions above written down, not improvised. A few practical notes on running it well:
- Ask to talk to the person who lived through implementation, not just the executive sponsor. The COO who approved the purchase has a different answer than the ops manager who fielded the support tickets. If you can get both, get both.
- Ask about volume and complexity relative to your own operation. A reference running 200 units a month at a single-rooftop store tells you less about what will happen at a five-store group doing 2,000 evaluations a week than a reference closer to your actual shape.
- Don’t let the vendor sit in on the call. If the sales rep insists on joining “just to help with context,” that’s worth noting on its own. A reference who’s comfortable talking without the vendor present is a stronger signal than one who isn’t.
- Ask what happens when something breaks now, in production, not during the pilot. Escalation paths and support commitments tend to erode after the contract is signed; a reference eighteen months in will tell you what that actually looks like.
This is one piece of a broader evaluation. Our AI vendor checklist for dealership operations covers the rest of the process, including DMS integration questions, security requirements, and how to structure a pilot; how to evaluate AI vendors for dealership operations walks through scoring vendors against your own KPIs rather than the vendor’s. If explainability matters to your compliance team specifically, why an AI document review tool needs a real audit trail covers the questions to ask about how decisions get logged and defended after the fact. And if the pitch is “the AI decides, a human reviews the exceptions,” human-in-the-loop AI for dealership operations explains what that phrase should actually mean in practice, since vendors use it loosely.
FAQ
Why are vendor-provided references less reliable than peer-sourced ones?
A vendor selects the references most likely to give a positive account, which is a rational thing for a sales process to do but makes the sample useless for predicting your own outcome. A peer contact reached independently, through a 20 Group or industry connection, has no incentive to filter the answer, so what you hear is closer to the unedited version.
What questions get the most honest answers in a reference call?
Specific questions about what broke during implementation and how it got resolved beat general satisfaction questions every time, because a specific question is harder to answer with a rehearsed line. “What’s one thing you wish you’d asked before signing” is usually the single most useful question in the whole call.
Once you’ve done the reference legwork and you’re ready to see what an AI operations deployment looks like end to end, Deskflow walks through how the process, the exceptions, and the audit trail actually work in production.