Judging what you are told.
QUESTIONS TO ASK
What specifically does it do? What was it tested on? What does it get wrong, and how often? What happens when it is uncertain? Who is accountable when it is wrong? What data does it use, and where does that data go?
WARNING SIGNS
Claims of understanding or intelligence, rather than specific capabilities Accuracy figures with no description of the test conditions Demonstrations on carefully chosen examples Reluctance to discuss failure modes "AI-powered" with no explanation of what that means
WHAT A GOOD ANSWER SOUNDS LIKE
Specific about the task, honest about limitations, clear about where human review is needed.
TESTING IT YOURSELF
On your actual data, on your actual task, including difficult cases.
A demonstration proves the demonstration works.
THE MOST USEFUL QUESTION
"Show me where it fails."
A vendor who can answer that understands their product. One who cannot has not looked.
FOR ANY CONSEQUENTIAL DECISION
Ask how a wrong output would be detected.
If the answer is that it would not be, that is the problem.