Like for like.
WHAT TO KEEP CONSTANT
The task The prompt The input material
Comparing a well-crafted prompt on one tool against a lazy one on another proves nothing.
WHAT TO MEASURE
Output quality, judged consistently Editing required Errors introduced Speed Cost
WHAT TO BE AWARE OF
Recency bias. The tool you tried most recently seems best.
Novelty. A different style is not necessarily a better one.
HAVE SOMEONE ELSE JUDGE
Blind comparison, where the judge does not know which tool produced which output.
That removes most bias and is easy to arrange.
WHAT DIFFERENCES ACTUALLY MATTER
Accuracy on your domain How much editing the output needs Whether it admits uncertainty
WHAT DIFFERENCES MATTER LESS
Style, which prompting controls Interface, which you adapt to
THE DECISION
The one requiring least editing on your actual work.