A recent Cursor study reveals that AI coding benchmark scores, particularly on SWE-bench Pro, are inflated due to answer retrieval rather than actual reasoning, with top models relying heavily on existing fixes. This discrepancy can lead to misleading enterprise procurement decisions, highlighting the need for stricter evaluation standards to differentiate coding ability from retrieval skills.













