Claim vs. evidence
- Claim
- “Our model is 40% more accurate than anything else in the category.”
- Evidence reviewed
- Internal eval harness, 3 customer benchmark files, 90 days of production traces, no held-out test set
- Finding
- The 40% figure comes from a curated set that overlaps training data. On production traces the advantage is 6–11% and only on two of six task types.
- Confidence
- Moderate
- Investment impact
- Win-rate assumptions built on category-wide accuracy leadership are not supported. Re-underwrite the differentiation premium and budget for an independent evaluation.