Skip to content
Probatio

Work product

Sample AI diligence report

An abridged, composite version of the deliverable, showing the structure of the analysis and reporting.

Composite case study

Snapshot

AI investment diligence snapshot

24
Claims tested
7
Unsupported
5
Material findings
MODERATEModel performance
HIGHClaims vs. evidence
HIGHModel dependency
MODERATEData moat
CRITICALInference economics
MODERATEScalability
HIGHAI governance
LOWProduct adoption
LOWTeam & execution

Section 1

Claims tested against evidence

Each marketed claim is restated, matched to the evidence actually reviewed, and resolved into an investment impact.

Claim vs. evidence

Claim
The product is fully autonomous — agents complete workflows without human review.
Evidence reviewed
60 days of production traces, escalation queue exports, two operations interviews
Finding
31% of agent runs route to a human reviewer, concentrated in the two highest-value workflows.
Confidence
High
Investment impact
Labour savings in the model overstate realised automation. Adjust the efficiency case and treat review-rate reduction as a value creation milestone rather than a delivered capability.

Claim vs. evidence

Claim
Our models are trained on proprietary customer data no competitor can access.
Evidence reviewed
Data processing agreements, training pipeline review, dataset lineage sampling
Finding
Fine-tuning uses roughly 4% proprietary data; the remainder is public corpora. Two of five enterprise contracts prohibit training use.
Confidence
Moderate
Investment impact
The data moat is thinner and legally narrower than presented. Discount defensibility, and treat contractual training rights as a diligence condition and a post-close remediation item.

Section 2

Material findings

Findings are ranked by materiality to the thesis, not by technical interest.

Risk: HIGH
Finding
No held-out evaluation set exists; regression testing is manual and run at release by one engineer.
Evidence
Repository review, CI configuration, release checklist, engineering interviews
Investment impact
Quality regressions reach customers before detection, surfacing as churn and support cost rather than an engineering metric, and blocking safe model upgrades.
Recommendation
Fund an automated eval suite with a frozen held-out set in the first 90 days and gate provider migration on it.
Risk: MODERATE
Finding
Retrieval quality degrades on tenants above roughly 50k documents.
Evidence
Latency and relevance telemetry by tenant size, index configuration review
Investment impact
The largest prospects experience the worst product quality, capping the upmarket motion the growth case depends on.
Recommendation
Scope a retrieval re-architecture; hold enterprise ARR assumptions flat until it lands.

AI risk heatmap

Diligence areas scored across evidence quality, severity, likelihood and remediation difficulty. Colour carries the pattern; labels confirm it.

AreaEvidenceSeverityLikelihoodRemediation
Model & evaluation
HIGH
MOD
HIGH
LOW
Data & retrieval
MOD
LOW
MOD
MOD
Economics
HIGH
CRIT
MOD
HIGH
Scalability & reliability
MOD
HIGH
LOW
MOD
Governance & security
CRIT
MOD
MOD
LOW
Product & adoption
LOW
MOD
HIGH
MOD
LOWMODHIGHCRIT

Section 3

Dependency, defensibility and economics

What the company controls, what it rents, and what happens to margin as usage scales.

AI dependency map

Solid cyan = proprietary and owned. Dashed amber = an external dependency the target does not control.

Proprietary External
ProductUI · permissions · audit trailOWNEDAI Orchestratorrouting · tools · guardrails · evalsOWNEDModel Providerstwo frontier vendors · no tested fallbackEXTERNALRetrievalown chunking · tenant vector indexOWNEDCloudsingle-region managed infrastructureEXTERNALDatacustomer corpora · labelled outcomesOWNED

Inference economics

Revenue indexed to 100. Inference cost and residual gross margin as monthly token volume scales.

Monthly tokens processed (x). Margin compression at scale is a pricing-model question, not an infrastructure question.

AI moat matrix

Defensibility by stack layer. The filled block marks where the layer actually sits; block height rises with strength.

LayerLowModerateStrong
  • Model

    Third-party foundation models, no fine-tune ownership

    Low
  • Data

    Labelled outcome data accumulated per customer

    Strong
  • Workflow

    Embedded in approval chain, replicable in ~2 quarters

    Moderate
  • Integrations

    Six systems of record, each a switching cost

    Moderate
  • Customer Data

    Tenant-scoped history improves retrieval quality

    Strong
  • Brand

    Category awareness undifferentiated in buyer interviews

    Low
  • Distribution

    Two channel partners, concentration in one

    Moderate

Worked example

Inference economics

Technical observation
Inference cost per active account rises roughly linearly with usage because context windows are not trimmed and retrieval returns a fixed large payload per call.
Business consequence
Gross margin compresses precisely as the largest customers expand, so the accounts that drive the growth case are the least profitable to serve.
Investment implication
Underwrite margin as a band tied to usage, not a point estimate, and treat prompt and retrieval optimisation as a funded first-year workstream with a measurable cost-per-account target.

Section 4

Investment thesis matrix

Each thesis assumption is tested and assessed, so the committee can see which parts of the case the evidence supports.

Investment thesis matrix

Each thesis assumption tested against evidence rather than against the management narrative.

AssumptionEvidenceFindingAssessmentInvestment impact
AI accuracy is materially ahead of incumbentsInternal eval set, 3 customer benchmarks, no held-out test setAdvantage holds on two of six task types; eval set overlaps training dataPartially supportedNarrow the win-rate assumption; fund an independent evaluation before close
Gross margin expands with scale12 months of provider invoices mapped to usageInference cost grows faster than seat revenue on agentic workloadsNot supportedAssume margin compression until pricing moves to consumption
AI features drive expansion revenueProduct telemetry, 20-account cohort, expansion ledgerExpansion concentrated in accounts using two of five AI featuresSupported, narrowlyPrioritise activation of the two proven features in the value creation plan

Diligence reduces uncertainty and documents what the evidence supports. It does not guarantee that every risk will be identified, and we make no claims about investment outcomes.

Want this run on a live deal?

Bring the thesis, the data room and the timeline. We will tell you what evidence exists, what is missing, and what it means for the deal.