Skip to content
Probatio

Primary practice

AI Due Diligence

The core engagement. We establish what the AI actually does, what evidence supports the claims made about it, and what the answers mean for price, structure and the value creation plan.

Scope

What we examine

AI due diligence is not a code review with AI vocabulary attached. It is an evidence exercise aimed at the specific ways AI companies fail.

AI performance and claims

  • Evaluation methodology: what is measured, on what data, against what baseline, and whether the test set is genuinely held out.
  • Production behaviour versus demo behaviour, read from traces rather than from the pitch.
  • Failure modes, hallucination and escalation rates in the workflows that carry commercial weight.
  • Human-in-the-loop reality: how much of the claimed automation is actually automated.

Defensibility and economics

  • Model dependency: what is owned, what is rented, and what happens on provider deprecation or repricing.
  • Data moat: whether accumulated data measurably improves the product or is merely stored.
  • Inference economics and gross margin at realistic usage, not at pilot usage.
  • AI governance: policy, access controls, data handling, model change management, audit trail.

The frame

Eight questions, applied in order

These structure the workplan, the management sessions and the final report.

01

Does it work?

Measured performance against the task the buyer is actually paying for.

02

Is it real?

Claims traced to evidence: evals, logs, code, not demo footage.

03

Is it proprietary?

What is owned versus rented from a foundation model provider.

04

Is it scalable?

Behaviour under load, latency budgets, failure modes.

05

Is it economical?

Inference economics and gross margin at realistic usage.

06

Is it valuable?

Whether AI features drive adoption, retention and price.

07

Is it durable?

Data moat, switching costs, model-provider commoditisation.

08

Can the org execute?

Team depth, evaluation discipline, AI governance maturity.

How findings read

Observation, consequence, implication

Every finding is carried through to what it changes about the investment.

Technical observation
Prompt templates, model versions and retrieval parameters are edited directly in production with no versioning or rollback path.
Business consequence
Output quality changes cannot be attributed, reproduced or reversed, so quality incidents become customer-facing before anyone can explain them.
Investment implication
This is an operational risk with revenue consequences, not a tooling preference. Treat change-management controls as a first-90-days condition and discount the reliability narrative accordingly.

Claim vs. evidence

Claim
The data we collect makes the model better every quarter.
Evidence reviewed
Retraining history, eval results by cohort, data labelling pipeline review
Finding
Data volume grows; measured task performance has been flat across three retraining cycles because labels are not outcome-linked.
Confidence
High
Investment impact
The compounding data-moat argument is not currently supported. Value the business on distribution and workflow lock-in, and treat outcome labelling as a funded initiative.

AI dependency map

Solid cyan = proprietary and owned. Dashed amber = an external dependency the target does not control.

Proprietary External
ProductUI · permissions · audit trailOWNEDAI Orchestratorrouting · tools · guardrails · evalsOWNEDModel Providerstwo frontier vendors · no tested fallbackEXTERNALRetrievalown chunking · tenant vector indexOWNEDCloudsingle-region managed infrastructureEXTERNALDatacustomer corpora · labelled outcomesOWNED

Inference economics

Revenue indexed to 100. Inference cost and residual gross margin as monthly token volume scales.

Monthly tokens processed (x). Margin compression at scale is a pricing-model question, not an infrastructure question.

AI moat matrix

Defensibility by stack layer. The filled block marks where the layer actually sits; block height rises with strength.

LayerLowModerateStrong
  • Model

    Third-party foundation models, no fine-tune ownership

    Low
  • Data

    Labelled outcome data accumulated per customer

    Strong
  • Workflow

    Embedded in approval chain, replicable in ~2 quarters

    Moderate
  • Integrations

    Six systems of record, each a switching cost

    Moderate
  • Customer Data

    Tenant-scoped history improves retrieval quality

    Strong
  • Brand

    Category awareness undifferentiated in buyer interviews

    Low
  • Distribution

    Two channel partners, concentration in one

    Moderate

Delivery

What you receive

Investment committee memo

Ranked material findings, each with evidence, confidence and investment impact.

Evidence appendix

Claims-versus-evidence register, dependency map, moat matrix, economics model.

Post-close agenda

The specific remediation and value creation items the diligence surfaced, sequenced.

Timelines are set against your process. Diligence reduces uncertainty and documents what the evidence supports; it does not guarantee that every risk will be identified, and we do not make claims about investment outcomes.

Scope

What we assess

The full sixteen dimensions. AI performance, defensibility and economics lead; technical and product evidence is pulled in wherever it changes the AI conclusion.

  • AI

    Model architecture

    What produces the output, and how much of it the company controls.

  • AI

    Evaluation discipline

    Held-out sets, regression suites, and whether results are reproducible.

  • AI

    Training & fine-tuning

    What was trained, on what, and whether it measurably improved the task.

  • Technical

    Data pipeline

    Ingestion, chunking, freshness, tenancy isolation and lineage.

  • AI

    Data rights

    Licensing, customer terms and training-use permissions behind the corpus.

  • Technical

    Retrieval quality

    Relevance and degradation as tenant corpora grow.

  • AI

    Agent reliability

    Tool-call errors, escalation rates and observed autonomy.

  • AI

    Inference economics

    Cost per unit of value delivered, and margin at realistic usage.

  • Technical

    Latency & scalability

    Behaviour under load, latency budgets and failure modes.

  • AI

    Model dependency

    Supplier concentration, exit paths and tested fallbacks.

  • Technical

    Security posture

    Customer data in prompts, logs, embeddings and traces.

  • AI

    AI governance

    Policy, model change management, human review and audit trail.

  • Technical

    Observability

    Whether quality regressions are detected before customers find them.

  • Product

    Product adoption

    Marketed AI features versus features actually used.

  • Product

    Workflow fit

    Where the AI sits in the customer's process and what it displaces.

  • Technical

    Team & execution

    Depth of people who can actually change model behaviour.

Considering an AI investment?

Bring the thesis, the data room and the timeline. We will tell you what evidence exists, what is missing, and what it means for the deal.