Independent AI behavioural due diligence
A company says its AI is controlled, safe, accurate or differentiated. Kronaxis tests the behavioural claim rather than repeating the vendor's architecture story, and reports what survives an independent test and where it fails. It sits beside your technical, cyber and legal diligence, not instead of them.
Provable, not plausible. Grounded in published research, the Distinct Fields series, ten papers on Zenodo with a full replication package.
What we test
Does the system actually produce the behavioural change the vendor claims?
What else changes when the claimed property is moved?
How stable are the results across prompts, samples, models and domains?
Where does the system fail, and are those failures systematic?
Does the vendor's own evaluation overstate performance against independent readers and simple baselines?
Are the policy, safety or assurance claims supported by the mechanism actually deployed?
Where it fits a transaction
The first mandate
Select the AI claims that materially affect the transaction. Kronaxis freezes a test for each, evaluates it independently from available artefacts and outputs, and reports what survives.
What you receive
An executive opinion on each tested claim, with an evidence matrix: claim, test, result, confidence, limitation.
Where the system fails and how systematically, and a comparison against the vendor's reported evidence and against simple baselines. Plus the questions the buyer should require the target to answer before completion.
Indicative first mandate, depending on transaction scope and access to artefacts. Intentionally modular: it sits beside technical, cyber, legal and financial diligence rather than displacing them, and focuses on the behavioural and control claims those workstreams usually do not test.
Most AI diligence asks what model is used, how it is hosted and whether policies exist. We ask the more basic commercial question: does the system actually behave as represented, and can you independently verify it?
Boundaries, stated first
This is an assurance and diligence engagement, not a legal opinion or a regulatory certification. It evaluates the claims and artefacts supplied and does not assert behaviour outside that scope. Where a criterion is learned or judgement based, we report it as judgement rather than proof, and formal or cryptographic components are described as proven only where the mechanism supports it. We do not claim to read minds or steer people.
Before you complete
If you have an AI heavy deal where the capability matters to the price, we can scope three to five claims as a fixed fee workstream, fast, from the artefacts you already hold.
Start the conversation
A short message is enough. A person replies, not a bot, and we scope three to five claims as a fixed fee workstream.
Prefer email? Write to hello@kronaxis.co.uk.