AI Assurance · Coskew

Evaluation you can trace.

Coskew Assurance measures whether an AI system actually does what it claims: rigorous benchmarking, adversarial safety testing, and plain-language interpretability evidence. We test the system as deployed, not the version that was tested upstream. Every number resolves to an artifact. Nothing ships on a demo.

Same mean, same variance, different skew. The signal most evaluations miss.

ACL · EMNLP · NeurIPS · ICML
Peer-reviewed evaluation methodology
Generation ≠ verification
The verifier is never the author
Every claim → an artifact
Replayable, auditable trail

What we assess

Four ways we tell a working system from a convincing one

Scoped independently or as one engagement. We evaluate models and agentic pipelines you built, or ones you are deciding whether to trust, in the configuration you actually run: fine-tuned, wired to your tools, on your data.

Evaluation

Benchmarking that reflects the real task

Pre-registered rubrics frozen before we look at outputs, so the goalposts never move after the demo. Task-faithful metrics instead of vanity leaderboards, with the failure modes named, not averaged away.

Safety

Adversarial & red-team testing

Hard users and misuse probes driving the system through its real interface: jailbreaks, prompt injection, deception-relevant behavior, and the boundary cases a happy-path demo never touches. Structured findings, ranked by severity.

Interpretability

Interpretability evidence as a deliverable

A plain-language report a non-technical stakeholder can read: what a model learned, what changed after fine-tuning or a model swap, where it behaves unexpectedly, and what to do about it. Every feature and number traces to a representation-analysis finding, with scope limits stated, not hidden.

Verification

Traceability & claim auditing

We check that every claim in a report, deck, or model card resolves to a real artifact: no invented metric, no fabricated feature, no number without a path. If you are buying a system on the strength of a vendor's safety claims, we find out whether those claims still hold for what you were sold.

How we work

The methodology, borrowed from how we run research

These are not slogans; they are the invariants our evaluation pipeline enforces on every engagement.

i.

Pre-registered success criteria

The pass/fail bar is frozen in writing before anything is measured. It is the anti-p-hacking discipline from empirical research: we do not move the goalposts once we have seen the result.

ii.

No claim without an ingested source

Nothing is asserted unless it traces to a snapshot on disk. Market claims, safety findings, and pitch numbers all resolve to an artifact, and the ingestion timestamp precedes the citation.

iii.

Generation is never verification

The agent (or person) that produces a result is never the one that signs off on it. An independent adversarial verifier runs before any review, so a broken build or an untraceable number is caught before it reaches a decision.

iv.

A replayable audit trail

Every routing decision is logged with its reasoning and confidence. The whole evaluation can be replayed from the log, so a skeptical reviewer can retrace exactly how each conclusion was reached.

Who it's for

Two buyers, one through-line: trust before you deploy

Labs & safety-eval firms

You need interpretability and alignment evidence as a deliverable your own clients (labs, governments, boards) can read and defend. We produce the traceable report, you keep the relationship.

Interpretability reports Alignment audits Third-party evidence

Organisations running AI they did not build

A model triaging patients, an agent reading credit files, a pipeline answering clinical questions. You inherited safety claims about a system that has since been fine-tuned, swapped or rewired, and you have no way to check them. We do the checking, and we say plainly when the answer is no. Health and finance first, because that is where our team has worked.

Deployed models & agents Red-teaming Evidence for regulators

Put a number you can defend behind it.

Tell us what you are shipping, or what you are being asked to trust. We will tell you how we would evaluate it, and what it would take to make the evidence hold up. We usually start with a short pre-audit, one or two weeks, at little or no cost, so you can see what independent evidence looks like before committing to anything.