Is your AI meeting its objectives?
Get Assured AI tracks whether the AI your business runs satisfies its Ethical and Business Objectives - Management by Objectives, applied to your AI estate - and recommends how to remediate every gap through critic-based evaluation. It reads the execution logs your AI already produces, reconstructs each workflow, scores every objective green, amber or red, and hands you the cheapest fix first. Read-only. Inside your own tenant. Start with Yardstick.
Partnered with ollo in ANZ for co-building AI agents with your teams.
Management by Objectives, for AI.
Definition: set explicit Ethical and Business Objectives for every AI workflow you run, measure their satisfaction continuously from the telemetry the AI already produces, and remediate the worst gap first. Drucker's discipline, applied to the AI estate. An objective nobody measures is a hope; a measurement nobody acts on is a report. This is neither.
Twelve ethical values. Your business outcomes. The health of the AI itself.
A conversational wizard elicits your objectives once: which ethical values are priorities (with the worst case you must avoid and the best case you want promoted), which business and customer outcomes the AI exists to serve, and which engineering failure modes worry you. Nothing is invented on your behalf - the register records what you said, and gaps stay visible as open questions rather than being filled with plausible fictions.
All twelve, always assessed
Your named priorities carry your worst and best cases into every assessment; the rest are still measured. Fairness on a hiring workflow and sustainability on a support bot are different risks - the scoring knows the difference.
The value the AI was bought for
Failed-run rates, turnaround times, cost per query, customer-happiness signals - each objective grounded in a system that can actually measure it, so satisfaction is read, never estimated.
The engineering failure modes, named and scored
Hallucination (in text and tool calls), prompt injection, context rot, tool-call loops, latency bloat, silent degradation, unconfirmed side-effects, and the soundness of every answer against the evidence it was given: faithfulness, precision, recall, noise sensitivity and more, each its own tracked objective with its own traffic light per workflow.
From logs to objectives to remediation.
No agents installed, no code changed, nothing written back. Four steps, continuously.
Logs become a living map
Execution logs (from platforms such as Langfuse) are reconstructed into provenance graphs of every AI workflow actually running: prompts, tools, inputs, outputs, lineage. Shadow AI surfaces here, because anything that runs leaves traces.
Objectives, elicited once
The wizard captures your Ethical and Business Objectives as a register: priorities, worst and best cases, measures grounded in connected systems.
Critics judge satisfaction
Layered critics score every objective on every workflow: green, amber or red, with the evidence attached. Deterministic findings set a floor no judge can argue down.
Cheapest fix first
Each red or amber carries concrete recommendations. Your team acts; the next cycle of logs measures whether it worked.
Every objective resolves to one of three states.
Conformance (green): the evidence shows the objective satisfied. Gap (amber): evidence is missing, partial or drifting - the objective cannot yet be shown satisfied. Violation (red): the evidence shows it breached. Estate rollups take the worst state across workflows, so a green dashboard means every workflow earned it.
Scored by critics that check each other.
A single LLM grading your AI is one opinion grading another. Objective satisfaction is judged instead by layers of critics with different failure modes, arranged so none is trusted alone.
Rules that cannot be charmed
Pattern, statistical and schema critics: identifier and secret detection, readability scoring, cohort comparisons, tool-call validation against bound schemas. Their findings cap the score - an LLM can never talk a deterministic red back to green.
Judgement where rules cannot see
Bounded judge prompts review the same evidence per template and per trace: is this persona stereotyping, is this retention script manipulation, was this distress disclosure handled well. Judges raise concerns; they cannot lower floors.
Claims against evidence
Every answer and tool call decomposes into atomic claims; a verifier checks each against what the AI was actually given. Fabrication, injected instructions reaching actions, and answers left on the table become measured rates: hallucination, faithfulness, precision, recall.
The critics are themselves benchmarked.
Every critic is validated against gold-labelled evaluation sets before its verdicts count: a 200-example, 739-case benchmark with labelled violations for the soundness critics, and persona replays with planted traps (declined items, retractions, embedded injections) for the objectives wizard. In its latest run the deterministic tool-call critic caught 109 of 109 planted schema violations with zero false positives across 195 valid calls. When a critic cannot judge confidently, it abstains and says so - abstention is reported, never hidden.
Compliance is a byproduct of governing well.
The evidence that shows your Ethical and Business Objectives being satisfied is the same evidence the EU AI Act, ISO/IEC 42001, the NIST AI RMF and GDPR ask for. Track objectives well and the audit trail writes itself: EU AI Act risk classification per workflow, evidence attached to every score, and an append-only record a regulator can walk. Filing paperwork governs nothing; this governs, and the paperwork follows.
Built for the people who answer for the AI.
AI owners & operators
Every workflow actually running, its objectives, its lights, and the next cheapest fix - before an incident, rather than a post-mortem after one.
Risk & Compliance
Ethical objectives measured continuously instead of sampled annually, with the evidence for every framework already attached to every score.
Finance & the Board
Business objectives and their satisfaction on the same page as the ethical ones, so the investment case and the assurance case are one conversation.
Management by Objectives for AI, answered.
What is Management by Objectives for AI?
Setting explicit Ethical and Business Objectives for every AI workflow an organisation runs, measuring their satisfaction continuously from execution logs, and remediating the worst gap first. It applies Peter Drucker's Management by Objectives to the AI estate: objectives are elicited once through a guided wizard, then every workflow is scored against them, live.
What are Ethical Objectives for AI?
Twelve values assessed on every workflow: fairness, privacy, wellbeing, transparency, accountability, safety and reliability, security, human autonomy and oversight, contestability, inclusivity and accessibility, sustainability, and lawfulness. Your organisation names its priorities with the worst case to avoid and the best case to promote; every value is still assessed.
What is critic-based remediation?
Scoring by layered critics with different failure modes: deterministic critics set floors no judge can argue down, LLM judges add discernment where rules cannot see, and a prover-verifier critic checks every claim an AI makes against the evidence it was given. Each red or amber score carries concrete, cheapest-first recommendations. The platform recommends; your team decides and acts; the next logs measure whether it worked.
What data does it need?
Only the execution logs your AI already produces, read from observability platforms such as Langfuse. Read-only, inside your own tenant, and your data never trains a model.
How is satisfaction scored?
Every objective on every workflow resolves to a traffic light: green for conformance, amber for a gap in evidence or performance, red for violation. Estate views take the worst light across workflows, and no data ever shows as green - an unmeasured objective is amber, visibly.
Does this cover the EU AI Act and ISO/IEC 42001?
Compliance falls out of the same evidence. Objective satisfaction is tracked against the telemetry that the EU AI Act, ISO/IEC 42001, the NIST AI RMF and GDPR ask you to produce, including per-workflow EU AI Act risk classification, so the audit trail is a byproduct of managing well rather than a separate exercise.
See whether your AI is meeting its objectives.
Connect one log source and see your first objective scores within a week. Read-only, in your own tenant, with the evidence behind every light and the cheapest fix ranked first.
Book an objectives assessment