Management by Objectives For the AI your business runs

Is your AI meeting its objectives?

Get Assured AI tracks whether the AI your business runs satisfies its Ethical and Business Objectives - Management by Objectives, applied to your AI estate - and recommends how to remediate every gap through critic-based evaluation. It reads the execution logs your AI already produces, reconstructs each workflow, scores every objective green, amber or red, and hands you the cheapest fix first. Read-only. Inside your own tenant. Start with Yardstick.

Powered by ollo SOC 2 compliant Your data is never used to train models.

Partnered with ollo in ANZ for co-building AI agents with your teams.

The idea

Management by Objectives, for AI.

Definition: set explicit Ethical and Business Objectives for every AI workflow you run, measure their satisfaction continuously from the telemetry the AI already produces, and remediate the worst gap first. Drucker's discipline, applied to the AI estate. An objective nobody measures is a hope; a measurement nobody acts on is a report. This is neither.

The objectives

Twelve ethical values. Your business outcomes. The health of the AI itself.

A conversational wizard elicits your objectives once: which ethical values are priorities (with the worst case you must avoid and the best case you want promoted), which business and customer outcomes the AI exists to serve, and which engineering failure modes worry you. Nothing is invented on your behalf - the register records what you said, and gaps stay visible as open questions rather than being filled with plausible fictions.

FairnessPrivacyWellbeingTransparencyAccountabilitySafety & reliabilitySecurityHuman autonomy & oversightContestabilityInclusivity & accessibilitySustainabilityLawfulness
ETHICAL OBJECTIVES

All twelve, always assessed

Your named priorities carry your worst and best cases into every assessment; the rest are still measured. Fairness on a hiring workflow and sustainability on a support bot are different risks - the scoring knows the difference.

BUSINESS OBJECTIVES

The value the AI was bought for

Failed-run rates, turnaround times, cost per query, customer-happiness signals - each objective grounded in a system that can actually measure it, so satisfaction is read, never estimated.

PROMPT HEALTH

The engineering failure modes, named and scored

Hallucination (in text and tool calls), prompt injection, context rot, tool-call loops, latency bloat, silent degradation, unconfirmed side-effects, and the soundness of every answer against the evidence it was given: faithfulness, precision, recall, noise sensitivity and more, each its own tracked objective with its own traffic light per workflow.

How it works

From logs to objectives to remediation.

No agents installed, no code changed, nothing written back. Four steps, continuously.

01 · MINE

Logs become a living map

Execution logs (from platforms such as Langfuse) are reconstructed into provenance graphs of every AI workflow actually running: prompts, tools, inputs, outputs, lineage. Shadow AI surfaces here, because anything that runs leaves traces.

02 · SET

Objectives, elicited once

The wizard captures your Ethical and Business Objectives as a register: priorities, worst and best cases, measures grounded in connected systems.

03 · SCORE

Critics judge satisfaction

Layered critics score every objective on every workflow: green, amber or red, with the evidence attached. Deterministic findings set a floor no judge can argue down.

04 · REMEDIATE

Cheapest fix first

Each red or amber carries concrete recommendations. Your team acts; the next cycle of logs measures whether it worked.

The states

Every objective resolves to one of three states.

Conformance (green): the evidence shows the objective satisfied. Gap (amber): evidence is missing, partial or drifting - the objective cannot yet be shown satisfied. Violation (red): the evidence shows it breached. Estate rollups take the worst state across workflows, so a green dashboard means every workflow earned it.

Critic-based evaluation

Scored by critics that check each other.

A single LLM grading your AI is one opinion grading another. Objective satisfaction is judged instead by layers of critics with different failure modes, arranged so none is trusted alone.

LAYER 1 · DETERMINISTIC

Rules that cannot be charmed

Pattern, statistical and schema critics: identifier and secret detection, readability scoring, cohort comparisons, tool-call validation against bound schemas. Their findings cap the score - an LLM can never talk a deterministic red back to green.

LAYER 2 · LLM JUDGES

Judgement where rules cannot see

Bounded judge prompts review the same evidence per template and per trace: is this persona stereotyping, is this retention script manipulation, was this distress disclosure handled well. Judges raise concerns; they cannot lower floors.

LAYER 3 · PROVER / VERIFIER

Claims against evidence

Every answer and tool call decomposes into atomic claims; a verifier checks each against what the AI was actually given. Fabrication, injected instructions reaching actions, and answers left on the table become measured rates: hallucination, faithfulness, precision, recall.

The critics are themselves benchmarked.

Every critic is validated against gold-labelled evaluation sets before its verdicts count: a 200-example, 739-case benchmark with labelled violations for the soundness critics, and persona replays with planted traps (declined items, retractions, embedded injections) for the objectives wizard. In its latest run the deterministic tool-call critic caught 109 of 109 planted schema violations with zero false positives across 195 valid calls. When a critic cannot judge confidently, it abstains and says so - abstention is reported, never hidden.

Compliance

Compliance is a byproduct of governing well.

The evidence that shows your Ethical and Business Objectives being satisfied is the same evidence the EU AI Act, ISO/IEC 42001, the NIST AI RMF and GDPR ask for. Track objectives well and the audit trail writes itself: EU AI Act risk classification per workflow, evidence attached to every score, and an append-only record a regulator can walk. Filing paperwork governs nothing; this governs, and the paperwork follows.

Who it's for

Built for the people who answer for the AI.

AI owners & operators

Every workflow actually running, its objectives, its lights, and the next cheapest fix - before an incident, rather than a post-mortem after one.

Risk & Compliance

Ethical objectives measured continuously instead of sampled annually, with the evidence for every framework already attached to every score.

Finance & the Board

Business objectives and their satisfaction on the same page as the ethical ones, so the investment case and the assurance case are one conversation.

Questions

Management by Objectives for AI, answered.

What is Management by Objectives for AI?

Setting explicit Ethical and Business Objectives for every AI workflow an organisation runs, measuring their satisfaction continuously from execution logs, and remediating the worst gap first. It applies Peter Drucker's Management by Objectives to the AI estate: objectives are elicited once through a guided wizard, then every workflow is scored against them, live.

What are Ethical Objectives for AI?

Twelve values assessed on every workflow: fairness, privacy, wellbeing, transparency, accountability, safety and reliability, security, human autonomy and oversight, contestability, inclusivity and accessibility, sustainability, and lawfulness. Your organisation names its priorities with the worst case to avoid and the best case to promote; every value is still assessed.

What is critic-based remediation?

Scoring by layered critics with different failure modes: deterministic critics set floors no judge can argue down, LLM judges add discernment where rules cannot see, and a prover-verifier critic checks every claim an AI makes against the evidence it was given. Each red or amber score carries concrete, cheapest-first recommendations. The platform recommends; your team decides and acts; the next logs measure whether it worked.

What data does it need?

Only the execution logs your AI already produces, read from observability platforms such as Langfuse. Read-only, inside your own tenant, and your data never trains a model.

How is satisfaction scored?

Every objective on every workflow resolves to a traffic light: green for conformance, amber for a gap in evidence or performance, red for violation. Estate views take the worst light across workflows, and no data ever shows as green - an unmeasured objective is amber, visibly.

Does this cover the EU AI Act and ISO/IEC 42001?

Compliance falls out of the same evidence. Objective satisfaction is tracked against the telemetry that the EU AI Act, ISO/IEC 42001, the NIST AI RMF and GDPR ask you to produce, including per-workflow EU AI Act risk classification, so the audit trail is a byproduct of managing well rather than a separate exercise.

See whether your AI is meeting its objectives.

Connect one log source and see your first objective scores within a week. Read-only, in your own tenant, with the evidence behind every light and the cheapest fix ranked first.

Book an objectives assessment