Skip to content

Our approach

Evidence about how this actually performs, before anyone builds anything.

Most enterprise AI programmes are architected on vendor benchmarks and executive enthusiasm, then discover their real cost and quality in production. We invert that: measure first, make the policy executable, and let the evidence decide what gets built.

  1. Phase 01Two weeks

    Baseline the reality, not the roadmap

    We start by measuring what is already happening. Shadow usage, existing licences, the pilots nobody has retired, and the true unit cost of each. In parallel we run a frontier scan: what the current generation of models can reliably do for your specific workloads this quarter, tested against your data rather than a public leaderboard.

    • Telemetry on actual usage, including unsanctioned tools
    • Current spend decomposed to cost per task, not cost per seat
    • Frontier capability tested against your workloads
    • A candidate list ranked by value, risk, and feasibility
  2. Phase 02Two to three weeks

    Evidence before architecture

    Before anything gets built we construct the evaluation set — drawn from your real traffic, labelled with your subject-matter experts, and calibrated so that automated grading agrees with human judgment often enough to be trusted. This becomes the contract. Every later decision about models, prompting, retrieval, or fine-tuning is settled by running it rather than by argument.

    • Golden datasets built from production traffic
    • Automated grading calibrated against expert review
    • Failure modes and refusal behaviour specified up front
    • A measured quality floor no release may fall below
  3. Phase 03Three to four weeks

    Design for the bill you can defend

    With a quality floor established, cost becomes an engineering problem rather than a surprise. We route each request to the cheapest model that clears the bar, cache aggressively at the semantic layer, and treat context engineering as the primary lever — most runaway inference bills are retrieval problems wearing a model costume. Where volume justifies it, we distil a frontier model down to a small one you can run yourself.

    • Model routing and cascades against measured quality
    • Context and retrieval engineering before model upgrades
    • Semantic caching and prompt compression where they hold up
    • Distillation and small-model paths for high-volume work
  4. Phase 04Concurrent

    Make the policy executable

    Governance that lives in a document is governance nobody follows. We express your obligations and priorities as code and enforce them at the gateway every request passes through: what data may leave, which models are permitted for which classifications, where a human must sign, and what each team is allowed to spend. Approved paths are built to be easier than the shadow ones, because that is the only enforcement mechanism that reliably works.

    • Policy-as-code enforced at the model gateway
    • Classification-aware routing, redaction, and residency rules
    • Budget ceilings and per-team cost attribution
    • Audit evidence generated continuously as a by-product
  5. Phase 05Six to eight weeks

    Ship one workload properly

    One real workload into production, with the evaluation gates, the guardrails, the traces, and the cost telemetry all live. We would rather put a single defensible system in front of real users than ten pilots in front of a steering committee. The first one establishes the pattern, and the pattern is what makes the second and third cheap.

    • Production release behind evaluation and safety gates
    • Human oversight placed by cost of error, not by default
    • Full tracing, replay, and rollback from day one
    • Measured cost per outcome reported against the estimate
  6. Phase 06Ongoing

    Scale the pattern, transfer the practice

    Once one workload is running, the work turns to reuse and independence. Teams adopt the pattern instead of reinventing it, autonomy is extended into loops where it is genuinely justified, and your staff are trained to operate and extend all of it. We plan our own exit from the first week, because capability only your consultants can run is a liability wearing the costume of an asset.

    • Reusable patterns and an internal platform teams pull from
    • Agentic loops extended where the evidence supports them
    • Role-specific enablement and prompting standards rolled out
    • Documented handoff with paired operation before we leave

The working agreement

Both sides of the arrangement, written down.

Engagements go wrong when expectations are implicit. These are ours, stated plainly before anything is signed.

What we ask of you

  • Access to real data and real traffic, not a sanitized sample
  • A business owner who can decide, present in the working sessions
  • Risk, security, and legal engaged from week one rather than at the gate
  • Willingness to retire pilots the evidence does not support

What you can expect from us

  • A target cost per outcome quoted before we build, then measured
  • Named senior practitioners who stay on your engagement
  • Written reasoning behind every material recommendation
  • A clean exit, full handoff, and no proprietary lock-in

Questions & answers

The questions executives actually ask.

If yours is not here, ask it directly — we answer our own email.

By establishing a quality floor first and then finding the cheapest way to stay above it. In practice the savings come from four places: routing each request to the smallest model that passes evaluation, fixing retrieval so you stop paying to stuff irrelevant context into every prompt, caching at the semantic layer, and distilling high-volume workloads onto small models you can host. None of it requires accepting worse output, because the evaluation set is what decides.

Start with the baseline.

Two weeks, and you will know what your AI estate actually costs, what it returns, and which half of it should be switched off.

Prefer email? Write to hello@metacogni.com