Skip to content
ONLINE·BOOKING Q4 2026 ENGAGEMENTS·ONDINE v1.10.1·--:-- UTC
implementation

Production AI Platforms and Agentic Systems

I help you build an LLM platform or bring an AI agent into production, working with your team and your existing systems.

the two build offers

01

AI Platform Foundation

Shared model access with permissions, cost tracking, and the tools needed to operate it.

view the full scope

Design and implement the shared foundation your workloads need across hosted, customer-controlled and self-hosted models: routing, identity, observability, evaluations and day-two operations.

platform architecturehosted, customer-controlled and self-hosted strategymodel gatewayauth and access controlrouting, quotas and budgetsK8s + GPU inferencemodel onboardingobservabilityevaluation infrastructureaudit loggingcost and capacity modeldelivery roadmap
02

Agentic Production Sprint

Put an agent to work on a specific business process, with authorised actions, human approval, and recovery when a step fails.

view the full scope

Start from a human-and-code process or an agent prototype. Allocate each step to deterministic code, a structured LLM call, a bounded agent loop or human approval, then connect it through typed contracts and evaluations.

workflow observation and decompositionexecutor selection: code / LLM / agent / humantyped input and output contractsagent harness and tool brokerpermissions and human approvaldurable executionevaluation dataset and release gatestrace-to-eval loopfailure and recovery handlingcost and iteration limitsobservabilitydeployment architecture

Explore the architecture, layer by layer

A reference map connecting applications, agents, models and controls.

Detailed viewFull AI platform map
100%
evidence planeTracingstep-level spansEvalsregression gatesAudit trailappend-only, hash-chainedReplayexplain any decisionCost attributionper run · per workflowconsumersProducts & copilotsInternal workflowsBatch & enrichmentagentic planeAgentsbounded autonomy · memoryOrchestrationplanning · durable stateTool brokerMCP · fail-closedHuman gatesapproval on writescontrol planeGatewaykeys · quotas · budgetsGuardrailsin / out, enforcedSemantic cachecost · latencyRoutingfailover · fallbackserving & knowledgeAPI modelsSelf-hosted GPURetrievalpermission-awareModel registrysigned weightsfoundationIdentity & secretsCloud · on-prem · sovereignEgress control
Hover any block: the question I ask during an LLM and agent platform review.
From model to evidence

Tap a layer, then a component, to see the question asked during a review.

  1. Review question

    What is the agent allowed to decide alone, and where exactly does its autonomy stop?

Working together

01

30-min scoping call

Most builds start here. We establish the existing LLM platform or agent workflow, which offer fits, and the first verifiable milestone.

02

or · evidence review first

If a system already exists, the fixed-scope Enterprise AI Evidence Review establishes the architecture and the ranked gaps, and build scope is written against its findings.

I work with you directly, on a focused engagement or a multi-month assignment. We define the scope in our first conversation.

Selected work and contact

Further details

My approach, with a workflow example

Two build paths: a shared LLM foundation, or one agentic business workflow taken to production.

Independent of any model vendor or agent platform: choose what to buy, build or keep customer-controlled, then make one LLM or agent workflow bounded, recoverable and independently reconstructable.

// design discipline

From real process to reliable agentic workflow.

We do not bolt an agent onto a procedure. Each decision goes to the least autonomous sufficient mechanism, and authority grows only when evidence supports it.

  1. 01REAL PROCESSobserve work as performed
  2. 02BOUNDARIESisolate trust and consequence
  3. 03MINIMUM MECHANISMcode · LLM · agent · human
  4. 04EVIDENCEtest trajectory and recovery
  5. 05BOUNDED AUTHORITYdelegate without surrendering
EXAMPLE · LOCAL SLICE EXECUTEDAn inference incident becomes a bounded, verified and reversible effect.
signal
observe
diagnose
authorise
act
verify / recover

Investigation stays read-only. Execution policy grants a revocable capability. A receipt and a fresh target observation establish the outcome.

6 reliability scenarios · no cloud claim

One control plane connects identity, policy, evaluations and evidence across every inference model.

what gets built
shared control plane
01

trace to eval

Production failures and real traces become evaluation sets and regression gates that block a bad release.

02

agent tool boundaries

Permissions, policy enforcement, and tool-call control, applied before a call reaches the tool server.

03

human approval

Sensitive actions proposed by the agent wait for human approval, recorded next to the evidence it was based on.

04

retrieval assurance

Permission-aware retrieval and citation checks for workflows where an unsupported answer is a reportable problem.

05

audit and siem

Model calls, tool calls and approvals are written to an append-only, hash-chained log, with trace export in a format your buyer's security team can ingest.

06

model gateway

One controlled entry point for managed APIs and private models: routing, failover, cost attribution, quotas, and runtime limits.

07

private inference

Self-hosted or customer-VPC serving, when a hosted API cannot meet the residency or control requirement.

08

platform foundations

Identity, central tracing, and shared controls, so each product team does not rebuild the same guardrails.

deployment follows the requirement
01

Managed API

when its native controls meet the need

02

Independent LLM control layer

when products and providers share policy

03

Customer VPC or self-hosted

when residency, control or measured economics require it

when customer-controlled inference is the right call

customer data cannot leave an approved environment

enterprise customers require customer-VPC or BYOC deployment

you have explicit model and data residency requirements

several teams need the same controls instead of one set each

hosting economics start to matter at your volume

you have an ops team that can run it after handover

Context and limits of the work and measurements

Controlled Vauban evidence

Repeatable lab experiments against a live Vauban gateway. They are not client production evidence.

  • EXP-02 · execution-governance comparison
  • EXP-04 · gateway failover, budget and key-control measurements
when an additional platform layer is premature
  • a managed API and its native controls already meet the full requirement
  • usage is low or still unpredictable
  • there is no platform or ops team to own it
  • the workload is still experimental
  • "on-prem is safer" with no threat model behind it

Professional case studies are anonymized by sector. Controlled Vauban evidence is linked above with its scope and limits.

LLM Platform & Agent Assessment Readiness · For regulated or high-risk teams preparing controls and evidence for independent assessment.