AI Governance Consulting That Starts by Finding the Systems Nobody Registered

Almost every organisation running AI is running more of it than its own register shows: a model in a product team, an assistant embedded in a tool somebody bought, a script that quietly became a decision. AI governance consulting that starts anywhere other than the inventory is governance of a subset.

Start With an Inventory

Clutch 5.0GoodFirms 5.0Google 4.3Upwork 4.8

What This Team Is Accountable For, and What It Refers Out

AI governance is a market where the boundary between advice, engineering and certification is routinely blurred. Ours is stated before anything else.

What we do is technical. Find the AI systems, classify them by what they actually decide, build the evaluation and logging that produce evidence, design the human oversight points, and assemble the technical documentation an assessor or a regulator will ask for. It is engineering work, done by people who build these systems rather than by people who read about them.

What we do not do is certify. We are not a notified body and we are not an accredited certification body. We can produce the technical evidence that goes into a conformity file; somebody else signs it. Legal interpretation of your obligations is your counsel’s work, and where a governance programme needs a major firm’s opinion on the law, we will say so rather than improvise.

Our AI Governance Team, in Numbers

The team and delivery record behind our governance programmes.

60+

AI Engineers

50+

AI Solutions Delivered

80+

AI-Integrated Workflows

30+

Industries Served

95%

Client Retention

The Engagements, Each One Startable on Its Own

Discovery is a prerequisite for the rest in practice, though clients regularly try to start with the evaluation suites.

AI System Discovery and Model Inventory

We find what is actually running: models your teams built, models inside tools you bought, and the scripts that have quietly become decision-makers, each with an owner, a purpose and what it decides. You get a register somebody can be held to, and usually a gap between it and the previous one.

Risk Classification and Obligation Mapping

We sort the inventory by what each system decides about whom, then map each entry to the obligations that actually attach to it rather than to the whole regime. What comes out is a shortlist: most systems in most inventories carry light obligations, and knowing which do not is as valuable as knowing which do.

Evaluation Suites for Bias, Robustness and Drift

We build test sets and harnesses that run against your models on a schedule, producing evidence rather than assurances. You end up with a measurement over time, which is the only form of evidence that survives a follow-up question.

Human Oversight and Approval Gate Design

We design where a person has to be in the loop, what they see when they are, and what they are actually able to change. Oversight that cannot alter an outcome is documentation rather than control, so you get gates built into the system, plus an honest note about which existing ones are decorative.

Decision Logging and Data Lineage

We record what was decided, by which model version, on what input, and where that input came from, built to be retrieved on request rather than reconstructed under pressure. The result is a record that turns a difficult question into a query.

Generative and Agentic System Controls

We govern the newer surface: prompt handling, output constraints, action limits, spend caps and what stops a running system. These need different controls from a scoring model and are frequently governed as though they were the same, so what you get is controls matched to systems that act rather than to systems that predict.

Technical Documentation for a Conformity File

We assemble the technical evidence an assessor expects, in the structure they expect it. That produces the file contents, not the sign-off, which comes from a body we are not.

Governance Work, and What the Inventory Turned Up

The number worth looking for on each card is systems found against systems previously known. That gap is the argument for doing discovery before anything else.

Feedback From Clients After an Inventory

Published on platforms that verify an engagement before a review is posted.

Reviewed on Upwork
Our experience working with Aipxperts has been exceptionally satisfying. From start to finish, they handled the project with professionalism and responsibility. Communication was seamless, and they effectively addressed our requirements, delivering high-quality results on time. Their technical expertise was particularly impressive, as they effortlessly solved complex problems. We highly recommend Aipxperts for their outstanding service and dedication to client satisfaction.
Full-Stack Developer Needed for Angular 15 and NestJS ProjectVerified Upwork client
Reviewed on Clutch
Hardik was very helpful in advice and completing the work.
TomAustralia

How Deep the Governance Has to Go in Your Sector

Depth is set by what your AI systems decide about people, and that varies enormously between sectors that otherwise look similar.

Fintech and financial services

Systems that affect access to credit or pricing sit at the demanding end, and the obligation is usually explanation rather than accuracy. We most often build the explanation and bias evidence a credit or pricing model has to produce on request across fintech and financial services, because a model that is right and cannot be explained is a problem here in a way it is not elsewhere.

Healthcare administration

Anything touching triage, prioritisation or access to a service carries a heavy oversight requirement, and the person in the loop has to be able to override rather than merely observe. We usually start with those approval gates and the decision logging behind them for healthcare administration systems.

Education and EdTech

Systems affecting assessment, progression or admission, with a user base that may include children. We settle the consent position and the oversight position together rather than sequentially, and document both, for education and EdTech platforms.

Food delivery

Automated allocation and scoring that affects couriers puts these systems in the category regulators and works councils look at hardest. We usually build the drift monitoring and the appeal path first for food delivery platforms, because a small error rate applied to millions of assignments is a governance problem regardless of the model’s quality.

Energy and utilities

Models adjacent to operational systems, where the governance question is what a model is permitted to influence rather than what it predicts. We establish the boundary between advisory and controlling and enforce it in the pipeline for energy and utilities operators.

Automotive and dealer networks

Obligations arrive contractually down a supply chain rather than from a regulator, so the specification is a manufacturer’s requirement document that changes without notice. We map those requirements onto a single control set for automotive and dealer networks, so a change lands in one place rather than several.

Logistics and warehousing

Automated decisions affecting workers, routing and allocation. Where a system influences how people are scheduled or measured the oversight requirement is heavier than most operators expect, so we begin with the human oversight design and the decision log for logistics and warehousing.

Why Choose Aipxperts Over a Large Consultancy

Differences that matter when you are weighing us against a large consultancy, plus one honest referral. That referral is the most useful thing on this page for a certain kind of client.

01You get controls written by the engineers who ship the modelsGovernance written by people who do not build these systems produces policies engineers route around. Ours are designed by the same engineers, which makes them implementable and occasionally makes them less comfortable to read.

02You maintain one control set instead of three overlapping onesOrganisations answering to a regulation, a customer questionnaire and an internal policy end up maintaining three overlapping frameworks. One control set with multiple tags is less work to maintain and far easier to evidence.

03You find out which controls the system actually enforcesThe gap between a control that is documented and a control the system actually enforces is where governance programmes fail an inspection. We test which is which and we write down the difference rather than reporting the policy.

04You can answer a question about an old decision in minutesIf answering a question about a decision from four months ago requires an engineer and a week, the evidence does not really exist. Logging is designed so the answer is a query.

05You get told when a larger firm is the right callIf what you need is a legal opinion on how a regulation applies to your business, or a certification, or a board-level assurance signature, that is a law firm or an accredited body. We will say so on the first call. What we are good at is the technical layer underneath whatever they tell you.

The EU AI Act Obligations Live Right Now, and Those Deferred

The timetable moved in 2026 and a great deal of published advice has not caught up. Everything below was re-verified against primary reporting on 17 August 2026. If you are reading this substantially later, check the dates before relying on them.

ObligationStatus as at 17 August 2026
Transparency duties under Article 50Applying now. Took effect 2 August 2026 and were largely untouched by the Digital Omnibus
Marking of synthetic content under Article 50(2)Systems placed on the market before 2 August 2026 have until 2 December 2026
High-risk systems listed in Annex IIIDeferred to 2 December 2027, roughly sixteen months later than originally scheduled
High-risk systems under Annex IDeferred to 2 August 2028, one year later than originally scheduled

The Tooling Behind Evaluation and Evidence Work

The evaluation, logging, lineage and documentation tooling our engineers use on governance work. Where you already run model or experiment tracking, the controls are built into yours rather than beside it.

Evaluation and Testing

RagasDeepEvalweightsandbiasesWeights & BiasesmlflowMLflowand custom evaluation harnesses built on your own cases

Bias, Explainability and Interpretability

SHAPLIMECaptummodel cards and documented data lineage

Drift and Performance Monitoring

Evidently AIArizeWhyLabsFiddler AIArthur AI

Pipeline Gates and CI

githubactionsGitHub ActionsazuredevopsAzure DevOpsjenkinsJenkinsmlflowMLflow Model RegistryKubeflow

Logging, Lineage and Audit Evidence

opentelemetryOpenTelemetryprometheusPrometheusgrafanaGrafanadvcDVCimmutable decision logs

Policy and GRC Platforms We Integrate With

Credo AIHolistic AITrustibleMonitauribmIBM watsonx.governanceOneTrust AI GovernanceServiceNow AI Control TowerCollibra

Cloud and Deployment

amazonwebservicesAWSmicrosoftazureMicrosoft AzuregooglecloudGoogle ClouddockerDockerkubernetesKubernetesterraformTerraform

Where Our AI Governance Reviews Are Published

Held on Clutch, GoodFirms, Upwork and Google, across web, mobile and enterprise delivery since 2012.

Upwork4.8150 reviewsClutch5.012 reviewsGoogle4.335 reviewsGoodFirms5.05 reviews

Your Models, Your Evidence, and What This Team Can Sign

Governance work reads model behaviour, training data and decision logs. The controls below apply throughout, and the one clients most need stated is that we assemble evidence rather than certify it.

The regimes the control language comes from

The NDA is signed before discovery. Assessment runs inside your environment where your policy requires it, and model weights, training data and evaluation sets do not cross your boundary without written authorisation. Frameworks and deliverables produced during the engagement assign to you.

Your data and your model intellectual propertyTraining data, model weights, evaluation sets and decision logs stay in your environment. Access is read-only, named and time-boxed to the engagement, and nothing is retained afterwards unless you ask.What the evidence has to surviveIt is built assuming it will be read by somebody adversarial who was not in the room. That standard is higher than internal reporting and it is the standard that makes the difference at an inspection.Where the report can goYours to share with a regulator, an assessor, an insurer, a customer or a board. We attach no conditions to its circulation and we do not name clients alongside governance findings without written consent.On certification, and the answer clients want firstAipxperts is not a notified body, not an accredited certification body and not a conformity assessment body, and holds neither ISO 27001 nor SOC 2 itself. We produce technical evidence that goes into a conformity file. The sign-off comes from somebody else, and any supplier suggesting they can provide it is describing something that does not exist.

What actually applies right now

Verified 13 August 2026. These dates need a named review owner and a re-check every quarter, because the deferrals moved once already this year.

ObligationStatus as at 13 August 2026
Digital Omnibus on AIIn force since 27 July 2026, following publication in the Official Journal on 24 July.
Article 50 transparencyApplying since 2 August 2026. Not deferred by the Omnibus.
Article 50(2) machine-readable markingGrace period to 2 December 2026, but only for generative systems already on the market before 2 August 2026. Anything deployed on or after that date complies immediately, so for a new AI feature the duty is live now.
Annex III high-risk obligationsDeferred to 2 December 2027: biometrics, employment, credit scoring, education, critical infrastructure.
Annex I embedded-product high-riskDeferred to 2 August 2028: AI inside products already covered by sectoral safety legislation.

Read the third row twice. Deferral got the headlines and the transparency duty did not move, so the nearest live obligation for most teams is a disclosure requirement already in force.

The Steps of a Governance Engagement, and the Limit of Each One

Every stage names what it does not cover, because on governance work the overclaim is the risk.

01Discovery and inventoryFinding the AI systems, including the ones inside purchased tools and the scripts that became decision-makers. The limit is that we find what your environment and your people reveal: a system nobody mentions and nothing logs stays hidden, so the inventory is a floor rather than a certainty.02Classification by what each system decidesSorting by effect on people rather than by technical sophistication, which is the sorting that obligations actually follow. Classification is a technical judgement, and where it carries legal consequences your counsel confirms it.03Obligation mappingTagging each system to the specific duties that attach to it, from regulation, contract or internal policy. We map to obligations as written; interpreting an ambiguous one is legal work, and we will say when we have hit one.04Gap assessment against runtime realityWhat the documentation says the controls are, against what the systems actually enforce. We test what is testable: controls that depend on human behaviour are assessed by interview, which is weaker evidence, and we mark them as such.05Evaluation and logging buildTest suites, drift monitoring and decision logging built into the systems on a schedule. An evaluation suite measures what it was built to measure, so new failure modes need new tests, which is why the suite is owned rather than delivered.06Oversight and control implementationApproval gates, action limits and the mechanisms that let a person actually intervene. A gate only works if the person behind it has time and authority; where they have neither, we will say the control is nominal.07Evidence assembly and handoverThe documentation set, structured for the audience that will ask for it, with your team able to maintain it. This is assembly, not attestation: we do not sign it, and nobody should represent our involvement as approval.

What Comes Up on a Governance Scoping Call

Most callers need the answer on scope. The answer on what we cannot sign is the one that decides whether we are the right firm.

Share your project vision

Tell us what you want to build. A specialist, not a salesperson, replies.

PDF, DOC or image, up to 10MB. Optional.
My idea is confidential – happy to sign an NDA.

The inventory, without exception. Every other question depends on knowing which systems exist, and the gap between what an organisation runs and what it has registered is consistently larger than anyone in the room expects.

No. Not a notified body, not an accredited certification body. We build the technical evidence that goes into the file and somebody accredited signs it. This is worth establishing on the first call rather than at the end.

That is a legal question about your systems, your users and where they are, and your counsel should answer it. What we can do is establish what your systems actually are and what they decide, which is the input that question needs and which almost nobody has ready.

The dates moved and the work did not shrink. Discovery is the slow part and it is unaffected by any timetable. Separately, most of our clients are driven by customer questionnaires rather than by regulators, and those arrive this year.

System count first, then how much is undocumented, then how deep the evaluation work goes. An inventory alone is a contained engagement and is often the right place to start before committing to anything larger.

A policy is a statement of intent. Governance is evidence that the intent is enforced. The gap between the two is where programmes fail an inspection, and testing it is usually the fastest way to find out where you actually stand.

An audit tests controls against an obligation and reports findings. This builds the controls and the evidence in the first place. Some clients want both, and where they do, a different team runs the audit so nobody reviews their own work.

Yes, and most of them are. Purchased tools with embedded AI are frequently the largest category in an inventory and the least documented, because nobody procured them as an AI system.

Different controls from a scoring model and frequently governed as though they were the same. Prompt handling, output limits, action boundaries and spend caps do not appear in a framework written for predictive models, and that mismatch is one of the commonest gaps we find.

When the inventory turns up a handful of low-impact systems with clear owners and no external obligation attached. That happens, it is a legitimate outcome, and it is why an inventory on its own is a reasonable first step.

Start With the List of AI Systems You Think You Have

Send whatever register exists, however incomplete, along with the questionnaire or the obligation that prompted this. Back comes what an inventory would involve, roughly where the gaps usually sit in an organisation like yours, and whether an inventory on its own is the right starting point before anything larger.

Send Your AI Register

Governance Notes From Live Programmes

Inventory findings, obligation mapping and evidence design, written up by the engineers who ran the programmes.