Skip to main content

SERVICE LINE

Data, Analytics & AI

Warehousing, dashboards, master data and machine learning that answer questions an executive actually asks. Built on Nigerian data, in Nigerian languages, under governance you can explain to a regulator.

From records to decisions

The Data & AI practice is led by Dr. Yakubu Sale Karaye, a statistician by training who spent a decade building measurement systems for development programmes in northern Nigeria before joining Brilliant Esystems. His team of twenty-three data engineers, analysts and machine learning specialists works on a simple premise: most Nigerian institutions are not short of data, they are short of trustworthy data arriving early enough to change a decision. A state ministry may hold eleven years of records across four systems and still be unable to say how many facilities it funds. The work of this practice is to close that gap and then keep it closed.

The foundation is unglamorous. We build data warehouses and lakehouses with modelled dimensions and documented lineage, extract-transform-load pipelines that run on schedule and alert when they do not, master data management so that one facility, one taxpayer or one member of staff has exactly one identity across the estate, and data quality monitoring that measures completeness, validity, timeliness and consistency as tracked metrics rather than complaints. Only once that is in place do dashboards mean anything, and we are direct with clients who want to start with the dashboard.

On top of that foundation we build analytics people use: operational dashboards for managers, statutory returns generated rather than assembled, self-service models for analysts, and predictive work where it earns its keep — revenue leakage detection, collection propensity, patient no-show forecasting, stock-out prediction, equipment failure prediction. We also do a substantial amount of document intelligence, applying optical character recognition and layout models to the scanned registers, handwritten ledgers, forms and certificates that hold much of Nigeria’s public record, and natural language processing across Hausa and English for classifying complaints, routing correspondence and analysing feedback from citizens who do not write in English.

Data and AI capabilities

Six capabilities, delivered in the order that makes the later ones possible.

Data warehousing and pipelines

Foundation

Dimensional and lakehouse models on PostgreSQL, Oracle or cloud warehouses, with orchestrated ETL and ELT pipelines, documented lineage, and alerting when a load fails or arrives late.

Dashboards and reporting

Decision support

Executive, operational and statutory reporting in Power BI, Metabase or Apache Superset, designed around the decisions each audience makes rather than around the tables that happen to exist.

Master data management

One version of a thing

Entity resolution and golden-record management for citizens, taxpayers, facilities, staff, suppliers and assets, with stewardship workflow and survivorship rules that the business owns.

Data quality

Measured, not assumed

Profiling, rule-based validation, completeness and timeliness scorecards by source system, and remediation workflow that puts bad records back in front of the people who can correct them.

Predictive and machine learning

Where it earns its keep

Forecasting, propensity, anomaly detection and risk scoring, delivered with baseline comparison, holdout evaluation, drift monitoring and a documented retraining schedule.

Document intelligence and language

Hausa and English

OCR and layout extraction from scanned registers, forms and certificates, plus text classification, routing and summarisation across Hausa and English correspondence and citizen feedback.

The data maturity roadmap

We assess every client against these five stages and are honest about which one they are in. Attempting stage four from stage one is the commonest reason analytics programmes are abandoned after eighteen months.

1

Stage 1 — Ad hoc

Data lives in operational systems and spreadsheets. Reporting is manual, slow and inconsistent between departments. The first intervention is a source inventory, a data quality baseline and a single agreed definition of the five or six measures the organisation argues about most.

2

Stage 2 — Consolidated

A warehouse exists, core sources load on schedule, and one set of numbers is published from one place. Manual reporting effort typically falls by half. Governance begins: named data owners, a business glossary and a change process for measure definitions.

3

Stage 3 — Governed and self-service

Master data is managed, quality is measured and trending, and trained analysts in the business build their own views on a curated semantic layer. The central team moves from producing reports to maintaining a platform, which is the point at which the model becomes affordable.

4

Stage 4 — Predictive

Historic data is deep and clean enough to model. Forecasting, anomaly detection and risk scoring go into production with monitoring, and predictions are wired into an operational workflow so that somebody actually acts on them.

5

Stage 5 — Embedded and adaptive

Analytics is part of the operating rhythm: models trigger workflow, decisions are logged and evaluated, model drift is monitored automatically, and the organisation runs deliberate experiments to test whether interventions work.

Typical use cases by sector

Selected engagements and the measure that changed. Outcomes are illustrative of typical engagements and depend heavily on data quality at the outset.

Sector Use case Approach Measurable outcome
Government and public sector Internally generated revenue leakage detection Assessment and collection data matched against bank settlement, with anomaly scoring on collection points Unattributed receipts reduced by 81% in two quarters
Financial services Early warning on loan portfolio deterioration Behavioural scoring on transaction and repayment patterns, refreshed nightly with drift monitoring Non-performing exposure identified 46 days earlier on average
Health Outpatient attendance and stock forecasting Time-series forecasting by clinic and by commodity, feeding the procurement calendar Essential-medicine stock-outs down 34% across nine facilities
Education Enrolment and facility funding reconciliation Master data resolution across four legacy registers with document intelligence on scanned returns Duplicate facility records reduced from 2,180 to 41
Telecommunications Network fault prediction and field crew routing Alarm-sequence modelling on site telemetry with dispatch optimisation Repeat site visits reduced by 27%
Agriculture Input distribution targeting and yield estimation Enumerator mobile data combined with satellite indices and Hausa-language feedback classification Input wastage reduced by 19% in the pilot local government areas

Practice metrics

23

Data specialists

Engineers, analysts and ML practitioners

47

Warehouses in production

Built and under support

2.1m

Pages digitised

Document intelligence, cumulative

96%

Pipeline reliability

Scheduled loads completing on time, FY2025

Frequently asked questions

Our data is a mess. Should we clean it before starting?
No — start, but start with profiling rather than with a dashboard. In the first three weeks we measure how bad the data actually is, source by source and field by field, and produce a quality baseline. That report usually reframes the programme: it shows which measures can be trusted today, which need remediation at source, and which are simply not recoverable. Cleaning data in a warehouse without fixing the operational system that produces it is a treadmill, and we will say so.
Do we need cloud infrastructure for analytics?
Not necessarily. Plenty of Nigerian institutions run perfectly good warehouses on PostgreSQL in our Kano data centre or on their own hardware, and for steady, predictable workloads that is often cheaper than a cloud warehouse once egress and licensing are modelled. Cloud makes sense for bursty workloads, large-scale model training and where elastic scale genuinely matters. We will price both and show the working.
How do you handle Hausa-language text?
We use multilingual transformer models fine-tuned on client-specific corpora, with a human-in-the-loop review stage for the first several thousand records so that accuracy can be measured rather than asserted. Typical classification accuracy on well-defined categories reaches the high eighties to low nineties per cent after tuning. Where accuracy is below the level a process needs, we route the uncertain cases to a person rather than pretending the model is better than it is.
Will an AI model replace our staff?
In the work we do, no — it changes what they spend the day on. Document intelligence removes the typing, not the officer who checks the record. Anomaly detection produces a shortlist, not a decision. Our contracts for public-sector clients require a named human decision-maker for any outcome affecting an individual, and we design the workflow so that overriding the model is easy and is logged.
How do you stop a model degrading after go-live?
Every production model ships with monitoring for input drift, output distribution shift and performance against a holdout or delayed-label set, plus an agreed retraining schedule and thresholds that trigger retraining early. Model versions, training data snapshots and evaluation results are all recorded, so that a decision made eight months ago can be explained with the model that actually made it.
Can we start small?
Most clients should. A six to eight week proof of value on one high-stakes question — with real data, a working pipeline and a dashboard people use — costs a fraction of a full programme and tells you far more about feasibility than a strategy document. If it does not work, you have lost eight weeks and learned something true. Worked examples sit at /solutions/business-intelligence-suite, /industries/government and /industries/agriculture.

Responsible AI and data protection

Every analytics and machine learning engagement runs under our responsible-AI standard. Before a model is built we record its purpose, the lawful basis for processing, the population affected, and the harm that a wrong output could cause. Models that affect individuals — eligibility, enforcement, credit, employment — require a documented human decision-maker, an appeal route, and bias testing across the protected characteristics that are meaningful in the Nigerian context, including geography, gender and language.

Personal data is minimised, pseudonymised in development and test environments, and retained only for the period the client’s retention schedule permits. Processing is governed by a data processing agreement, and data subject to Nigerian residency obligations under the Nigeria Data Protection Act stays in our Nigerian facilities. We do not use client data to train models for other clients, ever, and that restriction is written into every contract. Questions to [email protected] or [email protected].

Pick one question that matters and let us answer it properly

An eight-week proof of value with your real data will tell you more than any strategy document. Dr. Yakubu Karaye’s team will scope it, price it and tell you plainly if the data will not support the question.