Catching overpriced and unusual US medical and hospital claims before they're paid.
An end-to-end pipeline for US health insurance, powered by three AI engines trained in-house: one filters and cleans the data, one detects anomalies in claims, and one recommends what reviewers should do next.
High-level design
Flow design
The problem
US policyholders submit claims for medical equipment and hospital bills to be reimbursed. Claims teams had no reliable reference for what an item or service should cost, so inflated prices, mismatched items and unusual claims were hard to spot, and checking every claim by hand was slow.
How it's built
- Collection engine: gathers medical-equipment and hospital pricing data from public US websites on a schedule.
- Data lake: stores the raw public data alongside a small set of the client's private internal data, kept separate and access-controlled.
- AI engine 1, data filtering (trained in-house): cleans, de-duplicates and standardises product names, units and prices, and filters out bad or incomplete records.
- Database: the cleaned reference data is loaded into a structured database that the other engines query.
- AI engine 2, anomaly detection (trained in-house): checks each equipment or hospital claim against the reference data and flags line items or claims that look unusual, such as prices far above market or items that don't match.
- AI engine 3, recommendations (trained in-house): suggests the next step for each claim to the reviewer, so investigators know where to look first.
- Human review: flagged claims go to investigators, so people make the final decision on every claim.
- Morning report: every morning the business automatically receives a report of the previous day's flagged claims and the recommended next steps.
What shipped
A complete pipeline, from data collection through cleaning, storage, anomaly detection and recommendations to an automatic morning report, that turns raw public data into a daily check on US medical-equipment and hospital claims.
Clean handoff
Each stage runs as its own engine, and the three AI engines are trained and updated separately, so data sources, cleaning rules, detection thresholds and recommendations can each be improved without rebuilding the whole pipeline.
Talk about a build like this →