Home / Work / US medical claims anomaly detection
Case study, ML · US health insurance

Catching overpriced and unusual US medical and hospital claims before they're paid.

An end-to-end pipeline for US health insurance, powered by three AI engines trained in-house: one filters and cleans the data, one detects anomalies in claims, and one recommends what reviewers should do next.

US medical claims anomaly detection: illustrative screen
What it looks like · illustrative screen with sample data

High-level design

US medical claims anomaly detection: high-level architecture
Architecture by zone · components marked AI are AI engines

Flow design

US medical claims anomaly detection: step-by-step flow by actor
Who does what, step by step · one lane per actor

The problem

US policyholders submit claims for medical equipment and hospital bills to be reimbursed. Claims teams had no reliable reference for what an item or service should cost, so inflated prices, mismatched items and unusual claims were hard to spot, and checking every claim by hand was slow.

How it's built

  • Collection engine: gathers medical-equipment and hospital pricing data from public US websites on a schedule.
  • Data lake: stores the raw public data alongside a small set of the client's private internal data, kept separate and access-controlled.
  • AI engine 1, data filtering (trained in-house): cleans, de-duplicates and standardises product names, units and prices, and filters out bad or incomplete records.
  • Database: the cleaned reference data is loaded into a structured database that the other engines query.
  • AI engine 2, anomaly detection (trained in-house): checks each equipment or hospital claim against the reference data and flags line items or claims that look unusual, such as prices far above market or items that don't match.
  • AI engine 3, recommendations (trained in-house): suggests the next step for each claim to the reviewer, so investigators know where to look first.
  • Human review: flagged claims go to investigators, so people make the final decision on every claim.
  • Morning report: every morning the business automatically receives a report of the previous day's flagged claims and the recommended next steps.

What shipped

A complete pipeline, from data collection through cleaning, storage, anomaly detection and recommendations to an automatic morning report, that turns raw public data into a daily check on US medical-equipment and hospital claims.

Clean handoff

Each stage runs as its own engine, and the three AI engines are trained and updated separately, so data sources, cleaning rules, detection thresholds and recommendations can each be improved without rebuilding the whole pipeline.

Talk about a build like this →