Templefield Technologies
Data Science / Distributed ML

End-to-End Record Linkage Pipeline

An end-to-end machine learning pipeline for linking internal records with external data sources

Industry
Public Administration
Client
Stiftung Zentrale Stelle Verpackungsregister, Osnabrück
Engagement period
June 2023 – December 2023
Delivery mode
Hybrid
End-to-end record linkage pipeline

Situation

The public-sector client managed over 1 petabyte of data on companies subject to legal obligations. To pursue fraudulent activity, they needed the ability to identify these companies within large, heterogeneous external data sources — a task far beyond manual matching.

Objective

Templefield Technologies was engaged to engineer a comprehensive, distributed machine learning system for record linkage that could reliably match internal records against external sources at petabyte scale.

Approach

  • Data preprocessing, standardization, and attribute transformation across heterogeneous sources
  • Record grouping through clustering algorithms to reduce the matching space
  • Distributed XGBoost classification training on Spark
  • Model performance tracking and production serving via MLflow
  • Kubernetes-native Spark deployment with workflow orchestration in Apache Airflow

Results

The delivered system enabled the client to automatically and efficiently identify matching companies across diverse data sources, streamlining the manual review process for the prosecution of fraudulent entities and increasing fraud detection by 12%.

Technology stack

PySparkXGBoostPythonKubernetesApache AirflowMLflowHadoopAmazon S3SQLDocker
All work