Data Science · Flagship

Real-Time Recommendation Engine at Scale

A real-time, collaborative-filtering recommendation engine that serves personalized results under production-like load.

PythonApache KafkaRedisPySpark

Streaming, shopping and news services rely on recommendations, but a model trained once a week cannot react to what a user clicked a minute ago, and one that works on a small data set often collapses under real traffic. This project builds a recommendation engine that learns from live events and serves personalized results quickly under production-like load.

The engine combines collaborative filtering, which learns from the behaviour of similar users, with content-based signals, which handle new items and new users with little history. Events such as views and purchases stream through Apache Kafka, where a PySpark streaming job updates user and item features continuously. Recent results and popular items are cached in Redis so that the serving layer can answer within milliseconds, and a fallback strategy covers a cold start. Ranking quality is evaluated offline with measures such as precision, recall and NDCG, and online through an A/B test framework that splits traffic between the current and the new model and reports which one users engage with more. Load tests show latency and throughput as traffic grows.

You will learn streaming data pipelines, recommendation algorithms, caching, evaluation and A/B testing. The Project Reference Guide documents the architecture and measurements, and the Reference Implementation includes the pipeline, the serving API and the test tools.