Your free ticket to P99 CONF is waiting — 60+ (fully virtual) engineering talks on all things performance (Sponsored)P99 CONF is the technical conference for anyone who obsesses over high-performance, low-latency applications. Leading engineers from today’s most impressive gamechangers will be sharing 60+ talks on topics like Rust, Go, Zig, distributed data systems, Kubernetes, and AI/ML. Sign up to get 30-day access to the complete O’Reilly library & learning platform, free books, and a chance to win 1 of 500 free swag packs! Join 30K of your peers for an unprecedented opportunity to learn from experts like Chip Huyen (author of the O’Reilly AI Engineering book), Madelyn Olson (Valkey co-creator) & Carl Lerche (Tokio creator) and more – for free, from anywhere. Bonus: Registrants get immediate 30-day access to the complete O’Reilly library, and attendees can enter to win 1 of 500 free swag packs. Recommendations are one of the most important pieces of the Netflix experience. If you’ve watched Netflix, you would have noticed the Netflix homepage recommending a bunch of movies and TV shows. More often than not, they happen to be stuff that might be of interest to you. They also keep changing with time based on what you’ve watched. The original Netflix recommendations engine was built on top of 1000s of hand‑crafted features across users, items, and interactions. It also used special architecture for various functions that contributed to the task of making recommendations. Though a lot of changes have happened to this stack over the years, it’s still pretty complex. It’s costly to onboard new use cases. At the same time, large language models (LLMs) have made great strides. They have progressed a lot in their ability to recommend things to a user. Due to their broader world knowledge and understanding of language, LLMs do a good job of figuring out relationships between different things. They can also make recommendations using natural language, which is a plus point. However, you cannot just pick an LLM off-the-shelf to generate recommendations for a company like Netflix. This is why the Netflix engineering team built GenRec. The basic idea behind GenRec is to use an LLM to make better sense of a user’s viewing history to assign scores to the various movies and TV shows streaming on Netflix. However, you can’t just give a general-purpose LLM a list of movies or shows that a user has watched and expect it to come up with interesting recommendations. Netflix has to make a bunch of adjustments to the base model to make it work with their content and member behavior. In this article, we will look at how GenRec was built. Here’s what we will cover:
Disclaimer: This post is based on publicly shared details from various sources. References at the end. Please comment if you notice any inaccuracies. The Basic Question with RecommendationsA recommendation system has to answer a basic question: What should appear first in the list of recommendations? The answer to this question is the key to deciding which items to show a particular person and in what order. On Netflix, those items can include movies, series, games, and other content types. The order of the recommendations is super important. A user’s attention span is limited. A viewer might check only a few titles before making a choice. If they don’t find something interesting in that time, they might even leave the platform to check elsewhere. It doesn’t matter if something that the user might like is present in the catalog. You can say that the recommendation system has failed to do its job. The ordering of the recommendations is handled by a component known as the ranker. It assigns scores to various items based on how suitable they appear for a particular user in a particular situation. The better an item’s score, the higher it appears on the recommendation list. To decide this score, a ranker needs a lot of information. Here are some examples:
Netflix’s GenRec is capable of ranking the full content catalog. It can also work with a smaller set of content choices that are provided to it. You can use the top-K ranking, which means the highest-ranked K items. Here, K is the number of items you need. The goal goes beyond predicting what the user will click or play next. Netflix wants to recommend content that can satisfy a user so that they keep coming back for more. This means that a user playing a movie is a great indicator that the user was interested in that recommendation. But it doesn’t guarantee that the recommendation was valuable. AI’s Next Bottleneck Is Deployment. (Sponsored) |