In Plain SightContents ↓
Explainer 001 · Machine learning

Netflix GenRec,
from first principles.

A viewing history gives a streaming service clues about what to recommend. Follow one small movie catalog to see how those clues become scores, how a model learns to improve them, and how to judge the result.

Basic ML is enoughFour small experimentsPaper v2 · August 2026
01 / The problem

Choosing which movie comes first

Imagine opening a streaming service after watching several space mysteries. It has three movies it could show you: The Last Orbit, another space mystery; Harbor Lights, a coastal drama; and A Little Louder, a warm comedy. All three fit on the shelf, but one must appear first. Putting the comedy first might make sense for someone else. Your recent viewing gives the service a reason to try a different order.

We will follow this small catalog through the lesson. The movies, viewers, and demo numbers are invented teaching examples, not Netflix data or a replica of its model. Keeping the catalog small lets us see what changes when a system represents a viewer, learns from a choice, and uses what it learned.

For now, assume the service has already selected these candidates. A ranker assigns each one a score, then sorts them. Finding candidates is a separate job; a later reranking stage might adjust their order for diversity or other concerns. Here we will concentrate on how the scores arise. [4]

Our first version gives the viewer a pair of numbers and gives each movie a pair too. A list of numbers like this is a vector. To score a movie, multiply the viewer’s first number by the movie’s first number, do the same for the second pair, and add the two results. This operation is a dot product. It lets the same movie receive different scores for different viewers. [2]

In the example below, Viewer A has recently watched space mysteries. Switch to Viewer B, whose recent viewing favors comedies, and watch the order reverse. The catalog stays fixed; only the viewer’s two numbers change.

Experiment 01Invented movies & numbers

The same movies for two viewers

Viewer A: imagine a recent run of space mysteries.

Viewer vector: [0.9, 0.2]

  1. 1
    The Last OrbitAn imagined space mystery
    0.83score
  2. 2
    Harbor LightsAn imagined coastal drama
    0.55score
  3. 3
    A Little LouderAn imagined warm comedy
    0.27score

Top score: 0.9 × 0.9 + 0.2 × 0.1 = 0.83.

These are compatibility scores, not percentages or measured viewing probabilities. The viewer vectors are hand-picked—not inferred from the descriptions.

Viewer A puts The Last Orbit first.

For Viewer A, The Last Orbit comes first. Its score tells us it outranks the other two movies under our rule; it does not tell us how likely the viewer is to watch it. Sorting these scores is easy. The harder part is choosing numbers that reflect something useful about a person and a movie, rather than assigning them by hand as we just did.

GenRec scores catalog items instead of generating recommendation prose. Textual histories, metadata, and context replace thousands of engineered features. [1]

02 / The representation

Turning viewing history into numbers

Our viewer’s history is richer than a pair of numbers. It contains particular titles in a particular order, and a movie description carries more information than a genre label. We need a way to turn that input into a representation the scoring system can use.

An embedding is a learned numeric representation of something, such as a viewer or a movie. Learning can make useful matches score higher without requiring us to name every coordinate. Our two-number vectors are convenient to inspect, but a learned coordinate need not mean “likes space mysteries” or “likes comedy.” [2] In our example, the hand-picked vectors let us demonstrate ranking while leaving the work of producing them unfinished.

A language model offers a way to process a textual viewing history. It splits text into tokens, pieces that are represented numerically. Transformer layers then produce hidden states: internal vectors that depend on the surrounding context. A task-specific output component, called a head, turns those states into predictions. Predicting another token is one possible task; classification is another. Using a language model therefore does not require asking it to write a sentence. [6]

Text → tokensSplit the viewing history into pieces.
Numeric statesRepresent the input in context.
A task headProduce the task’s predictions.

In GenRec, the LLM produces a pooled user/context representation. A jointly trained head combines it with learned item embeddings; catalog scores determine the ranking. [1]

“Pooled” means that information from the sequence is combined into a representation for the next component to use. For our viewer, the idea is to turn the history into numbers suitable for scoring the shelf. That representation need not be an English summary of their taste, and pooling does not imply a particular averaging rule.

How closely does the dot-product toy match?

The toy illustrates how two numeric representations can produce a score. It does not specify the paper’s pooling method or scoring formula. Its coordinates are assigned by hand rather than learned from text.

This gives us a route from history to scores, but not yet a reason to trust those scores. A model could produce a vector for the viewer and still put an unwanted movie first. To improve its predictions, it needs examples of what happened after histories like this one.

03 / The learning signal

Learning from the movie they chose

Suppose our viewer chooses A Little Louder despite their recent run of space mysteries. We can turn that event into a training example: show the model only the earlier history, ask it to predict the choice, and compare its prediction with the recorded answer. The chosen movie is the target. Keeping it out of the input prevents the model from simply reading the answer.

The ranker needs feedback more precise than “the order was wrong.” A loss gives a numerical penalty to a prediction, so learning has something to reduce. For this exercise, we first convert scores into probabilities, then penalize the model for assigning too little probability to the observed choice.

The raw scores used here are called logits. Softmax exponentiates each logit and divides by the total, producing probabilities that sum to one without changing the order. [5] In our three-movie exercise, these probabilities describe competing predictions of the next choice. Increasing one movie’s share leaves less for the others. They are not independent estimates of how much the viewer likes each movie, nor does normalization alone make them accurate forecasts.

If the model gives the chosen movie very little probability, we want a large penalty. If it gives that movie more probability, we want a smaller one. The negative logarithm supplies this behavior: for target probability p, the loss is −ln(p). This is the single-target multiclass loss used in our exercise, rather than the full binary log-loss equation. [3]

The next demo uses the same movies but a separate set of starting logits, chosen to make the error visible. It initially assigns A Little Louder only 13.3% probability. Take one learning step and compare the target probability with the loss. The first rises as the second falls.

Experiment 02Real arithmetic, tiny model

Learning from A Little Louder

Observed target: A Little Louder

Last Orbit
59.8%
Harbor
26.9%
A Little Louder
13.3%
Target probability13.3%
Loss · lower is better2.014
Steps taken0

The target is currently the least likely choice.

This updates three logits directly using a gradient, not a full neural network. A neural model would instead update weights that produce its scores. Repeating one example shows the mechanism, not generalization.

See the update, and weight this example

zᵢ ← zᵢ − 0.5 × w × (pᵢ − yᵢ)

y is 1 for the chosen movie, 0 otherwise. The learning rate is 0.5. Weight w scales this example’s loss and gradient; it is not a serving-time score bonus.

Weighted loss: 2.014. Set the weight to zero and a step changes nothing.

A learning step uses a gradient, which describes how changing the adjustable numbers would change the loss. Here we adjust the logits directly. In a neural model, training would adjust the weights that produced the scores, so a change could affect predictions for other histories too. The optional calculation shows the update, but the visible behavior is enough to follow the lesson: the recorded choice pushes the prediction away from its initial mistake.

Repeating this one example makes the model better at this one example. It cannot tell us whether the model will help our viewer tomorrow, or help another viewer with a different history. That requires learning across varied examples and evaluating predictions on choices the model did not practice on.

04 / The training recipe

Adapting to the catalog and its viewers

One choice also leaves much unresolved about our viewer. Choosing a comedy tonight does not erase their interest in space mysteries. A useful training set would have to include many such histories and outcomes, rather than teaching the model that every viewer who watches The Last Orbit must want A Little Louder next.

There are two kinds of preparation to distinguish. A model needs familiarity with the setting in which it will work, including the information used to describe movies and viewing. It also needs practice making the particular prediction we care about. As the catalog and viewing patterns change, an old collection of examples may become less representative of the choices it now faces.

GenRec’s Phase 1 adapts an open-source LLM to Netflix data. Phase 2 frequently refreshes ranking specialization. [1]

In our example, language practice concerns the text describing the movies and their context; ranking practice concerns which movie should come ahead of the others. These tasks give learning different kinds of feedback. A good prediction of text is not, by itself, evidence that the viewer’s next choice will appear near the top of the shelf.

Training combines ranking and language objectives. Separate reward models weight ranking examples toward satisfaction and business goals—not full reinforcement learning. [1]

Weighting an example changes how strongly it contributes to learning. In the earlier demo’s optional controls, doubling the example weight doubles that step’s adjustment; setting it to zero leaves the logits unchanged. It does not add a bonus to the comedy’s score when the shelf is displayed. This distinction matters because the choice of training feedback shapes later predictions, while a recorded choice remains an imperfect description of what the viewer enjoyed.

Once training has produced usable weights, the service can apply them to a new history without repeating the learning exercise for every request. Our viewer opens the app; the system must now compute the shelf’s order using the model it has already trained.

05 / The serving path

Using the model without writing a response

During training, our exercise had a recorded answer and used a loss to change the model. When serving a recommendation, the viewer has not made their next choice yet. The system has the current history and candidate movies, and needs to produce scores. This is a prediction step, not another round of learning from an answer.

A language model processes its input context in a stage called prefill. For text generation, autoregressive decoding follows: it produces output tokens sequentially, with each step using the output so far. A task that can obtain its answer from the input representation does not need that continuation. [7] Our shelf needs an ordering of existing movies, rather than a sentence describing the ordering.

Use the stepper to add output tokens to the text-generation track. The scoring track stays finished while the text track continues. This illustrates the extra dependency in writing a response; the chips do not measure how long either task takes.

Experiment 03Stages, not a speed benchmark

Scoring a shelf versus writing a response

A scoring task
Input readScores ready
A text-generation task
Input readNext output token…

Input processed. The text task has generated 0 of 3 illustrative output tokens.

The chips show dependencies, not elapsed time. Each output chip is one invented token; no real tokenizer or language model runs here.

GenRec serving reads the context in one prefill pass, without token-by-token decoding. Context compression reduced about 5,000 tokens to 1,700 with negligible offline degradation. [1]

Skipping output generation still leaves input processing to do. In our running example, the history must become a representation before the scoring component can compare movies. The serving path explains how an order can be produced, but it does not establish that the order is better. For that, we need to test predictions against outcomes rather than inspect the machinery alone.

06 / The evidence

Measuring whether the order improved

Return to the viewer who chose A Little Louder. For training, we used the probability assigned to that movie to compute a loss. For evaluation, we can ask where it appeared in the ranked shelf. These are different measurements: the loss gives learning a penalty to reduce, while an evaluation metric describes an aspect of the resulting predictions.

Reciprocal rank gives the relevant result a value of one divided by its position. First place earns 1, second earns ½, and third earns ⅓. Averaging these values over evaluation examples gives mean reciprocal rank (MRR). It rewards putting the relevant result nearer the top. [8]

Imagine that the relevant movie below is A Little Louder. Move it from third place to first and the reciprocal rank rises from one-third to one. This is just one evaluation example; a mean would require more. The measurement tells us where a recorded target landed, not whether the viewer enjoyed the film after pressing play.

Experiment 04One invented evaluation example

Where the chosen movie appears

  1. 1Other
  2. 2Other
  3. 3Relevant
Reciprocal rank · higher is better1 ÷ 3 = 0.333

This is one reciprocal rank, not a mean. An offline score asks about a recorded target, not the entire experience of watching.

A reported improvement also needs a baseline and a scale. For an invented baseline MRR of 0.250, a 1.6% relative lift gives 0.254, an absolute increase of 0.004, not 0.016. Relative lift multiplies the baseline by a proportional change; it does not add that percentage directly to the metric. This baseline is only an arithmetic illustration.

Against the production baseline, GenRec reported approximately 40× fewer Phase-2 labeled examples and +1.6% relative MRR. Earlier foundation training is excluded from that data comparison. [1]

An offline comparison uses recorded examples, much as our exercise used the viewer’s known choice. A live comparison asks a different question: what happens when people actually encounter the system’s recommendations? It can assess behavior under the tested experience, rather than only how well a model orders historical targets.

A four-week A/B test on approximately 10% of traffic reported +0.006% relative improvement in the core online metric. Tests covered selected batch-compute surfaces, not every recommendation. Statistical significance is not the same as a large practical effect. [1]

The offline and online numbers describe different measures under different conditions, so their sizes are not directly comparable. For our tiny shelf, putting the chosen comedy first would improve reciprocal rank. Establishing that the new shelf improves the viewing experience would require evidence about that experience. Neither the shape of the model nor a successful learning step can stand in for that test.

Follow the sources.

Citations marked [1] identify summaries of the paper; the other sources cover the foundations. The imaginary catalog and runnable demos are original teaching material, not Netflix data, a replica model, or a performance simulation.

  1. [1]GenRec paper v2 — Li et al.
  2. [2]Google: embeddings and matrix factorization
  3. [3]Google: log loss
  4. [4]Google: recommendation systems overview
  5. [5]Google: softmax
  6. [6]Hugging Face: transformer task heads
  7. [7]Hugging Face: prefill and decoding
  8. [8]NIST: Mean reciprocal rank

Paper version 2. Sources checked October 1, 2026.