Skip to content
Krishna Panjiyar
MLBackend

Two-Stage Movie Recommender

Retrieval plus re-ranking over MovieLens, reaching 87% Precision@10 and 0.82 NDCG@10, served on AWS Fargate.

  • Python
  • PyTorch
  • Hugging Face
  • OpenAI API
  • FastAPI
  • Docker
  • AWS Fargate

Results

  • 87%

    Precision@10 (v2)

  • 0.82

    NDCG@10 (v2)

  • 0.78

    Recall@20 (v2)

  • 0.85

    precision@10 (v1, SVD + cosine)

v1 is public: TruncatedSVD plus item-item cosine similarity on MovieLens ml-latest-small, 0.86 RMSE on an 80/20 split, Flask and Gunicorn API, Docker, pytest in GitHub Actions, and preprocessing runtime cut by 30%.

Problem

Recommend movies a user will actually like from the MovieLens ratings data, and let people ask for recommendations in plain language instead of picking from filters.

Version 1 used TruncatedSVD and item-item cosine similarity. Version 2 set out to rank better by adding text embeddings and a dedicated re-ranking stage, without running the heavier scorer over the whole catalog.

My role

Solo project. I built both versions end to end: data preparation, the models, offline evaluation, the API, and deployment.

Architecture

Architecture diagram: v2 request path: cheap retrieval narrows the catalog, then a scoring layer re-ranks a short list.Architecture diagram: v2 request path: cheap retrieval narrows the catalog, then a scoring layer re-ranks a short list.
v2 request path: cheap retrieval narrows the catalog, then a scoring layer re-ranks a short list. Open full size (opens in a new tab)Open full size (opens in a new tab)
Diagram source (Mermaid)
flowchart TD
  Q["User request or natural-language query"] --> LLM["LLM query refinement (OpenAI API)"]
  LLM --> R["Stage 1: Retrieval"]
  MF["Matrix factorization vectors"] --> R
  TE["Transformer text embeddings (Hugging Face)"] --> R
  R -->|"top 100 by cosine similarity"| S["Stage 2: Pointwise scoring layer"]
  S -->|"final top 10"| API["FastAPI service on AWS Fargate"]
  API --> U["Ranked recommendations"]

Key decisions and tradeoffs

Two stages instead of one big model
Retrieval with vectors and cosine similarity is cheap enough to run against the whole catalog. The more expensive scoring layer then only sees 100 candidates, so ranking quality improves without paying that cost on every movie.
Combine collaborative and content signals
Matrix factorization captures what similar users rated. Transformer text embeddings capture what a movie is about, which can help titles with few ratings. Retrieval uses both.
Language model only at the edge
The LLM refines the user's query before retrieval. It does not rank. Ranking stays in models that are measured offline with Precision, NDCG, and Recall.
Measure ranking, not just rating error
v1 reported RMSE and precision@10. v2 adds NDCG@10 and Recall@20 because the product question is whether the right movies appear near the top of the list.

What I'd improve next

  • Publish the v2 code with a reproducible evaluation script so every metric on this page can be rerun.
  • Add a short screen recording of a natural-language query flowing through both stages.
  • Track online signals such as clicks next to the offline metrics.