N=1 Personalization: Replacing 25 Cohorts with User-Level Ranking
From cohort-based defaults to true user-level personalization at Cars24 UAE — +150% Click Recall@50, +16% buyer conversion, and solving cold-start for 35% of the user base.
Contents
At Cars24 UAE, we had 25 cohorts and ~10,000 listed cars at any given time. Every user in the same cohort — say, “mid-segment SUV buyers in Dubai” — saw the exact same catalog ranking. Two users who both liked SUVs but disagreed on everything else (one wanted a 2022 Tucson under 80k AED, the other wanted a 2019 Pajero under 50k AED) got identical recommendations.
This is the cohort problem. You cluster users into segments, train a ranking model per segment, and serve the segment’s ranking to everyone in it. It works until you realize that the differences within a cohort often matter more than the differences between cohorts.
I joined Cars24 as a Data Scientist in January 2024 and rebuilt the recommendation engine from scratch for the UAE geography — data sourcing through production deployment, A/B testing, 100% rollout, dashboarding, and ongoing maintenance. The system I describe here went live on 12 September 2024 for 100% of UAE users across all platforms (app, mobile web, desktop). It was later adopted in Australia, Thailand, and India, where I collaborated with local DS teams to port and adapt it.
The results after full rollout:
- Click Recall@50: 12% → 30% (+150%)
- Buyer conversion (U2BI): +16%
- View-to-buy intent (V2Bi): +21%
- Cold-start coverage: 65% → 100% of relevant users
This post is about how the system works, why each component exists, and where it breaks.
1. The Problem: Why Cohorts Stop Working#
Cars24’s existing personalization stack was ~3 years old when I started. It was a two-stage system: unsupervised clustering to assign users to one of 25 demand cohorts, followed by supervised ranking models per cohort. The clustering used clickstream data — impressions, clicks, searches, filters, wishlists, gallery views, inspection report views — to segment users. The ranking models then optimized car ordering within each cohort to maximize booking likelihood.
The system had been iterated on multiple times and wasn’t bad. But it had three structural problems:
Granularity ceiling. Within any cohort, user preferences diverged enough that the cohort-level ranking was a compromise for everyone. The “mid-segment SUV” cohort contained users who wanted a budget family car and users who wanted a premium compact SUV. No single ranking satisfies both.
Coverage gap. The cohort system personalized ~65% of users — those with enough clickstream history to be assigned a cohort confidently. The remaining 35% got default sorting. These weren’t marginal users — they included every new user and every low-engagement returner. They were exactly the users who needed personalization most, because they hadn’t yet found what they were looking for.
No user-level signal fusion. The cohort assignment was a hard partition. A user’s search for “Honda Civic 2021” and their filter for “under 60k AED” were used for clustering but lost in the aggregation. The ranking model saw the cohort label, not the individual signals.
The fix wasn’t to make more cohorts. It was to eliminate cohorts entirely and rank at the individual user level — N=1 personalization.
2. Architecture Overview#
The system is a two-stage retrieval-and-ranking pipeline. The first stage generates a broad candidate set using learned user and item embeddings. The second stage re-ranks those candidates with a feature-rich gradient-boosted model. Scores are published asynchronously and consumed at listing render time.
graph LR
subgraph "Data Layer"
A["Clickstream\n(GA4)"] --> FE
B["Search & Filter\nImpressions"] --> FE
C["Car Attributes\n& Inventory"] --> FE
D["Demand Signals\n(I2V, V2BI, STR)"] --> FE
end
FE["Feature\nEngineering"] --> TT["Two-Tower\nRetrieval Model"]
TT --> HNSW["HNSW Index\n(ANN)"]
HNSW -->|"Top-K candidates"| LGB["LightGBM\nRe-ranker"]
LGB --> PS["PubSub\nScore Publishing"]
PS --> LR["Listing\nRender"]
class TT amber
class HNSW violet
class LGB blue
class PS greenWhy two stages instead of one?
Scale. Cars24 UAE had ~10,000 active listings and 10M+ monthly sessions. Scoring every user against every car with a full-featured model at request time doesn’t fit in a 100ms latency budget. The Two-Tower model produces embeddings offline (or near-real-time), and HNSW retrieval runs in sub-millisecond time. Only the top-K candidates hit the expensive re-ranker.
Feature richness. The Two-Tower model operates on features that can be encoded into dense embeddings — user behavior patterns, car attributes. The LightGBM re-ranker operates on features that don’t embed well but matter enormously for ranking: real-time demand signals, price competitiveness, inventory freshness, position in the sales funnel. Two stages let each model do what it’s best at.
Decoupled iteration. I could improve the retrieval model and the ranker independently. A better embedding didn’t require retraining the ranker, and vice versa. This mattered when I was the only DS on the UAE team — parallel experiment tracks weren’t an option; sequential, decoupled iterations were.
3. Two-Tower Retrieval: Learning User-Car Affinity#
The Two-Tower architecture (also called a dual-encoder) learns separate embedding representations for users and items, then scores relevance as the similarity between the two embeddings in a shared latent space. It’s the same architecture behind YouTube’s recommendation candidate generation and most large-scale retrieval systems.
Why Two-Tower for Used Cars#
The used-car marketplace has a property that makes collaborative filtering and matrix factorization awkward: inventory turns over completely. A car listed today may be sold tomorrow. There’s no long-lived item catalog for a user-item interaction matrix to accumulate signal on. Every car is effectively a new item.
The Two-Tower model handles this gracefully because it learns to embed attributes, not specific item IDs. A 2022 Hyundai Tucson in white with 30,000 km and a price of 72,000 AED gets an embedding based on its features — make, model, year, body type, color, odometer, price, condition score. When that car sells and a similar one is listed, the new car lands in a similar region of the embedding space without needing any interaction history.
The Two Towers#
graph TD
subgraph "User Tower"
U1["Click history\n(make/model/price distribution)"] --> UE["User\nEmbedding"]
U2["Search query tokens"] --> UE
U3["Filter selections\n(body type, price range, year)"] --> UE
U4["Session features\n(time on site, pages viewed)"] --> UE
U5["Recency-weighted\ninteraction sequence"] --> UE
end
subgraph "Item Tower"
I1["Car attributes\n(make, model, year, body type)"] --> IE["Item\nEmbedding"]
I2["Price & market position\n(vs. segment median)"] --> IE
I3["Condition & inspection\nscore"] --> IE
I4["Demand signals\n(I2V, STR, BI rate)"] --> IE
I5["Listing age &\ninventory freshness"] --> IE
end
UE --> SIM["Cosine Similarity\nin Shared Space"]
IE --> SIM
SIM --> SC["Relevance\nScore"]
class UE amber
class IE blue
class SIM greenThe user tower encodes behavioral signals into a dense vector. For users with click history (≥3 clicks in the past 30 days), the primary features are the distribution of their clicked cars’ attributes — what makes, models, body types, price ranges, and age ranges they’ve engaged with, weighted by recency and depth of engagement (a gallery view + inspection report view counts more than an impression). For cold-start users (more on this in Section 5), search query tokens and filter selections substitute for click history.
The item tower encodes car attributes and demand signals. The key insight is including demand-side features (I2V, STR, BI rate) in the item embedding — these capture “cars that similar users have shown interest in” without requiring explicit collaborative filtering. A car with high I2V (impressions-to-views) among users whose embeddings are similar to the current user gets a relevance boost, even if this specific user hasn’t seen it.
HNSW: Approximate Nearest Neighbor Serving#
Once both towers are trained, every car in the active inventory gets an item embedding, and these embeddings are indexed in an HNSW (Hierarchical Navigable Small World) graph for approximate nearest neighbor search. At serving time, the user embedding is computed, and HNSW returns the top-K most similar item embeddings in sub-millisecond time.
HNSW was the right index choice over alternatives (LSH, IVF, brute-force) for our scale: ~10,000 items is small enough that HNSW’s memory overhead is negligible, but large enough that brute-force cosine similarity at request time would eat into the latency budget when multiplied across concurrent sessions. HNSW gave us consistent sub-millisecond retrieval with recall well above 95%.
The index is rebuilt on a batch cadence as inventory changes. Because cars are listed and sold daily, the index needs to stay fresh — stale embeddings for sold cars are the most common failure mode. A car that no longer exists ranking in someone’s top 10 is worse than no personalization at all.
4. LightGBM Re-ranker: Precision Ranking#
The Two-Tower model retrieves a broad candidate set optimized for relevance. The LightGBM re-ranker takes those candidates and re-orders them using features that are too complex or too dynamic for the embedding model.
Feature Categories#
The re-ranker operates on a richer feature set organized into five categories:
Car attributes. Odometer reading, odometer-per-year, price relative to segment median (the “proportion RFC of bought price”), odometer deviation from make-model median. These capture whether a specific car is a good or bad deal relative to comparable inventory.
Demand signals. Impressions-to-views (I2V) at 15-day and 90-day windows, view-to-booking-initiated (V2BI) at both windows, sell-through rate (STR) at 90 days, BI per listing per day. These features encode the platform’s collective intelligence — if many users view a car but nobody books it, something is wrong with the listing (price too high, condition concern, bad photos).
Market science. Active listing count for the same make-model-year, listing count for competitors (similar body type and price band), price deviation from the competitive set. A car priced 10% below its competitive set is a value proposition the ranker should surface.
Car condition. Post-refurb inspection score, specific defect flags. Users care about condition but don’t always filter for it — the ranker can implicitly learn that users who click on newer, lower-odometer cars are condition-sensitive.
Supply context. Mean similar-car listing count at the body type and price level. If there’s only one SUV under 50k AED on the platform, it should rank higher for SUV-preferring users regardless of other signals — scarcity matters.
The Ship Pipeline#
This is where most ML projects fail. A model that improves offline metrics but degrades the user experience in production is worse than no model. The ship pipeline is a sequence of gates, each of which must pass before the next:
graph LR
A["Offline Eval\nNDCG@K / MAP\nmust beat baseline"] --> B["Position-Bias\nCorrected Replay\non logged traffic"]
B --> C["Shadow\nDeployment\n(no user impact)"]
C --> D["Power-Analysed\nA/B Test\n+ SRM checks"]
D --> E["100% Rollout\n+ PSI Drift\nMonitoring"]
class A,B amber
class C violet
class D blue
class E greenOffline evaluation. NDCG@K and MAP on a held-out test set. The model must beat the current production baseline by a meaningful margin. This gate catches models that don’t learn anything useful.
Position-bias-corrected replay. This is the critical gate that most teams skip. Logged click data is biased — users click on items because they were shown at the top, not only because they’re relevant. If you train on raw click data, your model learns to reproduce the existing ranking, not improve it. Position-bias correction (inverse propensity scoring or a position feature that’s zeroed at inference) debiases the training signal. The replay step evaluates the new model’s ranking against what users actually engaged with, correcting for the fact that they only saw the old ranking.
Shadow deployment. The new model scores every request in production alongside the existing model, but only the existing model’s scores are served. This catches infrastructure issues — latency spikes, missing features, null scores — without affecting users.
Power-analysed A/B test. User-level randomization (not session-level — a user must see the same model across their entire journey). Sample ratio mismatch (SRM) checks run daily to detect randomization bugs. The test runs until statistical power is sufficient for the primary metric (U2BI), with secondary metrics monitored but not used for ship/no-ship decisions.
PSI drift monitoring. Post-launch, Population Stability Index tracks whether the feature distributions in production are drifting from the training distribution. A PSI spike means the model is scoring on data it wasn’t trained for — retrain trigger.
5. Cold-Start: The 35% Problem#
The user split in Cars24 UAE at a daily level looked like this:
- ~24,000 unique daily users
- ~12,000 with ≥10 seconds and ≥1 impression (“relevant users”)
- Of those, ~5,400 had ≥3 clicks in the past 30 days — the users the system was designed for
- ~2,500 had 1-2 clicks — enough signal to personalize partially
- ~4,080 had zero clicks — 34% of the relevant base
The cohort-based system couldn’t personalize the zero-click users at all. They got default sorting. The N1 system, originally designed for ≥3-click users, was experimentally expanded to all users — but without click history, the user tower had nothing to embed.
The Insight: Search and Filter as Proxy Signals#
The key observation was that intentful users — users who actually want to buy a car — almost always do one of two things before clicking: they search or they apply filters. A user who types “Honda Civic 2021” into the search bar and then filters by “under 60k AED” has told you an enormous amount about their preferences without clicking a single car.
These signals are weaker than clicks (a search for “Honda” doesn’t mean they’ll buy a Honda — maybe they’re comparing), but they’re dramatically stronger than nothing. And for cold-start users, “stronger than nothing” is the bar.
The implementation: search query tokens and filter-selection impressions are included as features in the user tower. For users with click history, these features are present but dominated by the stronger click signals. For zero-click users, they are the signal — the user embedding is constructed primarily from what they searched for and filtered by.
What Worked and What Didn’t#
For users with ≥3 clicks (45% of relevant base), the results were strong across the board:
| Metric | Cohort PZN | N1 PZN | Uplift |
|---|---|---|---|
| Click Recall@50 | 12% | 30% | +150% |
| BI Recall@50 | — | 33% | (new metric) |
| U2BI | baseline | +16% | |
| Time on site | baseline | +6% |
For users with 1-2 clicks (21% of relevant base):
- Clicks per user improved by 13% (1.16 → 1.30)
- Time spent increased by 33% (28 → 33 minutes)
- Conversion signal too sparse to measure reliably
For zero-click users (34% of relevant base):
- Time spent improved by 13% (22% for brand-new users)
- Clicks per user: no improvement
That last line is the honest result. The search/filter proxy gets the right category of cars in front of zero-click users (hence the time spent improvement — they’re engaging longer with the content), but it doesn’t make them click. The likely reason: zero-click users aren’t failing to find relevant cars. They’re casually browsing, comparing prices, or waiting for better deals. Personalization solves a discovery problem; these users may have a decision problem that personalization can’t fix.
The 35% cold-start problem is solved in the sense that 100% of relevant users now receive personalized rankings. It’s unsolved in the sense that personalized rankings don’t move the needle for users who aren’t ready to engage.
6. Results: What the Numbers Mean#
After full rollout on 12 September 2024 (100% of UAE users, all platforms), the system ran for several weeks before the numbers stabilized. Here’s what we measured and what each metric actually captures.
Click Recall@50#
What it measures: Of all the cars a user eventually clicked on (anywhere in the catalog), what fraction appeared in the top 50 recommendations from our model?
Why it matters: This is the metric that best captures “did we understand what this user wanted?” A high Click Recall@50 means the model successfully predicted which cars the user would find interesting, and surfaced them early — before the user had to scroll past 50 generic listings.
Result: 12% → 30% (+150%). Under the cohort system, only 12% of eventual clicks were on cars in the top 50. The N1 system surfaced 30%. This means the user finds what they want 2.5x faster.
Buyer Conversion (U2BI)#
What it measures: User-to-booking-initiated. The fraction of users who initiate a booking (the first hard commitment step — choosing a car, entering details, proceeding to payment).
Result: +16% uplift over the cohort baseline. Faster discovery → faster decision → more bookings.
V2Bi (View-to-Buy Intent)#
What it measures: Of the cars that received a detail-page view, what fraction progressed to booking initiation?
Result: +21% uplift via the LightGBM re-ranker. This metric isolates the re-ranker’s contribution — the Two-Tower model gets the right cars into the candidate set, but the re-ranker determines the order, and order determines which cars get viewed.
The Cohort vs N1 Intuition#
The simplest way to understand the difference:
Imagine 6 users and 10 cars. Under cohort-based personalization, users A, B, and D are in cohort C1. They all see the same ranking: cars 3, 1, 2, 5, rest. Under N1, each user sees their own ranking: A sees 1, 2, 3, rest. B sees 1, 2, 5, rest. D sees 1, 3, rest. The rankings aren’t wildly different — the cohort wasn’t wrong about the broad category. But the order within that category is where the value lives, and cohorts can’t capture that.
7. What I’d Do Differently#
The Data Lag Problem#
The system personalized users based on interaction data that was 2 days old. User A’s recommendations on Tuesday reflected their clicks from Sunday — not Monday. In a marketplace where new inventory is listed daily and user intent shifts session-to-session, this is a meaningful gap.
The failure mode: a user browses Nissan Sunny and Hyundai Lancerx on day T-2. On day T-1, they shift to Sportage and Creta (maybe they got a salary raise, maybe they’re shopping for someone else). On day T0, the recommendations still push Sunny and Lancerx. The user has to manually search for what they actually want now.
The fix is reducing the data pull lag to same-day or half-day. The blocking constraint was the batch pipeline architecture — GA4 event data landed in Snowflake with a processing delay, and the feature engineering pipeline ran on a daily cron. Moving to a streaming architecture (Kafka → real-time feature updates) would solve this but required infrastructure investment beyond the DS team’s scope.
Recency Weighting#
Related to the data lag: when we did have interaction data, all interactions within the 30-day window were weighted equally. A click from 28 days ago counted the same as a click from yesterday. This is obviously wrong — user preferences shift, and recent actions are far more predictive. Exponential time decay on interaction weights (half-life of ~7 days) would improve the embedding quality for active users.
Quality Guardrails for Low-Click Users#
Users with exactly 3 clicks qualified for full personalization. But 3 clicks on 3 variants of the same car (e.g., three white Honda Civics) doesn’t mean the user wants a white Honda Civic — they might be comparing prices for the same model. A guardrail requiring clicks on at least 2 unique cars (different make-model combinations) would trade coverage for signal quality. I’d expect a small drop in the personalizable user count but a meaningful lift in recommendation accuracy for the remaining users.
Zero-Click Users Need Something Else#
The honest conclusion from Section 5: personalization doesn’t move clicks for users with zero click history. These users may need a different intervention entirely — better default assortment, more compelling pricing visibility, or proactive notifications when a car matching their search history drops in price. Personalization solves “I know what I want but can’t find it.” Zero-click users may be stuck on “I’m not sure I want to buy at all.”
Online Scoring#
The shipped system used batch-computed scores published via PubSub. The natural next step is online scoring — computing rankings at request time using the user’s in-session behavior. This captures within-session intent shifts (the user started browsing sedans but is now looking at SUVs) and eliminates the data lag entirely. The Two-Tower architecture supports this cleanly: recompute the user embedding with in-session features, query HNSW, re-rank. The infrastructure cost is the blocker, not the ML.
Summary#
The shift from 25 cohorts to N=1 personalization wasn’t a model upgrade — it was an architecture change. The cohort system worked by finding groups of similar users and giving them the same ranking. The N1 system works by learning what each individual user wants and ranking the entire catalog for them.
The Two-Tower model handles the core problem: representing user preferences and car attributes in a shared embedding space where relevance is a distance metric. The LightGBM re-ranker handles the nuances: demand dynamics, price competitiveness, inventory freshness — features that change daily and don’t embed well.
The cold-start solution — using search and filter signals as proxy features for the user embedding — extended coverage from 65% to 100% of users. It doesn’t make zero-click users click, but it puts the right category of cars in front of them, which is better than the default sorting they had before.
The ship pipeline — offline eval, position-bias correction, shadow deployment, power-analysed A/B, drift monitoring — is where most of the engineering effort went. A recommendation model is only as good as the infrastructure that validates it, deploys it, and catches when it degrades. Without the pipeline, the model is a notebook. With it, it’s a production system serving 10M+ sessions.
If I were building this again from scratch, the one thing I’d change is the data freshness architecture. Batch pipelines with 2-day lag are the ceiling on recommendation quality. Everything else — the model, the features, the cold-start hack — iterates. The data lag is structural.
I built this system at Cars24 UAE. I’m Aman Jain. If you’re working on recommendation systems, personalization, or cold-start problems, reach me at amanjain.codes.