Machine learning
Every model is a logistic regression trained by a hand-written mini-batch SGD loop inside Convex actions, fed by two families of self-trained 128-dimensional embeddings, and shipped through one champion-challenger gate.
There is no separate training service, notebook, feature store or GPU. Training data is the product's own tables: an impression log, a friendship graph, notification outcomes, daily activity snapshots. Training is a scheduled Convex action that reads those tables, packs the rows into Float32 matrices to fit the action's memory budget, runs gradient descent, and writes the weights back into a table. Serving is a query that reads the weights and computes a dot product. Every model's weights, training history and validation metrics are shown on the admin dashboard, which is where the models are debugged.
The models
| Model | Predicts | Features | Trains | Decides |
|---|---|---|---|---|
| Feed ranker | 6 heads: like, comment, feed dwell of 7 s or more, detail dwell of 3 s or more, fullscreen open, skip | 37 | Nightly at 03:00 | Order of the top 80 unseen candidates per session |
| Friend ranker | Would A and B form a healthy friendship | 12 | On demand from admin | Ranking of up to 200 friend-suggestion candidates |
| Notification ranker | 2 heads: opened within 1 h, churn | 12 type one-hots plus user state | Nightly | Gate on every proactive push |
| Churn risk | Inactive in the next 24 h | 11 | Weekly, Sunday 03:00 | Who gets the daily 18:00 win-back push |
| Post propensity | Posts in the next 7 days | 12 | Nightly cache refresh | Interpretation, and the "your circle is active" nudge |
| Retention importance | Returns next week, given surfaces used | 15 | Weekly | Which surface to experiment on next |
| Taste embeddings | Co-engagement similarity (item2vec) | 128-d | Online plus daily at 04:30 | The tasteSim feature, taste-matched fan-out, viral reseeding |
| Graph embeddings | Friendship-graph proximity (Node2Vec) | 128-d | Daily at 05:00, frozen every 14 days | The graphSim feature, shared-orbit suggestions, tengslakort |
| Dating Elo and power score | Desirability and deck order | 6 terms | Per swipe, replay on Monday, rescore daily | Swipe deck order and the daily pick |
The pages
- 5.1Feed rankerRetrieval by fan-out, then a six-head logistic regression over the top 80 candidates.
- 5.2EmbeddingsTaste vectors from item2vec and graph vectors from Node2Vec, both 128 dimensions.
- 5.3Friend suggestionsCandidate sources, the friend ranker, Gumbel noise, and the closeness score underneath.
- 5.4Growth modelsHook signals, the notification ranker, churn risk, post propensity and retention importance.
- 5.5DatingElo with weekly replay, the power score, and the daily pick.
Linear models
Logistic regression was chosen for three reasons. Its weights are readable: the admin dashboard shows every head's weights and training history, and when the skip head learned inverted signs, the fix was a sign constraint. It trains on up to a hundred thousand rows inside a single serverless action. And it ships through a simple gate: the challenger replaces the champion only if it beats both the seed baseline and the current champion on held-out log loss.
Extra expressiveness comes from features: hand-built crosses such as freshness times friend, embeddings for taste and graph position, and debiasing controls that absorb position and cold-start effects.