In this case study

Friend suggestions

Friend suggestions are a candidate generator and a ranker. Candidates come from the graph and from revealed interest. The ranker is a logistic regression over twelve features, trained on real friendships, and served with Gumbel noise so the list varies between visits.

Candidates

Up to 200 candidates are collected per user from these sources:

  • Friends of friends, prioritised through close friends.
  • School cohorts with overlapping years.
  • Three revealed-interest sources: profiles the viewer opened in the last 30 days, people who viewed the viewer in the last 14 days, and recent likers of the viewer's posts.

The friend ranker

The ranker predicts whether A and B would form a healthy friendship. Its twelve features:

FeatureWhat it measures
Mutual countNumber of mutual friends
Closeness-weighted mutualsMutuals weighted by how close each is
Closeness chainThe sum over mutuals M of closeness(A, M) times closeness(B, M)
Taste similaritySimilarity of the two taste vectors
Town matchIDF-weighted: sharing Reykjavík scores about 0.08, a village near 1.0
School cohortSame school with overlapping years
Age proximityHow close in age
Like history, A to BPast likes in one direction
Like history, B to APast likes in the other direction
Profile viewsWhether either has viewed the other
Candidate recencyHow recently the candidate appeared in a source
Frozen graph similaritySimilarity of the two frozen Node2Vec vectors

The town match is weighted by inverse document frequency because a plain same-town flag would match everyone in Reykjavík with everyone else. A shared village scores near 1.0 and a shared capital about 0.08.

Labels

  • Positives: friendships formed in the last year, in both directions.
  • Hard negatives: declined friend requests.
  • Soft negatives: three random non-friends per positive.

The model is retrained on demand from the admin dashboard rather than on a schedule.

Serving

At serving time, each score gets Gumbel noise at temperature 0.15. That is equivalent to sampling from a softmax over the scores rather than taking the top of the list, so the suggestions vary from visit to visit.

The closeness score

Underneath the friend ranker, the feed's fan-out and the connection cards is one score. Likes, comments and messages increment directional counters between friends. Closeness is the geometric mean of the log-compressed incoming and outgoing counts, normalised by the user's own maximum, which gives a number between zero and one per friend.

All counters decay by a factor of 0.9 every Monday.

The score is used in three places:

  • It orders friends for the feed's fan-out waves, so the closest friends see a post first.
  • It feeds the friend ranker through the closeness-weighted mutuals and the closeness chain.
  • It is calibrated weekly against a sample of mature friendships to produce the tengslakort score out of ten, the number shown on the connection card between two people.

Shared orbit

Profiles also show shared-orbit suggestions that do not go through the ranker at all: a vector search around a 30/70 blend of the viewer's and the profile's fresh graph vectors. It returns people whose vectors are near both.