Professional services / Recruiting platform

LinkedIn: Semantic matching of job postings and resumes

Company
LinkedIn
Country
United States
Adoption stage
In operation
Source published
Date basis
The date the source was published. It can differ from the date adoption started.
How the source was checked
Read the full source text

The work problem

LinkedIn hit the limits of Pensieve, its existing embedding stack, in the recommendations that connect more than a billion members with tens of millions of job postings. The company said Pensieve relied on smaller and less precise models and was tied to taxonomies that were hard to maintain by hand and to rigid upstream pipelines. Putting a large language model straight into the serving path, it said, brings heavy compute cost, a more complex deployment pipeline and the ongoing burden of adapting the model to domain data.

Technology and data

The company built JUDE and put it into production. It is an internal platform that turns job postings, member profiles and resumes into embeddings in a shared space. A single shared base model is paired with prompt templates for each input type and fine tuned in a two tower setup. Training runs first on relevance labels and then continues on actual application records. The team used LoRA, Flash Attention 2, bfloat16 and DeepSpeed ZeRO across multiple H100 machines. On the serving side it moved from a Lambda architecture to a Kappa architecture. Kafka change streams select only the items that changed for recomputation, and the results are written into the Venice key value store for online reads. The model itself is served as a gRPC service on Kubernetes GPU pods.

Results

LinkedIn said it added these embeddings to the second stage ranking models in both job recommendations and job search, replacing overlapping standardised features, and ramped traffic gradually. In online A/B tests it reported Qualified Applications up 2.07%, the ratio of dismissals to applies down 5.13%, and total job applications up 1.91%. The company added that this was the largest metric gain it had seen from a single model change in that half year. On cost, recomputing only changed items cut inference cost by up to three times, while hashing and deduplication reduced inference volume during initial backfill by about six times. At the 7 billion parameter scale it held response time under 300 milliseconds at the 95th percentile.

Limits and open questions

The source is an engineering blog written by the company itself and it does not state the exact name and size of the base model used. It does not say when the A/B tests ran, for how long, or at what traffic share. The metrics are relative changes only, with no absolute figures. No adoption month is given, so the record is ordered by the publication date of the post. In earlier records this was a pending candidate whose source had not been read, and in this round the source was read in full and the entry moved to the reviewed set.

Sources

Compiled from public sources. These are not results from ATF Works customers.

Read original (opens in a new tab)