Professional services / Recruiting platform
LinkedIn: Semantic matching of job postings and resumes
- Company
- Country
- United States
- Adoption stage
- In operation
- Source published
- Date basis
- The date the source was published. It can differ from the date adoption started.
- How the source was checked
- Read the full source text
The work problem
LinkedIn hit the limits of Pensieve, its existing embedding stack, in the recommendations that connect more than a billion members with tens of millions of job postings. The company said Pensieve relied on smaller and less precise models and was tied to taxonomies that were hard to maintain by hand and to rigid upstream pipelines. Putting a large language model straight into the serving path, it said, brings heavy compute cost, a more complex deployment pipeline and the ongoing burden of adapting the model to domain data.
Technology and data
The company built JUDE and put it into production. It is an internal platform that turns job postings, member profiles and resumes into embeddings in a shared space. A single shared base model is paired with prompt templates for each input type and fine tuned in a two tower setup. Training runs first on relevance labels and then continues on actual application records. The team used LoRA, Flash Attention 2, bfloat16 and DeepSpeed ZeRO across multiple H100 machines. On the serving side it moved from a Lambda architecture to a Kappa architecture. Kafka change streams select only the items that changed for recomputation, and the results are written into the Venice key value store for online reads. The model itself is served as a gRPC service on Kubernetes GPU pods.
Results
LinkedIn said it added these embeddings to the second stage ranking models in both job recommendations and job search, replacing overlapping standardised features, and ramped traffic gradually. In online A/B tests it reported Qualified Applications up 2.07%, the ratio of dismissals to applies down 5.13%, and total job applications up 1.91%. The company added that this was the largest metric gain it had seen from a single model change in that half year. On cost, recomputing only changed items cut inference cost by up to three times, while hashing and deduplication reduced inference volume during initial backfill by about six times. At the 7 billion parameter scale it held response time under 300 milliseconds at the 95th percentile.
Limits and open questions
The source is an engineering blog written by the company itself and it does not state the exact name and size of the base model used. It does not say when the A/B tests ran, for how long, or at what traffic share. The metrics are relative changes only, with no absolute figures. No adoption month is given, so the record is ordered by the publication date of the post. In earlier records this was a pending candidate whose source had not been read, and in this round the source was read in full and the entry moved to the reviewed set.
Sources
- JUDE: LLM based representation learning for LinkedIn job recommendationslinkedin.com, Accessed
Compiled from public sources. These are not results from ATF Works customers.