Learn from your own latents and not from tokens: A sample-complexity theory

Published
Source
arXiv
Paper number
269
Field
Machine Learning
arXiv ID
2605.27734

Key points

  • Predictor: A network that takes a tuple of latent values as input and tries to predict the distribution of its sibling latent values.
  • Clusterer: A network that compresses the output of the predictor into a discrete code or cluster identity, and that code becomes the input to the next layer.
  • Teacher as a lifter: The EMA teacher gradually incorporates the learned latent values into its own targets, which shifts the prediction task from token level to latent level during training.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)