Learn from your own latents and not from tokens: A sample-complexity theory
- Published
- Source
- arXiv
- Paper number
- 269
- Field
- Machine Learning
- arXiv ID
- 2605.27734
Key points
- Predictor: A network that takes a tuple of latent values as input and tries to predict the distribution of its sibling latent values.
- Clusterer: A network that compresses the output of the predictor into a discrete code or cluster identity, and that code becomes the input to the next layer.
- Teacher as a lifter: The EMA teacher gradually incorporates the learned latent values into its own targets, which shifts the prediction task from token level to latent level during training.
Paper links
External research summaries. These are not HDATF publications or measured product results.