Context-weighted Discrete Flow Matching

Published
Source
arXiv
Paper number
715
Field
Machine Learning
arXiv ID
2607.21427

Key points

  • It confirms experimentally that the more filled neighbors a token has, the lower its prediction entropy becomes.
  • It proposes a context-weighted sampler that reflects local context density in sampling without any extra training.
  • A Scaled Cross-Entropy loss reweights the training signal and lowers OpenWebText generation perplexity by 63%.
  • On molecule generation with QM9, it increases the number of valid molecules by 2.8x and the number of novel molecules by 1.9x.
  • It keeps quality similar to semi-autoregressive block-diffusion methods while preserving the flexibility of order-free generation.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)