NCP-ArchPreview Technical Report: Moving towards Latent Space Language Models through Next Concept Prediction
- Published
- Source
- arXiv
- Paper number
- 1085
- Field
- LLMs / NLP
- arXiv ID
- 2609.10715
Key points
- The method builds a concept vocabulary from internal representations, trains next-concept prediction together with next-token prediction, and feeds predicted concepts back into token generation.
- The 8.9B model reached the final pretraining loss of the OLMo-3-7B baseline with 51.3% of its total training tokens, and after full training scored 2.45 points higher on the downstream task average.
- Against a separate 8.9B-sized baseline, it approached a similar training loss using 85% of the baseline compute, showing the gains are not explained by model size alone.
- It presents domain adaptation that tunes only a 17M-parameter concept module, and feeding concept representations into the DFlash2 draft generator raised mean acceptance length by 4.17%, suggesting benefits for speculative decoding.
- This architecture preview was evaluated at standard context lengths without long-context training, so whether the same gains hold in long-document settings needs further verification.
Paper links
External research summaries. These are not HDATF publications or measured product results.