Concurrent Image Understanding and Generation: Self-Correcting Coupled Markov Jump Processes

Published
Source
arXiv
Paper number
655
Field
Machine Learning
arXiv ID
2607.13188

Key points

  • It generates text and images simultaneously in a single decoding process, with each side verifying the other's decisions.
  • It proposes CO2Jump, a training-free sampler that outperforms prior methods on image editing and visual reasoning tasks.
  • It adds a remasking mechanism so that tokens once fixed can be hidden again when evidence changes, allowing the model to correct contradictions on its own.
  • It builds and plans to release three large multimodal datasets: JEdit-1M, JMaze-200K, and JNono-200K.
  • Performance improves steadily as the number of denoising steps increases, showing that the benefits of cross-modal coupling accumulate.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)