Measuring the Gap Between Human and LLM Research Ideas

Published
Source
arXiv
Paper number
558
Field
LLMs / NLP
arXiv ID
2607.01233

Key points

  • We introduce a two-axis taxonomy of research taste: seven opportunity patterns times six method paradigms, based on NSF, NIH, and DARPA.
  • LLMs score 47% to 64% on bridge opportunity versus humans at 12.1%, and 22% to 39% on synthesis and unification versus humans at 5.1%.
  • Human ideas have consistently higher normalized entropy on both axes, meaning they are more broadly distributed.
  • Cosine similarity between LLMs, 0.83, is higher than human LLM similarity, 0.72 to 0.78, showing model-to-model convergence.
  • LLMs emphasize the integrate and unify archetype at 34.2% versus humans at 2.35%, while humans focus on replace and decouple.
  • The key takeaway is that LLM ideas are reasonable but distributionally narrow, and the task is to align research taste.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)