Ask to Be Sure: Informative Interactions for Confident Multi-Turn LLM Recommendation

Published
Source
arXiv
Paper number
934
Field
Information Retrieval
arXiv ID
2608.15949

Key points

  • It computes uncertainty from the frequency and rank of movies in the recommendation list, and it uses how much the current response and the next user response reduce that uncertainty as the reward.
  • INSPIRED has 801 training dialogs and 99 test dialogs, while ReDial has 8,631 training dialogs and 1,036 test dialogs after filtering, and the simulated conversations run for up to five turns.
  • Averaged over three runs, turn-entropy DPO reaches 27.94% hit rate on INSPIRED simulated dialogs, which is higher than CollabLLM's reward at 26.60%, and on ReDial conversation-level entropy DPO reaches 32.83% versus 30.03%.
  • On ReDial, the average number of turns needed to reach the correct movie is 2.74 versus 2.86 for CollabLLM DPO, but this is not measured in real human conversations.
  • The evaluation depends on automatic recommendation metrics and LLM simulated users, sensitivity to the setting of five recommendation lists sampled five times is not validated, and direct recommendation Hit@1 remains low at 3.32% on INSPIRED and 2.24% on ReDial.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)