站内搜索 - AI4Meder

论文ICLR 2026 Poster2026 年trustworthy medical AI

ATPO：面向多轮医学对话的自适应树策略优化

ICLR 2026 Poster accepted paper at ICLR 2026. Effective information seeking in multi-turn medical dialogues is critical for accurate diagnosis, especially when dealing with incomplete information. Aligning Large Language Models (LLMs) for these interactive scenarios is challenging due to the uncertainty inherent in user-agent interactions, which we formulate as a Hierarchical Markov Decision Process (H-MDP). While conventional Reinforcement Learning (RL) methods like Group Relative Policy Optimization (GRPO) struggle with long-horizon credit assignment and Proximal Policy Optimization (PPO) suffers from unstable value estimation in this context, we propose a novel uncertainty-aware Adaptive Tree Policy Optimization (ATPO) algorithm. Our method adaptively allocates the rollout budget to states with high uncertainty, quantified by a composite metric of Bellman error and action-value variance.

医学影像计算临床语言智能可信、安全、公平与隐私论文 Reinforcement Learning (RL)Large Language Models (LLMs)查看论文详情

搜索医学 AI 论文与资源

1 条结果

ATPO：面向多轮医学对话的自适应树策略优化