AI4Meder 站内搜索

搜索医学 AI 论文与资源

按论文、数据资源、技术竞赛、投稿截止日期和课程资源检索社区内容，快速进入对应详情页。

17 条结果

输入关键词或点击标签，按论文、数据资源、竞赛截止日期、征稿与课程缩小范围。标签：medical LLM

论文ICLR 2026 Poster2026 年medical LLM agent

大语言模型能否匹配系统综述的结论？

ICLR 2026 Poster accepted paper at ICLR 2026. Systematic reviews (SR), in which experts summarize and analyze evidence across individual studies to provide insights on a specialized topic, are a cornerstone for evidence-based clinical decision-making, research, and policy. Given the exponential growth of scientific articles, there is growing interest in using large language models (LLMs) to automate SR generation. However, the ability of LLMs to critically assess evidence and reason across multiple documents to provide recommendations at the same proficiency as domain experts remains poorly characterized. We therefore ask: **Can LLMs match the conclusions of systematic reviews written by clinical experts when given access to the same studies?** To explore this question, we present MedEvidence, a benchmark pairing findings from 100 medical SRs with the studies they are based on.

医学影像计算临床语言智能论文 Benchmarks Multi-document Reasoning Medical AI 查看论文详情

论文ICLR 2026 Poster2026 年medical LLM agent

GALAX：面向精准医疗中可解释强化引导子图推理的图增强语言模型

ICLR 2026 Poster accepted paper at ICLR 2026. In precision medicine, quantitative multi-omic features, topological context, and textual biological knowledge play vital roles in identifying disease-critical signaling pathways and targets, guiding the discovery of novel therapeutics and effective treatment strategies. Existing pipelines capture only one or two of these—numerical omics ignore topological context, text-centric LLMs lack quantitative grounded reasoning, and graph-only models underuse rich node semantics and the generalization power of LLMs—thereby limiting mechanistic interpretability. Although Process Reward Models (PRMs) aim to guide reasoning in LLMs, they remain limited by coarse step definitions, unreliable intermediate evaluation, and vulnerability to reward hacking with added computational cost. These gaps motivate jointly integrating quantitative multi-omic signals, topological structure with node annotations, and literature-scale text via LLMs, using subgraph reasoning as the principle bridge linking numeric evidence, topological knowledge and language context.

医学影像计算临床语言智能论文 Reinforcement Learning Large Language Model (LLM)Text-Numeric Graph (TNG)查看论文详情

论文ICLR 2026 Poster2026 年medical LLM agent

Doctor-R1：通过体验式 Agent 强化学习掌握临床问诊

ICLR 2026 Poster accepted paper at ICLR 2026. The professionalism of a human doctor in outpatient service depends on two core abilities: the ability to make accurate medical decisions and the medical consultation skill to conduct strategic, empathetic patient inquiry. Existing Large Language Models (LLMs) have achieved remarkable accuracy on medical decision-making benchmarks. However, they often lack the ability to conduct the strategic and empathetic consultation, which is essential for real-world clinical scenarios. To address this gap, we propose Doctor-R1, an AI doctor agent trained to master both of the capabilities by ask high-yield questions and conduct strategic multi-turn inquiry to guide decision-making.

医学影像计算临床语言智能论文 Doctor Agent Clinical Inquiry Agentic Reinforcement Learning 查看论文详情

论文ICLR 2026 Poster2026 年medical LLM agent

AnesSuite：面向 LLM 麻醉学推理的综合基准与数据集套件

ICLR 2026 Poster accepted paper at ICLR 2026. The application of large language models (LLMs) in the medical field has garnered significant attention, yet their reasoning capabilities in more specialized domains like anesthesiology remain underexplored. To bridge this gap, we introduce AnesSuite, the first comprehensive dataset suite specifically designed for anesthesiology reasoning in LLMs. The suite features AnesBench, an evaluation benchmark tailored to assess anesthesiology-related reasoning across three levels: factual retrieval (System 1), hybrid reasoning (System 1.x), and complex decision-making (System 2). Alongside this benchmark, the suite includes three training datasets that provide an infrastructure for continued pre-training (CPT), supervised fine-tuning (SFT), and reinforcement learning with verifiable rewards (RLVR). Code/project link: https://github.com/MiliLab/AnesSuite

医学影像计算临床语言智能论文 Large language model Reasoning Anesthesiology 查看论文详情

论文ICLR 2026 Poster2026 年trustworthy medical AI

LiveClin：无泄漏的实时临床基准

ICLR 2026 Poster accepted paper at ICLR 2026. The reliability of medical LLM evaluation is critically undermined by data contamination and knowledge obsolescence, leading to inflated scores on static benchmarks. To address these challenges, we introduce LiveClin, a live benchmark designed for the approximating real-world clinical practice. Built from contemporary, peer-reviewed case reports and updated biannually, LiveClin ensures clinical currency and resists data contamination. Using a verified AI–human workflow involving 239 physicians, we transform authentic patient cases into complex, multimodal evaluation scenarios that span the entire clinical pathway. Code/project link: https://github.com/AQ-MedAI/LiveClin

医学影像计算医疗多模态临床语言智能论文 MultiModal Medical Benchmark ICLR 2026 查看论文详情

论文ICLR 2026 Poster2026 年medical LLM agent

KnowGuard：面向多轮临床推理的知识驱动拒答

ICLR 2026 Poster accepted paper at ICLR 2026. In clinical practice, physicians refrain from making decisions when patient information is insufficient. This behavior, known as abstention, is a critical safety mechanism preventing potentially harmful misdiagnoses. Recent investigations have reported the application of large language models (LLMs) in medical scenarios. However, existing LLMs struggle with the abstentions, frequently providing overconfident responses despite incomplete information. This limitation stems from conventional abstention methods relying solely on model self-assessments, which lack systematic strategies to identify knowledge boundaries with external medical evidences.

医学影像计算临床语言智能论文 multi-agent system 临床推理医学问答查看论文详情

论文ICLR 2026 Poster2026 年medical LLM agent

K-Prism：知识引导与提示融合的通用医学图像分割模型

ICLR 2026 Poster accepted paper at ICLR 2026. Medical image segmentation is fundamental to clinical decision-making, yet existing models remain fragmented. They are usually trained on single knowledge sources and specific to individual tasks, modalities, or organs. This fragmentation contrasts sharply with clinical practice, where experts seamlessly integrate diverse knowledge: anatomical priors from training, exemplar-based reasoning from reference cases, and iterative refinement through real-time interaction. We present $\textbf{K-Prism}$, a unified segmentation framework that mirrors this clinical flexibility by systematically integrating three knowledge paradigms: (i) $\textit{semantic priors}$ learned from annotated datasets, (ii) $\textit{in-context knowledge}$ from few-shot reference examples, and (iii) $\textit{interactive feedback}$ from user inputs like clicks or scribbles. Code/project link: https://github.com/bangwayne/K-Prism

医学影像计算论文 Medical Image Image Segmentation Universal Model Prompt Integration 查看论文详情

论文Submitted to ICLR 20262025 年可信、安全、公平与隐私

MedSentry：理解并缓解医学 LLM 多 Agent 系统中的安全风险

评估医疗 LLM 多智能体系统在攻击场景下的安全风险，比较多种多智能体拓扑，并提出检测与纠正机制以缓解恶意智能体造成的医疗安全问题。

可信、安全、公平与隐私医疗 AI 论文会议论文查看论文详情

数据资源Chinese medical instruction and dialogue textChinese medical instruction-tuning datasetAbout 140K medical SFT examples; see Hugging Face card开放访问

HuatuoGPT2-SFT-GPT4-140K 医学指令数据集

HuatuoGPT2-SFT-GPT4-140K is a Chinese medical supervised fine-tuning dataset containing medical instruction-style conversations and GPT-4-assisted responses. It is useful for Chinese medical assistant alignment and medical LLM instruction tuning.

临床语言智能数据集 medical SFT Chinese medical LLM instruction tuning medical_llm_agent 查看数据资源

数据资源Chinese medical question-answer textChinese medical QA corpusAbout 26 million medical QA pairs开放访问

Huatuo-26M：大规模中文医学问答数据集

Huatuo-26M is a large-scale Chinese medical question-answering dataset with about 26 million QA pairs collected for medical language modeling and medical dialogue research. It is suitable for Chinese medical LLM pretraining, fine-tuning, and QA system development.

临床语言智能数据集 Chinese medical QA large-scale corpus LLM training medical_llm_agent 查看数据资源

数据资源medical exam question-answer textmedical exam QA benchmarkUSMLE, Mainland China, and Taiwan exam-style QA splits; see repository开放访问

MedQA：含美国、中国大陆与台湾拆分的医学考试问答数据集

MedQA is a medical examination question answering benchmark with English and Chinese medical licensing-style question sets, including mainland China and Taiwan variants. It is widely used for medical QA and medical reasoning evaluation.

临床语言智能数据集 medical QA Chinese exam QA USMLE LLM evaluation 查看数据资源

数据资源Chinese medical exam and QA textChinese medical LLM evaluation benchmarkMultiple Chinese medical exam and benchmark splits; see Hugging Face card开放访问

CMB：中文医学基准

CMB is a comprehensive Chinese medical benchmark for evaluating medical large language models on medical exams, reasoning, and clinical knowledge questions. It is suited for Chinese medical QA, LLM evaluation, and instruction-following assessment.

临床语言智能数据集 Chinese medical benchmark medical LLM exam QA medical_llm_agent 查看数据资源

数据资源TextLLM benchmarkBenchmark and leaderboard开放访问

MedHELM 医学 LLM 评测基准

Medical LLM benchmark and leaderboard intended to broaden coverage beyond single medical QA datasets.

benchmark leaderboard medical LLM 查看数据资源

数据资源Text and medical imagesModelMedGemma / MedSigLIP model family开放访问

MedGemma / MedSigLIP 医学 AI 模型

Google Health AI Developer Foundations open model resources for medical text and medical image understanding, including MedGemma 1.5 resources.

medical LLM medical VLM open model 查看数据资源

技术竞赛报名入口公开，赛程未来阶段仍开放（2026-05-03 核验）sleep apnea detection and medical large-model applicationssleep monitoring signals and medical LLM applications截止北京时间 2026-08-07

京东健康·全球医疗 AI 创新大赛

京东健康全球医疗 AI 创新大赛公开页面显示赛事聚焦睡眠监测智能算法与医疗大模型创新应用两个方向，面向全球高校、科研机构、企业和个人开放报名，赛程含 6.17-8.7 初赛、后续复赛和 9.21 决赛。

临床语言智能竞赛 JD Health medical LLM sleep apnea medical AI competition 查看竞赛详情

征稿与合作ICTAI 2026截止北京时间 2026-06-30会议征稿

ICTAI 2026 征稿

CCF-Deadlines lists ICTAI 2026 with papers due 2026-06-30 AoE and conference dates 2026-11-02 to 2026-11-04 in Boca Raton. ICTAI is relevant to medical AI tools, clinical decision-support systems, medical LLM agents, and deployable AI techniques for healthcare workflows.

可信、安全、公平与隐私征稿 CCF C ICTAI AI tools clinical decision support 查看征稿详情

征稿与合作EMNLP 2026截止北京时间 2026-05-25会议征稿

EMNLP 2026 征稿

CCF-Deadlines lists EMNLP 2026 with papers due 2026-05-25 UTC-12 and conference dates 2026-10-24 to 2026-10-29 in Budapest. EMNLP is relevant to clinical NLP, biomedical language models, medical text mining, EHR note understanding, and safe medical LLM evaluation.

临床语言智能可信、安全、公平与隐私征稿 CCF B EMNLP clinical NLP 查看征稿详情