用“战舰”游戏训练AI提问,MIT与哈佛研究揭示推理策略关键
2026年,人工智能代理热潮持续升温,但这些半自主程序在高风险领域如医疗诊断和科学发现中,往往难以在不确定环境下提出有效问题。麻省理工学院计算机科学与人工智能实验室(CSAIL)与哈佛大学工程与应用科学学院的研究人员,通过经典猜谜游戏“战舰”对此进行探索。他们设计了“协作战舰”游戏,让人类玩家扮演“队长”提问,队友“观察员”回答,并构建了“BattleshipQA”数据集。随后测试了GPT-5等大模型和Llama 4 Scout等小模型。结果显示,大模型能击败人类,但小模型表现不佳。关键在于模型不擅长提出有用问题。研究者引入蒙特卡洛推理策略,模型据此评估各选项概率,从而提出更高效的问题。Llama 4 Scout的胜率从8%提升至82%,甚至超越GPT-5,而其成本仅为后者的1%。此外,通过将问题转换为验证代码,模型回答准确率平均提升15%。研究表明,赋予AI代理“世界模型”使其能预测和模拟环境,是提升提问能力的关键。 #AI #人工智能 #机器学习 #提问策略 #MIT #哈佛 #战舰游戏 #推理模型 #大模型
2026年,人工智能代理热潮持续升温,但这些半自主程序在高风险领域如医疗诊断和科学发现中,往往难以在不确定环境下提出有效问题。麻省理工学院计算机科学与人工智能实验室(CSAIL)与哈佛大学工程与应用科学学院的研究人员,通过经典猜谜游戏“战舰”对此进行探索。他们设计了“协作战舰”游戏,让人类玩家扮演“队长”提问,队友“观察员”回答,并构建了“BattleshipQA”数据集。随后测试了GPT-5等大模型和Llama 4 Scout等小模型。结果显示,大模型能击败人类,但小模型表现不佳。关键在于模型不擅长提出有用问题。研究者引入蒙特卡洛推理策略,模型据此评估各选项概率,从而提出更高效的问题。Llama 4 Scout的胜率从8%提升至82%,甚至超越GPT-5,而其成本仅为后者的1%。此外,通过将问题转换为验证代码,模型回答准确率平均提升15%。研究表明,赋予AI代理“世界模型”使其能预测和模拟环境,是提升提问能力的关键。 #AI #人工智能 #机器学习 #提问策略 #MIT #哈佛 #战舰游戏 #推理模型 #大模型
Deep Dive into Qiushi's Latest Edition: How to Build Future Industries? What are China's Fault-lines in Basic Research? How are LLMs Impacting Knowledge? What's Driving Japan's 'Neo-Militarism'? - Tracking People's Daily
Manoj Kewalramani Jun 03, 2026 5 1 Share Hi folks, In today’s edition, since I found little interesting in the paper, I am offering breakdowns of interesting articles in the latest edition of the Qiushi journal. I am covering four articles below: The first one is by the journal’s editorial department on China’s approach to future industries. The second is by Zhang Jun on the importance of basic/foundational research for cultivating future industries. He is quite critical in this regard. Third, there’s a piece by Sun Maosong talking about the impact of LLMs and generative AI on knowledge production. Finally, I am briefly discussi
Manoj Kewalramani Jun 03, 2026 5 1 Share Hi folks, In today’s edition, since I found little interesting in the paper, I am offering breakdowns of interesting articles in the latest edition of the Qiushi journal. I am covering four articles below: The first one is by the journal’s editorial department on China’s approach to future industries. The second is by Zhang Jun on the importance of basic/foundational research for cultivating future industries. He is quite critical in this regard. Third, there’s a piece by Sun Maosong talking about the impact of LLMs and generative AI on knowledge production. Finally, I am briefly discussi
ng the foreign affairs article in the journal, which deals with Japan’s so-called “neo-militarism”. I am also offering some takeaways from my perspective on some these articles. I hope you enjoy this format. If it works, I might continue this for future editions of the journal. In the latest edition of Qiushi, the lead article was an excerpt of a speech that Xi Jinping delivered during the 24th collective study session of the Politburo in January 2026. The speech focussed on cultivating future industries. Sinocism has the full translation of Xi’s article. The journal’s editorial department also has an interesting take on future industries, which I am sharing below. Instead of the full text, I am covering some of the key points. “Developing future industries cannot mean trying to cover everything and exerting force equally in all directions; it is imperative to clarify the direction and accurately identify the priorities. The Fourth Plenary Session of the 20th CPC Central Committee proposed that quantum technology, biomanufacturing, hydrogen energy and nuclear fusion energy, brain-computer interfaces, embodied intelligence, and sixth-generation mobile communications be promoted as new points of economic growth . In this important speech, General Secretary Xi Jinping explicitly identified these fields as the main areas of focus for the development of future industries for China during the 15th FYP period. Only by gaining a deep understanding of the development dynamics, industrial ecosystems, and prospective pathways of these fields can we better grasp the priorities, apply precisely targeted policies, and drive significant progress in the development of future industries during the 15th FYP. 发展未来产业,不能面面俱到、平均用力,必须明确方向、找准重点。党的二十届四中全会提出,要推动量子科技、生物制造、氢能和核聚变能、脑机接口、具身智能、第六代移动通信等成为新的经济增长点。在这篇重要讲话中,习近平总书记明确把这些领域作为“十五五”时期我国未来产业发展的主攻方向。只有深入认识这些领域的发展态势、产业生态与前景路径,才能更好把握重点、精准施策,推动未来产业发展在“十五五”时期取得明显进展. Why have these six major future industries been chosen as the main areas of focus? This was not done arbitrarily, but was a scientific and prudent choice made on the basis of an accurate assessment of factors such as global trends in science and technology, national strategic needs, and the laws governing industrial transformation. In terms of global trends in science and technology, all six of these industries represent frontier directions of the technological revolution and industrial transformation, and stand at critical junctures for technological breakthroughs. According to relevant research, nuclear fusion energy, sixth-generation mobile communications, and brain-computer interfaces are currently in the embryonic stage of development; hydrogen energy, quantum computing and precision measurement, and humanoid robots are in the growth stage; and quantum communications, new types of pharmaceutical manufacturing, and the like have already entered the expansion stage. The tiered distribution of these six industries chosen by China balances short-term breakthroughs with medium- and long-term development. 发展未来产业,不能面面俱到、平均用力,必须明确方向、找准重点。党的二十届四中全会提出,要推动量子科技、生物制造、氢能和核聚变能、脑机接口、具身智能、第六代移动通信等成为新的经济增长点。在这篇重要讲话中,习近平总书记明确把这些领域作为“十五五”时期我国未来产业发展的主攻方向。只有深入认识这些领域的发展态势、产业生态与前景路径,才能更好把握重点、精准施策,推动未来产业发展在“十五五”时期取得明显进展. In terms of national strategic needs, these six industries span the domains of future manufacturing, future information, future energy, and future health, and are characterised by high technological content, high added value, and low resource consumption.
跨模型激活迁移在多跳推理中遭遇负面结果
一篇来自 arXiv 的学术论文报告了关于跨模型激活迁移(Cross-Model Activation Transfer)在 Pythia 多跳推理设置中的负面研究结果。该研究由 Peiyan Zhang 等人完成,论文标题明确指出在 Pythia 多跳环境下,尝试将激活模式从一个模型迁移到另一个模型未能取得预期效果。实验表明,即使模型结构相似,激活迁移在多跳推理任务中会导致性能显著下降,无法实现有效的知识转移。这一发现对当前流行的模型激活工程(如激活增强、编辑)提出了挑战,提示研究者需谨慎对待跨模型迁移的适用性。该论文已提交 arXiv,等待 DOI 注册,其完整内容可供公开访问。 #AI #机器学习 #跨模型迁移 #多跳推理 #负面结果 #arXiv #大模型 #激活工程 #Pythia
一篇来自 arXiv 的学术论文报告了关于跨模型激活迁移(Cross-Model Activation Transfer)在 Pythia 多跳推理设置中的负面研究结果。该研究由 Peiyan Zhang 等人完成,论文标题明确指出在 Pythia 多跳环境下,尝试将激活模式从一个模型迁移到另一个模型未能取得预期效果。实验表明,即使模型结构相似,激活迁移在多跳推理任务中会导致性能显著下降,无法实现有效的知识转移。这一发现对当前流行的模型激活工程(如激活增强、编辑)提出了挑战,提示研究者需谨慎对待跨模型迁移的适用性。该论文已提交 arXiv,等待 DOI 注册,其完整内容可供公开访问。 #AI #机器学习 #跨模型迁移 #多跳推理 #负面结果 #arXiv #大模型 #激活工程 #Pythia
Claude Code 推出动态工作流
Claude Code 正从单一代码助手升级为可编排的 Agent 工作台。其最新推出的 workflows(工作流)功能,核心在于让 Claude 不再局限于同一上下文窗口内的“想完再做”,而是能根据任务动态生成执行框架。该框架支持任务拆分、子 Agent 派发、并行处理、交叉验证与循环迭代,使多个 AI 实体可以像团队一样协同作业。这一变化意味着开发者能借助 Claude Code 处理更复杂的多步骤任务,提升自动化效率与代码质量。目前该功能已逐步开放,有望推动 AI 编程工具向更高级的自主协作方向演进。 #Claude #AI #工作流 #Agent #代码助手 #动态编排 #编程工具
Claude Code 正从单一代码助手升级为可编排的 Agent 工作台。其最新推出的 workflows(工作流)功能,核心在于让 Claude 不再局限于同一上下文窗口内的“想完再做”,而是能根据任务动态生成执行框架。该框架支持任务拆分、子 Agent 派发、并行处理、交叉验证与循环迭代,使多个 AI 实体可以像团队一样协同作业。这一变化意味着开发者能借助 Claude Code 处理更复杂的多步骤任务,提升自动化效率与代码质量。目前该功能已逐步开放,有望推动 AI 编程工具向更高级的自主协作方向演进。 #Claude #AI #工作流 #Agent #代码助手 #动态编排 #编程工具
LLM Agent 不确定性感知澄清机制
来自 ICML 2026 的提交论文提出一种新颖方法,通过信息增益指标驱动 LLM Agent 在推理过程中主动进行不确定性感知的澄清。该方法使得 Agent 能够在面对模糊或不完整输入时,判断何时以及如何向用户或环境提出澄清问题,从而减少错误决策。研究通过实验验证了基于信息增益的澄清策略相较于基线方法在任务完成准确率和效率上的显著提升,为构建更可靠、更具交互性的智能 Agent 提供了重要理论基础。论文由 Mengyi Deng 等七位作者共同完成,预印本已在 arXiv 发布。 #LLM #AIAgent #信息增益 #不确定性 #ICML2026 #arXiv #人工智能 #机器学习 #自然语言处理
来自 ICML 2026 的提交论文提出一种新颖方法,通过信息增益指标驱动 LLM Agent 在推理过程中主动进行不确定性感知的澄清。该方法使得 Agent 能够在面对模糊或不完整输入时,判断何时以及如何向用户或环境提出澄清问题,从而减少错误决策。研究通过实验验证了基于信息增益的澄清策略相较于基线方法在任务完成准确率和效率上的显著提升,为构建更可靠、更具交互性的智能 Agent 提供了重要理论基础。论文由 Mengyi Deng 等七位作者共同完成,预印本已在 arXiv 发布。 #LLM #AIAgent #信息增益 #不确定性 #ICML2026 #arXiv #人工智能 #机器学习 #自然语言处理
Neural Radiated-Noise Fields 论文提出水下航行器噪声三维预测新方法
一项最新研究提出“神经辐射噪声场”(Neural Radiated-Noise Fields)技术,用于在三维场景中预测无人水下航行器(UUV)的辐射噪声频谱。该论文由 Yan Wu 等作者完成,已在 arXiv 预印本平台发布。传统水下噪声预测依赖经验模型或简化仿真,难以精确捕捉复杂三维环境中的声场分布。新方法借鉴神经辐射场(NeRF)思想,利用神经网络隐式编码水下声源及传播路径,可动态预测不同方位和距离下的噪声频谱。实验表明,该模型能有效重构辐射噪声空间分布,对提升 UUV 隐身性能、海洋监测及声呐探测仿真具有重要意义。论文提供了完整的理论框架与初步验证结果,为后续水声领域深度学习应用开辟了新方向。 #论文 #水下航行器 #噪声预测 #神经网络 #三维场景 #声学 #AI #水声技术
一项最新研究提出“神经辐射噪声场”(Neural Radiated-Noise Fields)技术,用于在三维场景中预测无人水下航行器(UUV)的辐射噪声频谱。该论文由 Yan Wu 等作者完成,已在 arXiv 预印本平台发布。传统水下噪声预测依赖经验模型或简化仿真,难以精确捕捉复杂三维环境中的声场分布。新方法借鉴神经辐射场(NeRF)思想,利用神经网络隐式编码水下声源及传播路径,可动态预测不同方位和距离下的噪声频谱。实验表明,该模型能有效重构辐射噪声空间分布,对提升 UUV 隐身性能、海洋监测及声呐探测仿真具有重要意义。论文提供了完整的理论框架与初步验证结果,为后续水声领域深度学习应用开辟了新方向。 #论文 #水下航行器 #噪声预测 #神经网络 #三维场景 #声学 #AI #水声技术
研究提出AI文本检测新框架
一篇来自arXiv的论文《Your AI Text is not Mine》重新定义了在现实假设下AI生成文本检测的方法与评估标准。该研究由Nils Dycke等四位学者完成,于2026年6月提交。传统AI文本检测多聚焦于区分人类与机器写作,但作者指出,不同AI模型、不同提示策略生成的文本存在显著差异,简单二分类无法满足实际需求。论文提出更精细的检测框架,强调需要区分“你的AI文本”和“我的AI文本”,即在多模型、多场景下实现归因与验证。研究还构建了新的评估基准,模拟真实环境中的文本混淆、改写等干扰因素,为提升检测鲁棒性提供了新路径。这项工作对学术诚信、内容审核及AI安全治理具有重要参考价值。 #AI #文本检测 #人工智能 #学术论文 #安全 #arXiv #深度学习
一篇来自arXiv的论文《Your AI Text is not Mine》重新定义了在现实假设下AI生成文本检测的方法与评估标准。该研究由Nils Dycke等四位学者完成,于2026年6月提交。传统AI文本检测多聚焦于区分人类与机器写作,但作者指出,不同AI模型、不同提示策略生成的文本存在显著差异,简单二分类无法满足实际需求。论文提出更精细的检测框架,强调需要区分“你的AI文本”和“我的AI文本”,即在多模型、多场景下实现归因与验证。研究还构建了新的评估基准,模拟真实环境中的文本混淆、改写等干扰因素,为提升检测鲁棒性提供了新路径。这项工作对学术诚信、内容审核及AI安全治理具有重要参考价值。 #AI #文本检测 #人工智能 #学术论文 #安全 #arXiv #深度学习