When do prophets profit in prediction markets?
Focus to learn more arXiv-issued DOI via DataCite (pending registration) Submission history From: Haifeng Xu [ view email ] [v1] Tue, 7 Jul 2026 11:41:46 UTC (311 KB) Full-text links: Access Paper: View a PDF of the paper titled When do prophets profit in prediction markets?, by Anri Gu and 4 other authors View PDF HTML (experimental) TeX Source view license Current browse context: < prev | next > new | recent | 2026-07 Change to browse by: cs References & Citations NASA ADS Google Scholar Semantic Scholar export BibTeX citation Loading... BibTeX formatted citation × loading... Data provided by: Bookmark Bibliographic Tools Bibliographic and Citation Tools Bibliographic Explorer Toggle Bibliographic Explorer ( What is the Explorer? ) Connected Papers Toggle Connected Papers ( What is Connected Papers? ) Litmaps Toggle Litma
Focus to learn more arXiv-issued DOI via DataCite (pending registration) Submission history From: Haifeng Xu [ view email ] [v1] Tue, 7 Jul 2026 11:41:46 UTC (311 KB) Full-text links: Access Paper: View a PDF of the paper titled When do prophets profit in prediction markets?, by Anri Gu and 4 other authors View PDF HTML (experimental) TeX Source view license Current browse context: < prev | next > new | recent | 2026-07 Change to browse by: cs References & Citations NASA ADS Google Scholar Semantic Scholar export BibTeX citation Loading... BibTeX formatted citation × loading... Data provided by: Bookmark Bibliographic Tools Bibliographic and Citation Tools Bibliographic Explorer Toggle Bibliographic Explorer ( What is the Explorer? ) Connected Papers Toggle Connected Papers ( What is Connected Papers? ) Litmaps Toggle Litma
ps ( What is Litmaps? ) Toggle scite Smart Citations ( What are Smart Citations? ) Code, Data, Media Code, Data and Media Associated with this Article alphaXiv Toggle alphaXiv ( What is alphaXiv? ) Links to Code Toggle CatalyzeX Code Finder for Papers ( What is CatalyzeX? ) DagsHub Toggle DagsHub ( What is DagsHub? ) GotitPub Toggle ( What is GotitPub? ) Huggingface Toggle Hugging Face ( What is Huggingface? ) ScienceCast Toggle ScienceCast ( What is ScienceCast? ) Demos Demos Replicate Toggle Replicate ( What is Replicate? ) Spaces Toggle Hugging Face Spaces ( What is Spaces? ) Spaces Toggle ( What is ? ) Related Papers Recommenders and Search Tools Link to Influence Flower Influence Flower ( What are Influence Flowers? ) Core recommender toggle CORE Recommender ( What is CORE? ) Author Venue Institution Topic About arXivLabs arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs .
奖励密度启发式算法提升动态多车辆路径规划性能
近日,一篇题为《Reward-Density Heuristic for Dynamic Multi-Vehicle Routing: Performance and Computational Efficiency》的论文在arXiv预印本平台上线。该论文由Manish Kolachalam和Rani Malhotra共同撰写,针对动态多车辆路径问题提出了一种基于奖励密度的启发式算法,并系统评估了其性能与计算效率。动态多车辆路径问题在物流配送、紧急响应等场景中具有重要应用,该研究为实时路径优化提供了新的算法思路。论文目前处于等待DOI注册阶段,全文可通过arXiv网站获取。 #动态多车辆路径 #启发式算法 #奖励密度 #性能评估 #计算效率 #arXiv #预印本 #物流优化 #算法研究
近日,一篇题为《Reward-Density Heuristic for Dynamic Multi-Vehicle Routing: Performance and Computational Efficiency》的论文在arXiv预印本平台上线。该论文由Manish Kolachalam和Rani Malhotra共同撰写,针对动态多车辆路径问题提出了一种基于奖励密度的启发式算法,并系统评估了其性能与计算效率。动态多车辆路径问题在物流配送、紧急响应等场景中具有重要应用,该研究为实时路径优化提供了新的算法思路。论文目前处于等待DOI注册阶段,全文可通过arXiv网站获取。 #动态多车辆路径 #启发式算法 #奖励密度 #性能评估 #计算效率 #arXiv #预印本 #物流优化 #算法研究
PolyWorkBench:多语言长周期LLM代理性能评估基准发布
来自香港理工大学等机构的研究团队在arXiv上提交了一篇论文,提出了一个新的基准测试PolyWorkBench,专门用于评估大型语言模型在多语言、长周期任务中的代理能力。该基准模拟了需要长期规划和持续交互的复杂场景,涵盖多种语言环境,旨在填补现有LLM评估中缺乏多语言长周期任务标准的空白。PolyWorkBench通过一系列精心设计的任务,测试模型在理解、记忆、推理和多步执行等方面的综合表现,为研究多语言AI代理的可靠性和实用性提供了重要参考工具。该基准的发布有望推动LLM在全球化、多语言实际应用中的发展。 #PolyWorkBench #LLM #多语言 #AI代理 #基准测试 #学术研究 #arXiv
来自香港理工大学等机构的研究团队在arXiv上提交了一篇论文,提出了一个新的基准测试PolyWorkBench,专门用于评估大型语言模型在多语言、长周期任务中的代理能力。该基准模拟了需要长期规划和持续交互的复杂场景,涵盖多种语言环境,旨在填补现有LLM评估中缺乏多语言长周期任务标准的空白。PolyWorkBench通过一系列精心设计的任务,测试模型在理解、记忆、推理和多步执行等方面的综合表现,为研究多语言AI代理的可靠性和实用性提供了重要参考工具。该基准的发布有望推动LLM在全球化、多语言实际应用中的发展。 #PolyWorkBench #LLM #多语言 #AI代理 #基准测试 #学术研究 #arXiv
Information Limits and Attractor Dynamics in Economies of Frontier LLM Agents: A Pre
Focus to learn more arXiv-issued DOI via DataCite (pending registration) Submission history From: Cheng Qian [ view email ] [v1] Tue, 7 Jul 2026 08:39:24 UTC (35 KB) Full-text links: Access Paper: View a PDF of the paper titled Information Limits and Attractor Dynamics in Economies of Frontier LLM Agents: A Pre-Registered Test, by Cheng Qian View PDF HTML (experimental) TeX Source view license Current browse context: < prev | next > new | recent | 2026-07 Change to browse by: cs References & Citations NASA ADS Google Scholar Semantic Scholar export BibTeX citation Loading... BibTeX formatted citation × loading... Data provided by: Bookmark Bibliographic Tools Bibliographic and Citation Tools Bibliographic Explorer Toggle Bibliographic Explorer ( What is the Explorer? ) Connected Papers
Focus to learn more arXiv-issued DOI via DataCite (pending registration) Submission history From: Cheng Qian [ view email ] [v1] Tue, 7 Jul 2026 08:39:24 UTC (35 KB) Full-text links: Access Paper: View a PDF of the paper titled Information Limits and Attractor Dynamics in Economies of Frontier LLM Agents: A Pre-Registered Test, by Cheng Qian View PDF HTML (experimental) TeX Source view license Current browse context: < prev | next > new | recent | 2026-07 Change to browse by: cs References & Citations NASA ADS Google Scholar Semantic Scholar export BibTeX citation Loading... BibTeX formatted citation × loading... Data provided by: Bookmark Bibliographic Tools Bibliographic and Citation Tools Bibliographic Explorer Toggle Bibliographic Explorer ( What is the Explorer? ) Connected Papers
Toggle Connected Papers ( What is Connected Papers? ) Litmaps Toggle Litmaps ( What is Litmaps? ) Toggle scite Smart Citations ( What are Smart Citations? ) Code, Data, Media Code, Data and Media Associated with this Article alphaXiv Toggle alphaXiv ( What is alphaXiv? ) Links to Code Toggle CatalyzeX Code Finder for Papers ( What is CatalyzeX? ) DagsHub Toggle DagsHub ( What is DagsHub? ) GotitPub Toggle ( What is GotitPub? ) Huggingface Toggle Hugging Face ( What is Huggingface? ) ScienceCast Toggle ScienceCast ( What is ScienceCast? ) Demos Demos Replicate Toggle Replicate ( What is Replicate? ) Spaces Toggle Hugging Face Spaces ( What is Spaces? ) Spaces Toggle ( What is ? ) Related Papers Recommenders and Search Tools Link to Influence Flower Influence Flower ( What are Influence Flowers? ) Core recommender toggle CORE Recommender ( What is CORE? ) Author Venue Institution Topic About arXivLabs arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs .
AgoraSim:一种混合代理基建模框架在arXiv发布
近日,一篇题为"AgoraSim: A Hybrid Agent-Based Modeling Framework"的论文在arXiv预印本平台正式发布。论文由Chung-Chi Chen撰写,于2026年7月7日提交。该研究提出了一种名为AgoraSim的混合代理基建模框架,为复杂系统模拟提供了新的思路。论文目前提供PDF、HTML及TeX源码等多种格式,其DOI正在通过DataCite注册中。该论文归属于计算机科学领域,感兴趣的读者可通过arXiv免费获取全文。 #arXiv #代理基建模 #混合模型 #学术论文 #计算机科学 #AgoraSim #2026
近日,一篇题为"AgoraSim: A Hybrid Agent-Based Modeling Framework"的论文在arXiv预印本平台正式发布。论文由Chung-Chi Chen撰写,于2026年7月7日提交。该研究提出了一种名为AgoraSim的混合代理基建模框架,为复杂系统模拟提供了新的思路。论文目前提供PDF、HTML及TeX源码等多种格式,其DOI正在通过DataCite注册中。该论文归属于计算机科学领域,感兴趣的读者可通过arXiv免费获取全文。 #arXiv #代理基建模 #混合模型 #学术论文 #计算机科学 #AgoraSim #2026
Auto-DSM Under the Lens: A Black-Box Evaluation Framework for LLM-Based DSM Generation
Focus to learn more arXiv-issued DOI via DataCite (pending registration) Submission history From: Theo Hofman [ view email ] [v1] Tue, 7 Jul 2026 08:17:49 UTC (1,631 KB) Full-text links: Access Paper: View a PDF of the paper titled Auto-DSM Under the Lens: A Black-Box Evaluation Framework for LLM-Based DSM Generation, by Niels Potters and 1 other authors View PDF HTML (experimental) TeX Source view license Current browse context: < prev | next > new | recent | 2026-07 Change to browse by: cs eess References & Citations NASA ADS Google Scholar Semantic Scholar export BibTeX citation Loading... BibTeX formatted citation × loading... Data provided by: Bookmark Bibliographic Tools Bibliographic and Citation Tools Bibliographic Explorer Toggle Bibliographic Explorer ( What is the Explorer?
Focus to learn more arXiv-issued DOI via DataCite (pending registration) Submission history From: Theo Hofman [ view email ] [v1] Tue, 7 Jul 2026 08:17:49 UTC (1,631 KB) Full-text links: Access Paper: View a PDF of the paper titled Auto-DSM Under the Lens: A Black-Box Evaluation Framework for LLM-Based DSM Generation, by Niels Potters and 1 other authors View PDF HTML (experimental) TeX Source view license Current browse context: < prev | next > new | recent | 2026-07 Change to browse by: cs eess References & Citations NASA ADS Google Scholar Semantic Scholar export BibTeX citation Loading... BibTeX formatted citation × loading... Data provided by: Bookmark Bibliographic Tools Bibliographic and Citation Tools Bibliographic Explorer Toggle Bibliographic Explorer ( What is the Explorer?
) Connected Papers Toggle Connected Papers ( What is Connected Papers? ) Litmaps Toggle Litmaps ( What is Litmaps? ) Toggle scite Smart Citations ( What are Smart Citations? ) Code, Data, Media Code, Data and Media Associated with this Article alphaXiv Toggle alphaXiv ( What is alphaXiv? ) Links to Code Toggle CatalyzeX Code Finder for Papers ( What is CatalyzeX? ) DagsHub Toggle DagsHub ( What is DagsHub? ) GotitPub Toggle ( What is GotitPub? ) Huggingface Toggle Hugging Face ( What is Huggingface? ) ScienceCast Toggle ScienceCast ( What is ScienceCast? ) Demos Demos Replicate Toggle Replicate ( What is Replicate? ) Spaces Toggle Hugging Face Spaces ( What is Spaces? ) Spaces Toggle ( What is ? ) Related Papers Recommenders and Search Tools Link to Influence Flower Influence Flower ( What are Influence Flowers? ) Core recommender toggle CORE Recommender ( What is CORE? ) Author Venue Institution Topic About arXivLabs arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs .
整合知识图谱与多语言学术语料,推动SSH领域自适应LLM研究
一项最新研究提出通过整合知识图谱和多语言学术语料库,构建适用于社会科学与人文学科(SSH)的领域自适应大语言模型。该论文题为《Integrating knowledge graphs and multilingual scholarly corpora for domain-adaptive LLMs in SSH》,由Adam Faci等六位研究人员共同完成,于2026年7月7日提交至arXiv预印本平台。论文目前提供PDF全文和HTML预览,相关数字对象标识符(DOI)正在通过DataCite登记中。该研究属于计算机科学范畴,旨在利用结构化知识图谱与多语种学术文献,提升大语言模型在特定学科的理解与推理能力,为跨学科人工智能应用提供了新方向。 #知识图谱 #多语言语料库 #大语言模型 #领域自适应 #SSH #arXiv #学术研究 #人工智能 #跨学科
一项最新研究提出通过整合知识图谱和多语言学术语料库,构建适用于社会科学与人文学科(SSH)的领域自适应大语言模型。该论文题为《Integrating knowledge graphs and multilingual scholarly corpora for domain-adaptive LLMs in SSH》,由Adam Faci等六位研究人员共同完成,于2026年7月7日提交至arXiv预印本平台。论文目前提供PDF全文和HTML预览,相关数字对象标识符(DOI)正在通过DataCite登记中。该研究属于计算机科学范畴,旨在利用结构化知识图谱与多语种学术文献,提升大语言模型在特定学科的理解与推理能力,为跨学科人工智能应用提供了新方向。 #知识图谱 #多语言语料库 #大语言模型 #领域自适应 #SSH #arXiv #学术研究 #人工智能 #跨学科
SearchEyes论文提出多模态深度搜索智能新方法,通过搜索世界模拟实现前沿搜索能力
近日,一篇题为《SearchEyes: Towards Frontier Multimodal Deep Search Intelligence via Search World Simulation》的论文提交至预印本平台arXiv。该论文由Zhengbo Jiao等18位研究人员共同撰写,提出了一种通过模拟搜索世界来训练和评估多模态深度搜索智能的创新框架。研究聚焦于融合文本、图像等多种信息模态,在模拟环境中提升模型对复杂查询的深度理解与精准检索能力。该工作旨在突破传统搜索方法的局限,为下一代智能检索系统提供新思路,引起了学术界的广泛关注。 #多模态搜索 #深度搜索 #人工智能 #SearchEyes #arXiv #学术论文 #前沿科技 #智能检索
近日,一篇题为《SearchEyes: Towards Frontier Multimodal Deep Search Intelligence via Search World Simulation》的论文提交至预印本平台arXiv。该论文由Zhengbo Jiao等18位研究人员共同撰写,提出了一种通过模拟搜索世界来训练和评估多模态深度搜索智能的创新框架。研究聚焦于融合文本、图像等多种信息模态,在模拟环境中提升模型对复杂查询的深度理解与精准检索能力。该工作旨在突破传统搜索方法的局限,为下一代智能检索系统提供新思路,引起了学术界的广泛关注。 #多模态搜索 #深度搜索 #人工智能 #SearchEyes #arXiv #学术论文 #前沿科技 #智能检索
PCBWorld 基准环境发布,推进引擎驱动的PCB设计自动化
近日,研究团队在arXiv预印本平台发表论文《PCBWorld: A Benchmark Environment for Engine-Grounded PCB Design Automation》。该论文由Hyungseok Song等六位作者联合提出,旨在为基于引擎的印刷电路板(PCB)设计自动化系统建立一个标准的测试和评估环境。PCBWorld专注于为自动化算法提供统一的性能衡量基准,从而帮助研究者在布局、布线及整体设计流程中比较不同算法效果。这一基准的发布有望加速电子设计自动化(EDA)领域的技术迭代,尤其是在引入AI和优化引擎后,提升PCB设计效率与质量。 #PCBWorld #PCB设计 #自动化 #基准测试 #EDA #电子设计 #人工智能 #arXiv
近日,研究团队在arXiv预印本平台发表论文《PCBWorld: A Benchmark Environment for Engine-Grounded PCB Design Automation》。该论文由Hyungseok Song等六位作者联合提出,旨在为基于引擎的印刷电路板(PCB)设计自动化系统建立一个标准的测试和评估环境。PCBWorld专注于为自动化算法提供统一的性能衡量基准,从而帮助研究者在布局、布线及整体设计流程中比较不同算法效果。这一基准的发布有望加速电子设计自动化(EDA)领域的技术迭代,尤其是在引入AI和优化引擎后,提升PCB设计效率与质量。 #PCBWorld #PCB设计 #自动化 #基准测试 #EDA #电子设计 #人工智能 #arXiv
优势加权排序法揭示潜在抑郁严重程度 助力二元抑郁检测
研究人员提出一种名为优势加权排序(Advantage-weighting Ranking)的方法,用于改进二元抑郁检测任务。该方法通过挖掘抑郁症状的潜在严重程度信息,将排序学习机制引入传统二分类模型,从而提升检测的敏感性和特异性。论文由Manning Gao及其合作者撰写,于2026年7月提交至arXiv预印本平台。该研究为抑郁症的自动筛查提供了新的技术路径,有望帮助早期识别抑郁患者,推动心理健康评估的智能化发展。传统方法常忽视抑郁程度差异,而新方法通过优势加权对严重程度进行有效建模,使模型性能进一步提升。 #抑郁检测 #优势加权 #排序学习 #机器学习 #心理健康 #arXiv #AI #深度学习
研究人员提出一种名为优势加权排序(Advantage-weighting Ranking)的方法,用于改进二元抑郁检测任务。该方法通过挖掘抑郁症状的潜在严重程度信息,将排序学习机制引入传统二分类模型,从而提升检测的敏感性和特异性。论文由Manning Gao及其合作者撰写,于2026年7月提交至arXiv预印本平台。该研究为抑郁症的自动筛查提供了新的技术路径,有望帮助早期识别抑郁患者,推动心理健康评估的智能化发展。传统方法常忽视抑郁程度差异,而新方法通过优势加权对严重程度进行有效建模,使模型性能进一步提升。 #抑郁检测 #优势加权 #排序学习 #机器学习 #心理健康 #arXiv #AI #深度学习
StateFuse:面向多智能体系统的确定性冲突保持记忆模型
一项题为《StateFuse: Deterministic Conflict-Preserving Memory for Multi-Agent Systems》的研究论文近日在arXiv预印本平台发布。该论文由Sergey Volkov等三位学者共同完成,提出了StateFuse模型,这是一种专为多智能体系统设计的确定性冲突保持记忆机制。论文强调该模型能够在多智能体交互中确保记忆操作的确定性和冲突的有效保留,从而提升系统的鲁棒性。该论文于2026年7月7日提交,目前可通过arXiv获取PDF和HTML格式的全文。 #StateFuse #多智能体系统 #arXiv #预印本 #论文 #AI #机器学习 #确定性记忆 #冲突保留
一项题为《StateFuse: Deterministic Conflict-Preserving Memory for Multi-Agent Systems》的研究论文近日在arXiv预印本平台发布。该论文由Sergey Volkov等三位学者共同完成,提出了StateFuse模型,这是一种专为多智能体系统设计的确定性冲突保持记忆机制。论文强调该模型能够在多智能体交互中确保记忆操作的确定性和冲突的有效保留,从而提升系统的鲁棒性。该论文于2026年7月7日提交,目前可通过arXiv获取PDF和HTML格式的全文。 #StateFuse #多智能体系统 #arXiv #预印本 #论文 #AI #机器学习 #确定性记忆 #冲突保留
新研究提出TurnOPD方法,提升长时域智能体在线策略蒸馏效率
据arXiv预印本,来自Yuhang Zhou等研究者的论文《TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training》提出了一种创新的在线策略蒸馏框架。该方法通过引入“回合感知”(Turn-Aware)机制,解决了传统在线策略蒸馏在长时域任务中样本效率低、训练不稳定的问题。具体而言,TurnOPD能够利用回合内状态的时间相关性,动态调整知识迁移策略,从而在复杂多步决策场景下显著加速智能体训练过程,并提升最终策略性能。该研究为将蒸馏技术应用于机器人控制、游戏AI等需要长期规划的领域提供了新思路。 #强化学习 #策略蒸馏 #长时域 #智能体 #知识蒸馏 #深度学习 #AI机器人 #样本效率
据arXiv预印本,来自Yuhang Zhou等研究者的论文《TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training》提出了一种创新的在线策略蒸馏框架。该方法通过引入“回合感知”(Turn-Aware)机制,解决了传统在线策略蒸馏在长时域任务中样本效率低、训练不稳定的问题。具体而言,TurnOPD能够利用回合内状态的时间相关性,动态调整知识迁移策略,从而在复杂多步决策场景下显著加速智能体训练过程,并提升最终策略性能。该研究为将蒸馏技术应用于机器人控制、游戏AI等需要长期规划的领域提供了新思路。 #强化学习 #策略蒸馏 #长时域 #智能体 #知识蒸馏 #深度学习 #AI机器人 #样本效率
arXiv论文提出标题特定激活引导控制工具使用
来自arXiv预印本平台的信息显示,一篇题为《Controlling Tool Use with Heading-Specific Activation Steering》的研究论文已被提交。该论文由Yuqi Chen等四位作者共同完成,提交版本为v1,时间显示为2026年7月7日。根据论文标题推断,该研究提出了一种通过标题特定(heading‑specific)的激活引导技术来实现对工具使用的精确控制。目前论文已在计算机科学(cs)类别下公开,并提供PDF和HTML预览版本供研究人员免费阅读。该研究有望为提升人工智能系统在工具调用时的可控性和安全性提供新的理论思路。 #arXiv #论文 #计算机科学 #工具控制 #激活引导 #AI #新研究
来自arXiv预印本平台的信息显示,一篇题为《Controlling Tool Use with Heading-Specific Activation Steering》的研究论文已被提交。该论文由Yuqi Chen等四位作者共同完成,提交版本为v1,时间显示为2026年7月7日。根据论文标题推断,该研究提出了一种通过标题特定(heading‑specific)的激活引导技术来实现对工具使用的精确控制。目前论文已在计算机科学(cs)类别下公开,并提供PDF和HTML预览版本供研究人员免费阅读。该研究有望为提升人工智能系统在工具调用时的可控性和安全性提供新的理论思路。 #arXiv #论文 #计算机科学 #工具控制 #激活引导 #AI #新研究
超越排行榜:LLM智能体在工具使用、规划与推理中的失败综合研究发布
近日,一篇题为《Beyond the Leaderboard: A Synthesis of Tool-Use, Planning, and Reasoning Failures in Large Language Model Agents》的论文在arXiv平台提交。该论文由Wael Albayaydh等人撰写,旨在超越传统排行榜的单一评估维度,系统综合了大型语言模型(LLM)智能体在工具使用、规划制定以及推理过程三个核心环节中的失败模式。通过对现有研究的归纳与梳理,论文全面揭示了这些智能体在执行复杂任务时常见的瓶颈与不足,为理解其局限性提供了整体视角,并对未来提升LLM智能体的鲁棒性和实用性具有重要参考价值。 #LLM #智能体 #工具使用 #规划 #推理 #失败分析 #arXiv #AI #大模型
近日,一篇题为《Beyond the Leaderboard: A Synthesis of Tool-Use, Planning, and Reasoning Failures in Large Language Model Agents》的论文在arXiv平台提交。该论文由Wael Albayaydh等人撰写,旨在超越传统排行榜的单一评估维度,系统综合了大型语言模型(LLM)智能体在工具使用、规划制定以及推理过程三个核心环节中的失败模式。通过对现有研究的归纳与梳理,论文全面揭示了这些智能体在执行复杂任务时常见的瓶颈与不足,为理解其局限性提供了整体视角,并对未来提升LLM智能体的鲁棒性和实用性具有重要参考价值。 #LLM #智能体 #工具使用 #规划 #推理 #失败分析 #arXiv #AI #大模型