SkillOpt: Executive Strategy for Self-Evolving Agent Skills
https://arxiv.org/abs/2605.23904
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs .
https://arxiv.org/abs/2605.23904
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs .
arXiv.org
SkillOpt: Executive Strategy for Self-Evolving Agent Skills
Agent skills today are hand-crafted, generated one-shot, or evolved through loosely controlled self-revision, none of which behaves like a deep-learning optimizer for the skill, and none of which...
arXiv推出开放协作框架ArXivLabs
https://arxiv.org/abs/2605.23899
学术预印本平台arXiv近期推出了名为ArXivLabs的协作框架。该框架旨在为个人与机构研究者提供一个直接在arXiv网站上开发、测试并共享新功能的平台。ArXivLabs强调开放、社区协作、追求卓越以及用户数据隐私等核心价值观。平台表示,所有希望参与的合作伙伴都必须认同并遵守这些原则。此举旨在通过社区驱动的创新,持续为arXiv的全球学术用户群增添价值与功能。 #arXiv #学术平台 #开放科学 #协作创新 #数据隐私 #科研工具 #学术社区
https://arxiv.org/abs/2605.23899
学术预印本平台arXiv近期推出了名为ArXivLabs的协作框架。该框架旨在为个人与机构研究者提供一个直接在arXiv网站上开发、测试并共享新功能的平台。ArXivLabs强调开放、社区协作、追求卓越以及用户数据隐私等核心价值观。平台表示,所有希望参与的合作伙伴都必须认同并遵守这些原则。此举旨在通过社区驱动的创新,持续为arXiv的全球学术用户群增添价值与功能。 #arXiv #学术平台 #开放科学 #协作创新 #数据隐私 #科研工具 #学术社区
arXiv.org
From Raw Experience to Skill Consumption: A Systematic Study of...
Language agents increasingly improve by reusing \emph{skills} -- structured procedural artifacts distilled from past experience. In particular, \emph{domain-level} and \emph{model-generated}...
AI编程流程优化
一位开发者分享了他如何简化日益复杂的AI辅助编程工作流。他指出,当前许多“AI编程流程”过度依赖大语言模型执行本应由确定性代码完成的任务,这不仅浪费算力,还因LLM的不确定性导致需要人工反复校验,反而降低了效率。为此,他放弃了复杂的命令/代理/技能体系,转而采用极简的Pi Agent工具,并将重复性任务(如运行SonarQube检查、处理代码审查)封装成确定性的可调用扩展模块。这种“模块化确定性构建块”的方法,使得关键步骤能稳定、准确地自动执行,显著减少了token消耗,并让开发者能从繁琐的流程监督中解放出来。他强调,工具应适配个人工作流,盲目采用他人复杂的现成方案可能适得其反,有时回归简洁、可靠的脚本是更优选择。 #AI编程 #开发工具 #效率优化 #PiAgent #开源 #工作流 #软件工程 #技术实践
一位开发者分享了他如何简化日益复杂的AI辅助编程工作流。他指出,当前许多“AI编程流程”过度依赖大语言模型执行本应由确定性代码完成的任务,这不仅浪费算力,还因LLM的不确定性导致需要人工反复校验,反而降低了效率。为此,他放弃了复杂的命令/代理/技能体系,转而采用极简的Pi Agent工具,并将重复性任务(如运行SonarQube检查、处理代码审查)封装成确定性的可调用扩展模块。这种“模块化确定性构建块”的方法,使得关键步骤能稳定、准确地自动执行,显著减少了token消耗,并让开发者能从繁琐的流程监督中解放出来。他强调,工具应适配个人工作流,盲目采用他人复杂的现成方案可能适得其反,有时回归简洁、可靠的脚本是更优选择。 #AI编程 #开发工具 #效率优化 #PiAgent #开源 #工作流 #软件工程 #技术实践
arXiv 推出 Labs 框架,促进学术社区协作创新
https://arxiv.org/abs/2605.23898
知名学术预印本平台 arXiv 宣布推出 arXivLabs 框架,旨在鼓励学术社群直接在其平台上协作开发与分享新功能。该框架面向个人研究者及机构组织开放,秉持开放、社区、卓越与用户数据隐私的核心价值观。arXiv 表示将只与坚守这些价值观的伙伴合作。此举旨在通过降低技术门槛,激发社区成员的创造力,共同丰富平台功能,提升全球科研人员的使用体验与工作效率。对于能为学术社区增添价值的项目构想,arXiv 已开放进一步了解的通道。 #arXiv #科研工具 #开放科学 #学术社区 #技术协作 #学术出版 #数据隐私
https://arxiv.org/abs/2605.23898
知名学术预印本平台 arXiv 宣布推出 arXivLabs 框架,旨在鼓励学术社群直接在其平台上协作开发与分享新功能。该框架面向个人研究者及机构组织开放,秉持开放、社区、卓越与用户数据隐私的核心价值观。arXiv 表示将只与坚守这些价值观的伙伴合作。此举旨在通过降低技术门槛,激发社区成员的创造力,共同丰富平台功能,提升全球科研人员的使用体验与工作效率。对于能为学术社区增添价值的项目构想,arXiv 已开放进一步了解的通道。 #arXiv #科研工具 #开放科学 #学术社区 #技术协作 #学术出版 #数据隐私
arXiv.org
SPACENUM: Revisiting Spatial Numerical Understanding in VLMs
Vision-Language Models (VLMs) are increasingly deployed in embodied environments, where they need produce numerical outputs such as action magnitudes and spatial coordinates. Although these...
提出对抗子空间对齐方法,增强多模态知识编辑鲁棒性
https://arxiv.org/abs/2605.23780
一项最新研究提出了名为“对抗子空间对齐”的新方法,旨在实现超越传统二元编辑模式的、更为鲁棒的多模态知识编辑。多模态知识编辑旨在不重新训练整个模型的情况下,精准修正或更新大语言模型中的特定知识,而该方法通过引入对抗性子空间对齐技术,旨在提升编辑过程在复杂场景下的稳定性和有效性,解决了现有编辑方法可能面临的脆弱性问题。该研究为维护和更新日益庞大的多模态AI模型的知识库提供了新的技术思路。 #人工智能 #机器学习 #多模态 #知识编辑 #鲁棒学习 #对抗训练 #学术研究 #AI
https://arxiv.org/abs/2605.23780
一项最新研究提出了名为“对抗子空间对齐”的新方法,旨在实现超越传统二元编辑模式的、更为鲁棒的多模态知识编辑。多模态知识编辑旨在不重新训练整个模型的情况下,精准修正或更新大语言模型中的特定知识,而该方法通过引入对抗性子空间对齐技术,旨在提升编辑过程在复杂场景下的稳定性和有效性,解决了现有编辑方法可能面临的脆弱性问题。该研究为维护和更新日益庞大的多模态AI模型的知识库提供了新的技术思路。 #人工智能 #机器学习 #多模态 #知识编辑 #鲁棒学习 #对抗训练 #学术研究 #AI
arXiv.org
Beyond Binary Edits Robust Multimodal Knowledge Editing with...
Multimodal large language models (MLLMs) need efficient mechanisms to update knowledge without degrading existing capabilities. While intrinsic multimodal knowledge editing achieves strong...
arXiv推出协作框架,支持程序验证研究
https://arxiv.org/abs/2605.23772
学术预印本平台arXiv发布了名为“arXivLabs”的协作框架,旨在邀请研究社区成员共同开发与分享网站新功能。该平台鼓励个人及组织在坚持其开放、社区、卓越与用户数据隐私等核心价值观的前提下参与创新。arXiv表示,此举措旨在通过众包方式为社区增添价值,具体项目已开放申请。 #arXiv #程序验证 #开源 #学术 #计算机科学 #软件工程 #研究工具
https://arxiv.org/abs/2605.23772
学术预印本平台arXiv发布了名为“arXivLabs”的协作框架,旨在邀请研究社区成员共同开发与分享网站新功能。该平台鼓励个人及组织在坚持其开放、社区、卓越与用户数据隐私等核心价值观的前提下参与创新。arXiv表示,此举措旨在通过众包方式为社区增添价值,具体项目已开放申请。 #arXiv #程序验证 #开源 #学术 #计算机科学 #软件工程 #研究工具
arXiv.org
Agentic Proving for Program Verification
Agentic systems have recently emerged as state-of-the-art approaches for automated theorem proving in formal mathematics. To assess how far these capabilities extend to program verification, we...
MemAudit:基于因果归因和结构异常检测的代理记忆审计工具发布
https://arxiv.org/abs/2605.23723
arXivLabs是一个支持在arXiv网站上协作开发与分享新功能的框架,秉持开放、社区、卓越及用户数据隐私的价值观。近期,与该框架相关的MemAudit工具被提出,旨在对AI代理中可能被毒化的记忆进行事后审计。该工具结合因果归因和结构异常检测技术,通过分析记忆数据中的因果关系和结构偏差,识别潜在的恶意干扰或数据毒化,从而增强AI系统的安全性和可靠性。这一进展反映了随着AI代理学习系统日益复杂,对其记忆完整性和安全审计的需求不断增长,MemAudit为开发者提供了有效的检测和修正手段
https://arxiv.org/abs/2605.23723
arXivLabs是一个支持在arXiv网站上协作开发与分享新功能的框架,秉持开放、社区、卓越及用户数据隐私的价值观。近期,与该框架相关的MemAudit工具被提出,旨在对AI代理中可能被毒化的记忆进行事后审计。该工具结合因果归因和结构异常检测技术,通过分析记忆数据中的因果关系和结构偏差,识别潜在的恶意干扰或数据毒化,从而增强AI系统的安全性和可靠性。这一进展反映了随着AI代理学习系统日益复杂,对其记忆完整性和安全审计的需求不断增长,MemAudit为开发者提供了有效的检测和修正手段
arXiv.org
MemAudit: Post-hoc Auditing of Poisoned Agent Memory via Causal...
Large language model agents increasingly rely on persistent memory to store past interactions, retrieve relevant demonstrations, and improve long-horizon task execution. However, this memory...
研究提出共享强化学习策略打造无限可追踪人格游戏NPC
https://arxiv.org/abs/2605.23652
近期,一项新研究提出了一种基于共享强化学习策略的方法,名为“Persona-Traceable Shared RL Policies”,旨在创建可扩展的游戏代理。该方法允许开发者使用单一策略控制无限数量的非玩家角色(NPC),同时确保每个NPC的人格特征可追踪,从而解决传统游戏AI中NPC行为多样性不足和开发成本高昂的挑战。研究论文通过arXiv平台及其arXivLabs协作框架分享,强调了开放、社区导向和数据隐私的科研价值观。该方法的潜在影响包括提升游戏开发效率、丰富游戏体验,并推动人工智能在互动娱乐领域的创新应用。 #游戏AI #强化学习 #NPC #人工智能 #arXiv #科技新闻 #共享策略 #研究突破
https://arxiv.org/abs/2605.23652
近期,一项新研究提出了一种基于共享强化学习策略的方法,名为“Persona-Traceable Shared RL Policies”,旨在创建可扩展的游戏代理。该方法允许开发者使用单一策略控制无限数量的非玩家角色(NPC),同时确保每个NPC的人格特征可追踪,从而解决传统游戏AI中NPC行为多样性不足和开发成本高昂的挑战。研究论文通过arXiv平台及其arXivLabs协作框架分享,强调了开放、社区导向和数据隐私的科研价值观。该方法的潜在影响包括提升游戏开发效率、丰富游戏体验,并推动人工智能在互动娱乐领域的创新应用。 #游戏AI #强化学习 #NPC #人工智能 #arXiv #科技新闻 #共享策略 #研究突破
arXiv.org
One Policy, Infinite NPCs: Persona-Traceable Shared RL Policies...
On a 300-persona life-simulation benchmark, pcsp achieves compositional zero-shot persona identification up to 17x above chance, Spearman rho approx 0.73 semantic-behavioral alignment, and 22x...
Solving the Aircraft Disassembly Scheduling Problem
https://arxiv.org/abs/2605.23592
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs .
https://arxiv.org/abs/2605.23592
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs .
arXiv.org
Solving the Aircraft Disassembly Scheduling Problem
Dismantling aircrafts reaching their end of life is a complex endeavour that is necessary in terms of sustainability but yields small income margins for air transport companies. An efficient...
Co-ReAct:为ReAct智能体引入步骤级协作者评分标准
https://arxiv.org/abs/2605.23590
一种名为Co-ReAct的新方法被提出,旨在提升ReAct类型智能体的性能。该方法创新性地将结构化评分标准作为步骤级协作者整合到智能体的推理过程中,从而引导其进行更规范、更可靠的逐步推理。此项目目前依托于arXivLabs平台进行开发和展示。 #CoReAct #AI智能体 #机器学习 #人工智能 #学术研究 #arXiv #技术前沿
https://arxiv.org/abs/2605.23590
一种名为Co-ReAct的新方法被提出,旨在提升ReAct类型智能体的性能。该方法创新性地将结构化评分标准作为步骤级协作者整合到智能体的推理过程中,从而引导其进行更规范、更可靠的逐步推理。此项目目前依托于arXivLabs平台进行开发和展示。 #CoReAct #AI智能体 #机器学习 #人工智能 #学术研究 #arXiv #技术前沿
arXiv.org
Co-ReAct: Rubrics as Step-Level Collaborators for ReAct Agents
ReAct-style agents for search-intensive, multi-step reasoning tasks rely largely on their own internal judgment to decide what evidence to seek, which reasoning or action step to take next, and...
研究创新结合约束规划与动态规划求解部分车间调度问题
https://arxiv.org/abs/2605.23569
一项新的学术研究探讨了在部分车间调度问题中整合约束规划与动态规划两种方法的可能性。该研究针对一个具体的生产调度案例,提出了一种混合求解策略,旨在充分利用约束规划在建模灵活性和动态规划在特定结构问题上的效率优势,以期提升复杂调度场景下的求解性能。这种方法论上的融合尝试,为优化算法领域提供了新的思路,对理论研究和实际工业应用中的排产优化均具有潜在的参考价值。 #运筹学 #算法 #调度优化 #学术研究 #AI应用
https://arxiv.org/abs/2605.23569
一项新的学术研究探讨了在部分车间调度问题中整合约束规划与动态规划两种方法的可能性。该研究针对一个具体的生产调度案例,提出了一种混合求解策略,旨在充分利用约束规划在建模灵活性和动态规划在特定结构问题上的效率优势,以期提升复杂调度场景下的求解性能。这种方法论上的融合尝试,为优化算法领域提供了新的思路,对理论研究和实际工业应用中的排产优化均具有潜在的参考价值。 #运筹学 #算法 #调度优化 #学术研究 #AI应用
arXiv.org
CP or DP? Why Not Both: A Case Study in the Partial Shop Scheduling Problem
Dynamic Programming (DP) and Constraint Programming (CP) are well-established paradigms for solving combinatorial optimization problems. Usually, these two approaches are used separately. This...
EDGE-OPD: Internalizing Privileged Context with Evidence Guided On-Policy Distillation
https://arxiv.org/abs/2605.23493
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs .
https://arxiv.org/abs/2605.23493
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs .
arXiv.org
EDGE-OPD: Internalizing Privileged Context with Evidence Guided...
On-Policy Distillation (OPD) has gained wide attraction as an LLM post-training paradigm due to its effectiveness in improving capabilities without introducing model distribution drift, and...
arXiv 推出 arXivLabs 框架,促进社区合作与创新
https://arxiv.org/abs/2605.23414
arXiv 最近发布了 arXivLabs 框架,这是一个全新平台,旨在允许研究人员和开发者直接在 arXiv 网站上开发并共享新特性。该框架强调开放性、社区参与、卓越标准和用户数据隐私等核心价值观,只有认同并遵守这些原则的合作伙伴才能参与项目开发。通过这一举措,arXiv 希望鼓励社区成员提出创新想法,以增强平台功能,推动学术交流和科技进展。arXiv 表示,这不仅是为了提升用户体验,还致力于构建更协作的科研生态系统,预计将吸引更多机构和个人加入,共同探索未来科研工具的发展方向。 #arXiv #科研 #开放科学 #社区合作 #科技新闻 #AI #框架开发 #创新平台
https://arxiv.org/abs/2605.23414
arXiv 最近发布了 arXivLabs 框架,这是一个全新平台,旨在允许研究人员和开发者直接在 arXiv 网站上开发并共享新特性。该框架强调开放性、社区参与、卓越标准和用户数据隐私等核心价值观,只有认同并遵守这些原则的合作伙伴才能参与项目开发。通过这一举措,arXiv 希望鼓励社区成员提出创新想法,以增强平台功能,推动学术交流和科技进展。arXiv 表示,这不仅是为了提升用户体验,还致力于构建更协作的科研生态系统,预计将吸引更多机构和个人加入,共同探索未来科研工具的发展方向。 #arXiv #科研 #开放科学 #社区合作 #科技新闻 #AI #框架开发 #创新平台
arXiv.org
When Planning Fails Despite Correct Execution: On Epistemic...
LLM-based multi-agent systems can fail even when planned actions are executed correctly because agents may misjudge their knowledge when evaluating plan feasibility, a phenomenon we term epistemic...
新研究提出人机协同多智能体呼吸机决策支持系统
https://arxiv.org/abs/2605.23320
一篇发布于学术平台的论文提出了一种名为“Human-in-the-Loop Multi-Agent Ventilator Decision Support with Contextual Bandit Preference Learning”的新方法。该研究旨在结合人类专家的判断与多智能体系统的协同分析,通过一种名为“情境强盗偏好学习”的技术,为临床呼吸机参数设置提供智能决策支持。该工作的目标是提升危重症治疗场景下关键决策的准确性与效率。 该论文通过arXiv平台发布。arXivLabs是一个开放框架,允许研究者直接在arXiv网站上协作开发和分享新功能,其核心价值包括开放、协作、卓越和保护用户数据隐私。 #arXiv #医学研究 #呼吸机 #人工智能 #决策支持系统 #医疗AI #临床智能
https://arxiv.org/abs/2605.23320
一篇发布于学术平台的论文提出了一种名为“Human-in-the-Loop Multi-Agent Ventilator Decision Support with Contextual Bandit Preference Learning”的新方法。该研究旨在结合人类专家的判断与多智能体系统的协同分析,通过一种名为“情境强盗偏好学习”的技术,为临床呼吸机参数设置提供智能决策支持。该工作的目标是提升危重症治疗场景下关键决策的准确性与效率。 该论文通过arXiv平台发布。arXivLabs是一个开放框架,允许研究者直接在arXiv网站上协作开发和分享新功能,其核心价值包括开放、协作、卓越和保护用户数据隐私。 #arXiv #医学研究 #呼吸机 #人工智能 #决策支持系统 #医疗AI #临床智能
arXiv.org
Human-in-the-Loop Multi-Agent Ventilator Decision Support with...
Ventilator decision support requires sequential decisions that track evolving physiology and disease trajectories while respecting safety boundaries and clinician specific tuning styles. Rule...
arXivLabs 推出 DART 语义恢复技术,助力结构化工具代理优化
https://arxiv.org/abs/2605.23311
arXivLabs 作为 arXiv 平台的协作框架,长期秉持开放性、社区卓越和用户数据隐私的核心价值观,支持协作者直接开发和共享新功能。近期,该框架引入了 DART 技术,专注于结构化工具代理的语义可恢复性。DART 旨在提升代理工具在复杂场景下的语义理解和恢复能力,从而增强用户体验和社区协作效率。通过这一技术,arXiv 进一步推动学术资源共享和工具创新,为全球研究者提供更可靠的支撑。此举不仅强化了 arXiv 对开放科学的承诺,还预计吸引更多合作伙伴加入,共同促进技术生态的繁荣发展。 #arXiv #DART #语义可恢复性 #工具代理 #AI #学术技术 #开放科学 #社区协作
https://arxiv.org/abs/2605.23311
arXivLabs 作为 arXiv 平台的协作框架,长期秉持开放性、社区卓越和用户数据隐私的核心价值观,支持协作者直接开发和共享新功能。近期,该框架引入了 DART 技术,专注于结构化工具代理的语义可恢复性。DART 旨在提升代理工具在复杂场景下的语义理解和恢复能力,从而增强用户体验和社区协作效率。通过这一技术,arXiv 进一步推动学术资源共享和工具创新,为全球研究者提供更可靠的支撑。此举不仅强化了 arXiv 对开放科学的承诺,还预计吸引更多合作伙伴加入,共同促进技术生态的繁荣发展。 #arXiv #DART #语义可恢复性 #工具代理 #AI #学术技术 #开放科学 #社区协作
arXiv.org
DART: Semantic Recoverability for Structured Tool Agents
When a structured tool agent fails mid-execution, the runtime faces a dilemma: replaying the entire task is safe but wasteful, while restoring from a local checkpoint is efficient but can leave...
Ontological Knowledge Blocks: Executable Compliance and Profile-Based Validation for Trustworthy AI Systems
https://arxiv.org/abs/2605.23297
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs .
https://arxiv.org/abs/2605.23297
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs .
arXiv.org
Ontological Knowledge Blocks: Executable Compliance and...
AI-enabled services deployed in critical digital infrastructure are subject to governance obligations spanning transparency, accountability, fairness, and traceability. Compliance today remains...
arXivLabs 框架发布,推动开放协作与创新
https://arxiv.org/abs/2605.23296
近日,学术预印本平台 arXiv 正式推出 arXivLabs 框架,旨在通过开放协作推动平台功能创新。该框架设计为一个协作环境,允许个人和组织直接在 arXiv 网站上参与新特性的创建与分享。arXiv 强调,所有参与者需认同其核心价值观,包括开放性、社区精神、卓越追求以及用户数据隐私保护。平台承诺只与遵守这些原则的伙伴合作。通过 arXivLabs,arXiv 希望激发社区创新,鼓励研究人员提出有价值的项目想法,从而增强平台服务,推动开放科学的发展。这一举措预计将提升学术交流的效率,并扩大 arXiv 在全球研究社区中的影响力,进一步促进学术资源的共享与协作。 #arXiv #arXivLabs #开放科学 #学术协作 #科技创新 #数据隐私 #社区创新 #开源框架
https://arxiv.org/abs/2605.23296
近日,学术预印本平台 arXiv 正式推出 arXivLabs 框架,旨在通过开放协作推动平台功能创新。该框架设计为一个协作环境,允许个人和组织直接在 arXiv 网站上参与新特性的创建与分享。arXiv 强调,所有参与者需认同其核心价值观,包括开放性、社区精神、卓越追求以及用户数据隐私保护。平台承诺只与遵守这些原则的伙伴合作。通过 arXivLabs,arXiv 希望激发社区创新,鼓励研究人员提出有价值的项目想法,从而增强平台服务,推动开放科学的发展。这一举措预计将提升学术交流的效率,并扩大 arXiv 在全球研究社区中的影响力,进一步促进学术资源的共享与协作。 #arXiv #arXivLabs #开放科学 #学术协作 #科技创新 #数据隐私 #社区创新 #开源框架
arXiv.org
Parallel Context Compaction for Long-Horizon LLM Agent Serving
Long-horizon LLM agents accumulate growing conversation histories that eventually exceed the model's context window. Context compaction via LLM-based summarization keeps the conversation bounded,...
arXivLabs 推广开放框架,邀全球学者共建学术平台
https://arxiv.org/abs/2605.23262
arXiv 近期积极推广其创新的 arXivLabs 框架,这是一个专为学术社区设计的协作平台,允许研究人员和开发者直接在 arXiv 网站上开发和分享新功能。该框架严格遵循开放、社区、卓越和用户数据隐私的核心价值观,要求所有合作的个人或组织必须遵守这些原则。通过 arXivLabs,arXiv 旨在激发全球学者的创造力,鼓励他们提出有价值的项目以增强平台功能,从而促进学术资源共享、推动开放科学的发展,并加强研究社区的协作与创新。 #arXiv #arXivLabs #开放科学 #学术社区 #数据隐私 #研究 #科技新闻 #知识工作
https://arxiv.org/abs/2605.23262
arXiv 近期积极推广其创新的 arXivLabs 框架,这是一个专为学术社区设计的协作平台,允许研究人员和开发者直接在 arXiv 网站上开发和分享新功能。该框架严格遵循开放、社区、卓越和用户数据隐私的核心价值观,要求所有合作的个人或组织必须遵守这些原则。通过 arXivLabs,arXiv 旨在激发全球学者的创造力,鼓励他们提出有价值的项目以增强平台功能,从而促进学术资源共享、推动开放科学的发展,并加强研究社区的协作与创新。 #arXiv #arXivLabs #开放科学 #学术社区 #数据隐私 #研究 #科技新闻 #知识工作
arXiv.org
Design and Report Benchmarks for Knowledge Work
The development of LLM agents has led to a growing body of work on knowledge-work AI, including coding, research, and healthcare. However, current knowledge-work evaluation and benchmark design...
GENSTRAT
https://arxiv.org/abs/2605.23238
一项名为GENSTRAT的前沿研究提出,旨在为大型语言模型的战略推理能力建立系统性的科学框架。该研究发表于预印本平台arXiv,其依托的arXivLabs是一个鼓励开放协作的框架,致力于推动社区创新,同时恪守开放、卓越与用户数据隐私等核心价值。该研究的目标是超越当前模型在复杂决策与博弈场景中的表现,深入探索并科学化其战略推理的内在机制,这一方向对于提升AI的自主决策与协作能力具有关键意义。 #人工智能 #大语言模型 #战略推理 #学术研究 #arXiv #GENSTRAT #AI科学
https://arxiv.org/abs/2605.23238
一项名为GENSTRAT的前沿研究提出,旨在为大型语言模型的战略推理能力建立系统性的科学框架。该研究发表于预印本平台arXiv,其依托的arXivLabs是一个鼓励开放协作的框架,致力于推动社区创新,同时恪守开放、卓越与用户数据隐私等核心价值。该研究的目标是超越当前模型在复杂决策与博弈场景中的表现,深入探索并科学化其战略推理的内在机制,这一方向对于提升AI的自主决策与协作能力具有关键意义。 #人工智能 #大语言模型 #战略推理 #学术研究 #arXiv #GENSTRAT #AI科学
arXiv.org
GENSTRAT: Toward a Science of Strategic Reasoning in Large Language Models
Large language models (LLMs) are increasingly deployed as economic agents in marketplaces, auctions, and bidding settings. Anticipating their behavior in any specific deployment is hard. Existing...
arXiv 推出 Labs 协作框架,助力学术社区功能共建
https://arxiv.org/abs/2605.23218
专注于预印本服务的学术平台 arXiv 推出了名为“arXivLabs”的协作框架。该框架旨在搭建一个开放的合作层,允许全球的研究人员、开发者以及相关机构直接在其网站上协作开发与分享新功能,从而更灵活地响应学术社区的需求。arXiv 强调,所有参与 Labs 的合作方都需认同其开放、社区、卓越与用户数据隐私的核心价值观,并鼓励有想法的社区成员通过该项目为平台增添价值。 #学术 #开放科学 #科研平台 #技术协作 #用户参与
https://arxiv.org/abs/2605.23218
专注于预印本服务的学术平台 arXiv 推出了名为“arXivLabs”的协作框架。该框架旨在搭建一个开放的合作层,允许全球的研究人员、开发者以及相关机构直接在其网站上协作开发与分享新功能,从而更灵活地响应学术社区的需求。arXiv 强调,所有参与 Labs 的合作方都需认同其开放、社区、卓越与用户数据隐私的核心价值观,并鼓励有想法的社区成员通过该项目为平台增添价值。 #学术 #开放科学 #科研平台 #技术协作 #用户参与
arXiv.org
Foundation Protocol: A Coordination Layer for Agentic Society
Autonomous agents are moving from tools into a layer of social infrastructure: they browse, purchase, deploy software, manage systems, and increasingly interact with one another. As these systems...