Measuring Intelligence Beyond Human Scale
Focus to learn more arXiv-issued DOI via DataCite (pending registration) Submission history From: Elad Hazan [ view email ] [v1] Wed, 8 Jul 2026 06:19:10 UTC (40 KB) Full-text links: Access Paper: View a PDF of the paper titled Measuring Intelligence Beyond Human Scale, by Jerry Han and 7 other authors View PDF TeX Source view license Current browse context: < prev | next > new | recent | 2026-07 Change to browse by: cs References & Citations NASA ADS Google Scholar Semantic Scholar export BibTeX citation Loading... BibTeX formatted citation × loading... Data provided by: Bookmark Bibliographic Tools Bibliographic and Citation Tools Bibliographic Explorer Toggle Bibliographic Explorer ( What is the Explorer? ) Connected Papers Toggle Connected Papers ( What is Connected Papers? ) Litmaps Toggle Litmaps ( What is Litmaps? ) Toggl
Focus to learn more arXiv-issued DOI via DataCite (pending registration) Submission history From: Elad Hazan [ view email ] [v1] Wed, 8 Jul 2026 06:19:10 UTC (40 KB) Full-text links: Access Paper: View a PDF of the paper titled Measuring Intelligence Beyond Human Scale, by Jerry Han and 7 other authors View PDF TeX Source view license Current browse context: < prev | next > new | recent | 2026-07 Change to browse by: cs References & Citations NASA ADS Google Scholar Semantic Scholar export BibTeX citation Loading... BibTeX formatted citation × loading... Data provided by: Bookmark Bibliographic Tools Bibliographic and Citation Tools Bibliographic Explorer Toggle Bibliographic Explorer ( What is the Explorer? ) Connected Papers Toggle Connected Papers ( What is Connected Papers? ) Litmaps Toggle Litmaps ( What is Litmaps? ) Toggl
e scite Smart Citations ( What are Smart Citations? ) Code, Data, Media Code, Data and Media Associated with this Article alphaXiv Toggle alphaXiv ( What is alphaXiv? ) Links to Code Toggle CatalyzeX Code Finder for Papers ( What is CatalyzeX? ) DagsHub Toggle DagsHub ( What is DagsHub? ) GotitPub Toggle ( What is GotitPub? ) Huggingface Toggle Hugging Face ( What is Huggingface? ) ScienceCast Toggle ScienceCast ( What is ScienceCast? ) Demos Demos Replicate Toggle Replicate ( What is Replicate? ) Spaces Toggle Hugging Face Spaces ( What is Spaces? ) Spaces Toggle ( What is ? ) Related Papers Recommenders and Search Tools Link to Influence Flower Influence Flower ( What are Influence Flowers? ) Core recommender toggle CORE Recommender ( What is CORE? ) Author Venue Institution Topic About arXivLabs arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs .
学习社会规范可提升动态人机协调的兼容性
近日,一篇题为《Learning social norms enhances compatibility in dynamic human-AI coordination》的论文在arXiv预印本平台发布。该论文由Yi Yang等学者撰写,指出在动态人机协作环境中,AI系统通过学习并遵循社会规范,能够显著提升与人类合作的兼容性和效率。研究认为,社会规范的学习有助于智能体更好地理解人类意图、预测行为变化,从而在复杂交互中做出更协调的决策。这一成果对自动驾驶、智能机器人以及多智能体系统等领域具有潜在应用价值。论文目前已在arXiv上公开,供相关领域研究者参考。 #人机协调 #社会规范 #AI #论文 #arXiv #动态协作 #多智能体
近日,一篇题为《Learning social norms enhances compatibility in dynamic human-AI coordination》的论文在arXiv预印本平台发布。该论文由Yi Yang等学者撰写,指出在动态人机协作环境中,AI系统通过学习并遵循社会规范,能够显著提升与人类合作的兼容性和效率。研究认为,社会规范的学习有助于智能体更好地理解人类意图、预测行为变化,从而在复杂交互中做出更协调的决策。这一成果对自动驾驶、智能机器人以及多智能体系统等领域具有潜在应用价值。论文目前已在arXiv上公开,供相关领域研究者参考。 #人机协调 #社会规范 #AI #论文 #arXiv #动态协作 #多智能体
Large Behavior Model: A Promptable Digital Twin of the Retail Customer
Focus to learn more arXiv-issued DOI via DataCite (pending registration) Submission history From: Krittin Pachtrachai PhD [ view email ] [v1] Wed, 8 Jul 2026 04:31:18 UTC (242 KB) Full-text links: Access Paper: View a PDF of the paper titled Large Behavior Model: A Promptable Digital Twin of the Retail Customer, by Wachiravit Modecrua and 2 other authors View PDF HTML (experimental) TeX Source view license Current browse context: < prev | next > new | recent | 2026-07 Change to browse by: cs References & Citations NASA ADS Google Scholar Semantic Scholar export BibTeX citation Loading... BibTeX formatted citation × loading... Data provided by: Bookmark Bibliographic Tools Bibliographic and Citation Tools Bibliographic Explorer Toggle Bibliographic Explorer ( What is the Explorer? ) Connected Papers T
Focus to learn more arXiv-issued DOI via DataCite (pending registration) Submission history From: Krittin Pachtrachai PhD [ view email ] [v1] Wed, 8 Jul 2026 04:31:18 UTC (242 KB) Full-text links: Access Paper: View a PDF of the paper titled Large Behavior Model: A Promptable Digital Twin of the Retail Customer, by Wachiravit Modecrua and 2 other authors View PDF HTML (experimental) TeX Source view license Current browse context: < prev | next > new | recent | 2026-07 Change to browse by: cs References & Citations NASA ADS Google Scholar Semantic Scholar export BibTeX citation Loading... BibTeX formatted citation × loading... Data provided by: Bookmark Bibliographic Tools Bibliographic and Citation Tools Bibliographic Explorer Toggle Bibliographic Explorer ( What is the Explorer? ) Connected Papers T
oggle Connected Papers ( What is Connected Papers? ) Litmaps Toggle Litmaps ( What is Litmaps? ) Toggle scite Smart Citations ( What are Smart Citations? ) Code, Data, Media Code, Data and Media Associated with this Article alphaXiv Toggle alphaXiv ( What is alphaXiv? ) Links to Code Toggle CatalyzeX Code Finder for Papers ( What is CatalyzeX? ) DagsHub Toggle DagsHub ( What is DagsHub? ) GotitPub Toggle ( What is GotitPub? ) Huggingface Toggle Hugging Face ( What is Huggingface? ) ScienceCast Toggle ScienceCast ( What is ScienceCast? ) Demos Demos Replicate Toggle Replicate ( What is Replicate? ) Spaces Toggle Hugging Face Spaces ( What is Spaces? ) Spaces Toggle ( What is ? ) Related Papers Recommenders and Search Tools Link to Influence Flower Influence Flower ( What are Influence Flowers? ) Core recommender toggle CORE Recommender ( What is CORE? ) Author Venue Institution Topic About arXivLabs arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs .
紧凑世界模型中的空间关系基础化
来自Yufeng Wang等作者的研究论文《Grounding Spatial Relations in a Compact World Model: Instruction Leakage and a Goal-Free Dynamics Fix》于2026年7月8日提交至arXiv预印本平台。该工作深入探讨了在紧凑世界模型中如何实现空间关系的有效基础化,并重点分析了指令泄漏问题——即模型在训练过程中意外依赖与任务无关的指令信息。为克服这一缺陷,作者提出了一种无目标动力学修复方法,旨在让模型在不依赖明确目标信号的情况下,仅通过动力学特征习得稳固的空间表征。研究属于计算机科学领域,相关工作有望提升AI系统在复杂环境中的空间推理与导航能力,为构建更鲁棒的世界模型提供了新思路。 #空间关系 #紧凑世界模型 #指令泄漏 #无目标动力学 #人工智能 #计算机科学 #arXiv #研究论文
来自Yufeng Wang等作者的研究论文《Grounding Spatial Relations in a Compact World Model: Instruction Leakage and a Goal-Free Dynamics Fix》于2026年7月8日提交至arXiv预印本平台。该工作深入探讨了在紧凑世界模型中如何实现空间关系的有效基础化,并重点分析了指令泄漏问题——即模型在训练过程中意外依赖与任务无关的指令信息。为克服这一缺陷,作者提出了一种无目标动力学修复方法,旨在让模型在不依赖明确目标信号的情况下,仅通过动力学特征习得稳固的空间表征。研究属于计算机科学领域,相关工作有望提升AI系统在复杂环境中的空间推理与导航能力,为构建更鲁棒的世界模型提供了新思路。 #空间关系 #紧凑世界模型 #指令泄漏 #无目标动力学 #人工智能 #计算机科学 #arXiv #研究论文
新论文揭示"驾驭效应":编排设计决定企业代理AI的Token经济学
一篇名为《The Harness Effect: How Orchestration Design Sets the Token Economics of Enterprise Agentic AI》的学术论文于2026年7月8日提交至预印本平台arXiv。该论文由Muayad Sayed Ali等32位研究者共同撰写,深入探讨了在企业级多智能体系统中,编排设计(即任务分配与流程协调机制)如何从根本上影响Token的消耗与经济效益。论文提出了“Harness Effect”(驾驭效应)概念,指出不同的编排策略会显著改变AI代理的Token使用模式,进而影响企业AI部署的成本与效率。这一研究成果为优化企业级AI代理的经济模型提供了关键理论指导,有望推动更高效、更经济的AI系统设计。 #论文 #AI代理 #Token经济学 #编排设计 #企业AI #arXiv #多智能体
一篇名为《The Harness Effect: How Orchestration Design Sets the Token Economics of Enterprise Agentic AI》的学术论文于2026年7月8日提交至预印本平台arXiv。该论文由Muayad Sayed Ali等32位研究者共同撰写,深入探讨了在企业级多智能体系统中,编排设计(即任务分配与流程协调机制)如何从根本上影响Token的消耗与经济效益。论文提出了“Harness Effect”(驾驭效应)概念,指出不同的编排策略会显著改变AI代理的Token使用模式,进而影响企业AI部署的成本与效率。这一研究成果为优化企业级AI代理的经济模型提供了关键理论指导,有望推动更高效、更经济的AI系统设计。 #论文 #AI代理 #Token经济学 #编排设计 #企业AI #arXiv #多智能体
SageMath 增强 LLM 智能体
一项预印本研究在 arXiv 公开,标题为《Evaluating SageMath-Augmented LLM Agents for Computational and Experimental Mathematics》。该工作由 Pavel Snopov 等人完成,旨在系统评估将开源数学软件 SageMath 与大语言模型(LLM)相结合所形成的智能体,在计算与实验数学任务上的表现。研究通过设计基准测试,分析了此类增强智能体在符号计算、代数推导与数学实验等场景中的能力与局限性,探讨了利用专业数学工具补齐 LLM 在精确计算与形式化验证方面短板的可能性,为未来 AI 辅助数学研究提供了新的思路与参考。 #arXiv #SageMath #LLM #大模型 #计算数学 #实验数学 #AI #机器学习 #预印本
一项预印本研究在 arXiv 公开,标题为《Evaluating SageMath-Augmented LLM Agents for Computational and Experimental Mathematics》。该工作由 Pavel Snopov 等人完成,旨在系统评估将开源数学软件 SageMath 与大语言模型(LLM)相结合所形成的智能体,在计算与实验数学任务上的表现。研究通过设计基准测试,分析了此类增强智能体在符号计算、代数推导与数学实验等场景中的能力与局限性,探讨了利用专业数学工具补齐 LLM 在精确计算与形式化验证方面短板的可能性,为未来 AI 辅助数学研究提供了新的思路与参考。 #arXiv #SageMath #LLM #大模型 #计算数学 #实验数学 #AI #机器学习 #预印本
成本效益代理工具助力ARC-AGI
近日,Kabir Moghe等人在arXiv预印本平台提交了一篇题为《Cost-Effective Agent Harnesses for Abstract Reasoning and Generalization on ARC-AGI-1》的论文。该研究聚焦ARC-AGI-1这一衡量AI抽象推理能力的关键基准,提出了一种高性价比的代理工具,旨在以较低计算成本提升模型的推理与泛化性能。通过优化代理架构与策略,研究团队在资源受限条件下实现了更高效的抽象问题解决,为人工智能向通用智能发展提供了新思路,并展示了降低大模型依赖度的可能性。 #人工智能 #抽象推理 #泛化 #ARCAGI #成本效益 #代理工具 #预印本 #AI研究
近日,Kabir Moghe等人在arXiv预印本平台提交了一篇题为《Cost-Effective Agent Harnesses for Abstract Reasoning and Generalization on ARC-AGI-1》的论文。该研究聚焦ARC-AGI-1这一衡量AI抽象推理能力的关键基准,提出了一种高性价比的代理工具,旨在以较低计算成本提升模型的推理与泛化性能。通过优化代理架构与策略,研究团队在资源受限条件下实现了更高效的抽象问题解决,为人工智能向通用智能发展提供了新思路,并展示了降低大模型依赖度的可能性。 #人工智能 #抽象推理 #泛化 #ARCAGI #成本效益 #代理工具 #预印本 #AI研究
QANTIS: 硬件校准的顺序POMDP信念更新方法在IBM Heron上实现
据arXiv预印本显示,Bayram Yuksel Eker等研究人员提交了题为“QANTIS: Hardware-Calibrated Sequential POMDP Belief Updates on IBM Heron”的论文。该研究提出一种硬件校准的序列部分可观察马尔可夫决策过程(POMDP)信念更新框架,旨在提升量子系统的状态估计与决策能力。论文基于IBM Heron量子处理器平台进行实验验证,展示了该方法在噪声环境下对量子态的实时校准与更新能力。这一工作将经典控制理论与量子硬件特性相结合,为量子计算中的自适应控制与误差缓解提供了新思路,有望推动量子处理器在实际应用中的可靠性提升。 #量子计算 #POMDP #IBMHeron #arXiv #论文 #计算机科学 #硬件校准 #人工智能
据arXiv预印本显示,Bayram Yuksel Eker等研究人员提交了题为“QANTIS: Hardware-Calibrated Sequential POMDP Belief Updates on IBM Heron”的论文。该研究提出一种硬件校准的序列部分可观察马尔可夫决策过程(POMDP)信念更新框架,旨在提升量子系统的状态估计与决策能力。论文基于IBM Heron量子处理器平台进行实验验证,展示了该方法在噪声环境下对量子态的实时校准与更新能力。这一工作将经典控制理论与量子硬件特性相结合,为量子计算中的自适应控制与误差缓解提供了新思路,有望推动量子处理器在实际应用中的可靠性提升。 #量子计算 #POMDP #IBMHeron #arXiv #论文 #计算机科学 #硬件校准 #人工智能
LLM-powered reasoning in agent
Important: e-prints posted on arXiv are not peer-reviewed by arXiv; they should not be relied upon without context to guide clinical practice or health-related behavior and should not be reported in news media as established information without consulting multiple experts in the field. Focus to learn more arXiv-issued DOI via DataCite (pending registration) Submission history From: Sifat Moon [ view email ] [v1] Tue, 7 Jul 2026 19:39:01 UTC (1,373 KB) Full-text links: Access Paper: View a PDF of the paper titled LLM-powered reasoning in agent-based modeling, by Sifat Afroj Moon and 7 other authors View PDF HTML (experimental) TeX Source view license Current browse context: < prev | next > new | recent | 2026-07 Change to browse by: cs References & Citations NASA ADS Google Scholar Semantic Scholar export BibTeX citation Loading... BibTeX f
Important: e-prints posted on arXiv are not peer-reviewed by arXiv; they should not be relied upon without context to guide clinical practice or health-related behavior and should not be reported in news media as established information without consulting multiple experts in the field. Focus to learn more arXiv-issued DOI via DataCite (pending registration) Submission history From: Sifat Moon [ view email ] [v1] Tue, 7 Jul 2026 19:39:01 UTC (1,373 KB) Full-text links: Access Paper: View a PDF of the paper titled LLM-powered reasoning in agent-based modeling, by Sifat Afroj Moon and 7 other authors View PDF HTML (experimental) TeX Source view license Current browse context: < prev | next > new | recent | 2026-07 Change to browse by: cs References & Citations NASA ADS Google Scholar Semantic Scholar export BibTeX citation Loading... BibTeX f
ormatted citation × loading... Data provided by: Bookmark Bibliographic Tools Bibliographic and Citation Tools Bibliographic Explorer Toggle Bibliographic Explorer ( What is the Explorer? ) Connected Papers Toggle Connected Papers ( What is Connected Papers? ) Litmaps Toggle Litmaps ( What is Litmaps? ) Toggle scite Smart Citations ( What are Smart Citations? ) Code, Data, Media Code, Data and Media Associated with this Article alphaXiv Toggle alphaXiv ( What is alphaXiv? ) Links to Code Toggle CatalyzeX Code Finder for Papers ( What is CatalyzeX? ) DagsHub Toggle DagsHub ( What is DagsHub? ) GotitPub Toggle ( What is GotitPub? ) Huggingface Toggle Hugging Face ( What is Huggingface? ) ScienceCast Toggle ScienceCast ( What is ScienceCast? ) Demos Demos Replicate Toggle Replicate ( What is Replicate? ) Spaces Toggle Hugging Face Spaces ( What is Spaces? ) Spaces Toggle ( What is ? ) Related Papers Recommenders and Search Tools Link to Influence Flower Influence Flower ( What are Influence Flowers? ) Core recommender toggle CORE Recommender ( What is CORE? ) Author Venue Institution Topic About arXivLabs arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs .
When Does In-Context Search Help? A Sampling-Complexity Theory of Reflection
Focus to learn more arXiv-issued DOI via DataCite (pending registration) Submission history From: Yotam Wolf [ view email ] [v1] Tue, 7 Jul 2026 18:36:04 UTC (746 KB) Full-text links: Access Paper: View a PDF of the paper titled When Does In-Context Search Help? A Sampling-Complexity Theory of Reflection-Driven Reasoning, by Yotam Wolf and 2 other authors View PDF HTML (experimental) TeX Source view license Current browse context: < prev | next > new | recent | 2026-07 Change to browse by: cs References & Citations NASA ADS Google Scholar Semantic Scholar export BibTeX citation Loading... BibTeX formatted citation × loading... Data provided by: Bookmark Bibliographic Tools Bibliographic and Citation Tools Bibliographic Explorer Toggle Bibliographic Explorer ( What is the Explorer? ) Connected P
Focus to learn more arXiv-issued DOI via DataCite (pending registration) Submission history From: Yotam Wolf [ view email ] [v1] Tue, 7 Jul 2026 18:36:04 UTC (746 KB) Full-text links: Access Paper: View a PDF of the paper titled When Does In-Context Search Help? A Sampling-Complexity Theory of Reflection-Driven Reasoning, by Yotam Wolf and 2 other authors View PDF HTML (experimental) TeX Source view license Current browse context: < prev | next > new | recent | 2026-07 Change to browse by: cs References & Citations NASA ADS Google Scholar Semantic Scholar export BibTeX citation Loading... BibTeX formatted citation × loading... Data provided by: Bookmark Bibliographic Tools Bibliographic and Citation Tools Bibliographic Explorer Toggle Bibliographic Explorer ( What is the Explorer? ) Connected P
apers Toggle Connected Papers ( What is Connected Papers? ) Litmaps Toggle Litmaps ( What is Litmaps? ) Toggle scite Smart Citations ( What are Smart Citations? ) Code, Data, Media Code, Data and Media Associated with this Article alphaXiv Toggle alphaXiv ( What is alphaXiv? ) Links to Code Toggle CatalyzeX Code Finder for Papers ( What is CatalyzeX? ) DagsHub Toggle DagsHub ( What is DagsHub? ) GotitPub Toggle ( What is GotitPub? ) Huggingface Toggle Hugging Face ( What is Huggingface? ) ScienceCast Toggle ScienceCast ( What is ScienceCast? ) Demos Demos Replicate Toggle Replicate ( What is Replicate? ) Spaces Toggle Hugging Face Spaces ( What is Spaces? ) Spaces Toggle ( What is ? ) Related Papers Recommenders and Search Tools Link to Influence Flower Influence Flower ( What are Influence Flowers? ) Core recommender toggle CORE Recommender ( What is CORE? ) Author Venue Institution Topic About arXivLabs arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs .
AgentLens:基于生产评估的编码智能体轨迹审查方法被提出
近日,一篇题为《AgentLens: Production-Assessed Trajectory Reviews for Coding Agent Evaluation》的论文在arXiv预印本平台发布。该论文由Andrey Podivilov等七位作者共同完成,提出了一种名为AgentLens的编码智能体评估新方法。与传统基准测试不同,AgentLens利用智能体在真实生产环境中的任务执行轨迹进行审查与评估,旨在更客观、准确地衡量编码智能体的实际能力。该方法有望为智能体的研发与部署提供更有效的反馈,推动编码智能体评估的标准化。目前论文已开放PDF全文下载。 #AgentLens #编码智能体 #评估方法 #生产评估 #轨迹审查 #arXiv #AI #大模型
近日,一篇题为《AgentLens: Production-Assessed Trajectory Reviews for Coding Agent Evaluation》的论文在arXiv预印本平台发布。该论文由Andrey Podivilov等七位作者共同完成,提出了一种名为AgentLens的编码智能体评估新方法。与传统基准测试不同,AgentLens利用智能体在真实生产环境中的任务执行轨迹进行审查与评估,旨在更客观、准确地衡量编码智能体的实际能力。该方法有望为智能体的研发与部署提供更有效的反馈,推动编码智能体评估的标准化。目前论文已开放PDF全文下载。 #AgentLens #编码智能体 #评估方法 #生产评估 #轨迹审查 #arXiv #AI #大模型
学术“人性化”工具消除AI写作痕迹引担忧
一款名为Humanizer的学术工具于6月20日发布,旨在个性化修改由AI生成的研究论文语气,抹去明显的机器写作痕迹。该工具由明尼苏达大学研究员Jie Ding开发,专为论文和基金申请书设计,核心原则是“在不随意化文风的前提下消除AI特有印记”。支持者认为它能帮助非母语学者提高写作效率,但批评者担忧这会诱使更多研究者违规使用AI且不披露,损害学术诚信。AI检测平台Pangram的CEO表示该工具并不高级,大部分处理过的文本仍能被其平台捕获,并表示正开发专门针对性的升级版本。开发者Ding强调工具仅是编辑辅助,不解除使用者披露AI协助的义务,并已修改说明以避免误导。这一争议凸显出AI辅助学术写作与规范之间的紧张关系。 #AI #学术写作 #人性化工具 #人工智能 #学术诚信 #科技争议 #大模型
一款名为Humanizer的学术工具于6月20日发布,旨在个性化修改由AI生成的研究论文语气,抹去明显的机器写作痕迹。该工具由明尼苏达大学研究员Jie Ding开发,专为论文和基金申请书设计,核心原则是“在不随意化文风的前提下消除AI特有印记”。支持者认为它能帮助非母语学者提高写作效率,但批评者担忧这会诱使更多研究者违规使用AI且不披露,损害学术诚信。AI检测平台Pangram的CEO表示该工具并不高级,大部分处理过的文本仍能被其平台捕获,并表示正开发专门针对性的升级版本。开发者Ding强调工具仅是编辑辅助,不解除使用者披露AI协助的义务,并已修改说明以避免误导。这一争议凸显出AI辅助学术写作与规范之间的紧张关系。 #AI #学术写作 #人性化工具 #人工智能 #学术诚信 #科技争议 #大模型
世界杯AI预测对决
由联想与中国移动咪咕联合发起的“世界杯人机预测对决”项目,对12款国产大模型进行了评测。在全部92场比赛中,机器整体预测准确率达64%,高于人类的53.8%。中国移动“九天”框架以64次预测正确居首,联想“天禧AI”和阿里“通义千问”并列第二。然而,当比赛爆出冷门时,算法便暴露了局限:西班牙对阵佛得角的比赛,多数模型预测西班牙胜,结果却是0-0闷平;佛得角随后逼平乌拉圭和沙特,以不败战绩出线,四款模型未能预见这一“童话”。DeepSeek、Kimi等也曾错误预测荷兰、德国晋级,而两队均在32强出局。专家指出,高准确率不代表机器真正理解足球,这一测试更像一场公开的算法压力考试。 #AI #世界杯 #大模型 #预测 #人工智能 #足球 #联想 #中国移动 #机器学习 #比赛分析
由联想与中国移动咪咕联合发起的“世界杯人机预测对决”项目,对12款国产大模型进行了评测。在全部92场比赛中,机器整体预测准确率达64%,高于人类的53.8%。中国移动“九天”框架以64次预测正确居首,联想“天禧AI”和阿里“通义千问”并列第二。然而,当比赛爆出冷门时,算法便暴露了局限:西班牙对阵佛得角的比赛,多数模型预测西班牙胜,结果却是0-0闷平;佛得角随后逼平乌拉圭和沙特,以不败战绩出线,四款模型未能预见这一“童话”。DeepSeek、Kimi等也曾错误预测荷兰、德国晋级,而两队均在32强出局。专家指出,高准确率不代表机器真正理解足球,这一测试更像一场公开的算法压力考试。 #AI #世界杯 #大模型 #预测 #人工智能 #足球 #联想 #中国移动 #机器学习 #比赛分析