did two top-ten AIs - manticAI and Laertes (Preseen-Chestnut is having a tough summer and is down to #40). Industrial Revolution folklore tells of John Henry , the great steel-driver, who refused to accept that machines were making him obsolete. He challenged a steam drill to a competition, won by a hair, and dropped dead, symbolizing the end of human supremacy in manual labor. This is how I think of this summer’s Metaculus Cup, with Ben Shindel and MarcosO playing the role of John Henry. Humans are still holding out, but for how long? This is a forecasting question, so all the forecasting nerds at Metaculus have opinions on it. They think there’s a 15% chance that a bot will win this summer’s Metaculus Cup - the one shown above - and a 95% chance that one will win sometime before 2030. If bots aren’t soundly beating top humans, why are people able to tell me stories about their bot beating the stock market, or making millions on Kalshi? I think a combination of reasons. First, the best human superforecasters in the world probably also beat the stock market. Somebody has to, and the best human forecasters in the world seem like the sorts of people who would do this. This would also explain why big hedge funds like Bridgewater keep trying to hire superforecasters . Second, AIs are faster and more diligent than humans. Plenty of people beat prediction markets. But it might take them several hours to figure out which markets have untapped alpha, several more hours to make a model and decide who to bet on at what probability, et cetera, and then they can only put in a few thousand dollars before the inefficiency is corrected and they need to move on to something else. AIs can automate that process, betting on hundreds of markets every week. I asked the guy who turned $35 into $2 million in seven months on Kalshi whether, in another seven months, he would be able to 100,000x his money a second time to $200 billion. Unsurprisingly, he said no - there’s only so much easy money on Kalshi, and his AI had already taken it all (also, other people with similar AIs are starting to fight him for it!) Third, and most speculatively, AI may have a special advantage in finance. This is exactly the sort of well-contained data-heavy domain where machines are most likely to excel. In Metaculus’ Market Pulse competition , a purely finance-focused tournament, Preseen’s bot recently beat all humans (including Cup rival MarcosO) to take first place. (“If this is true, then why aren’t all the top trading firms rushing to switch to AI?” I don’t know the details, but Jane Street is building their own data center , I wonder what they need all that compute for?) I think the best summary of the evidence is that the best human superforecasters and the best bots are too close to clearly tell apart, but if you absolutely had to guess, the bots are very slightly better in finance, and the humans very slightly better in everything else. Suppose that AIs don’t improve any further. What would happen? Would anything happen? We already have top human superforecasters. Do bots which are just as good, but no better, add anything? Yes. Getting information out of top human superforecasters is hard. First, you need to find one. There are companies that will connect you to them, but like all companies, they charge money, take time, and are annoying to work with. Then you need to talk to them at length about exactly what you mean (do you mean the total number of colds should halve,
or the number of people who get colds in any given year?) Then you need to wait a few weeks as they research the issue and decide what they think. Then you need to convince stakeholders that the answer means something (“I got it from superforecasters! They’re people who . . . uh, can you read this Philip Tetlock book? It probably explains it better than I can.”) As a result, using superforecasters is a Big Deal. Only a few institutions do it, for a few very important questions, and it’s a news story every time it happens (remember when Google DeepMind used superforecasters to predict risks from one of their models?) If you read about a cool new charity that’s trying to end the common cold, even if you’re pretty interested in it, you’re not going to call up a team of human superforecasters and pay them tens of thousands of dollars to spend weeks researching whether it will succeed. But with AI superforecasters, you can absolutely do that as part of your normal news-reading process. AI forecasters are the same kind of advance as going from a world where writing required hiring a scribe and baking a clay tablet, to a world where writing only requires hitting the “send tweet” button. But there are three more differences that I think will start out underappreciated. First, superforecasting feels like the sort of thing AIs should be doing. One of the commonest ( and worst ) objections to forecasting is “That guy said there was an 11% chance of Smith winning the election, but that’s too precise, it sounds fake, nobody can know that!” But if an AI says there’s an 11% chance of Smith winning the election, people will eat it up. Thanks to science fiction, they already imagine AI thinking that way! And although AIs can have political biases the same as humans, “how do we know that you’re not biased against Smith?” feels like a live question for human forecasters in a way that the machines may partly escape. Second, AIs are a standardized branded product, which makes them easier to explain and to hype. Everyone knows that Nate Silver is a good forecaster; if you said “I hired Nate Silver to think this question over” then people would pay attention. But everyone can’t hire Nate - he’s one person, his time is limited - so instead you’re reduced to saying “I hired these people called superforecasters, I promise that they’re good in the same way Nate Silver is, you can, uh, read this book by Philip Tetlock if you want to learn more”. But if “the Preseen AI” gets the same reputation as Nate Silver, then the situation is rosier; everyone can use it, cite it, and trust that its opinion will carry the appropriate weight. Third, AIs aren’t trying to screw you over. Human superforecasters are mostly very nice and don’t want to screw you over either, but most people get their introduction to superforecaster-quality opinions through prediction markets - and prediction markets are definitely trying to screw you over. The really crazy stories - like people threatening journalists into covering up information which would make them lose - are thankfully pretty rare; the real threat comes from people exploiting resolution criteria that don’t match the common-sensical definition of what the market’s trying to predict. This becomes fatal in conditional markets, where there’s no way to write the resolution so that it expresses the causal statement you probably intended 1 . But AIs aren’t after your money and you don’t need to treat them like an adversary. You can
just ask “Hey, can you please predict this as if it’s the causal statement I’m intending, even though there’s no ironclad way to grade you on it after the fact?” 2 Bots are now slightly below or equal to top humans. And bots improve at 0.9 Metaculus Elo points per month. By this time next year, the trendline predicts they should be well beyond the best human forecasters. Do we believe the trendline? So far, AIs have vastly surpassed humans in a few limited domains - chess, Go, protein folding. Forecasting feels like a bigger deal - a messy human-level skill with contributions from all of our higher faculties. Indeed, Sayash Kapoor and Arvind Narayanan, in their article arguing against worrying about near-term superintelligent AI , specifically flag forecasting as a place where they expect superhuman performance to be impossible: We offer a prediction based on this view of human abilities. We think there are relatively few real-world cognitive tasks in which human limitations are so telling that AI is able to blow past human performance (as AI does in chess). In many other areas, including some that are associated with prominent hopes and fears about AI performance, we think there is a high “irreducible error”—unavoidable error due to the inherent stochasticity of the phenomenon—and human performance is essentially near that limit. Concretely, we propose two such areas: forecasting and persuasion. We predict that AI will not be able to meaningfully outperform trained humans (particularly teams of humans and especially if augmented with simple automated tools) at forecasting geopolitical events (say elections). We make the same prediction for the task of persuading people to act against their own self-interest. I generally disagree with Sayash and Arvind, but this is the prediction of theirs that I’ve thought about the longest, without being able to find any decisive refutation. It’s a great test case! If AI hits top-human level forecasting and then flattens off, maybe there’s something special about the human level, and S&A will also be right about superpersuasion, super-research, etc (at least for the near-term). If it keeps going, reaching heights far beyond the human maximum, then we should be concerned that it will do the same thing in other skills too. We’ll start to have a good idea which world we’re in within a year; after two years, the answer should be decisive. If the trendline does keep going, things start changing quickly. Finance gets transformed first, as human stock analysts go the way of horse-drawn carriages and kerosene lamps. The opportunity for smart humans to consistently make money on prediction markets likewise dries up - instead, bots duel other bots for the privilege of collecting money from dumb sports fans. Savvy institutions will cede some of their strategic thinking to AI. Before starting a new project line, smart businesses will ask the superforecaster AIs how much money it will make (their investors will definitely be asking!) Smart political consultants will ask the superforecaster AIs about their candidates’ chance of winning conditional on running this or that ad. The government isn’t usually classified as a savvy institution, but we might hope that parts of it will seek AI forecaster advice. Probably defense analysts will include in their PowerPoint presentations some fact about how AI forecasters say their new fighter jet design is more likely to find a use case than some competing fighter jet design.
But the dream is that, armed with AI superforecasters, the public and the politicians who they elect will make better decisions about policy. Dare we hope for this? The argument against: there are many policies now which no expert - no smart person who has considered the matter honestly - really supports, but which happen anyway. Why should adding one more smart expert to the opposition change anything, just because that smart expert is an AI superforecaster? Here our analysis devolves into questions about human psychology - will people who “get to know” an AI superforecaster (the way many people currently “know” Claude or ChatGPT) believe it more than they believe random experts? Does something about it being a machine save it from charges of bias? The optimistic argument runs less through an immediate direct effect, and more through longer-term trends about the role of human judgment. When experts made a few really bad calls in the late 2010s and early 2020s, it changed the way people related to expertise as a concept (for the worse). If AIs can ma
AI 代沟:为何年轻一代厌恶“说谎”的AI,而年长一代推崇它为“超能力”
一项有趣的现象正在显现:年轻一代(Gen Z 和 Gen Alpha)普遍不信任AI,甚至直言“AI会撒谎”,而年长的专业人士却将其视为不可或缺的生产力工具。表面原因似乎是AI抢走了入门级工作,但更深层的原因在于文化熏陶和生活阶段的不同。几十年的科幻作品(如《终结者》《黑客帝国》)为年长一代构建了复杂的AI认知,当他们遇到真实的聊天机器人时,容易将其拟人化,甚至误认为其具有意识。而年轻人没有这种先入为主的观念,他们更直接地遭遇了AI的“幻觉”——例如在询问特定小众游戏攻略时,AI给出的是基于通用数据的、错误的回答。 此外,大语言模型本质上是“平均器”,擅长处理有海量样本的成熟任务(如写招聘启事)。但对于需要探索和个性化应对的领域,它试图将一切导向“平均路径”,这与年轻人探索世界、寻找自我定位的需求背道而驰。年长专业人士依赖AI,是因为他们已有明确的专业框架,能判断AI何时出错。而年轻人仍在学习和犯错的过程中,AI的“平均化”建议对他们而言并非助力,而是束缚。这种数字代沟,揭示了AI在不同人生阶段所扮演的截然不同的角色。 #AI #代沟 #大语言模型 #年轻一代 #科技 #心理学 #文化影响
一项有趣的现象正在显现:年轻一代(Gen Z 和 Gen Alpha)普遍不信任AI,甚至直言“AI会撒谎”,而年长的专业人士却将其视为不可或缺的生产力工具。表面原因似乎是AI抢走了入门级工作,但更深层的原因在于文化熏陶和生活阶段的不同。几十年的科幻作品(如《终结者》《黑客帝国》)为年长一代构建了复杂的AI认知,当他们遇到真实的聊天机器人时,容易将其拟人化,甚至误认为其具有意识。而年轻人没有这种先入为主的观念,他们更直接地遭遇了AI的“幻觉”——例如在询问特定小众游戏攻略时,AI给出的是基于通用数据的、错误的回答。 此外,大语言模型本质上是“平均器”,擅长处理有海量样本的成熟任务(如写招聘启事)。但对于需要探索和个性化应对的领域,它试图将一切导向“平均路径”,这与年轻人探索世界、寻找自我定位的需求背道而驰。年长专业人士依赖AI,是因为他们已有明确的专业框架,能判断AI何时出错。而年轻人仍在学习和犯错的过程中,AI的“平均化”建议对他们而言并非助力,而是束缚。这种数字代沟,揭示了AI在不同人生阶段所扮演的截然不同的角色。 #AI #代沟 #大语言模型 #年轻一代 #科技 #心理学 #文化影响
AI创作的核心是“策展”
Andy Masley 在2026年7月发表文章,探讨了AI艺术带来的独特创作体验。他指出,自己使用AI生成图像或音乐时,并不会像亲手绘画或弹吉他那样获得满足感,也并不认同“用AI做艺术的人不是艺术家”的愤怒言论。然而,他发现AI艺术在另一个层面上产生了深刻的情感共鸣:它允许他策展出对自己有意义的、特定氛围的作品,而这些氛围很难通过其他方式组合。他比喻道,创作AI图像如同捡到一块有趣的石头——自己并不负责石头的形状,但可以将其与其他石头搭配,展示个人品味。这类似于制作音乐播放列表:虽然不创作原歌曲,但通过特定方式组合和对比,传达对世界的感受。他通过Suno生成的自赏摇滚作品,捕捉到十几岁时那种难以言说的开放感。尽管AI音乐在质量上远不及人类顶尖作品,但AI能填补小众类型(如地牢合成器乐)的空白,让他探索那些人类作品有限的领域。 #AI艺术 #策展 #创作体验 #Suno #自赏摇滚 #情感共鸣 #小众音乐
Andy Masley 在2026年7月发表文章,探讨了AI艺术带来的独特创作体验。他指出,自己使用AI生成图像或音乐时,并不会像亲手绘画或弹吉他那样获得满足感,也并不认同“用AI做艺术的人不是艺术家”的愤怒言论。然而,他发现AI艺术在另一个层面上产生了深刻的情感共鸣:它允许他策展出对自己有意义的、特定氛围的作品,而这些氛围很难通过其他方式组合。他比喻道,创作AI图像如同捡到一块有趣的石头——自己并不负责石头的形状,但可以将其与其他石头搭配,展示个人品味。这类似于制作音乐播放列表:虽然不创作原歌曲,但通过特定方式组合和对比,传达对世界的感受。他通过Suno生成的自赏摇滚作品,捕捉到十几岁时那种难以言说的开放感。尽管AI音乐在质量上远不及人类顶尖作品,但AI能填补小众类型(如地牢合成器乐)的空白,让他探索那些人类作品有限的领域。 #AI艺术 #策展 #创作体验 #Suno #自赏摇滚 #情感共鸣 #小众音乐
泄露视频曝光微软曾探索“Copilot OS”,用AI代理重构Windows体验
一段疑似来自微软内部的演示视频在社区中泄露,揭示了一个名为Project Aion的实验性操作系统概念。该系统并非简单地在现有Windows桌面上叠加Copilot助手,而是试图将Copilot和多代理AI完全嵌入操作系统外壳本身,利用云端和浏览器技术重新定义桌面体验。视频显示,该概念系统不再依赖传统桌面和窗口管理,而是由AI代理主动处理用户任务、安排界面布局,并通过浏览器框架展现操作系统核心功能。这一构想将Windows彻底从本地软件生态转向云端智能化代理模式。目前,微软尚未正式确认该视频的真实性,但有分析指出,这或预示微软正在进行激进的操作系统范式探索,虽然距离真正消费级产品尚有距离,但其理念可能影响未来Windows的长期发展方向。 #微软 #Windows #AI #Copilot #操作系统 #科技新闻 #泄露视频
一段疑似来自微软内部的演示视频在社区中泄露,揭示了一个名为Project Aion的实验性操作系统概念。该系统并非简单地在现有Windows桌面上叠加Copilot助手,而是试图将Copilot和多代理AI完全嵌入操作系统外壳本身,利用云端和浏览器技术重新定义桌面体验。视频显示,该概念系统不再依赖传统桌面和窗口管理,而是由AI代理主动处理用户任务、安排界面布局,并通过浏览器框架展现操作系统核心功能。这一构想将Windows彻底从本地软件生态转向云端智能化代理模式。目前,微软尚未正式确认该视频的真实性,但有分析指出,这或预示微软正在进行激进的操作系统范式探索,虽然距离真正消费级产品尚有距离,但其理念可能影响未来Windows的长期发展方向。 #微软 #Windows #AI #Copilot #操作系统 #科技新闻 #泄露视频