AI知识库 @ai521
364 subscribers
24.2K photos
45 videos
20 files
925 links
@ai521 专注分享最实用的AI内容

🤖 AI教程(新手到进阶)
🧠 AI知识科普(大模型 / 提示词 / 自动化)
📰 AI资讯更新(每日最新AI动态)
📚 AI实战技巧(写作 / 绘画 / 编程 / 赚钱)
🔧 最新AI工具推荐

每天更新AI干货
长期做一个真正有价值的AI频道
Download Telegram
The AI Superforecasters Are Here

Scott Alexander Jul 02, 2026 217 108 29 Share The annual prediction market conference was earlier this month. This was the year prediction markets went from an obscure hobby to a multi-billion dollar industry; from semi-illegal to having the President’s son as an advisor. I can’t remember if anyone talked about any of that. It didn’t even register. All eyes were on the AI superforecasters. I met an AI superforecaster startup founder who told me his AI had turned $35 into $2 million on Kalshi over seven months. I met another who said they were beating the stock market by 25% with a market-neutral portfolio - of course this could be luck, but they’d beaten Kalshi and Polymarket by similar margins. In fact, I believe all of these people. The extending-lines-on-graphs community has long predicted that AIs would beat the best human forecasters sometime
in 2026 - 2027. What did you expect the bots-finally-beat-humans-at-predicting-the-future moment to look like? Vibes? Papers? Essays? In retrospect, sure: it will look like AIs making crazy profits on prediction markets and beating the stock market by some comfortable amount. But what happens next? Before getting into details, what exactly are we talking about? An AI superforecaster is an AI - usually a frontier model like ChatGPT or Claude - which has been modified to be good at forecasting. This usually means a “scaffold” - a program that handholds it through a long research process with various prompts, tools, advice about when to create subagents, etc. The overall experience is a lot like using any other AI, but slower and more expensive, because it’s doing more work. This might make more sense with an example. FutureSearch - the company that claims to be beating the stock market - kindly offered to let me try their AI superforecaster and write about it here. For a test question - some Silicon Valley philanthropists recently started a project to end respiratory infections like the common cold . I decided to ask about their chances of success. Since forecasters need very precise questions, I asked how likely it was that the rate of colds would be cut in half by 2040: By two minutes in, the AI had deployed three subagents, read 16 websites, and (at the exact moment I took this screenshot) was “investigating the scalability of ASHRAE Standard 241 air cleaning technology for widespread residential adoption by 2040.” After five minutes, it had its answer: the chance of US respiratory infections halving by 2040 was 7% : The agents’ reasoning centers on a tough conjunctive chain of requirements that all must succeed simultaneously: 1: The biology is brutal. 200+ cold-causing viruses, 150+ rhinovirus serotypes alone — 50 years of vaccine efforts have failed, and experts as recently as late 2025 didn’t expect a cold vaccine within 5 years. During COVID lockdowns with unprecedented behavioral changes, rhinovirus incidence dropped only modestly and rebounded immediately. 2: Intercept’s own timeline is tight. Their roadmap projects commercialization 5–7+ years out, leaving little room to clear Phase 3 trials, FDA approval, and achieve mass deployment by 2040 — all for under $500M (likely far too little for that scope). 3: Adoption is a wildcard. For a mild illness, compliance with regular nasal sprays or prophylactics is expected to be low. Air-cleaning infrastructure faces steep retrofit costs and a mixed evidence base ( a 2025 RCT of germicidal UV found no significant reduction in infections). 4: Measurement may be impossible. The US lacks routine population-wide cold surveillance, and the 2020–2026 baseline is distorted by pandemic-era behavior — making a verified >50% reduction hard to demonstrate even if real progress occurs. The initiative is genuinely novel and credible, but halving cold frequency in 14 years without transformative AI acceleration is an extraordinarily high bar. You could follow up by forecasting a softer threshold (e.g., 25% reduction by 2035), or by ranking the specific Intercept technical approaches (nasal sprays vs. vaccines vs. air cleaning) by their likelihood of meaningful impact. …and, in case I wanted to learn more, cited its 212 sources. The forecast had taken five minutes and cost me $8 in credits. But is it true? Ideally we would wait until 2040 and see. For now, I started by comparing its answer to another
superforecaster AI. Preseen is the company that claimed to 100,000x their seed money on Kalshi. Here’s their answer : 8.8% compared to FutureSearch’s 7%, not bad! Are either of these true? I asked a human superforecaster to predict this question, to see if she got the same as the AI. She said that depending on an ambiguity in the wording, she would give it 5-10%. Again, not bad! Of course, it would be even better to do the same experiment at scale and figure out how AIs compare to humans once and for all. But measuring forecasting ability is hard. You can’t say something like “it gets 85% of questions right”, because that depends entirely on question difficulty. If the questions are things like “will the sun rise tomorrow morning”, then even a 100% hit rate is unimpressive. Instead, we can only match different forecasters against each other and determine who is better or worse. Any anchoring in an absolute space will come from the inclusion of groups whose predictive abilities we intuitively understand (eg the average member of the public, CIA analysts, etc). The forecasting website Metaculus matches AIs against humans and each other on a common metric. Here are their results over time : The Metaculus Community Prediction is a “wisdom of crowds” style aggregation of all the forecasters on Metaculus. The Metaculus Pro Forecasters are top professional superforecasters. This graph makes it look like - as of May 2026 when Gemini 3.1 was state of the art - AI was approaching the Community Prediction. This is no mean feat, but it’s still far from the professional superforecaster level. But in a recent blog post , Metaculus adds context. The graph above only measures out-of-the-box brand-name AIs like GPT and Claude. It doesn’t count forecasting-focused scaffolds like FutureSearch. A different investigation by Metaculus finds that these efforts are “worth 9 months of base model progress”, eg a well-scaffolded AI today is already as good at forecasting as base models will be in nine months. If you extend the dotted green line on the graph to July 2026, then add nine months for the extra scaffolding, it looks like the best AIs should be around 31, compared to top pro forecasters’ 36. So in theory, the absolute best forecasters in the world are still beating the top AIs, but the margin of victory is less than the graph suggests, and we should expect human-AI parity in about six months. But the claim that scaffolded AIs are nine months behind base models is itself ~9 months old. Several people in the field told me that they thought this underestimated true progress. Claims by the AI startups themselves may be treated skeptically, but even a few top human superforecasters said they were no longer confident they could beat the bots. Seems like time for a head-to-head matchup. The Metaculus Cup - the World Cup of forecasting! - is on the case. Once a season, top humans and AIs compete on about fifty questions like “Who will win the upcoming Nepali elections?” and “Will the US attack Iran?” Here are the winners of the most recent tournament: Humans took the top two spots, but Preseen’s AI came in third. Every forecasting competition involves a heavy dose of luck, so realistically at this point humans and AIs are in a statistical dead heat. We can confirm by looking at the intermediate results of the ongoing summer Metaculus Cup: Of humans who placed in the top ten during spring, 2/10 - benshindel and MarcosO - repeated their performance in summer. So
did two top-ten AIs - manticAI and Laertes (Preseen-Chestnut is having a tough summer and is down to #40). Industrial Revolution folklore tells of John Henry , the great steel-driver, who refused to accept that machines were making him obsolete. He challenged a steam drill to a competition, won by a hair, and dropped dead, symbolizing the end of human supremacy in manual labor. This is how I think of this summer’s Metaculus Cup, with Ben Shindel and MarcosO playing the role of John Henry. Humans are still holding out, but for how long? This is a forecasting question, so all the forecasting nerds at Metaculus have opinions on it. They think there’s a 15% chance that a bot will win this summer’s Metaculus Cup - the one shown above - and a 95% chance that one will win sometime before 2030. If bots aren’t soundly beating top humans, why are people able to tell me stories about their bot beating the stock market, or making millions on Kalshi? I think a combination of reasons. First, the best human superforecasters in the world probably also beat the stock market. Somebody has to, and the best human forecasters in the world seem like the sorts of people who would do this. This would also explain why big hedge funds like Bridgewater keep trying to hire superforecasters . Second, AIs are faster and more diligent than humans. Plenty of people beat prediction markets. But it might take them several hours to figure out which markets have untapped alpha, several more hours to make a model and decide who to bet on at what probability, et cetera, and then they can only put in a few thousand dollars before the inefficiency is corrected and they need to move on to something else. AIs can automate that process, betting on hundreds of markets every week. I asked the guy who turned $35 into $2 million in seven months on Kalshi whether, in another seven months, he would be able to 100,000x his money a second time to $200 billion. Unsurprisingly, he said no - there’s only so much easy money on Kalshi, and his AI had already taken it all (also, other people with similar AIs are starting to fight him for it!) Third, and most speculatively, AI may have a special advantage in finance. This is exactly the sort of well-contained data-heavy domain where machines are most likely to excel. In Metaculus’ Market Pulse competition , a purely finance-focused tournament, Preseen’s bot recently beat all humans (including Cup rival MarcosO) to take first place. (“If this is true, then why aren’t all the top trading firms rushing to switch to AI?” I don’t know the details, but Jane Street is building their own data center , I wonder what they need all that compute for?) I think the best summary of the evidence is that the best human superforecasters and the best bots are too close to clearly tell apart, but if you absolutely had to guess, the bots are very slightly better in finance, and the humans very slightly better in everything else. Suppose that AIs don’t improve any further. What would happen? Would anything happen? We already have top human superforecasters. Do bots which are just as good, but no better, add anything? Yes. Getting information out of top human superforecasters is hard. First, you need to find one. There are companies that will connect you to them, but like all companies, they charge money, take time, and are annoying to work with. Then you need to talk to them at length about exactly what you mean (do you mean the total number of colds should halve,
or the number of people who get colds in any given year?) Then you need to wait a few weeks as they research the issue and decide what they think. Then you need to convince stakeholders that the answer means something (“I got it from superforecasters! They’re people who . . . uh, can you read this Philip Tetlock book? It probably explains it better than I can.”) As a result, using superforecasters is a Big Deal. Only a few institutions do it, for a few very important questions, and it’s a news story every time it happens (remember when Google DeepMind used superforecasters to predict risks from one of their models?) If you read about a cool new charity that’s trying to end the common cold, even if you’re pretty interested in it, you’re not going to call up a team of human superforecasters and pay them tens of thousands of dollars to spend weeks researching whether it will succeed. But with AI superforecasters, you can absolutely do that as part of your normal news-reading process. AI forecasters are the same kind of advance as going from a world where writing required hiring a scribe and baking a clay tablet, to a world where writing only requires hitting the “send tweet” button. But there are three more differences that I think will start out underappreciated. First, superforecasting feels like the sort of thing AIs should be doing. One of the commonest ( and worst ) objections to forecasting is “That guy said there was an 11% chance of Smith winning the election, but that’s too precise, it sounds fake, nobody can know that!” But if an AI says there’s an 11% chance of Smith winning the election, people will eat it up. Thanks to science fiction, they already imagine AI thinking that way! And although AIs can have political biases the same as humans, “how do we know that you’re not biased against Smith?” feels like a live question for human forecasters in a way that the machines may partly escape. Second, AIs are a standardized branded product, which makes them easier to explain and to hype. Everyone knows that Nate Silver is a good forecaster; if you said “I hired Nate Silver to think this question over” then people would pay attention. But everyone can’t hire Nate - he’s one person, his time is limited - so instead you’re reduced to saying “I hired these people called superforecasters, I promise that they’re good in the same way Nate Silver is, you can, uh, read this book by Philip Tetlock if you want to learn more”. But if “the Preseen AI” gets the same reputation as Nate Silver, then the situation is rosier; everyone can use it, cite it, and trust that its opinion will carry the appropriate weight. Third, AIs aren’t trying to screw you over. Human superforecasters are mostly very nice and don’t want to screw you over either, but most people get their introduction to superforecaster-quality opinions through prediction markets - and prediction markets are definitely trying to screw you over. The really crazy stories - like people threatening journalists into covering up information which would make them lose - are thankfully pretty rare; the real threat comes from people exploiting resolution criteria that don’t match the common-sensical definition of what the market’s trying to predict. This becomes fatal in conditional markets, where there’s no way to write the resolution so that it expresses the causal statement you probably intended 1 . But AIs aren’t after your money and you don’t need to treat them like an adversary. You can
just ask “Hey, can you please predict this as if it’s the causal statement I’m intending, even though there’s no ironclad way to grade you on it after the fact?” 2 Bots are now slightly below or equal to top humans. And bots improve at 0.9 Metaculus Elo points per month. By this time next year, the trendline predicts they should be well beyond the best human forecasters. Do we believe the trendline? So far, AIs have vastly surpassed humans in a few limited domains - chess, Go, protein folding. Forecasting feels like a bigger deal - a messy human-level skill with contributions from all of our higher faculties. Indeed, Sayash Kapoor and Arvind Narayanan, in their article arguing against worrying about near-term superintelligent AI , specifically flag forecasting as a place where they expect superhuman performance to be impossible: We offer a prediction based on this view of human abilities. We think there are relatively few real-world cognitive tasks in which human limitations are so telling that AI is able to blow past human performance (as AI does in chess). In many other areas, including some that are associated with prominent hopes and fears about AI performance, we think there is a high “irreducible error”—unavoidable error due to the inherent stochasticity of the phenomenon—and human performance is essentially near that limit. Concretely, we propose two such areas: forecasting and persuasion. We predict that AI will not be able to meaningfully outperform trained humans (particularly teams of humans and especially if augmented with simple automated tools) at forecasting geopolitical events (say elections). We make the same prediction for the task of persuading people to act against their own self-interest. I generally disagree with Sayash and Arvind, but this is the prediction of theirs that I’ve thought about the longest, without being able to find any decisive refutation. It’s a great test case! If AI hits top-human level forecasting and then flattens off, maybe there’s something special about the human level, and S&A will also be right about superpersuasion, super-research, etc (at least for the near-term). If it keeps going, reaching heights far beyond the human maximum, then we should be concerned that it will do the same thing in other skills too. We’ll start to have a good idea which world we’re in within a year; after two years, the answer should be decisive. If the trendline does keep going, things start changing quickly. Finance gets transformed first, as human stock analysts go the way of horse-drawn carriages and kerosene lamps. The opportunity for smart humans to consistently make money on prediction markets likewise dries up - instead, bots duel other bots for the privilege of collecting money from dumb sports fans. Savvy institutions will cede some of their strategic thinking to AI. Before starting a new project line, smart businesses will ask the superforecaster AIs how much money it will make (their investors will definitely be asking!) Smart political consultants will ask the superforecaster AIs about their candidates’ chance of winning conditional on running this or that ad. The government isn’t usually classified as a savvy institution, but we might hope that parts of it will seek AI forecaster advice. Probably defense analysts will include in their PowerPoint presentations some fact about how AI forecasters say their new fighter jet design is more likely to find a use case than some competing fighter jet design.
But the dream is that, armed with AI superforecasters, the public and the politicians who they elect will make better decisions about policy. Dare we hope for this? The argument against: there are many policies now which no expert - no smart person who has considered the matter honestly - really supports, but which happen anyway. Why should adding one more smart expert to the opposition change anything, just because that smart expert is an AI superforecaster? Here our analysis devolves into questions about human psychology - will people who “get to know” an AI superforecaster (the way many people currently “know” Claude or ChatGPT) believe it more than they believe random experts? Does something about it being a machine save it from charges of bias? The optimistic argument runs less through an immediate direct effect, and more through longer-term trends about the role of human judgment. When experts made a few really bad calls in the late 2010s and early 2020s, it changed the way people related to expertise as a concept (for the worse). If AIs can ma
AI驱动汽车租赁自动化管理平台

一篇深度分析文章指出,一款结合人工智能、自动化与数据驱动工具的智能汽车租赁软件平台正重塑行业。该平台旨在简化运营流程、提升客户体验、降低成本并增加利润。文章总结出定义优秀汽车租赁软件的七大关键特性,涵盖AI算法优化车辆调度、智能预测维护需求、自动化合同处理智能客服等,实现车队管理的全面自动化和智能化。该平台已成为租赁企业数字化转型的核心工具,帮助企业在激烈竞争中提升效率与盈利能力。 #AI #汽车租赁 #自动化 #车队管理 #软件平台 #数据分析 #客户体验 #降本增效 #数字化转型
Telegram必备的搜索引擎,极搜JISOU帮你精准找到,想要的群组、频道、视频、音乐

👉 t.me/jisou2?start=a_8247614025
人类编辑与AI辅助

一位开发者创建了一本人类编辑、AI辅助的复古风格杂志,专门用于评测YouTube上的视频内容。该杂志旨在应对当前大量机器自动上传的视频浪潮,并探讨AI辅助媒体在审核这些内容时的责任与义务——当AI辅助的杂志评论“垃圾内容浪潮”时,它应该向读者提供怎样的承诺?项目强调人工把关与AI工具的结合,试图在算法泛滥的时代为内容质量把关。 #AI辅助 #人类编辑 #YouTube评测 #复古杂志 #内容审核 #垃圾信息 #AI责任 #科技新闻 #独立媒体
AI 代沟:为何年轻一代厌恶“说谎”的AI,而年长一代推崇它为“超能力”

一项有趣的现象正在显现:年轻一代(Gen Z 和 Gen Alpha)普遍不信任AI,甚至直言“AI会撒谎”,而年长的专业人士却将其视为不可或缺的生产力工具。表面原因似乎是AI抢走了入门级工作,但更深层的原因在于文化熏陶和生活阶段的不同。几十年的科幻作品(如《终结者》《黑客帝国》)为年长一代构建了复杂的AI认知,当他们遇到真实的聊天机器人时,容易将其拟人化,甚至误认为其具有意识。而年轻人没有这种先入为主的观念,他们更直接地遭遇了AI的“幻觉”——例如在询问特定小众游戏攻略时,AI给出的是基于通用数据的、错误的回答。 此外,大语言模型本质上是“平均器”,擅长处理有海量样本的成熟任务(如写招聘启事)。但对于需要探索和个性化应对的领域,它试图将一切导向“平均路径”,这与年轻人探索世界、寻找自我定位的需求背道而驰。年长专业人士依赖AI,是因为他们已有明确的专业框架,能判断AI何时出错。而年轻人仍在学习和犯错的过程中,AI的“平均化”建议对他们而言并非助力,而是束缚。这种数字代沟,揭示了AI在不同人生阶段所扮演的截然不同的角色。 #AI #代沟 #大语言模型 #年轻一代 #科技 #心理学 #文化影响
AI创作的核心是“策展”

Andy Masley 在2026年7月发表文章,探讨了AI艺术带来的独特创作体验。他指出,自己使用AI生成图像或音乐时,并不会像亲手绘画或弹吉他那样获得满足感,也并不认同“用AI做艺术的人不是艺术家”的愤怒言论。然而,他发现AI艺术在另一个层面上产生了深刻的情感共鸣:它允许他策展出对自己有意义的、特定氛围的作品,而这些氛围很难通过其他方式组合。他比喻道,创作AI图像如同捡到一块有趣的石头——自己并不负责石头的形状,但可以将其与其他石头搭配,展示个人品味。这类似于制作音乐播放列表:虽然不创作原歌曲,但通过特定方式组合和对比,传达对世界的感受。他通过Suno生成的自赏摇滚作品,捕捉到十几岁时那种难以言说的开放感。尽管AI音乐在质量上远不及人类顶尖作品,但AI能填补小众类型(如地牢合成器乐)的空白,让他探索那些人类作品有限的领域。 #AI艺术 #策展 #创作体验 #Suno #自赏摇滚 #情感共鸣 #小众音乐