AI知识库 @ai521
362 subscribers
24.2K photos
45 videos
20 files
925 links
@ai521 专注分享最实用的AI内容

🤖 AI教程(新手到进阶)
🧠 AI知识科普(大模型 / 提示词 / 自动化)
📰 AI资讯更新(每日最新AI动态)
📚 AI实战技巧(写作 / 绘画 / 编程 / 赚钱)
🔧 最新AI工具推荐

每天更新AI干货
长期做一个真正有价值的AI频道
Download Telegram
Xbox 大规模宕机,光盘游戏也无法运行

微软 Xbox 平台自周日晚上起发生大规模服务中断,不仅影响了数字版游戏的在线游玩,还意外地阻止了玩家运行光盘版游戏。此次宕机持续数小时,导致全球玩家无法正常登录账户、启动游戏或进行联机。即便是安装光盘游戏,系统也会因无法验证许可证或联网权限而报错。目前微软尚未公布具体原因及修复时间,只表示正在排查问题。此次故障再次凸显云游戏和数字授权对实体游戏体验的潜在影响。 #Xbox #微软 #宕机 #游戏 #光盘游戏 #数字游戏 #服务中断 #技术故障
韩国企业引入海外AI模型遭遇Token成本高昂难题,三星实施严格配额制

随着韩国大型企业加速将海外大模型引入业务流程,以三星、SK集团为代表的巨头正面临AI词元(Token)使用成本快速攀升的挑战。为控制开支,三星在向员工开放Claude、Gemini和ChatGPT等美国模型时,实施了严格的Token配额制度,员工需证明AI提升工作效率才能申请更高额度,此举引发部分员工不满。尽管Claude因编程能力在韩国开发群体中广受欢迎,但半导体等数据敏感行业对AI部署仍极为谨慎,SK海力士至今主要依赖自研模型,而三星芯片部门基本禁止使用外部大模型。此外,中小企业因缺乏管理体系,只能通过公布Token消耗排名等方式隐性限制使用。目前,Naver、Nexon等公司已大规模部署Claude,但整体成本压力正倒逼企业探索AI应用与运营成本的平衡之道。 #韩国 #AI #大模型 #Token #三星 #企业应用 #成本控制
主流AI模型政治倾向测试结果一致偏左

一个名为Unslop的小型研究实验室近日公布了一项实验,对16个主流AI大模型进行了政治倾向测试。测试采用经典的Political Compass问卷,涉及62个经济与社会政策问题。每个模型分别进行了标准版、反向措辞版和乱序版共数十次运行。结果显示,除xAI的Grok模型表现不稳定外,包括GPT系列、Claude系列、Gemini、Llama、DeepSeek、Qwen、Kimi、GLM等在内的所有模型均一致落在“自由意志左派”象限,展现出支持社会平等、反资本主义等进步倾向。研究作者坦言测试并非完全严谨,但结果反映了当前主流AI在政治价值观上存在系统性偏向,与驱动其发展的资本主义背景形成反差。 #AI #大模型 #政治倾向 #研究 #GPT #Claude #Grok #DeepSeek #科技新闻
Mealyu AI营养健身教练上线,拍照即可分析餐食

Mealyu是一款基于AI的营养与健身教练应用,用户只需拍摄餐食照片,系统即可快速估算热量、蛋白质、碳水化合物、脂肪及纤维等营养成分,并给出质量评分。应用还提供主动式AI教练功能,根据用户目标实时提醒饮食与补水调整,自动计算每日营养目标,并预测达成减重、增肌或身体重塑等目标的具体日期。此外,Mealyu能根据用户水平、场地和设备生成个性化训练方案,支持进度追踪与周期性报告。目前该应用提供免费基础版和高级订阅版(约12.42欧元/月或149欧元/年),高级版可解锁无限照片分析、24/7 AI教练及个性化训练课程。自上线以来,应用已收获超过2300条用户评价,平均评分4.9星,多名用户反馈其帮助显著提升了饮食管理和训练效果。 #AI #营养 #健身 #健康 #科技 #应用 #个性化 #饮食管理 #运动
超过三成新播客由AI生成,内容质量堪忧

据播客搜索引擎Listen Notes最新数据显示,目前超过30%的新增播客节目是由AI自动生成的低质量内容。这些播客通常缺乏真实人物参与和深度思考,内容空洞、重复性高,被业内称为“AI垃圾”。分析指出,随着AI语音合成和文案生成工具普及,大量创作者利用低成本方式批量生产播客,导致平台内容质量整体下滑,用户发现优质内容的难度增加。Listen Notes呼吁平台加强审核机制,同时提醒听众注意辨别AI生成内容。 #AI #播客 #内容质量 #ListenNotes #科技趋势 #数据报告 #智能语音
Our toolchain assumes one human writer, AI agents break that illusion

Christopher Meiklejohn One Writer Our tools assume one writer, and assume that writer is a human. Nothing computes what a change reads and writes at runtime, so the only known fix is brute force priced for organizations. 27 Jul 2026 In this blog post, I discuss three days in July 2026 when a single agent session ran away from me, and what those three days revealed about the concurrency assumptions buried in our development tooling. Six weeks ago I wrote The Test Suite Was the Incident my test suite had grown a pile of shared data nobody owned, every pull request paid to rebuild it, and the resulting failures had nothing to do with the changes under review. That cost me about $180 in one night. I got a worse one. It lasted three days, and in one twenty-four-hour stretch of it I burned through an entire Codex 20x
max plan. This post is not really about that, though. It is about a property of our tooling that the three days made impossible to ignore. Nearly every layer of this assumes one writer, and assumes that writer is a human. Git hands you a conflict and waits. Code review assumes somebody reads. A migration sequence assumes somebody is assigning the order. Each of those protocols terminates in a person. That is fine while there is exactly one, and while they are, in fact, a person. The agent runtime turns out to be on that list too, which I did not expect. It spawned eighty-four workers into a single checkout without being able to say what any one of them would read or write, and that is the same question git can’t answer about a diff. That assumption was invisible for forty years because nothing ever bound it at my scale. Agents break both halves at once: there are many of them, and not one of them is the person the protocol was waiting for. Git is the partial exception, and I will come to why the exception does not help. Nothing in the stack detects the violation when it happens. It gets caught later, somewhere else, attributed to the wrong change, and paid for at full price. Some context for readers arriving fresh. Zabriskie is a social app for live-music fans, and it’s also a deliberate experiment: I’m building a real, deployed, actually-used application almost entirely with AI agents (agents that wrote the features, agents that wrote the tests guarding those features, and agents that now open most of the pull requests), in order to find out what that’s like and, more usefully, where it breaks. I’ve written almost none of the code. That framing matters, because several things below look like obvious mistakes and are. I let agents design a migration scheme with only another agent reviewing it, I stopped reading most diffs, and I let sixty-four pull requests go up in a single day, all of which a careful engineer would tell you not to do and would be right about. But the point of running an experiment at the extreme is to find the walls. That week I found several at once. A migration, throughout, is a versioned SQL file that changes the database schema. CI is the automated checking that runs on every proposed change: build the app, spin up a fresh database, run the tests. Here is the shape of the three days. Treat these numbers as texture, not as evidence, for a reason I will get to. | | 24 Jul | 25 Jul | 26 Jul | | --- | --- | --- | --- | | pull requests opened | 19 | 64 | 28 | | pull requests merged | 20 | 50 | 30 | | incidents logged | 0 | 6 | 301 | Days in that table are UTC; clock times in the narrative below are Eastern (where I was). The burst ran past midnight UTC: 301 incidents on the 26th plus 60 more before 1 AM on the 27th, so 361 for the burst. Every incident count below is scoped to that window, and 253 rows is the total before 24 July. I offer that comparison as a sense of scale and not as a baseline, for the following reason. Those incidents exist because a standing instruction tells agents to log their own mistakes, and ten minutes into the worst night I tightened that instruction. The log therefore measures reported failures. Look at the daily series and it gets worse: there are days that week with ten and twenty merged pull requests and zero logged incidents, which at any real failure rate means nobody was logging rather than nothing broke. In short, I can’t give you a trustworthy baseline. What follows rests on
mechanism and on a few dated, checkable events, not on 361. The Session Late Saturday night the queue jammed. Sixty-four pull requests had gone up that day, ` was red, and nineteen open pull requests were stuck behind a suite that could not tell me which of them was broken. At 12:31 AM I opened a session with Codex and complained that CI was wasting too much money. Codex read that and hired a workforce. Over seventeen hours that session made 81 spawn calls, producing 74 direct children; those children spawned 10 more, for 84 threads. Then 358 calls waiting on them, 96 listing them, 85 sending follow-up work, and 13 interrupting them, which comes to six hundred and thirty-three tool calls of pure management overhead (spawn, wait, list, follow up, interrupt) against 54 messages from me over the same seventeen hours. All 84 ran in the same checked-out copy of the repository. Trivial parallelization, at scale, with a coordinator that only managed agents, on a toolchain built for one writer. The rest of this post is about why that combination is so much worse than it sounds. I asked afterward why it had spawned anything, since I never requested it. A language model asked why it did something produces plausible narrative, not introspection, and I apply that same skepticism to the self-reported incidents above. So the quotes below are not evidence, and I’m not offering them as any. I had already reached the same conclusion from the tool-call counts before I asked. I print them because they state that conclusion more plainly than I did, and because there is something worth looking at in a system that can describe the failure this precisely and could not avoid it. I chose to spawn multiple agents even though you did not ask me to. The environment permitted proactive delegation, but that was permission, not a requirement. I treated the availability of agents as a reason to use them instead of first asking whether they would reduce time, cost, or risk. It then diagnosed why the task was a bad candidate: A PR queue is largely serial because every merge changes ` , which changes the integration state of every remaining PR. The work needed one authoritative coordinator moving through the queue in order. Parallel agents could not independently merge overlapping PRs without continuously invalidating one another’s assumptions. On what those 358 waits were doing: The agents’ outputs created additional coordination work for me. I had to read their findings, compare conflicting recommendations, inspect supporting evidence, and decide how to combine them. The agents therefore generated work for the coordinator instead of reliably removing work from the critical path. A system with no notion of its own concurrency was handed a concurrency primitive and used it, reasonably, on a task the substrate could not support. What that substrate is, and why nothing in it objected, is the rest of this post. One number deserves care. I do not know how many of those 84 threads wrote to the checkout instead of reading it; from Codex’s own account many were investigating. A reader is harmless and a writer is not, and nothing in the toolchain drew the distinction or could. That’s the same missing primitive one level up: not even the agent runtime knew which of its children were writers. The Substrate Start with the layer that did anticipate this, because it is the one people reach for. Branches are optimistic concurrency control; worktrees go further, giving each writer a
physically separate checkout so that two agents can hold two versions of the tree at once. That works. However, git’s isolation stops at the edge of the source tree. A worktree gives an agent its own files. It doesn’t give it its own database, port range, mock server, or position in the migration sequence. Everything below the filesystem is shared and singular. That assumption was invisible to me because it never bound me. It has bound large organizations for a very long time, however, and over the last two decades they built the response to it: merge queues, hermetic builds, database-per-test, trunk-based development, automated culprit-finding, and whole infrastructure teams whose only job is to keep the thing moving. So what is new here isn’t the concurrency. It’s that a solo developer now operates in the regime that used to require an infrastructure organization, with none of the infrastructure and no headcount to build it. Agents removed the cap on my arrival rate. They didn’t hand me Google’s build system. The organizational fix is worth naming, because it isn’t subtle. Large companies make every developer’s environment a faithful copy of the one that builds and ships the product, either by moving development onto cloud machines provisioned from the same definition as CI, or by reproducing that environment in miniature on the laptop: a container per service, a database per test, a toolchain pinned to a lockfile. Somebody owns that, full time. The cost of getting it wrong is paid once, by a platform team, instead of every day by everyone else. I don’t have a platform team, and neither does anyone else building alone with agents. An independent developer can now reach a working prototype absurdly fast, faster than at any point in my career, and then hit the wall this whole post is describing: one environment, artifacts that assume one writer, and no organization to absorb the difference. Getting off the ground is close to solved. Staying up is not, and I don’t have a good answer for it. That’s the same argument I made a few months ago about civil engineering coming at it from the other side. When construction gets cheap, design is what fails, and the design here is the substrate: what runs where, what is isolated from what, and what has to be serialized. An agent isn’t going to do that for you, and it won’t show up in the prototype. Partial isolation is then its own trap. A worktree gives you the feeling of a private workspace (clean tree, own branch, no file collisions), and it reads as properly parallel right up until two private workspaces write the same database row. Nothing errors. Nothing warns. The contention surfaces later, somewhere else, as a red X on an unrelated pull request. Worse, the isolation below git is advisory, and agents have to choose it. Mine routinely don’t. One incident reads “PR 1868 isolated E2E attempt fell back to shared ports”: the agent tried to isolate its test environment, isolation failed, and nothing stopped the run. Worse still, a worktree is cut from a commit and stays there, so it’s isolated in the past and decays with every merge. When I changed one of the agent gates in July (they’re tracked scripts in the repository, not .git/hooks ` , so every worktree carries its own copy pinned to the commit it was cut from) the fix merged at 1:48 AM, and four hours later 73 of 74 worktrees were still running the old one. A stale worktree can’t detect its own staleness, so it goes green about a world that
no longer exists, and the error is deferred to the only actor holding current ` , which is CI. Convergence and Invariant Preservation One framing before the specifics, because it unifies them. Every merge mechanism here is built to terminate, and none of them is built to preserve invariants. Git is honest about the cases it cannot decide: on a textual conflict it halts and asks you. The trouble is the far more common case, where it doesn’t halt, produces a merge confidently, and the invariant it never knew about is now false. That distinction is the oldest lesson in the replicated-data literature, and it has a canonical counterexample. Take a replicated map where each field merges independently under its own perfectly reasonable rule. One field holds a person’s name; another holds the length of that name. Two replicas concurrently write different names. Each field converges exactly as specified, and the result is a record whose ` came from one replica and whose ` came from the other, with the invariant tying them now false. Nothing merged incorrectly. The composition of correct local merges is simply not a correct global merge. Closing that gap is the point of work like Balegas and colleagues’ Indigo which enforces application invariants over eventually consistent stores, and which I have written about before Some of what follows has that shape. The first case does not, and I will not dress it up. The Sealed Prefix This project has 1,388 migrations, and rebuilding a test database from all of them on each of eight parallel test machines is as slow as it sounds. So, in response to my complaining for weeks that CI was too expensive, an agent froze a database snapshot into the repository for CI to restore, applying only what came after. It’s a good optimization: it saves 13 to 18 minutes of machine time per run. It merged at 5:01 AM Eastern on the third day. Fifty-four minutes later the first pull request failed, because its migration no longer sorted after the newly frozen prefix. Then another. By late morning a single incident covers four at once: “Four queued PRs carried migrations older than the sealed CI baseline suffix.” The diagnosis here is ordinary, and a critic will say so. Other agents had already opened pull requests that appended migrations; one agent then sealed a new prefix underneath them; those two streams of work conflicted; and nothing anywhere in the system noticed until CI rejected the queued pull requests several hours later. That’s parallel work colliding and finding out late, not a subtle invariant bug. Caching a prefix of an ordered log made the collision loud and retroactive, but it did not invent the collision. So the loop closes: I complained CI was expensive, an agent made CI cheaper, and the mechanism became a new source of CI failures against the whole queue. Merge Functions and Runtime State Git detects conflicts over lines of text; my conflicts live in shared runtime state. From that week: “Song-call E2E suites deleted each other’s shared pending call.” Two test files, no overlapping lines, both green alone, and the row one of them depends on is the row the other deletes. The collision is in the database at runtime, not in the diff. Those two files are the ` and the ` . They share no lines, each merges cleanly, each is individually green, and the invariant binding them is false the moment they land together. No improvement to git’s merge algorithm catches that, because the property is not a property of any
file: it is a property of the composition, and git has no representation of the composition to check. Verification Under Composition Which gives the sharpest version of the problem: green(A on base) and green(B on base) does not imply green(merge of A and B). The test suite is the only thing in my pipeline that checks the invariant at all, since git checks text and timestamps check ordering. So the tests are my invariant checker. I do run them on the composition. Pull requests rebase onto current ` and merge in sequence, each one retested before it lands, and ` itself is checked after land. That catches the break. It just catches it late, and only by paying the serial tax: rebase, retest, merge, next. That tax is exactly the queue-wide cost the sealed prefix made visible on the morning when one seal invalidated everything queued behind it. Serialization is the honest answer, and I already pay it. Speculative merge queues try to buy the same guarantee back with parallelism: build the candidate futures, main+A ` , main+A+B ` , main+A+B+C ` , test them concurrently, and discard a failure from the middle. OpenStack’s Zuul has gated that way since 2012; GitHub’s merge queue ships a version; and bors which lands Rust, does the cheaper batch-and-bisect variant. It works, and it works by brute force, which I will come back to. Shared Environments One port range, one development database, one mock server. I do have a script assigning each worktree its own ports by hashing the directory name (the right idea), and it hashes into 99 slots. With dozens of worktrees on disk the birthday math makes a collision effectively certain, and two agents get handed the same port whenever the colliding pair happens to be running at once. Undivided Work Trivial parallelism doesn’t work on a problem that was never split into non-conflicting units. The agents were pointed at one jammed queue and found their own boundaries, which were mostly the same boundaries. What looked like a coordinator was a process manager: spawn, wait, list, follow up, six hundred and thirty-three tool calls of overhead, all of it without any deep context about which pieces of work actually interacted. A coordinator that only manages agents can’t divide work it does not understand, so the children rediscover the same failures, edit the same new scripts, and invalidate one another’s assumptions. Codex said as much afterward: the agents generated work for the coordinator instead of removing work from the critical path. Classified by root cause, 245 of the 361 entries collapse into ten systemic problems, 73 are one-off product bugs, and 43 I could not classify. Here are the ten, because a taxonomy whose majority is invisible isn’t a taxonomy: | systemic root cause | incidents | | --- | --- | | tooling the session was writing that same day | 87 | | no per-agent workspace isolation | 33 | | errors skipped, swallowed, or reported as success | 32 | | affected-spec selection far too broad | 20 | | migration order versus the sealed baseline | 16 | | tests leaning on shared fixture state | 16 | | recovery work that would not converge | 12 | | CI cost structure | 12 | | evidence produced against a base that moved | 10 | | verification run without its dependencies | 7 | I would defend the shape of that table rather than the ratio: workers under a shallow coordinator rediscover the same things, often enough to dominate a log. Look down the column and the majority is exactly that failure. The top row