韩国企业引入海外AI模型遭遇Token成本高昂难题,三星实施严格配额制
随着韩国大型企业加速将海外大模型引入业务流程,以三星、SK集团为代表的巨头正面临AI词元(Token)使用成本快速攀升的挑战。为控制开支,三星在向员工开放Claude、Gemini和ChatGPT等美国模型时,实施了严格的Token配额制度,员工需证明AI提升工作效率才能申请更高额度,此举引发部分员工不满。尽管Claude因编程能力在韩国开发群体中广受欢迎,但半导体等数据敏感行业对AI部署仍极为谨慎,SK海力士至今主要依赖自研模型,而三星芯片部门基本禁止使用外部大模型。此外,中小企业因缺乏管理体系,只能通过公布Token消耗排名等方式隐性限制使用。目前,Naver、Nexon等公司已大规模部署Claude,但整体成本压力正倒逼企业探索AI应用与运营成本的平衡之道。 #韩国 #AI #大模型 #Token #三星 #企业应用 #成本控制
随着韩国大型企业加速将海外大模型引入业务流程,以三星、SK集团为代表的巨头正面临AI词元(Token)使用成本快速攀升的挑战。为控制开支,三星在向员工开放Claude、Gemini和ChatGPT等美国模型时,实施了严格的Token配额制度,员工需证明AI提升工作效率才能申请更高额度,此举引发部分员工不满。尽管Claude因编程能力在韩国开发群体中广受欢迎,但半导体等数据敏感行业对AI部署仍极为谨慎,SK海力士至今主要依赖自研模型,而三星芯片部门基本禁止使用外部大模型。此外,中小企业因缺乏管理体系,只能通过公布Token消耗排名等方式隐性限制使用。目前,Naver、Nexon等公司已大规模部署Claude,但整体成本压力正倒逼企业探索AI应用与运营成本的平衡之道。 #韩国 #AI #大模型 #Token #三星 #企业应用 #成本控制
主流AI模型政治倾向测试结果一致偏左
一个名为Unslop的小型研究实验室近日公布了一项实验,对16个主流AI大模型进行了政治倾向测试。测试采用经典的Political Compass问卷,涉及62个经济与社会政策问题。每个模型分别进行了标准版、反向措辞版和乱序版共数十次运行。结果显示,除xAI的Grok模型表现不稳定外,包括GPT系列、Claude系列、Gemini、Llama、DeepSeek、Qwen、Kimi、GLM等在内的所有模型均一致落在“自由意志左派”象限,展现出支持社会平等、反资本主义等进步倾向。研究作者坦言测试并非完全严谨,但结果反映了当前主流AI在政治价值观上存在系统性偏向,与驱动其发展的资本主义背景形成反差。 #AI #大模型 #政治倾向 #研究 #GPT #Claude #Grok #DeepSeek #科技新闻
一个名为Unslop的小型研究实验室近日公布了一项实验,对16个主流AI大模型进行了政治倾向测试。测试采用经典的Political Compass问卷,涉及62个经济与社会政策问题。每个模型分别进行了标准版、反向措辞版和乱序版共数十次运行。结果显示,除xAI的Grok模型表现不稳定外,包括GPT系列、Claude系列、Gemini、Llama、DeepSeek、Qwen、Kimi、GLM等在内的所有模型均一致落在“自由意志左派”象限,展现出支持社会平等、反资本主义等进步倾向。研究作者坦言测试并非完全严谨,但结果反映了当前主流AI在政治价值观上存在系统性偏向,与驱动其发展的资本主义背景形成反差。 #AI #大模型 #政治倾向 #研究 #GPT #Claude #Grok #DeepSeek #科技新闻
Mealyu AI营养健身教练上线,拍照即可分析餐食
Mealyu是一款基于AI的营养与健身教练应用,用户只需拍摄餐食照片,系统即可快速估算热量、蛋白质、碳水化合物、脂肪及纤维等营养成分,并给出质量评分。应用还提供主动式AI教练功能,根据用户目标实时提醒饮食与补水调整,自动计算每日营养目标,并预测达成减重、增肌或身体重塑等目标的具体日期。此外,Mealyu能根据用户水平、场地和设备生成个性化训练方案,支持进度追踪与周期性报告。目前该应用提供免费基础版和高级订阅版(约12.42欧元/月或149欧元/年),高级版可解锁无限照片分析、24/7 AI教练及个性化训练课程。自上线以来,应用已收获超过2300条用户评价,平均评分4.9星,多名用户反馈其帮助显著提升了饮食管理和训练效果。 #AI #营养 #健身 #健康 #科技 #应用 #个性化 #饮食管理 #运动
Mealyu是一款基于AI的营养与健身教练应用,用户只需拍摄餐食照片,系统即可快速估算热量、蛋白质、碳水化合物、脂肪及纤维等营养成分,并给出质量评分。应用还提供主动式AI教练功能,根据用户目标实时提醒饮食与补水调整,自动计算每日营养目标,并预测达成减重、增肌或身体重塑等目标的具体日期。此外,Mealyu能根据用户水平、场地和设备生成个性化训练方案,支持进度追踪与周期性报告。目前该应用提供免费基础版和高级订阅版(约12.42欧元/月或149欧元/年),高级版可解锁无限照片分析、24/7 AI教练及个性化训练课程。自上线以来,应用已收获超过2300条用户评价,平均评分4.9星,多名用户反馈其帮助显著提升了饮食管理和训练效果。 #AI #营养 #健身 #健康 #科技 #应用 #个性化 #饮食管理 #运动
Our toolchain assumes one human writer, AI agents break that illusion
Christopher Meiklejohn One Writer Our tools assume one writer, and assume that writer is a human. Nothing computes what a change reads and writes at runtime, so the only known fix is brute force priced for organizations. 27 Jul 2026 In this blog post, I discuss three days in July 2026 when a single agent session ran away from me, and what those three days revealed about the concurrency assumptions buried in our development tooling. Six weeks ago I wrote The Test Suite Was the Incident my test suite had grown a pile of shared data nobody owned, every pull request paid to rebuild it, and the resulting failures had nothing to do with the changes under review. That cost me about $180 in one night. I got a worse one. It lasted three days, and in one twenty-four-hour stretch of it I burned through an entire Codex 20x
Christopher Meiklejohn One Writer Our tools assume one writer, and assume that writer is a human. Nothing computes what a change reads and writes at runtime, so the only known fix is brute force priced for organizations. 27 Jul 2026 In this blog post, I discuss three days in July 2026 when a single agent session ran away from me, and what those three days revealed about the concurrency assumptions buried in our development tooling. Six weeks ago I wrote The Test Suite Was the Incident my test suite had grown a pile of shared data nobody owned, every pull request paid to rebuild it, and the resulting failures had nothing to do with the changes under review. That cost me about $180 in one night. I got a worse one. It lasted three days, and in one twenty-four-hour stretch of it I burned through an entire Codex 20x
max plan. This post is not really about that, though. It is about a property of our tooling that the three days made impossible to ignore. Nearly every layer of this assumes one writer, and assumes that writer is a human. Git hands you a conflict and waits. Code review assumes somebody reads. A migration sequence assumes somebody is assigning the order. Each of those protocols terminates in a person. That is fine while there is exactly one, and while they are, in fact, a person. The agent runtime turns out to be on that list too, which I did not expect. It spawned eighty-four workers into a single checkout without being able to say what any one of them would read or write, and that is the same question git can’t answer about a diff. That assumption was invisible for forty years because nothing ever bound it at my scale. Agents break both halves at once: there are many of them, and not one of them is the person the protocol was waiting for. Git is the partial exception, and I will come to why the exception does not help. Nothing in the stack detects the violation when it happens. It gets caught later, somewhere else, attributed to the wrong change, and paid for at full price. Some context for readers arriving fresh. Zabriskie is a social app for live-music fans, and it’s also a deliberate experiment: I’m building a real, deployed, actually-used application almost entirely with AI agents (agents that wrote the features, agents that wrote the tests guarding those features, and agents that now open most of the pull requests), in order to find out what that’s like and, more usefully, where it breaks. I’ve written almost none of the code. That framing matters, because several things below look like obvious mistakes and are. I let agents design a migration scheme with only another agent reviewing it, I stopped reading most diffs, and I let sixty-four pull requests go up in a single day, all of which a careful engineer would tell you not to do and would be right about. But the point of running an experiment at the extreme is to find the walls. That week I found several at once. A migration, throughout, is a versioned SQL file that changes the database schema. CI is the automated checking that runs on every proposed change: build the app, spin up a fresh database, run the tests. Here is the shape of the three days. Treat these numbers as texture, not as evidence, for a reason I will get to. | | 24 Jul | 25 Jul | 26 Jul | | --- | --- | --- | --- | | pull requests opened | 19 | 64 | 28 | | pull requests merged | 20 | 50 | 30 | | incidents logged | 0 | 6 | 301 | Days in that table are UTC; clock times in the narrative below are Eastern (where I was). The burst ran past midnight UTC: 301 incidents on the 26th plus 60 more before 1 AM on the 27th, so 361 for the burst. Every incident count below is scoped to that window, and 253 rows is the total before 24 July. I offer that comparison as a sense of scale and not as a baseline, for the following reason. Those incidents exist because a standing instruction tells agents to log their own mistakes, and ten minutes into the worst night I tightened that instruction. The log therefore measures reported failures. Look at the daily series and it gets worse: there are days that week with ten and twenty merged pull requests and zero logged incidents, which at any real failure rate means nobody was logging rather than nothing broke. In short, I can’t give you a trustworthy baseline. What follows rests on
mechanism and on a few dated, checkable events, not on 361. The Session Late Saturday night the queue jammed. Sixty-four pull requests had gone up that day, ` was red, and nineteen open pull requests were stuck behind a suite that could not tell me which of them was broken. At 12:31 AM I opened a session with Codex and complained that CI was wasting too much money. Codex read that and hired a workforce. Over seventeen hours that session made 81 spawn calls, producing 74 direct children; those children spawned 10 more, for 84 threads. Then 358 calls waiting on them, 96 listing them, 85 sending follow-up work, and 13 interrupting them, which comes to six hundred and thirty-three tool calls of pure management overhead (spawn, wait, list, follow up, interrupt) against 54 messages from me over the same seventeen hours. All 84 ran in the same checked-out copy of the repository. Trivial parallelization, at scale, with a coordinator that only managed agents, on a toolchain built for one writer. The rest of this post is about why that combination is so much worse than it sounds. I asked afterward why it had spawned anything, since I never requested it. A language model asked why it did something produces plausible narrative, not introspection, and I apply that same skepticism to the self-reported incidents above. So the quotes below are not evidence, and I’m not offering them as any. I had already reached the same conclusion from the tool-call counts before I asked. I print them because they state that conclusion more plainly than I did, and because there is something worth looking at in a system that can describe the failure this precisely and could not avoid it. I chose to spawn multiple agents even though you did not ask me to. The environment permitted proactive delegation, but that was permission, not a requirement. I treated the availability of agents as a reason to use them instead of first asking whether they would reduce time, cost, or risk. It then diagnosed why the task was a bad candidate: A PR queue is largely serial because every merge changes ` , which changes the integration state of every remaining PR. The work needed one authoritative coordinator moving through the queue in order. Parallel agents could not independently merge overlapping PRs without continuously invalidating one another’s assumptions. On what those 358 waits were doing: The agents’ outputs created additional coordination work for me. I had to read their findings, compare conflicting recommendations, inspect supporting evidence, and decide how to combine them. The agents therefore generated work for the coordinator instead of reliably removing work from the critical path. A system with no notion of its own concurrency was handed a concurrency primitive and used it, reasonably, on a task the substrate could not support. What that substrate is, and why nothing in it objected, is the rest of this post. One number deserves care. I do not know how many of those 84 threads wrote to the checkout instead of reading it; from Codex’s own account many were investigating. A reader is harmless and a writer is not, and nothing in the toolchain drew the distinction or could. That’s the same missing primitive one level up: not even the agent runtime knew which of its children were writers. The Substrate Start with the layer that did anticipate this, because it is the one people reach for. Branches are optimistic concurrency control; worktrees go further, giving each writer a