GPT-5.6 成为微软 365 Copilot 首选模型,办公效率再升级
微软宣布,GPT-5.6 已正式成为 Microsoft 365 Copilot 的默认搭载模型,为 Word、Excel、PowerPoint 以及 Chat 和协作(Cowork)功能提供更强大的 AI 支持。升级后的 Copilot 在文本生成、数据分析、幻灯片制作与实时协作中表现出更快的响应速度和更高的内容质量,帮助用户大幅提升工作效率。微软表示,GPT-5.6 的引入标志着企业级 AI 助手进入新阶段,未来将逐步向更多计划开放。 #微软 #GPT5.6 #Copilot #AI #办公效率 #智能助手 #科技新闻
微软宣布,GPT-5.6 已正式成为 Microsoft 365 Copilot 的默认搭载模型,为 Word、Excel、PowerPoint 以及 Chat 和协作(Cowork)功能提供更强大的 AI 支持。升级后的 Copilot 在文本生成、数据分析、幻灯片制作与实时协作中表现出更快的响应速度和更高的内容质量,帮助用户大幅提升工作效率。微软表示,GPT-5.6 的引入标志着企业级 AI 助手进入新阶段,未来将逐步向更多计划开放。 #微软 #GPT5.6 #Copilot #AI #办公效率 #智能助手 #科技新闻
AI复盘200期播客:押中美光180%涨幅,错过Cursor 60亿收购
Limitless Podcast 在200期纪念中,将全部文字稿输入AI分析,验证了早期投资判断。节目组曾押注“算力即国家安全”,美光股价因此上涨180%,SK海力士成为韩国市值最大公司,三星利润超越英伟达。但遗憾错过Cursor,后者被SpaceX以600亿美元收购。AI还发现,行业焦点从宏大概念(如AGI、超级智能)转向实际应用,提及频率分别下降90%和54%,而能源瓶颈成为新天花板。内存制造商平均回报达153%,最佳模拟投资组合回报4倍,Valor Atomics因解决数据中心能源短缺获得13倍回报。未来三大趋势:太空训练模型、垂直领域AI、边缘端本地模型。看多与看空言论比例为2.8:1,反映极度乐观态度。 #AI #播客 #投资 #算力 #内存 #能源 #科技趋势
Limitless Podcast 在200期纪念中,将全部文字稿输入AI分析,验证了早期投资判断。节目组曾押注“算力即国家安全”,美光股价因此上涨180%,SK海力士成为韩国市值最大公司,三星利润超越英伟达。但遗憾错过Cursor,后者被SpaceX以600亿美元收购。AI还发现,行业焦点从宏大概念(如AGI、超级智能)转向实际应用,提及频率分别下降90%和54%,而能源瓶颈成为新天花板。内存制造商平均回报达153%,最佳模拟投资组合回报4倍,Valor Atomics因解决数据中心能源短缺获得13倍回报。未来三大趋势:太空训练模型、垂直领域AI、边缘端本地模型。看多与看空言论比例为2.8:1,反映极度乐观态度。 #AI #播客 #投资 #算力 #内存 #能源 #科技趋势
GPT5.6引爆价格战成本暴跌| AI快讯详情
据God of Prompt称:Luna每百万词$1,Sonnet5入门$2,Terra较GPT5.5便宜一半。 GPT-5.6的发布将人工智能竞争从能力竞赛转向价格战。根据God of Prompt的讨论,Luna每百万令牌仅需1美元,Sonnet 5入门版为2美元,Terra价格比GPT-5.5低一半。这标志着智能成本大幅下降,为企业提供更实惠的AI接入机会。 价格下降反映了AI基础设施的成熟,硬件和算法改进带来的效率提升让提供商在不牺牲质量的情况下降低费率。Luna的定价针对数据处理高量用户,Terra的策略激励成本敏感企业迁移。 降低的智能成本解锁了将AI嵌入SaaS产品的规模化盈利策略。初创公司可推出AI驱动工具,促进与大型企业的竞争。未来展望显示AI基础模型进一步商品化,2027年前价格持续下降。 GPT-5.6的发布促使竞争对手通过激进定价争夺推理服务市场份额。 小企业获得先进模型的实惠接入,能自动化流程并有效竞争。 潜在风险包括性能不一致,企业可通过试点测试缓解。 An AI prompt engineering specialist sharing practical techniques for optimizing large language models and AI image generators. The content features prompt design strategies, AI tool tutorials, and creative applications of generative AI for both beginners and advanced users.
据God of Prompt称:Luna每百万词$1,Sonnet5入门$2,Terra较GPT5.5便宜一半。 GPT-5.6的发布将人工智能竞争从能力竞赛转向价格战。根据God of Prompt的讨论,Luna每百万令牌仅需1美元,Sonnet 5入门版为2美元,Terra价格比GPT-5.5低一半。这标志着智能成本大幅下降,为企业提供更实惠的AI接入机会。 价格下降反映了AI基础设施的成熟,硬件和算法改进带来的效率提升让提供商在不牺牲质量的情况下降低费率。Luna的定价针对数据处理高量用户,Terra的策略激励成本敏感企业迁移。 降低的智能成本解锁了将AI嵌入SaaS产品的规模化盈利策略。初创公司可推出AI驱动工具,促进与大型企业的竞争。未来展望显示AI基础模型进一步商品化,2027年前价格持续下降。 GPT-5.6的发布促使竞争对手通过激进定价争夺推理服务市场份额。 小企业获得先进模型的实惠接入,能自动化流程并有效竞争。 潜在风险包括性能不一致,企业可通过试点测试缓解。 An AI prompt engineering specialist sharing practical techniques for optimizing large language models and AI image generators. The content features prompt design strategies, AI tool tutorials, and creative applications of generative AI for both beginners and advanced users.
Show HN: LocalClip – A local-first AI video clipper for Mac
LocalClip local Studio it works 100% local · private · free forever One recording. Weeks of content. Stop uploading hours of video. LocalClip is an AI video studio that runs directly on your Mac — it finds your best moments and turns them into vertical clips with subtitles, titles and hashtags. No cloud. No uploads. No per-minute pricing. Download for Mac how it works Free · Apple Silicon · runs on your GPU (mlx) localhost:8000 · LocalClip Live Q&A — Tuesday stream (1h 42m) up to 8 clips of 20–60s Creating clips Queued Subtitles Best moments 4 Creating clips 0:38 this is the part nobody tells you The part nobody tells you advice #real 0:27 I tried it for 30 days straight 30-day experiment results challenge #results 0:45 the biggest mistake beginners make Biggest beginner mistake tips #howto Rendering clip 4… Why loca
LocalClip local Studio it works 100% local · private · free forever One recording. Weeks of content. Stop uploading hours of video. LocalClip is an AI video studio that runs directly on your Mac — it finds your best moments and turns them into vertical clips with subtitles, titles and hashtags. No cloud. No uploads. No per-minute pricing. Download for Mac how it works Free · Apple Silicon · runs on your GPU (mlx) localhost:8000 · LocalClip Live Q&A — Tuesday stream (1h 42m) up to 8 clips of 20–60s Creating clips Queued Subtitles Best moments 4 Creating clips 0:38 this is the part nobody tells you The part nobody tells you advice #real 0:27 I tried it for 30 days straight 30-day experiment results challenge #results 0:45 the biggest mistake beginners make Biggest beginner mistake tips #howto Rendering clip 4… Why loca
l Stop uploading hours of video Cloud tools make you upload multi-GB files, wait in a queue and pay for every processed minute — because they pay for the GPUs. LocalClip runs on the hardware you already own. Faster processing No upload, no queue. Transcription and rendering start instantly on your Mac's GPU (mlx / Apple Silicon). Privacy by design Your footage never touches a server. Nothing to upload, nothing stored in someone else's cloud. Fixed cost No per-minute billing, no per-clip credits. Because it runs on your device, we don't pay for your compute. Unlimited clips Render as many clips from as many videos as you want. Your hardware is the only limit — and it's idle anyway. AI Studio The clipper is just the beginning LocalClip is a local video-understanding engine. It watches your recording once — and turns it into everything you need to publish. 🔴 Live stream🎙️ Podcast💻 Zoom / Meet🖥️ Webinar🎬 Any video AI understands your content Transcribes, detects the best moments — 100% on your Mac Viral clips Today Word-by-word subtitles Today Titles & hashtags Today Chapters Soon Summary & blog post Soon LinkedIn / X posts Soon How it works From long video to clips in 3 steps Drop a file, let it run, post the results. That's the whole workflow. Drop your video A recorded video or a saved livestream — pick a file, no upload needed. AI does the work It transcribes, finds the strongest moments and cuts vertical 9:16 clips with word-by-word subtitles. Post everywhere The difference LocalClip vs cloud tools Same output. None of the uploading, waiting or paying per minute. LocalClip Cloud tools Runs on your GPU (free)— Video stays private— No upload wait— Fixed cost — no per-minute billing— Word-by-word subtitles✓ Finish a live. Have a week of content. Get the beta for macOS (Apple Silicon). Drop your email and we'll send you the download + updates. Get the DMG ⬇ Download for Mac (Apple Silicon) Free · ~615 MB · unsigned beta — right-click → Open the first time ⚠️ LocalClip needs a Mac with Apple Silicon (M1 or newer). Your Mac looks like Intel, so it won't run — write to us if you need an Intel build. How to install 1. Open the DMG and drag LocalClip to Applications. Easiest (no Terminal) Double-click LocalClip → on the “Apple could not verify…” warning, open System Settings → Privacy & Security → scroll down → “Open Anyway” → confirm with your password / Touch ID. Or with Terminal — Run this once, then open it normally: xattr -dr /Applications/ Unsigned beta — nothing is uploaded or collected. Ideas, problems or questions? Tell me what you'd improve or what broke. I read everything and reply. Send LocalClip Your personal AI video studio that never uploads your content. Why local · AI Studio · How it works · Updates · Feedback Created byLou Alcala
Show HN: Sell your unused AI Credits or buy Claude credits for 50% off
Second Hand Tokens AI tokens at 50% off the price Developers buy more tokens than they use. You buy the leftovers at 50% off. Get API key your unused AI credits Copy ` from openai import OpenAI client = OpenAI( base_url=" api_key="YOUR_API_KEY", ) Standard response = ( model="claude-sonnet-4-6", messages=[{"role": "user", "content": "Hello!"}], ) print( [0]. ) ` ! : Claude Haiku 4.5 logo Claude Sonnet 4.6 Anthropic −50% claude-sonnet-4-6 Input / 1M tokens $3.00$1.50 Output / 1M tokens $15.00$7.50 Get API key ! : Claude Haiku 4.5 logo Claude Opus 4.6 Anthropic −50% claude-opus-4-6 Input / 1M tokens $5.00$2.50 Output / 1M tokens $25.00$12.50 Get API key ! : Claude Haiku 4.5 logo Claude Haiku 4.5 Anthropic −50% claude-haiku-4-5 Input / 1M tokens $1.
Second Hand Tokens AI tokens at 50% off the price Developers buy more tokens than they use. You buy the leftovers at 50% off. Get API key your unused AI credits Copy ` from openai import OpenAI client = OpenAI( base_url=" api_key="YOUR_API_KEY", ) Standard response = ( model="claude-sonnet-4-6", messages=[{"role": "user", "content": "Hello!"}], ) print( [0]. ) ` ! : Claude Haiku 4.5 logo Claude Sonnet 4.6 Anthropic −50% claude-sonnet-4-6 Input / 1M tokens $3.00$1.50 Output / 1M tokens $15.00$7.50 Get API key ! : Claude Haiku 4.5 logo Claude Opus 4.6 Anthropic −50% claude-opus-4-6 Input / 1M tokens $5.00$2.50 Output / 1M tokens $25.00$12.50 Get API key ! : Claude Haiku 4.5 logo Claude Haiku 4.5 Anthropic −50% claude-haiku-4-5 Input / 1M tokens $1.
00$0.500 Output / 1M tokens $5.00$2.50 Get API key ! : Llama 3 70B Instruct logo Llama 4 Maverick 17B Meta −50% llama-4-maverick-17b Input / 1M tokens $0.240$0.120 Output / 1M tokens $0.970$0.485 Get API key ! : Llama 3 70B Instruct logo Llama 3 70B Instruct Meta −50% llama-3-70b-instruct Input / 1M tokens $2.65$1.32 Output / 1M tokens $3.50$1.75 Get API key ! : DeepSeek R1 logo DeepSeek V3.2 DeepSeek −50% deepseek-v3.2 Input / 1M tokens $0.580$0.290 Output / 1M tokens $1.68$0.840 Get API key ! : DeepSeek R1 logo DeepSeek R1 DeepSeek −50% deepseek-r1 Input / 1M tokens $1.35$0.675 Output / 1M tokens $5.40$2.70 Get API key How it works Three steps to cheaper AI tokens. 01 Pick a model Choose from Claude, GPT-4o, Gemini, and more. See live second-hand pricing. 02 Get your API key We route your requests through pre-funded seller accounts. Drop-in compatible with standard SDKs. 03 Pay half price You're billed at 50% of retail. Sellers recover unused credits. Everyone wins. Second Hand Tokens © 2026
Building a real-time AI tutor for 5-year
We set out to build the first AI tutor to teach math and reading to kids ages 4-9. For AI to actually teach a five-year-old, pedagogy must be baked into the engineering. A child can't wait for a slow reply, can't read a chat interface, and can't unhear anything a model gets wrong. We wanted to share some of the learnings that shaped our architectural decisions building a real-time AI tutor. A 2-second pause in conversation feels different to a child than to a developer, or even to an adult on the phone speaking to an automated agent. A couple of seconds is enough for a child's attention to wander and for learning to stop. Good teachers manage this without pausing to think. They acknowledge a child immediately, even when they hold the answer back to let the child work. Teaching is matching the right approach to the curre
We set out to build the first AI tutor to teach math and reading to kids ages 4-9. For AI to actually teach a five-year-old, pedagogy must be baked into the engineering. A child can't wait for a slow reply, can't read a chat interface, and can't unhear anything a model gets wrong. We wanted to share some of the learnings that shaped our architectural decisions building a real-time AI tutor. A 2-second pause in conversation feels different to a child than to a developer, or even to an adult on the phone speaking to an automated agent. A couple of seconds is enough for a child's attention to wander and for learning to stop. Good teachers manage this without pausing to think. They acknowledge a child immediately, even when they hold the answer back to let the child work. Teaching is matching the right approach to the curre
nt moment, and most approaches aren't answers. When we set out to build an AI tutor for children ages 4-9, we wanted to build a tutor that actually teaches and not just a chatbot that responds quickly. We knew the constraint underneath would be hard, and that it wasn't optional: sub-second response on every turn. Most agents trade off speed for quality through reasoning budgets. Our architecture has to ground the tutor in pedagogy and respond to the child in real-time. A teacher is constantly deciding how to engage a student, whether to say something, draw on the whiteboard, play a game, or change topics entirely. The standard pattern for an agent today is a tool loop. The LLM outputs one or more tool calls, waits for them to execute, observes the results, and decides what to do next. So the straightforward way to build a teaching agent is to make a tool for each action a teacher could take. But the tool loop has a latency problem. Frontier models take 2–3 seconds to produce their first token, then decode at around 30 tokens per second. Our actions average a few dozen tokens. Add round-trip latency and audio playback, and a standard loop means 3-4 seconds of downtime between each sentence or change on the screen. In one of our earlier playtests we watched it happen in real time. A six-year-old boy waited for the agent to think, then asked: Why is he not doing anything? When is this starting. It's boring. Another child in the same round of playtests figured out she only needed to pay attention part of the time and could still keep up. Latency had taught her to tune the tutor out. That was also the moment she stopped learning. The convenient fix would be a smaller, faster model. That's where a scope problem shows up. Teaching is a broad task. A tutor might pick between dozens of actions in a single lesson, and the hardest call is often to withhold the answer and give a hint, ask a smaller question, or let the child struggle just enough that the insight is theirs when it lands. Smaller models struggled to follow instructions across that breadth. An early version of our agent that used one was responsive yet constantly giving the answer away. Every time it did, it took away the moment where the learning happens. So we built a custom harness to balance instruction following, latency, and a flexible action space. The model streams multiple actions in a single response. An interpreter parses and executes each action while the model is still generating the next ones. The child only has to wait for the first action about 30 tokens in, not for the whole response to complete. Separating generation from execution buys us two more things. We can change which actions are available depending on the situation. For instance, when a question is on screen the agent gets instructions and options for scaffolding rather than answering. And we can validate each action without a latency hit on the happy path. Only if the stream produces an invalid action do we interrupt and re-generate, otherwise execution never pauses. None of this is free. Owning the loop means we've had to build our own observability and tracing instead of leaning on a framework. And we're swimming against the current: frontier models are heavily post-trained on the tool-use pattern. If future models get fast enough, our harness is designed to be replaced by the simpler loop. Lesson: Agent frameworks are building toward background work, where the tradeoff between speed and thinking is
easy. Real-time learning sits at the other extreme. Teaching at conversation speed means owning the loop ourselves. A real teacher both reflects on what a student just did and anticipates what they'll do next. Teach the same lesson a hundred times and you see the patterns. But you also know this child, where they've been stalling, what excites them, what's likely to trip them up today. You start the lesson with a plan and adjust it on the fly. We call the agent that interacts with the child the converser . Our early experiments showed that a smaller action space led to better instruction following, so we built a second agent, the planner , to review the conversation against the lesson's objectives and manage the converser's context. The first version ran synchronously, which of course was too slow. Plans that expired after a fixed number of turns weren't reliable. Neither was having the converser ask for a new plan. What worked was an asynchronous planner that runs while the child is thinking or talking, the same way a teacher reflects and anticipates in the gaps of a conversation. Those gaps are where the judgment calls get made: challenge the child or let them succeed, stay on the concept or move on. A teacher makes them on intuition; a model has to reason its way there, and running async is what buys it the time. Async also means two agents running at once, both reading and writing shared state without coordinating. So we store every turn, every tap, and every UI update as an immutable event on an append-only log. Either agent reads and appends without waiting on the other. That trajectory format enables another kind of anticipation. Whenever the converser asks a closed-ended question (i.e. coming up with a fill in the blank question, playing I Spy, completing an equation etc.), the harness hypothesizes the child's likely answers and pre-generates a response to each one on its own branch, forked from the trajectory. When the child answers, we match it to a branch and play the response without waiting on a fresh model call. The tradeoff is cost, and the occasional miscall. The planner runs on a more capable, more expensive model, and it runs on every turn. And a prediction is still a prediction. Sometimes a child who was ready to be pushed gets handed an easy win instead. It's harder to evaluate whether a converser error was a mistake or downstream of a faulty plan. We don't yet have a clean signal for when to trust the plan versus what's happening live in the moment. Lesson: The child interacts with the app in real time while the agents run in discrete generations, so leverage the time the child thinks or talks. Let the planner reflect the past and anticipate the future while the converser handles the present, and the slow pedagogical reasoning happens concurrently with the real-time exchange. When the next move is predictable, generate it before the child even answers. Most AI products build guardrails in serial with a model call or agent turn. A user won't notice when the token stream goes through a content filter and a developer is willing to wait for a CLI tool call to be auto-reviewed. There's nowhere to hide in a real-time conversation with a five-year-old. Nor is there an undo: a child can't unhear what the tutor said. The safety system has to gate any action, on every turn. Our safety classifier is an LLM that takes ~500-1000ms to run. Waiting to run the converser until that check completes adds a second of delay to every
turn that we can't afford. Here’s another advantage of decoupling generation from execution in our harness. The safety classifier blocks execution without blocking generation. As soon as the child finishes speaking, we dispatch both the classifier and a small model to generate the converser's first action in parallel. That model reacts quickly with an eager response that mirrors or acknowledges what the child said ("you like dinosaurs! me too"). While a rules-based check would be faster and cheaper, it wouldn't survive the ways a five-year-old actually talks. Every category we add to the safety policy adds tokens and requires re-tuning a non-deterministic classifier. Sometimes a transcription error spooks the classifier and triggers a false positive. We review these cases and use them to improve how the agent understands the child. By the time that eager action has generated, the classifier has usually returned safe. That check unblocks the converser to generate while the eager action executes. The child hears one continuous turn despite the multiple model calls. But the harder problem than latency is what to do when that reflexive action is the wrong choice. Mirroring is great for everyday conversation with a child. Other times it's the opposite of what pedagogy suggests. Take a child who mentions, mid-lesson, that a classmate called them a bad name. The same reflex that turns "I like dinosaurs" into "you like dinosaurs! me too" would echo the mean name back to the child. So whenever the safety classifier flags the child's turn, we throw out the eager action. The converser is handled different guidance for this turn: don't repeat the name, acknowledge it must have felt bad, and suggest speaking to a grown-up. Note: Our safety systems are governed by policies developed with child-development experts. How our safety system works in detail will be a separate article in and of itself. Lesson: Gate execution on the safety check rather than generation to prevent a latency hit. Replace the reflexive response with guidance tailored to the child's situation whenever the check fails. You can't build an AI tutor for children by picking the right model, prompting it, and calling it a day. Building an AI tutor for children requires much more. It's about engineering a real-time system that gives enough time to be both factually and pedagogically right — the time to withhold an answer when the answer diminishes learning, the time to choose the next action before the child finishes their thought, the time to second-guess a quick reflex before the child hears it. These pieces seem small in isolation, but the magic isn’t until you have all these pieces work together that you a tutor that thinks ahead, recovers gracefully, and feels like it's with the child, and not catching up to them. If these problems sound interesting to you, or if you want to learn more about what it actually takes to build a real-time AI tutor for kids, we'd love to talk. We're hiring.