Anthropic Art:生成满背景手绘编辑插画的 Agent Skill
这是一个非官方 Agent Skill,可把抽象主题转换成简洁视觉隐喻,并按已核验的配色、象牙白不规则承载形和粗黑手绘线条生成满背景编辑插画。仓库包含 Codex 安装方式、可复用提示词、风格规范及地球/太阳/月亮示例。
Anthropic Art: Full-Background Hand-Drawn Editorial Illustration Skill
An unofficial Agent Skill that turns abstract themes into simple visual metaphors and generates full-background editorial illustrations using a verified palette, irregular ivory carrier shapes, and bold hand-drawn lines. The repository includes Codex installation steps, reusable prompts, style rules, and earth/sun/moon examples.
https://github.com/HalfAI1102/anthropic-art
这是一个非官方 Agent Skill,可把抽象主题转换成简洁视觉隐喻,并按已核验的配色、象牙白不规则承载形和粗黑手绘线条生成满背景编辑插画。仓库包含 Codex 安装方式、可复用提示词、风格规范及地球/太阳/月亮示例。
Anthropic Art: Full-Background Hand-Drawn Editorial Illustration Skill
An unofficial Agent Skill that turns abstract themes into simple visual metaphors and generates full-background editorial illustrations using a verified palette, irregular ivory carrier shapes, and bold hand-drawn lines. The repository includes Codex installation steps, reusable prompts, style rules, and earth/sun/moon examples.
https://github.com/HalfAI1102/anthropic-art
GitHub
GitHub - HalfAI1102/anthropic-art: Generate full-background Anthropic-style editorial illustrations with verified palette and hand…
Generate full-background Anthropic-style editorial illustrations with verified palette and hand-drawn visual rules. - HalfAI1102/anthropic-art
Claude 5 模型的上下文工程新规则
Thariq 解释 Claude Code 为新一代模型删掉 80% 以上系统提示词却没有可测量的编码评测损失,并总结六个转变:规则转向模型判断、示例转向接口设计、一次性灌入转向渐进披露、重复指令转向简洁工具描述、CLAUDE.md 记忆转向自动记忆、简单规格转向丰富参考资料。
The New Rules of Context Engineering for Claude 5 Models
Thariq explains why Claude Code removed over 80% of its system prompt for newer models with no measurable loss on coding evaluations, replacing older habits with model judgement, expressive interfaces, progressive disclosure, simple tool descriptions, auto-memory, and rich references.
https://x.com/trq212/status/2080710971228918066
Thariq 解释 Claude Code 为新一代模型删掉 80% 以上系统提示词却没有可测量的编码评测损失,并总结六个转变:规则转向模型判断、示例转向接口设计、一次性灌入转向渐进披露、重复指令转向简洁工具描述、CLAUDE.md 记忆转向自动记忆、简单规格转向丰富参考资料。
The New Rules of Context Engineering for Claude 5 Models
Thariq explains why Claude Code removed over 80% of its system prompt for newer models with no measurable loss on coding evaluations, replacing older habits with model judgement, expressive interfaces, progressive disclosure, simple tool descriptions, auto-memory, and rich references.
https://x.com/trq212/status/2080710971228918066
X (formerly Twitter)
Thariq is on vacation (@trq212) on X
The new rules of context engineering for Claude 5 models
Kimi K3:开放前沿智能技术报告
Moonshot AI 的 47 页 Kimi K3 报告介绍一个开放的原生多模态 MoE 模型:2.8 万亿总参数、1040 亿激活参数、100 万 token 上下文,并结合 Kimi Delta Attention、Attention Residuals、Stable LatentMoE、多档推理强度的智能体强化学习和 3T 级训练/部署基础设施。报告称整体缩放效率较 Kimi K2 提升约 2.5 倍,并已开放完整模型权重。
Kimi K3: Open Frontier Intelligence — Technical Report
Moonshot AI's 47-page report presents an open native multimodal MoE model with 2.8T total parameters, 104B activated parameters, a 1M-token context window, Kimi Delta Attention, Attention Residuals, Stable LatentMoE, multi-effort agentic RL, and 3T-class training and serving infrastructure. It reports roughly 2.5× better overall scaling efficiency than Kimi K2 and releases the full model weights.
https://github.com/MoonshotAI/Kimi-K3/blob/main/k3_tech_report.pdf
Moonshot AI 的 47 页 Kimi K3 报告介绍一个开放的原生多模态 MoE 模型:2.8 万亿总参数、1040 亿激活参数、100 万 token 上下文,并结合 Kimi Delta Attention、Attention Residuals、Stable LatentMoE、多档推理强度的智能体强化学习和 3T 级训练/部署基础设施。报告称整体缩放效率较 Kimi K2 提升约 2.5 倍,并已开放完整模型权重。
Kimi K3: Open Frontier Intelligence — Technical Report
Moonshot AI's 47-page report presents an open native multimodal MoE model with 2.8T total parameters, 104B activated parameters, a 1M-token context window, Kimi Delta Attention, Attention Residuals, Stable LatentMoE, multi-effort agentic RL, and 3T-class training and serving infrastructure. It reports roughly 2.5× better overall scaling efficiency than Kimi K2 and releases the full model weights.
https://github.com/MoonshotAI/Kimi-K3/blob/main/k3_tech_report.pdf
GitHub
Kimi-K3/k3_tech_report.pdf at main · MoonshotAI/Kimi-K3
Open Frontier Intelligence. Contribute to MoonshotAI/Kimi-K3 development by creating an account on GitHub.
梁文锋四小时投资人会议实录(完整版,来源未核验)
这份由虎嗅依据网传 42 页语音转写 PDF 整理的记录,涵盖愿景、开源、定价、AGI、Scaling、国产芯片与组织等议题。公开 YouMind 文档当前在 Q6 处中断,且未获梁文锋、DeepSeek 或原始录音核验,请审慎参考。
Liang Wenfeng’s Four-Hour Investor Meeting Transcript (Unverified Source)
Compiled from a circulated 42-page speech-to-text PDF, it covers vision, open source, pricing, AGI, scaling, domestic chips, and organization. The public YouMind document currently cuts off during Q6 and has not been verified against the original recording or confirmed by Liang Wenfeng or DeepSeek.
https://youmind.com/d/qEXsZr2Z52kFGz
这份由虎嗅依据网传 42 页语音转写 PDF 整理的记录,涵盖愿景、开源、定价、AGI、Scaling、国产芯片与组织等议题。公开 YouMind 文档当前在 Q6 处中断,且未获梁文锋、DeepSeek 或原始录音核验,请审慎参考。
Liang Wenfeng’s Four-Hour Investor Meeting Transcript (Unverified Source)
Compiled from a circulated 42-page speech-to-text PDF, it covers vision, open source, pricing, AGI, scaling, domestic chips, and organization. The public YouMind document currently cuts off during Q6 and has not been verified against the original recording or confirmed by Liang Wenfeng or DeepSeek.
https://youmind.com/d/qEXsZr2Z52kFGz
YouMind
梁文锋四小时投资人会议实录(完整版)
以及我们公司的其他同事,我们一开始来做这个公司,初衷是没有想到说我最后要赚多少钱,要到资本市场上去,要上市,要怎么样的,所以我们是没有这个初衷的。 最开始的几十个人完全没有这么想过。如果他这么想,他就不会来。所以总体讲,我们是怀着一个对这个世界非常大的善意来做这个事情,然后我们觉得这是对人类有用的,这是一个金钱以外的…
从 GPT-2 到 Kimi K3:22,580 倍参数增长背后的架构演化
ali 用 22 张图和 17 组代码串联 GPT-2、线性注意力、DeltaNet、Gated DeltaNet、Kimi Delta Attention、混合 MLA/MoE、SiTU 与 Attention Residuals,说明 Kimi K3 的进步不只是扩大参数量,而是围绕记忆写入、遗忘、路由和检索逐步重构架构。
22580: From GPT2 to Kimi3, Explained
Ali traces the path from GPT-2 to Kimi K3 with 22 diagrams and 17 code examples, covering linear attention, DeltaNet, Gated DeltaNet, Kimi Delta Attention, hybrid MLA/MoE, SiTU, and Attention Residuals. The article argues that Kimi K3 is not blind scaling: each step adds a specific mechanism for memory, forgetting, routing, or retrieval.
https://x.com/waterloo_intern/status/2081762065392541951
ali 用 22 张图和 17 组代码串联 GPT-2、线性注意力、DeltaNet、Gated DeltaNet、Kimi Delta Attention、混合 MLA/MoE、SiTU 与 Attention Residuals,说明 Kimi K3 的进步不只是扩大参数量,而是围绕记忆写入、遗忘、路由和检索逐步重构架构。
22580: From GPT2 to Kimi3, Explained
Ali traces the path from GPT-2 to Kimi K3 with 22 diagrams and 17 code examples, covering linear attention, DeltaNet, Gated DeltaNet, Kimi Delta Attention, hybrid MLA/MoE, SiTU, and Attention Residuals. The article argues that Kimi K3 is not blind scaling: each step adds a specific mechanism for memory, forgetting, routing, or retrieval.
https://x.com/waterloo_intern/status/2081762065392541951
X (formerly Twitter)
ali (@waterloo_intern) on X
22580: From GPT2 to Kimi3, Explained
租用智能,拥有记忆:一份宣言
Nowledge Labs 认为,模型、Agent harness 与应用界面正在变成可随时替换的短期租赁品,真正持续增值的是团队在工作中积累的上下文与判断。文章提出 L1–L4 四级记忆模型:从单一工具记住个人,走向带来源签名、时效链与分布式托管的组织记忆层。
Rent the Intelligence. Own the Memory. A Manifesto.
Nowledge Labs argues that models, agent harnesses, and interfaces are short-term rentals, while portable, signed, versioned memory is the asset that compounds. The manifesto maps four levels of memory, from one tool remembering one user to an organization-wide memory layer with distributed custody.
https://x.com/wey_gu/status/2083118261601005612
Nowledge Labs 认为,模型、Agent harness 与应用界面正在变成可随时替换的短期租赁品,真正持续增值的是团队在工作中积累的上下文与判断。文章提出 L1–L4 四级记忆模型:从单一工具记住个人,走向带来源签名、时效链与分布式托管的组织记忆层。
Rent the Intelligence. Own the Memory. A Manifesto.
Nowledge Labs argues that models, agent harnesses, and interfaces are short-term rentals, while portable, signed, versioned memory is the asset that compounds. The manifesto maps four levels of memory, from one tool remembering one user to an organization-wide memory layer with distributed custody.
https://x.com/wey_gu/status/2083118261601005612
X (formerly Twitter)
Wey Gu 古思为 (@wey_gu) on X
Rent the Intelligence. Own the Memory. A Manifesto.
时空可组合性:可撤销副作用与响应式依赖的编程范式
论文将动态组件系统拆成时间可组合性与空间可组合性,用可撤销副作用追踪组件清理,用响应式 coeffect 管理依赖变化,并在 Cordis TypeScript 元框架中实现配置协调与事务式热更新;Koishi 4000+ 插件生态提供真实采用案例。本地归档包含 88 页原始 PDF、完整文字稿与关键页面。
A Programming Paradigm for Spatiotemporal Composability
The paper formalizes dynamic component composition through revertible effects and reactive coeffects, then implements the model in the Cordis TypeScript meta-framework with configuration reconciliation and transactional HMR. The local archive includes the complete 88-page PDF, full text transcript, and key rendered pages.
https://github.com/cordiverse/paper/blob/main/paper.pdf
论文将动态组件系统拆成时间可组合性与空间可组合性,用可撤销副作用追踪组件清理,用响应式 coeffect 管理依赖变化,并在 Cordis TypeScript 元框架中实现配置协调与事务式热更新;Koishi 4000+ 插件生态提供真实采用案例。本地归档包含 88 页原始 PDF、完整文字稿与关键页面。
A Programming Paradigm for Spatiotemporal Composability
The paper formalizes dynamic component composition through revertible effects and reactive coeffects, then implements the model in the Cordis TypeScript meta-framework with configuration reconciliation and transactional HMR. The local archive includes the complete 88-page PDF, full text transcript, and key rendered pages.
https://github.com/cordiverse/paper/blob/main/paper.pdf
GitHub
paper/paper.pdf at main · cordiverse/paper
A Programming Paradigm for Spatiotemporal Composability - cordiverse/paper
关于扩展定律的思考
模型扩展不等于只增加参数:数据、生命周期推理成本、MoE 激活参数与有效深度、任务类型以及后训练都会改变最优配置。作者以 Kaplan、Chinchilla、MoE 和 GLM-5.3 为线索,说明下一步最值得扩展的旋钮未必是模型规模。
Thoughts About Scaling Law
Model scaling is not just parameter growth. Data, lifetime inference cost, MoE activation and effective depth, task mix, and post-training all shift the optimum; Kaplan, Chinchilla, MoE research, and GLM-5.3 show that the next useful scaling dial may not be model size.
https://x.com/jietang/status/2089941544581403107
模型扩展不等于只增加参数:数据、生命周期推理成本、MoE 激活参数与有效深度、任务类型以及后训练都会改变最优配置。作者以 Kaplan、Chinchilla、MoE 和 GLM-5.3 为线索,说明下一步最值得扩展的旋钮未必是模型规模。
Thoughts About Scaling Law
Model scaling is not just parameter growth. Data, lifetime inference cost, MoE activation and effective depth, task mix, and post-training all shift the optimum; Kaplan, Chinchilla, MoE research, and GLM-5.3 show that the next useful scaling dial may not be model size.
https://x.com/jietang/status/2089941544581403107
任意规模的 Git
大规模 Git 托管的瓶颈在于 packfile 的随机访问、强一致性,以及复制与压缩成本。Cursor 回顾了 Spokes 的 3PC 架构,并介绍 Continuity 如何以 S3 WAL 为事实来源、NVMe 为热缓存,结合 CAS 与读取时校验,实现线性化写入、强一致读取和可弹性伸缩的副本数。
Git at Any Scale
Large-scale Git hosting is constrained by packfile access patterns, strong consistency, and the cost of replication and compaction. Cursor revisits Spokes and explains how Continuity uses an S3-backed WAL, NVMe hot caches, CAS, and read-time validation to deliver linearizable writes, consistent reads, and elastic replicas.
https://cursor.com/cn/blog/git-at-any-scale
大规模 Git 托管的瓶颈在于 packfile 的随机访问、强一致性,以及复制与压缩成本。Cursor 回顾了 Spokes 的 3PC 架构,并介绍 Continuity 如何以 S3 WAL 为事实来源、NVMe 为热缓存,结合 CAS 与读取时校验,实现线性化写入、强一致读取和可弹性伸缩的副本数。
Git at Any Scale
Large-scale Git hosting is constrained by packfile access patterns, strong consistency, and the cost of replication and compaction. Cursor revisits Spokes and explains how Continuity uses an S3-backed WAL, NVMe hot caches, CAS, and read-time validation to deliver linearizable writes, consistent reads, and elastic replicas.
https://cursor.com/cn/blog/git-at-any-scale
AI 工程技能地图
Andrew Ng 团队基于逾 10,000 条招聘信息、专家访谈与问卷,将核心能力归纳为构建与部署 AI 应用、软件工程基础、使用编程智能体和塑造构建方向。这些是所有开发者都需要的 AI 工程技能,而不只是“AI 工程师”的职位要求。
The AI Engineering Skills Map
Based on 10,000+ job postings, expert interviews, and surveys, Andrew Ng's team identifies four core capabilities: building and deploying AI applications, software engineering fundamentals, using coding agents, and shaping the build. These are broad skills every developer will need, not merely requirements for the title “AI Engineer.”
https://x.com/AndrewYNg/status/2088302050706686198
Andrew Ng 团队基于逾 10,000 条招聘信息、专家访谈与问卷,将核心能力归纳为构建与部署 AI 应用、软件工程基础、使用编程智能体和塑造构建方向。这些是所有开发者都需要的 AI 工程技能,而不只是“AI 工程师”的职位要求。
The AI Engineering Skills Map
Based on 10,000+ job postings, expert interviews, and surveys, Andrew Ng's team identifies four core capabilities: building and deploying AI applications, software engineering fundamentals, using coding agents, and shaping the build. These are broad skills every developer will need, not merely requirements for the title “AI Engineer.”
https://x.com/AndrewYNg/status/2088302050706686198
AI 工程技能地图:构建与部署 AI 应用
Andrew Ng 将“构建与部署 AI 应用”拆成六项能力:LLM 基础、数据 grounding、智能体系统、评估驱动开发、生产运行和机器学习基础。核心是通过评估/错误分析闭环、可观测性、安全治理与成本延迟优化,用不可预测的 AI 组件构建可靠系统。
AI Engineering Skills Map: Building and Deploying AI Applications
Andrew Ng breaks this capability into six areas: LLM foundations, grounding models with data, agentic systems, evaluation-driven development, production operations, and machine learning foundations. The goal is to turn unpredictable AI components into reliable systems through disciplined eval loops, observability, security, and cost/latency optimization.
https://x.com/AndrewYNg/status/2090840747738374568
Andrew Ng 将“构建与部署 AI 应用”拆成六项能力:LLM 基础、数据 grounding、智能体系统、评估驱动开发、生产运行和机器学习基础。核心是通过评估/错误分析闭环、可观测性、安全治理与成本延迟优化,用不可预测的 AI 组件构建可靠系统。
AI Engineering Skills Map: Building and Deploying AI Applications
Andrew Ng breaks this capability into six areas: LLM foundations, grounding models with data, agentic systems, evaluation-driven development, production operations, and machine learning foundations. The goal is to turn unpredictable AI components into reliable systems through disciplined eval loops, observability, security, and cost/latency optimization.
https://x.com/AndrewYNg/status/2090840747738374568
https://x.com/localhost_5173
刷到一个非常酷的博主,在做一些非常有 taste 的 小而美的工具,我认为这值得我真人分享到这里
https://kero.sh/ 和 https://waku.sh/ 我都很喜欢
刷到一个非常酷的博主,在做一些非常有 taste 的 小而美的工具,我认为这值得我真人分享到这里
https://kero.sh/ 和 https://waku.sh/ 我都很喜欢
Durable AgentHarness 设计
pi 的 Harness V2 规范把每个 prompt 建模为可持久化、可恢复的 operation,并以共享追加式对话树、并行 lanes、每 lane 操作日志和全局 facts 组织会话。它用“先写 intent、后追加 result”的协议覆盖崩溃恢复、工具调用、队列、压缩、导航、hooks/events/telemetry,以及 Memory/JSONL/SQLite 后端和确定性测试驱动。
Durable AgentHarness design
A normative design and implementation plan for pi's crash-recoverable AgentHarness: durable operations, parallel lanes over an append-only conversation tree, intent/result logging, deterministic effect boundaries, tool recovery, hooks/events/telemetry, compaction/navigation, and Memory/JSONL/SQLite storage.
https://github.com/earendil-works/pi/blob/harness-v2/j4/packages/agent/docs/harness-v2.md
pi 的 Harness V2 规范把每个 prompt 建模为可持久化、可恢复的 operation,并以共享追加式对话树、并行 lanes、每 lane 操作日志和全局 facts 组织会话。它用“先写 intent、后追加 result”的协议覆盖崩溃恢复、工具调用、队列、压缩、导航、hooks/events/telemetry,以及 Memory/JSONL/SQLite 后端和确定性测试驱动。
Durable AgentHarness design
A normative design and implementation plan for pi's crash-recoverable AgentHarness: durable operations, parallel lanes over an append-only conversation tree, intent/result logging, deterministic effect boundaries, tool recovery, hooks/events/telemetry, compaction/navigation, and Memory/JSONL/SQLite storage.
https://github.com/earendil-works/pi/blob/harness-v2/j4/packages/agent/docs/harness-v2.md
GitHub
pi/packages/agent/docs/harness-v2.md at harness-v2/j4 · earendil-works/pi
AI agent toolkit: unified LLM API, agent loop, TUI, coding agent CLI - earendil-works/pi
AgentHarness 实施规范
pi 当前 main 分支的 AgentHarness 规范以 immutable entries、mutable registers 和 append-only usage ledger 作为三类权威存储,并把完整 `op.state` register 作为可恢复的持久化程序计数器。文档完整定义 effect sandwich、lane 并发、对话树与 branch index、恢复/终止、hooks/events/telemetry、schema evolution、v3 兼容及确定性测试策略。
AgentHarness — implementation specification
The current implementation specification for pi's durable AgentHarness: immutable conversation entries, mutable typed registers holding total operation state, an append-only usage ledger, atomic effect-sandwich transactions, lane concurrency, crash recovery, schema evolution, and deterministic testing.
https://github.com/earendil-works/pi/blob/main/packages/agent/docs/harness.md
pi 当前 main 分支的 AgentHarness 规范以 immutable entries、mutable registers 和 append-only usage ledger 作为三类权威存储,并把完整 `op.state` register 作为可恢复的持久化程序计数器。文档完整定义 effect sandwich、lane 并发、对话树与 branch index、恢复/终止、hooks/events/telemetry、schema evolution、v3 兼容及确定性测试策略。
AgentHarness — implementation specification
The current implementation specification for pi's durable AgentHarness: immutable conversation entries, mutable typed registers holding total operation state, an append-only usage ledger, atomic effect-sandwich transactions, lane concurrency, crash recovery, schema evolution, and deterministic testing.
https://github.com/earendil-works/pi/blob/main/packages/agent/docs/harness.md
GitHub
pi/packages/agent/docs/harness.md at main · earendil-works/pi
AI agent toolkit: unified LLM API, agent loop, TUI, coding agent CLI - earendil-works/pi
📚 新书签 / New Bookmark
中文|Sam Altman:打造 OpenAI 与押注“不可能”
这场 78 分钟对谈涵盖 AI 采用中的人类与组织惯性、模型与算力的幂律基础设施、非共识研究下注、迭代部署与权力集中风险、OpenAI 的平台战略、YC 经验,以及从无产品的混乱摸索走向 GPT 的历程。完整可读文字稿、原生字幕与封面均已本地归档。
English | Sam Altman on Building OpenAI & Betting on the Impossible
A 78-minute conversation on adoption inertia, compute as a power-law infrastructure project, non-consensus research bets, iterative deployment, centralized-power risks, OpenAI's platform strategy, YC lessons, and the chaotic pre-product path to GPT. The complete readable transcript, native captions, and cover are archived locally.
https://www.youtube.com/watch?v=kG8AoExkX40
中文|Sam Altman:打造 OpenAI 与押注“不可能”
这场 78 分钟对谈涵盖 AI 采用中的人类与组织惯性、模型与算力的幂律基础设施、非共识研究下注、迭代部署与权力集中风险、OpenAI 的平台战略、YC 经验,以及从无产品的混乱摸索走向 GPT 的历程。完整可读文字稿、原生字幕与封面均已本地归档。
English | Sam Altman on Building OpenAI & Betting on the Impossible
A 78-minute conversation on adoption inertia, compute as a power-law infrastructure project, non-consensus research bets, iterative deployment, centralized-power risks, OpenAI's platform strategy, YC lessons, and the chaotic pre-product path to GPT. The complete readable transcript, native captions, and cover are archived locally.
https://www.youtube.com/watch?v=kG8AoExkX40
YouTube
Sam Altman on Building OpenAI & Betting on the Impossible
Sam Altman has spent his career at the intersection of startups, investing and artificial intelligence. He says he was fascinated by AI as a child in St. Louis, studied it in college and eventually helped start OpenAI in 2015 after concluding that the most…
从零构建你自己的 AI 编程智能体 Harness
Vercel Academy 的 38 节实作课程以 TeensyCode 为主线,从工具循环出发,逐步加入工具契约、执行级安全门、结构化提示词、沙箱抽象、上下文裁剪、子智能体委派、生命周期、人机审批、规划验证、产品表面与扩展机制。完整课程概览、11 个模块、全部逐课链接、封面和作者头像均已本地归档。
Build Your Own AI Coding Agent Harness
A 38-lesson Vercel Academy course that builds TeensyCode from a tool loop into a production-oriented coding-agent harness, covering tool contracts, execution safety, prompts, sandboxes, context pruning, subagent delegation, lifecycle, human approval, verification, surfaces, and extensibility. The full overview, all 11 modules and lesson links, cover, and author image are archived locally.
https://vercel.com/academy/build-ai-agent-harness
Vercel Academy 的 38 节实作课程以 TeensyCode 为主线,从工具循环出发,逐步加入工具契约、执行级安全门、结构化提示词、沙箱抽象、上下文裁剪、子智能体委派、生命周期、人机审批、规划验证、产品表面与扩展机制。完整课程概览、11 个模块、全部逐课链接、封面和作者头像均已本地归档。
Build Your Own AI Coding Agent Harness
A 38-lesson Vercel Academy course that builds TeensyCode from a tool loop into a production-oriented coding-agent harness, covering tool contracts, execution safety, prompts, sandboxes, context pruning, subagent delegation, lifecycle, human approval, verification, surfaces, and extensibility. The full overview, all 11 modules and lesson links, cover, and author image are archived locally.
https://vercel.com/academy/build-ai-agent-harness
Vercel
Build Your Own AI Coding Agent Harness | Vercel Academy
Build an AI coding agent harness from scratch using AI SDK, Vercel Sandbox, and just-bash. Covers the tool loop, tool design, system prompts, sandbox abstraction, context pruning, subagent delegation, lifecycle management, and extensibility.
《推理工程》:从 CUDA 到 Kubernetes 的生产 AI 指南
Philip Kiely 的 256 页《Inference Engineering》梳理运行时、基础设施与工具三层技术栈,覆盖用例和预算、模型架构、GPU 硬件、CUDA/框架/推理引擎、量化与推测解码、KV cache、并行与解耦、多模态 serving 及生产运维。完整图书介绍、章节路线、作者、FAQ、读者评价和 20 张原图已本地归档。
Inference Engineering
Philip Kiely's 256-page guide maps the runtime, infrastructure, and tooling layers of production AI inference—from use cases and budgets through model architecture, GPU hardware, CUDA/frameworks/engines, quantization, speculative decoding, KV-cache reuse, parallelism, disaggregation, multimodal serving, and production operations. The full landing page, chapter map, author bio, FAQ, reader reactions, and 20 original images are archived locally.
https://www.baseten.co/inference-engineering/
Philip Kiely 的 256 页《Inference Engineering》梳理运行时、基础设施与工具三层技术栈,覆盖用例和预算、模型架构、GPU 硬件、CUDA/框架/推理引擎、量化与推测解码、KV cache、并行与解耦、多模态 serving 及生产运维。完整图书介绍、章节路线、作者、FAQ、读者评价和 20 张原图已本地归档。
Inference Engineering
Philip Kiely's 256-page guide maps the runtime, infrastructure, and tooling layers of production AI inference—from use cases and budgets through model architecture, GPU hardware, CUDA/frameworks/engines, quantization, speculative decoding, KV-cache reuse, parallelism, disaggregation, multimodal serving, and production operations. The full landing page, chapter map, author bio, FAQ, reader reactions, and 20 original images are archived locally.
https://www.baseten.co/inference-engineering/
Baseten
Inference Engineering | Baseten Books
Inference Engineering by Philip Kiely is your guide to the hardware, software, techniques, and infrastructure required to run AI models in production.
从零开始的 AI 工程:503 节全栈开源课程
一套免费、MIT 许可的全栈 AI 工程课程:20 个阶段、约 320 小时,从数学与经典机器学习一路构建到 Transformer、LLM、智能体、群体系统、生产基础设施和安全对齐。每节课都交付可复用的 prompt、skill、agent 或 MCP server。
AI Engineering from Scratch: A 503-Lesson Open Curriculum
A free MIT-licensed, 20-phase curriculum spanning roughly 320 hours, from math and classical ML through transformers, LLMs, agents, swarms, production infrastructure, and alignment. Every lesson ships a reusable prompt, skill, agent, or MCP server.
https://github.com/rohitg00/ai-engineering-from-scratch
一套免费、MIT 许可的全栈 AI 工程课程:20 个阶段、约 320 小时,从数学与经典机器学习一路构建到 Transformer、LLM、智能体、群体系统、生产基础设施和安全对齐。每节课都交付可复用的 prompt、skill、agent 或 MCP server。
AI Engineering from Scratch: A 503-Lesson Open Curriculum
A free MIT-licensed, 20-phase curriculum spanning roughly 320 hours, from math and classical ML through transformers, LLMs, agents, swarms, production infrastructure, and alignment. Every lesson ships a reusable prompt, skill, agent, or MCP server.
https://github.com/rohitg00/ai-engineering-from-scratch
GitHub
GitHub - rohitg00/ai-engineering-from-scratch: Learn it. Build it. Ship it for others.
Learn it. Build it. Ship it for others. Contribute to rohitg00/ai-engineering-from-scratch development by creating an account on GitHub.
从零构建大语言模型:Sebastian Raschka 官方代码仓库
这本 Manning 图书的官方实践仓库用 PyTorch 从底层实现 GPT 风格 LLM:从分词、attention 和模型组装,到预训练、GPT-2 权重加载、分类微调与指令微调。七章主线、五个附录和持续更新的 bonus material 都配有可运行 Notebook 与精简代码。
Build a Large Language Model (From Scratch): Official Code Repository
Sebastian Raschka’s official Manning book repository implements a GPT-like LLM in PyTorch from tokenization and attention through pretraining, GPT-2 weight loading, classification finetuning, and instruction finetuning, with runnable notebooks, five appendices, exercises, and extensive evolving bonus material.
https://github.com/rasbt/LLMs-from-scratch
这本 Manning 图书的官方实践仓库用 PyTorch 从底层实现 GPT 风格 LLM:从分词、attention 和模型组装,到预训练、GPT-2 权重加载、分类微调与指令微调。七章主线、五个附录和持续更新的 bonus material 都配有可运行 Notebook 与精简代码。
Build a Large Language Model (From Scratch): Official Code Repository
Sebastian Raschka’s official Manning book repository implements a GPT-like LLM in PyTorch from tokenization and attention through pretraining, GPT-2 weight loading, classification finetuning, and instruction finetuning, with runnable notebooks, five appendices, exercises, and extensive evolving bonus material.
https://github.com/rasbt/LLMs-from-scratch
GitHub
GitHub - rasbt/LLMs-from-scratch: Implement a ChatGPT-like LLM in PyTorch from scratch, step by step
Implement a ChatGPT-like LLM in PyTorch from scratch, step by step - rasbt/LLMs-from-scratch
《从零构建大模型》社区中文翻译
MLNLP-World 维护的非官方中文学习仓库,把 Sebastian Raschka 的 LLMs-from-scratch 主线翻译为中文,覆盖安装、文本处理、attention、GPT、预训练、分类/指令微调与三个实作附录,并逐章标注译者和校对者。项目仍标记为 v0.1.0 / building;学习时应同步查阅更新更快的英文上游。
LLMs from Scratch: Community Chinese Translation
An unofficial MLNLP-World community translation of Sebastian Raschka’s LLMs-from-scratch learning repository, covering setup, text processing, attention, GPT construction, pretraining, classification/instruction finetuning, and three hands-on appendices with translator and reviewer attribution. It remains marked v0.1.0 / building, so readers should cross-check the faster-moving English upstream.
https://github.com/MLNLP-World/LLMs-from-scratch-CN
MLNLP-World 维护的非官方中文学习仓库,把 Sebastian Raschka 的 LLMs-from-scratch 主线翻译为中文,覆盖安装、文本处理、attention、GPT、预训练、分类/指令微调与三个实作附录,并逐章标注译者和校对者。项目仍标记为 v0.1.0 / building;学习时应同步查阅更新更快的英文上游。
LLMs from Scratch: Community Chinese Translation
An unofficial MLNLP-World community translation of Sebastian Raschka’s LLMs-from-scratch learning repository, covering setup, text processing, attention, GPT construction, pretraining, classification/instruction finetuning, and three hands-on appendices with translator and reviewer attribution. It remains marked v0.1.0 / building, so readers should cross-check the faster-moving English upstream.
https://github.com/MLNLP-World/LLMs-from-scratch-CN
GitHub
GitHub - MLNLP-World/LLMs-from-scratch-CN: LLMs-from-scratch项目中文翻译
LLMs-from-scratch项目中文翻译. Contribute to MLNLP-World/LLMs-from-scratch-CN development by creating an account on GitHub.
深入 vLLM:高吞吐 LLM 推理系统的完整剖面
Aleksa Gordic 从一次离线 generate 调用出发,逐层拆解 vLLM V1 的 engine core、调度器、PagedAttention、continuous batching、chunked prefill、prefix caching、guided/speculative decoding、P/D 解耦、多 GPU/多节点服务,以及 TTFT、ITL、吞吐和 roofline 评测。
Inside vLLM: Anatomy of a High-Throughput LLM Inference System
Aleksa Gordic traces vLLM V1 from a single offline generate call through the engine core, scheduler, PagedAttention, continuous batching, chunked prefill, prefix caching, guided/speculative decoding, disaggregated prefill/decode, multi-GPU and multi-node serving, and latency-throughput benchmarking.
https://vllm.ai/blog/2025-09-05-anatomy-of-vllm
Aleksa Gordic 从一次离线 generate 调用出发,逐层拆解 vLLM V1 的 engine core、调度器、PagedAttention、continuous batching、chunked prefill、prefix caching、guided/speculative decoding、P/D 解耦、多 GPU/多节点服务,以及 TTFT、ITL、吞吐和 roofline 评测。
Inside vLLM: Anatomy of a High-Throughput LLM Inference System
Aleksa Gordic traces vLLM V1 from a single offline generate call through the engine core, scheduler, PagedAttention, continuous batching, chunked prefill, prefix caching, guided/speculative decoding, disaggregated prefill/decode, multi-GPU and multi-node serving, and latency-throughput benchmarking.
https://vllm.ai/blog/2025-09-05-anatomy-of-vllm
vllm.ai
Inside vLLM: Anatomy of a High-Throughput LLM Inference System
How vLLM's inference engine works, covering PagedAttention, continuous batching, prefix caching, speculative decoding, multi-GPU serving, scheduling, and benchm