Cloudflare saved 100 TB of RAM with math and Rust π₯
Cloudflare reclaimed more than 100 TB of RAM globally in a Pingora-based consistent-hashing service. No new hardware. The win came from changing data representation and the algorithms around it.
β
Lessons:
β’ Measure retained memory, not only allocation rate.
β’ Check collection shape: duplicated keys, oversized buckets, and pointer-heavy graphs add up fast.
β’ Fix the model first: a smaller or more compact structure usually beats micro-optimizing a hot loop.
For high-cardinality caches, routing tables, or tenant maps, take a heap dump before reaching for another cache node. A few bytes per entry becomes expensive at fleet scale.
[ Blog ] :
https://blog.cloudflare.com/saving-100-tb-of-ram-with-math
γ°οΈγ°οΈγ°οΈγ°οΈγ°οΈγ°οΈ
#Performance #Memory #Rust
@ProgrammingTip
Cloudflare reclaimed more than 100 TB of RAM globally in a Pingora-based consistent-hashing service. No new hardware. The win came from changing data representation and the algorithms around it.
β’ Measure retained memory, not only allocation rate.
β’ Check collection shape: duplicated keys, oversized buckets, and pointer-heavy graphs add up fast.
β’ Fix the model first: a smaller or more compact structure usually beats micro-optimizing a hot loop.
For high-cardinality caches, routing tables, or tenant maps, take a heap dump before reaching for another cache node. A few bytes per entry becomes expensive at fleet scale.
[ Blog ] :
https://blog.cloudflare.com/saving-100-tb-of-ram-with-math
γ°οΈγ°οΈγ°οΈγ°οΈγ°οΈγ°οΈ
#Performance #Memory #Rust
@ProgrammingTip
Please open Telegram to view this post
VIEW IN TELEGRAM
Cloudflare Blog
Saving another 100TB of RAM with math (and Rust)
Cloudflare's global network is immense but not limitless. As we look for small ways to trim our resource usage, we sometimes get lucky and we can cut significantly more. Hereβs how we reduced one of our Pingora-based service's RAM usage with statistics.
Hex turns GPT-6 Astra analysis into visual reports π
Hex is using GPT-6 Astra to turn complex analysis into visual reports. The useful bit is not just asking a model to summarize a table. It is moving from an analysis request to a result people can inspect and share.
What this points to:
β’ Analysis as an artifact: teams need charts, assumptions, and outputs, not a chat answer pasted into Slack.
β’ Human review still matters: a clean report can hide bad joins, stale data, or a wrong metric definition.
β’ Tool context is the product: models get more useful when they operate inside the workspace where data and business logic already live.
For AI app builders, this is the bar: produce a result that can survive review, not just a plausible paragraph.
[ Read More ] :
https://openai.com/index/hex-gpt-6-astra
γ°γ°γ°γ°γ°γ°
#AI #LLM #GPT6 #Data
@ProgrammingTip
Hex is using GPT-6 Astra to turn complex analysis into visual reports. The useful bit is not just asking a model to summarize a table. It is moving from an analysis request to a result people can inspect and share.
What this points to:
β’ Analysis as an artifact: teams need charts, assumptions, and outputs, not a chat answer pasted into Slack.
β’ Human review still matters: a clean report can hide bad joins, stale data, or a wrong metric definition.
β’ Tool context is the product: models get more useful when they operate inside the workspace where data and business logic already live.
For AI app builders, this is the bar: produce a result that can survive review, not just a plausible paragraph.
[ Read More ] :
https://openai.com/index/hex-gpt-6-astra
γ°γ°γ°γ°γ°γ°
#AI #LLM #GPT6 #Data
@ProgrammingTip
OpenAI
Hex turns complex analysis into visual reports with GPTβ6 Astra
GPT-6 Astra helps Hexβs data agents turn answers into interactive visualizations that employees are proud to share.
GPT-6 Sol and GPT-6 Luna are in the API π
OpenAI shipped GPT-6 Sol and GPT-6 Luna for API developers, alongside their availability in Codex and ChatGPT.
This is a two-model release, so do not blindly swap your existing production model. Put both behind the same eval set first: tool calls, structured output, long-context retrieval, refusal behavior, and latency under your real prompt size.
What to do this week:
β’ Add Sol and Luna as versioned model options in your config.
β’ Run replay traffic against a fixed golden set.
β’ Log model ID, token use, tool errors, and task success separately.
A model migration is an engineering change, not a dropdown change.
[ Read More ] :
https://openai.com/index/introducing-gpt-6-sol-and-luna/
γ°γ°γ°γ°γ°γ°
#AI #OpenAI #LLM #API
@ProgrammingTip
OpenAI shipped GPT-6 Sol and GPT-6 Luna for API developers, alongside their availability in Codex and ChatGPT.
This is a two-model release, so do not blindly swap your existing production model. Put both behind the same eval set first: tool calls, structured output, long-context retrieval, refusal behavior, and latency under your real prompt size.
What to do this week:
β’ Add Sol and Luna as versioned model options in your config.
β’ Run replay traffic against a fixed golden set.
β’ Log model ID, token use, tool errors, and task success separately.
A model migration is an engineering change, not a dropdown change.
[ Read More ] :
https://openai.com/index/introducing-gpt-6-sol-and-luna/
γ°γ°γ°γ°γ°γ°
#AI #OpenAI #LLM #API
@ProgrammingTip
OpenAI
Introducing GPT-6 Sol and Luna
Meet GPT-6 Sol and Luna, two models that bring frontier intelligence to everyday work with different balances of capability and cost.
Claude Opus 5.5 is now on the Claude Platform π
Anthropic introduced Claude Opus 5.5, a lower-cost frontier model available in Claude Code and through the Claude Platform.
That matters if your agent workload has been split between a high-end model for hard tasks and cheaper models for everything else. A lower-cost Opus tier can change where that handoff happens.
What to check:
β’ Run your existing eval set, not a few cherry-picked prompts.
β’ Measure tool-call accuracy and recovery after a failed call.
β’ Compare total agent cost: tokens, retries, and human review time.
For Claude Code users, model choice is now a practical repo-level config decision, not just a benchmark chart.
[ Blog ] :
https://www.anthropic.com/claude-opus-5-5
γ°γ°γ°γ°γ°γ°
#AI #LLM #Claude #ClaudeCode
@ProgrammingTip
Anthropic introduced Claude Opus 5.5, a lower-cost frontier model available in Claude Code and through the Claude Platform.
That matters if your agent workload has been split between a high-end model for hard tasks and cheaper models for everything else. A lower-cost Opus tier can change where that handoff happens.
What to check:
β’ Run your existing eval set, not a few cherry-picked prompts.
β’ Measure tool-call accuracy and recovery after a failed call.
β’ Compare total agent cost: tokens, retries, and human review time.
For Claude Code users, model choice is now a practical repo-level config decision, not just a benchmark chart.
[ Blog ] :
https://www.anthropic.com/claude-opus-5-5
γ°γ°γ°γ°γ°γ°
#AI #LLM #Claude #ClaudeCode
@ProgrammingTip
Anthropic
Introducing Claude Opus 5.5
Claude Opus 5.5 leads in agentic coding and knowledge work, and costs 40% less to run than Opus 5 on typical workloads.
Claude Opus 5.5 is available in the API π
Anthropic released Claude Opus 5.5, and it is available through the Claude API.
The practical pitch is simple: lower typical token costs than Opus 5, plus a faster mode when response time matters more than squeezing out the last bit of reasoning.
What to check in your evals:
β’ Run the same tool-use and coding tasks against Opus 5.
β’ Measure latency separately for normal and faster mode.
β’ Track input and output tokens, not just the model's listed price.
A cheaper high-end model changes agent architecture decisions. Some workflows that needed routing to a smaller model may now fit under one stronger default.
[ Read More ] :
https://www.anthropic.com/claude/opus
γ°γ°γ°γ°γ°γ°
#AI #LLM #Claude #API
@ProgrammingTip
Anthropic released Claude Opus 5.5, and it is available through the Claude API.
The practical pitch is simple: lower typical token costs than Opus 5, plus a faster mode when response time matters more than squeezing out the last bit of reasoning.
What to check in your evals:
β’ Run the same tool-use and coding tasks against Opus 5.
β’ Measure latency separately for normal and faster mode.
β’ Track input and output tokens, not just the model's listed price.
A cheaper high-end model changes agent architecture decisions. Some workflows that needed routing to a smaller model may now fit under one stronger default.
[ Read More ] :
https://www.anthropic.com/claude/opus
γ°γ°γ°γ°γ°γ°
#AI #LLM #Claude #API
@ProgrammingTip
Anthropic
Claude Opus
Hybrid reasoning model built for serious coding and AI agents, featuring a 1M context window.
Claude Sonnet 5.5 is built for coding agents π
Anthropic released Claude Sonnet 5.5 for the Claude Platform. It targets the work developers actually hand to coding agents: navigating a repo, making changes across files, using tools, and checking the result.
What changed:
β’ Faster agentic work: aimed at shorter tool loops and less idle time while an agent investigates a codebase.
β’ Lower-cost option: positioned for teams that need to run coding tasks repeatedly, not just ask one-off questions.
β’ Production focus: Anthropic calls out software engineering and multi-step agent workflows directly.
If your agent spends more time calling tools than writing code, model latency and per-task cost matter as much as benchmark scores. Test it on a real issue queue, with your actual tool permissions.
[ Read More ] :
https://www.anthropic.com/claude-sonnet-5-5
γ°γ°γ°γ°γ°γ°
#AI #LLM #Claude #CodingAgents
@ProgrammingTip
Anthropic released Claude Sonnet 5.5 for the Claude Platform. It targets the work developers actually hand to coding agents: navigating a repo, making changes across files, using tools, and checking the result.
What changed:
β’ Faster agentic work: aimed at shorter tool loops and less idle time while an agent investigates a codebase.
β’ Lower-cost option: positioned for teams that need to run coding tasks repeatedly, not just ask one-off questions.
β’ Production focus: Anthropic calls out software engineering and multi-step agent workflows directly.
If your agent spends more time calling tools than writing code, model latency and per-task cost matter as much as benchmark scores. Test it on a real issue queue, with your actual tool permissions.
[ Read More ] :
https://www.anthropic.com/claude-sonnet-5-5
γ°γ°γ°γ°γ°γ°
#AI #LLM #Claude #CodingAgents
@ProgrammingTip
Anthropic
Introducing Claude Sonnet 5.5
Claude Sonnet 5.5 is a clear upgrade over Claude Sonnet 5, runs 30%+ faster, and costs up to 30% less for most work.
OpenAI introduced Dots, always-on agents π
OpenAI introduced Dots, its take on always-on agents.
This is not a one-prompt, one-answer workflow. The pitch is an agent that can stay active around work instead of waiting for you to reopen a chat and restate the task.
Why this matters:
β’ Long-running work: agents need durable context, not a pile of copied prompts.
β’ Real handoff points: a useful agent should surface decisions and results, not silently keep doing things.
β’ Agent ops: permissions, logs, cancellation, and cost limits become product features.
If you build agent workflows, the hard part is no longer getting a model to call a tool. It is making an autonomous process observable enough that somebody will trust it.
[ Read More ] :
https://openai.com/index/introducing-dots/
γ°γ°γ°γ°γ°γ°
#AI #Agents #OpenAI
@ProgrammingTip
OpenAI introduced Dots, its take on always-on agents.
This is not a one-prompt, one-answer workflow. The pitch is an agent that can stay active around work instead of waiting for you to reopen a chat and restate the task.
Why this matters:
β’ Long-running work: agents need durable context, not a pile of copied prompts.
β’ Real handoff points: a useful agent should surface decisions and results, not silently keep doing things.
β’ Agent ops: permissions, logs, cancellation, and cost limits become product features.
If you build agent workflows, the hard part is no longer getting a model to call a tool. It is making an autonomous process observable enough that somebody will trust it.
[ Read More ] :
https://openai.com/index/introducing-dots/
γ°γ°γ°γ°γ°γ°
#AI #Agents #OpenAI
@ProgrammingTip
OpenAI
Introducing dots
Dots by OpenAI are proactive assistants that can keep working across complex projects and everyday tasks. Learn how dots help you stay in control while work moves forward.
Cloudflare shipped an agentic CLI for its API β‘οΈ
Cloudflare launched
This is a practical alternative to collecting one-off curl commands in a wiki or maintaining small admin scripts for every service. The CLI gives humans and coding agents one command-line entry point for Cloudflare operations.
Where it fitsβ
:
β’ Inspect and change Cloudflare resources while debugging a service.
β’ Give an agent a constrained operational interface instead of raw dashboard access.
β’ Turn repeatable incident steps into checked-in commands and runbooks.
Do not hand an agent broad production credentials because it has a nice CLI. Use scoped tokens, separate environments, and audit the resulting changes.
[ Blog ] :
https://blog.cloudflare.com/cloudflare-cf-cli-launch
γ°οΈγ°οΈγ°οΈγ°οΈγ°οΈγ°οΈ
#Cloudflare #CLI #DevOps #Agents #LLM
@ProgrammingTip
Cloudflare launched
cf, an agentic CLI for its full API surface.This is a practical alternative to collecting one-off curl commands in a wiki or maintaining small admin scripts for every service. The CLI gives humans and coding agents one command-line entry point for Cloudflare operations.
Where it fits
β’ Inspect and change Cloudflare resources while debugging a service.
β’ Give an agent a constrained operational interface instead of raw dashboard access.
β’ Turn repeatable incident steps into checked-in commands and runbooks.
Do not hand an agent broad production credentials because it has a nice CLI. Use scoped tokens, separate environments, and audit the resulting changes.
[ Blog ] :
https://blog.cloudflare.com/cloudflare-cf-cli-launch
γ°οΈγ°οΈγ°οΈγ°οΈγ°οΈγ°οΈ
#Cloudflare #CLI #DevOps #Agents #LLM
@ProgrammingTip
Please open Telegram to view this post
VIEW IN TELEGRAM
Cloudflare Blog
Introducing cf: the agentic CLI for the entire Cloudflare API
We are releasing cf, our new command-line tool that mirrors the entire Cloudflare API and supports programmatic TypeScript configuration. We are also open-sourcing Forge, our internal SDK generator.
GPT-6 Astra gets an Ultrafast API tier π
OpenAI added an Ultrafast speed tier for GPT-6 Astra in Codex and the API.
The important bit is real-time work. Astra can now run Responses API workflows over WebSockets, which is a much better fit for interactive coding assistants, live agent status, and UI flows where waiting on a full request feels bad.
What to check:
β’ Responses API: use it for the agent loop and tool calls.
β’ WebSockets: keep one live connection instead of polling.
β’ Ultrafast tier: test it where latency matters more than squeezing every last token of quality.
If your app streams agent work to a browser, this is worth benchmarking against your current model setup.
[ Article ] :
https://community.openai.com/t/build-ultrafast-with-astra-in-codex-and-the-api/1402393
γ°γ°γ°γ°γ°γ°
#AI #OpenAI #API #LLM
@ProgrammingTip
OpenAI added an Ultrafast speed tier for GPT-6 Astra in Codex and the API.
The important bit is real-time work. Astra can now run Responses API workflows over WebSockets, which is a much better fit for interactive coding assistants, live agent status, and UI flows where waiting on a full request feels bad.
What to check:
β’ Responses API: use it for the agent loop and tool calls.
β’ WebSockets: keep one live connection instead of polling.
β’ Ultrafast tier: test it where latency matters more than squeezing every last token of quality.
If your app streams agent work to a browser, this is worth benchmarking against your current model setup.
[ Article ] :
https://community.openai.com/t/build-ultrafast-with-astra-in-codex-and-the-api/1402393
γ°γ°γ°γ°γ°γ°
#AI #OpenAI #API #LLM
@ProgrammingTip
OpenAI Developer Community
Build Ultrafast with Astra in Codex and the API
Our premium speed tier, Ultrafast offers up to 8x faster token generation (300 tokens per second) in Codex and up to 6x in the API. In Codex, Ultrafast generates code faster, so you can move more quickly from idea to code to iteration π β¦
AI is changing the developer career ladder π
GitHub's latest developer career advice is refreshingly practical: AI can write more of the first draft, but it cannot own the engineering outcome for you.
The skills to double down on:
β’ Systems thinking: understand the service, data flow, failure modes, and tradeoffs around the code.
β’ Judgment: spot the plausible-looking AI patch that breaks security, cost, or production behavior.
β’ Communication: turn a vague product request into constraints an agent, teammate, and reviewer can act on.
Using Copilot or an agent is becoming normal. Being the person who can frame the task, verify the output, and ship it safely is still the hard part.
[ Article ] :
https://github.blog/ai-and-ml/ai-is-rewriting-the-developer-career-ladder-heres-how-to-stand-out/
γ°γ°γ°γ°γ°γ°
#AI #GitHub #Copilot #Career
@ProgrammingTip
GitHub's latest developer career advice is refreshingly practical: AI can write more of the first draft, but it cannot own the engineering outcome for you.
The skills to double down on:
β’ Systems thinking: understand the service, data flow, failure modes, and tradeoffs around the code.
β’ Judgment: spot the plausible-looking AI patch that breaks security, cost, or production behavior.
β’ Communication: turn a vague product request into constraints an agent, teammate, and reviewer can act on.
Using Copilot or an agent is becoming normal. Being the person who can frame the task, verify the output, and ship it safely is still the hard part.
[ Article ] :
https://github.blog/ai-and-ml/ai-is-rewriting-the-developer-career-ladder-heres-how-to-stand-out/
γ°γ°γ°γ°γ°γ°
#AI #GitHub #Copilot #Career
@ProgrammingTip
The GitHub Blog
AI is changing developer work. Here are three skills to strengthen.
Learn three ways to get noticed and grow your career as AI reshapes how developers build software.
Cloudflare shipped a Web Search API β‘οΈ
Cloudflare introduced a Web Search API. Search can now be a service call in the app stack instead of a pile of scraped HTML, brittle selectors, and browser automation.π³οΈβπ
[ Read More ] :
https://developers.cloudflare.com/changelog/post/2026-10-02-introducing-web-search-api
γ°οΈγ°οΈγ°οΈγ°οΈγ°οΈγ°οΈ
#AI #Cloudflare #Agents
@ProgrammingTip
Cloudflare introduced a Web Search API. Search can now be a service call in the app stack instead of a pile of scraped HTML, brittle selectors, and browser automation.
[ Read More ] :
https://developers.cloudflare.com/changelog/post/2026-10-02-introducing-web-search-api
γ°οΈγ°οΈγ°οΈγ°οΈγ°οΈγ°οΈ
#AI #Cloudflare #Agents
@ProgrammingTip
Please open Telegram to view this post
VIEW IN TELEGRAM
Cloudflare Docs
Introducing Web Search API Β· Changelog
Search the web from your AI agents and applications through AI Gateway with Ceramic.ai, Exa, and Linkup.