ROI of AI-Assisted Development
One of the biggest questions around AI adoption is what business value does it actually bring?
I'm currently introducing AI into the SDLC across my teams, so measuring its real impact is something I'm really interested in.
Back in February 2026, DORA published a dedicated report on this topic: ROI of AI-assisted Software Development report.
Key takeaways:
๐ธ Measure your baseline first. Before evaluating the impact of AI, you need to understand where you're now. Without "before" metrics, it's impossible to measure "after".
๐ธ Expect a J-curve. Most organizations experience a temporary productivity drop before seeing benefits. Main reasons for that:
- The learning curve. Time to master new tools and workflows.
- The verification tax. Time to verify AI-generated code and establish trustworthiness of agents output.
- Pipeline adaptation. Scaling other downstream processes like testing and change approval to handle increased number of code changes.
๐ธ Invest in engineering practices. The quality of an internal platform, clear workflows, reliable test automation, and strong engineering culture become critical to produce predictable delivery. The rule
๐ธ Measure both positive and negative outcomes. Higher throughput is valuable only if it doesn't come with higher instability, more defects, or lower quality.
๐ธ Measure the following business values:
- Cost efficiency
- Productivity
- Developer experience
- User experience
- Business growth
The report provides guidance on how to measure these areas and combine them into an ROI calculation. It also includes an online calculator where you can check your own numbers.
๐ธ AI adoption takes time. According to the report, the average adoption journey takes around 8 months, in large enterprise rollout can take even 12โ18 months.
Overall, the report is great: no hype, just a practical and rational approach.
It also strongly aligns with my own view of AI adoption. You can't simply buy AI licenses for developers and expect the investment to pay for itself.
Successful AI adoption requires improvements across the entire SDLC: better development processes, a stronger engineering culture, higher test coverage, and investments in guardrails and workflows that keep software delivery stable and predictable.
#ai #engineering #ai4sdlc
One of the biggest questions around AI adoption is what business value does it actually bring?
I'm currently introducing AI into the SDLC across my teams, so measuring its real impact is something I'm really interested in.
Back in February 2026, DORA published a dedicated report on this topic: ROI of AI-assisted Software Development report.
Key takeaways:
๐ธ Measure your baseline first. Before evaluating the impact of AI, you need to understand where you're now. Without "before" metrics, it's impossible to measure "after".
๐ธ Expect a J-curve. Most organizations experience a temporary productivity drop before seeing benefits. Main reasons for that:
- The learning curve. Time to master new tools and workflows.
- The verification tax. Time to verify AI-generated code and establish trustworthiness of agents output.
- Pipeline adaptation. Scaling other downstream processes like testing and change approval to handle increased number of code changes.
๐ธ Invest in engineering practices. The quality of an internal platform, clear workflows, reliable test automation, and strong engineering culture become critical to produce predictable delivery. The rule
garbage in -> garbage out is still actual.๐ธ Measure both positive and negative outcomes. Higher throughput is valuable only if it doesn't come with higher instability, more defects, or lower quality.
๐ธ Measure the following business values:
- Cost efficiency
- Productivity
- Developer experience
- User experience
- Business growth
The report provides guidance on how to measure these areas and combine them into an ROI calculation. It also includes an online calculator where you can check your own numbers.
๐ธ AI adoption takes time. According to the report, the average adoption journey takes around 8 months, in large enterprise rollout can take even 12โ18 months.
Overall, the report is great: no hype, just a practical and rational approach.
It also strongly aligns with my own view of AI adoption. You can't simply buy AI licenses for developers and expect the investment to pay for itself.
Successful AI adoption requires improvements across the entire SDLC: better development processes, a stronger engineering culture, higher test coverage, and investments in guardrails and workflows that keep software delivery stable and predictable.
#ai #engineering #ai4sdlc
Google Cloud
DORA: ROI of AI-assisted Software Development
Unlock the true value of AI in dev, Learn to navigate the J-curve, translate metrics into financial outcomes, and reinvest recovered capacity
๐ฅ5โค1๐1
AI Adoption: What to Measure First
In the previous post, we looked at how DORA recommends measuring the ROI of AI adoption.
But what if you don't know the revenue generated by your features or other business metrics? Yet you're still expected to show the effectiveness of AI adoption.
Let's bring the DORA approach down to the engineering team level.
I would split AI adoption into two phases: AI adoption itself and getting business benefits. These phases have different goals, metrics and outcomes.
It doesn't make much sense to measure business impact of AI until the team has actually gone through the adoption phase.
What to measure to understand the AI adoption state:
๐ธ Active users. How many engineers actually use AI in their daily work?
๐ธ Monthly usage per user. How actively do engineers use AI tools? Token consumption or API cost per engineer can be a good indicator. AI Gateway solutions such as LiteLLM can help collect these metrics.
๐ธ AI-generated code ratio. What percentage of code is generated by AI? Important note: this metric measures only how widely AI is being used. It says nothing about whether AI is being used effectively or producing good results.
๐ธ AI-assisted feature ratio. How many features are developed with AI assistance? The metric counts if AI was used in design, implementation, testing, documentation, or code review.
๐ธ Engineering readiness. Is the engineering ecosystem ready for agents? Here I would use criteria similar to an Agent Readiness Framework: documentation,
Giving developers AI licenses does not mean AI has been adopted.
Not after one month. Not after three. Not after five.
AI adoption happens at the team level. The goal of all these metrics is to understand whether the teams actually change the way they work. And it requires training people, overcoming resistance, changing engineering practices, and adapting the development process.
In the next post, I'll share my approach to measuring AI impact during the second phase using only tools that almost every engineering team already has.
#ai #engineering #ai4sdlc
In the previous post, we looked at how DORA recommends measuring the ROI of AI adoption.
But what if you don't know the revenue generated by your features or other business metrics? Yet you're still expected to show the effectiveness of AI adoption.
Let's bring the DORA approach down to the engineering team level.
I would split AI adoption into two phases: AI adoption itself and getting business benefits. These phases have different goals, metrics and outcomes.
It doesn't make much sense to measure business impact of AI until the team has actually gone through the adoption phase.
What to measure to understand the AI adoption state:
๐ธ Active users. How many engineers actually use AI in their daily work?
๐ธ Monthly usage per user. How actively do engineers use AI tools? Token consumption or API cost per engineer can be a good indicator. AI Gateway solutions such as LiteLLM can help collect these metrics.
๐ธ AI-generated code ratio. What percentage of code is generated by AI? Important note: this metric measures only how widely AI is being used. It says nothing about whether AI is being used effectively or producing good results.
๐ธ AI-assisted feature ratio. How many features are developed with AI assistance? The metric counts if AI was used in design, implementation, testing, documentation, or code review.
๐ธ Engineering readiness. Is the engineering ecosystem ready for agents? Here I would use criteria similar to an Agent Readiness Framework: documentation,
AGENTS.md, test coverage, CI stability, guardrails, and so on.Giving developers AI licenses does not mean AI has been adopted.
Not after one month. Not after three. Not after five.
AI adoption happens at the team level. The goal of all these metrics is to understand whether the teams actually change the way they work. And it requires training people, overcoming resistance, changing engineering practices, and adapting the development process.
In the next post, I'll share my approach to measuring AI impact during the second phase using only tools that almost every engineering team already has.
#ai #engineering #ai4sdlc
๐ฅ3โค1๐1
AI Adoption: Measuring Feature Development
Let's assume your team has already completed AI adoption and now management expects to see positive business results.
The good news is that AI doesn't magically change your business metrics. You don't need to invent new KPIs. You just need to track how your existing ones change (or finally start collecting them).
I suggest focusing on quality metrics first, then gradually shift toward delivery speed and cost optimization.
Since I lead platform engineering teams, we have two major types of work: new features development and L4 support. These activities are different, so they should be measured differently. Let's start with product development.
What can be measured:
๐ธ Team Budget. The cost of the team in $, mandays, or FTEs.
๐ธ AI Cost ($). How much does your team spend on AI? Those jokes about "it being cheaper to hire a junior" may stop being jokes soon. :)
๐ธ Bug Density (defects/LOE). AI helps us deliver more code, but it can also introduce more defects. The first goal is to make sure quality doesn't get worse.
๐ธ Test Coverage (unit, integration, E2E). AI is very good at writing tests. Increasing test coverage across different levels is usually one of the easiest wins.
๐ธ % of Toil Budget. How much engineering effort goes into routine work such as CI maintenance, vulnerability fixes, upgrades, and similar operational tasks? AI should gradually reduce this type of work.
๐ธ Feature Delivery Rate (per sprint, release, or quarter). I'm personally skeptical about this metric. In R&D, one feature may take two days while another takes two months, so averages often tell you nothing. But for some teams it can be a useful indicator.
๐ธ Time to Market. The average time to implement and deliver a feature. This is also difficult to measure in R&D, but it can be useful for teams who develop customer-facing features.
So don't focus on measuring AI. Measure the results of your work: cost, quality, and time against your baseline (baseline is the metrics value before AI adoption). If you want to show the value of AI in the future, you need to start collecting those metrics today.
#ai #engineering #ai4sdlc
Let's assume your team has already completed AI adoption and now management expects to see positive business results.
The good news is that AI doesn't magically change your business metrics. You don't need to invent new KPIs. You just need to track how your existing ones change (or finally start collecting them).
I suggest focusing on quality metrics first, then gradually shift toward delivery speed and cost optimization.
Since I lead platform engineering teams, we have two major types of work: new features development and L4 support. These activities are different, so they should be measured differently. Let's start with product development.
What can be measured:
๐ธ Team Budget. The cost of the team in $, mandays, or FTEs.
๐ธ AI Cost ($). How much does your team spend on AI? Those jokes about "it being cheaper to hire a junior" may stop being jokes soon. :)
๐ธ Bug Density (defects/LOE). AI helps us deliver more code, but it can also introduce more defects. The first goal is to make sure quality doesn't get worse.
๐ธ Test Coverage (unit, integration, E2E). AI is very good at writing tests. Increasing test coverage across different levels is usually one of the easiest wins.
๐ธ % of Toil Budget. How much engineering effort goes into routine work such as CI maintenance, vulnerability fixes, upgrades, and similar operational tasks? AI should gradually reduce this type of work.
๐ธ Feature Delivery Rate (per sprint, release, or quarter). I'm personally skeptical about this metric. In R&D, one feature may take two days while another takes two months, so averages often tell you nothing. But for some teams it can be a useful indicator.
๐ธ Time to Market. The average time to implement and deliver a feature. This is also difficult to measure in R&D, but it can be useful for teams who develop customer-facing features.
So don't focus on measuring AI. Measure the results of your work: cost, quality, and time against your baseline (baseline is the metrics value before AI adoption). If you want to show the value of AI in the future, you need to start collecting those metrics today.
#ai #engineering #ai4sdlc
๐5โค2๐ฅ2
AI Adoption: Measuring Support
Continuing the previous post, let's look at how AI adoption can be measured for support teams.
The principles are exactly the same: cost, quality, and time.
Here are the metrics I'd track:
๐ธ Team Budget. Same as for development teams: the cost of the team in $, mandays, or FTEs.
๐ธ AI Cost ($). How much does the team spend on AI, including autonomous agents if you're using them.
๐ธ Incoming Load. The number of incoming tickets. At the beginning, this metric probably won't change much. But in the future it can show whether the overall quality is improving or getting worse. I recommend tracking it as a percentage change from the baseline before AI adoption.
๐ธ Backlog Size. The number of non-resolved tickets. This metric should always be viewed together with the team budget and SLA.
๐ธ Time to Resolution. How quickly customers receive a solution to their problem.
๐ธ Reopen Rate. Has the percentage of reopened tickets increased? I've seen exactly this happen when teams introduced AI-based ticket assessment. Resolution time went down, but the number of reopened tickets increased several times. A clear sign that support quality had dropped.
As you can see, none of these metrics are AI-specific. Most tracking systems can calculate them out of the box, while a few may require collecting and analyzing historical data.
And one final point: never look at these metrics in isolation. Resolution time is decreased -> reopen rate doubles, team budget decreased -> backlog increased -> resolution time increased.
So it's the combination of metrics that tells you whether AI is actually improving support process or not.
#ai #engineering #ai4sdlc
Continuing the previous post, let's look at how AI adoption can be measured for support teams.
The principles are exactly the same: cost, quality, and time.
Here are the metrics I'd track:
๐ธ Team Budget. Same as for development teams: the cost of the team in $, mandays, or FTEs.
๐ธ AI Cost ($). How much does the team spend on AI, including autonomous agents if you're using them.
๐ธ Incoming Load. The number of incoming tickets. At the beginning, this metric probably won't change much. But in the future it can show whether the overall quality is improving or getting worse. I recommend tracking it as a percentage change from the baseline before AI adoption.
๐ธ Backlog Size. The number of non-resolved tickets. This metric should always be viewed together with the team budget and SLA.
๐ธ Time to Resolution. How quickly customers receive a solution to their problem.
๐ธ Reopen Rate. Has the percentage of reopened tickets increased? I've seen exactly this happen when teams introduced AI-based ticket assessment. Resolution time went down, but the number of reopened tickets increased several times. A clear sign that support quality had dropped.
As you can see, none of these metrics are AI-specific. Most tracking systems can calculate them out of the box, while a few may require collecting and analyzing historical data.
And one final point: never look at these metrics in isolation. Resolution time is decreased -> reopen rate doubles, team budget decreased -> backlog increased -> resolution time increased.
So it's the combination of metrics that tells you whether AI is actually improving support process or not.
#ai #engineering #ai4sdlc
๐3๐ฅ2โค1
Agent Plugins Spec
Agent Plugins - a new specification intended to standardize how skills and MCP servers are packaged across different agent harnesses.
The spec is a result of collaboration between Cursor, Microsoft, OpenAI, Vercel, and AWS, and looks like an attempt to provide an alternative to the growing plugin ecosystem around Claude Code.
According to the spec, a plugin should have the following structure:
And that's basically where the specification ends.
There are no answers to questions like: How should users discover available plugins? How should plugins be distributed and installed? How to deliver a new version and roll it out to clients? All of that is left to each particular harness implementation.
In other words, Agent Plugins standardizes the package format, but not package management.
And this makes it difficult to operationalize across an organization. Moreover, Claude Code, for example, doesn't support it at all, and there is currently no obvious reason for Anthropic to adopt it.
Specifications are good. I actually really like them because they bring some order to the chaos of different integrations. But at the moment, Agent Plugins is far behind APM packages or the Claude Code plugins ecosystem.
Overall, it looks promising, but for now it's more something to watch than something you can build a sustainable process around. Let's see.
#ai #engineering
Agent Plugins - a new specification intended to standardize how skills and MCP servers are packaged across different agent harnesses.
The spec is a result of collaboration between Cursor, Microsoft, OpenAI, Vercel, and AWS, and looks like an attempt to provide an alternative to the growing plugin ecosystem around Claude Code.
According to the spec, a plugin should have the following structure:
my-plugin/
โโโ plugin.json
โโโ skills/
โ โโโ summarize/
โ โโโ SKILL.md
โ โโโ scripts/
โ โโโ references/
โโโ mcp.json
โโโ com.example.client/
โโโ hooks/
And that's basically where the specification ends.
There are no answers to questions like: How should users discover available plugins? How should plugins be distributed and installed? How to deliver a new version and roll it out to clients? All of that is left to each particular harness implementation.
In other words, Agent Plugins standardizes the package format, but not package management.
And this makes it difficult to operationalize across an organization. Moreover, Claude Code, for example, doesn't support it at all, and there is currently no obvious reason for Anthropic to adopt it.
Specifications are good. I actually really like them because they bring some order to the chaos of different integrations. But at the moment, Agent Plugins is far behind APM packages or the Claude Code plugins ecosystem.
Overall, it looks promising, but for now it's more something to watch than something you can build a sustainable process around. Let's see.
#ai #engineering
Agent Plugins
A portable package format for reusable components that extend AI agents.
๐3โค1๐ฅ1๐1
The AI-Driven Leader
This is how The AI-Driven Leader by Geoff Woods begins. An international bestseller with a catchy title and a good rating on Amazon. Actually, I had pretty high expectations about it.
But let's start from the beginning.
This is a book written by a non-engineer for non-engineers. It explains in simple terms what GenAI is, what risks it brings, and how leaders can integrate AI into their work.
If you boil it down, the whole idea is to use AI as a Thought Partner (I wrote about it here). To demonstrate that the approach works, he provides simple real-life stories and examples of specific prompts. If you don't know where to start with AI, the author suggest 5 main use cases: strategic thinking, decision-making, content creation, idea generation, and analysis.
And that's basically where the meaningful AI-related content ends.
The remaining 80% of the book is mostly basic leadership recommendations that I would summarize as "do the right things, don't do the wrong things". Ideas like "leaders should focus more on strategy and less on operations" don't really have much to do with AI.
Overall, I was left with mixed feelings.
On one hand, the book is quite entertaining, easy to read, it provides a simplified introduction to GenAI and explains what leaders can actually do with it.
On the other hand, it feels like a mix of generic leadership practices with AI that sometimes added just because it's a hot topic. At the same time, the author is extremely positive about the solutions AI produces. And as engineers, we know that something that looks like a convincing result at first glance can be AI slop on the second.
From my perspective, the book can be useful if you're only starting to use AI at work and don't know where to begin. Otherwise, you can safely skip this one and spend your time reading something more useful.
#booknook #ai #leadership
"Our humanity lies in our ability to think strategically, be creative, and communicate and collaborate to solve complex problems... AI presents a unique opportunity for us to reclaim these human strengths and have machines adjust to meet our needs."
This is how The AI-Driven Leader by Geoff Woods begins. An international bestseller with a catchy title and a good rating on Amazon. Actually, I had pretty high expectations about it.
But let's start from the beginning.
This is a book written by a non-engineer for non-engineers. It explains in simple terms what GenAI is, what risks it brings, and how leaders can integrate AI into their work.
If you boil it down, the whole idea is to use AI as a Thought Partner (I wrote about it here). To demonstrate that the approach works, he provides simple real-life stories and examples of specific prompts. If you don't know where to start with AI, the author suggest 5 main use cases: strategic thinking, decision-making, content creation, idea generation, and analysis.
And that's basically where the meaningful AI-related content ends.
The remaining 80% of the book is mostly basic leadership recommendations that I would summarize as "do the right things, don't do the wrong things". Ideas like "leaders should focus more on strategy and less on operations" don't really have much to do with AI.
Overall, I was left with mixed feelings.
On one hand, the book is quite entertaining, easy to read, it provides a simplified introduction to GenAI and explains what leaders can actually do with it.
On the other hand, it feels like a mix of generic leadership practices with AI that sometimes added just because it's a hot topic. At the same time, the author is extremely positive about the solutions AI produces. And as engineers, we know that something that looks like a convincing result at first glance can be AI slop on the second.
From my perspective, the book can be useful if you're only starting to use AI at work and don't know where to begin. Otherwise, you can safely skip this one and spend your time reading something more useful.
#booknook #ai #leadership
๐4๐ฅ2โค1๐1
The OpenAIโHugging Face Incident
A talk that sounds more like a detective story than just a presentation: "The OpenAIโHugging Face Incident" from OpenAI engineers at Black Hat USA.
On July 16, Hugging Face reported that they had been targeted by a cyberattack carried out by autonomous agents. A couple of days later, OpenAI realized that the attack had been caused by their agents eval runs. That awkward moment when the investigation leads back to you ๐ฒ.
The talk gives the full chronology of what happened at OpenAI, from the first attacks on their internal infrastructure in May to the Hugging Face incident and its remediation. It's really breathtaking, and I definitely recommend watching the full video.
But what impressed me most was why it happened.
Frontier models can be quite lazy. They usually prefer to find an answer on the Internet rather than solve the task on their own. That's why Internet access is usually disabled during evaluation runs.
But when the task is too complex, an agent can get stuck and start cheating.
That's what really happened at OpenAI.
One of the agents got stuck on its task and came up with an idea: if it could leave a note describing the problem in the environment, maybe another agent with another model could find it and help solve the task.
So the agent uploaded a file to an Artifactory instance that is shared across all OpenAI's infrastructure. Over time, Artifactory was turned into a message board for agents. They started using it to collaborate, share tasks, and exchange vulnerability exploits, including an Artifactory SSRF vulnerability that gave them access to the Internet.
Then agents decided to hack Hugging Face, because it might have data files hidden from the model that could contain the answer to the original task.
What makes it more interesting is that, according to the talk, the model understood there was a rule prohibiting attacks on external resources. But it evaluated the potential benefit as greater than the cost of breaking that rule.
At the end of the talk, the speakers highlighted that it was an unintentional attack. But such attacks will become real in the near future. So we, as an industry, should invest in continuous agentic red teaming, automated vulnerability detection and patching, and automated incident response.
And this is where I had another thought. How much of this is real, and how much is marketing for the next generation models and AI products? After all, it's quite convenient to first demonstrate a new class of threat and then offer the tools to defend against it. Anyway, the threat landscape is changing. Again. And the industry has to invent new ways to deal with it.
#ai #security #engineering
A talk that sounds more like a detective story than just a presentation: "The OpenAIโHugging Face Incident" from OpenAI engineers at Black Hat USA.
On July 16, Hugging Face reported that they had been targeted by a cyberattack carried out by autonomous agents. A couple of days later, OpenAI realized that the attack had been caused by their agents eval runs. That awkward moment when the investigation leads back to you ๐ฒ.
The talk gives the full chronology of what happened at OpenAI, from the first attacks on their internal infrastructure in May to the Hugging Face incident and its remediation. It's really breathtaking, and I definitely recommend watching the full video.
But what impressed me most was why it happened.
Frontier models can be quite lazy. They usually prefer to find an answer on the Internet rather than solve the task on their own. That's why Internet access is usually disabled during evaluation runs.
But when the task is too complex, an agent can get stuck and start cheating.
That's what really happened at OpenAI.
One of the agents got stuck on its task and came up with an idea: if it could leave a note describing the problem in the environment, maybe another agent with another model could find it and help solve the task.
So the agent uploaded a file to an Artifactory instance that is shared across all OpenAI's infrastructure. Over time, Artifactory was turned into a message board for agents. They started using it to collaborate, share tasks, and exchange vulnerability exploits, including an Artifactory SSRF vulnerability that gave them access to the Internet.
Then agents decided to hack Hugging Face, because it might have data files hidden from the model that could contain the answer to the original task.
What makes it more interesting is that, according to the talk, the model understood there was a rule prohibiting attacks on external resources. But it evaluated the potential benefit as greater than the cost of breaking that rule.
At the end of the talk, the speakers highlighted that it was an unintentional attack. But such attacks will become real in the near future. So we, as an industry, should invest in continuous agentic red teaming, automated vulnerability detection and patching, and automated incident response.
And this is where I had another thought. How much of this is real, and how much is marketing for the next generation models and AI products? After all, it's quite convenient to first demonstrate a new class of threat and then offer the tools to defend against it. Anyway, the threat landscape is changing. Again. And the industry has to invent new ways to deal with it.
#ai #security #engineering
YouTube
Black Hat USA 2026 | The 'Breaking' News: The OpenAIโHugging Face Incident
The 'Breaking' News: The OpenAIโHugging Face Incident - A Technical Reconstruction and Its Implications for AI
When AI Goes Rogue. The Incident That Changed Everything. An OpenAI evaluation agent broke out of its sandbox, infiltrated Hugging Face infrastructureโฆ
When AI Goes Rogue. The Incident That Changed Everything. An OpenAI evaluation agent broke out of its sandbox, infiltrated Hugging Face infrastructureโฆ
๐ฅ3โค2๐1
Loop Engineering
Over the past year, AI has been constantly bringing new terms and practices into our work. Yet another XX Engineering probably doesn't already surprise anyone.
Let's take a look at Loop Engineering, which is hyped right now.
The main idea is that an engineer no longer prompts an agent directly. Instead, they design a system that does it on its own. A loop here can be thought of as a recursive goal: you define the purpose, and the AI iterates until complete.
A typical loop consists of the following steps:
๐ธ Discovery: Define what should be done during the iteration. The key is letting the agent find its own work rather than providing it manually.
๐ธ Handoff. Move the task from the scheduling system into the hands of the agent that does the work.
๐ธ Verification. Check whether the result satisfies the goal. This is what prevents the loop from blindly moving forward.
๐ธ Persistence. Save the result somewhere it can survive between sessions and agents, e.g., PR, issue tracker, database, etc.
๐ธ Scheduling. Define when the loop should run. For example, an automated morning CI triage.
Building blocks for the loop:
๐ธ Automations. Automated procedures or triggers that start the loop.
๐ธ Worktrees. Built-in git mechanism for multiple independent working directories in one repo to allow agents to work independently.
๐ธ Skills. Project knowledge and instructions.
๐ธ Plugins\Connectors. Tools (usually MCPs) that give agents access to the outside world: issue trackers, databases, APIs, and other systems.
๐ธ Subagents. Separation of tasks and responsibilities. For example, one agent does the work while another verifies it.
๐ธ Memory. Long-term storage for task state, execution results, and lessons learned.
The overall concept is pretty cool, but it requires very strong engineering discipline.
The cost of an incorrectly designed loop is much higher than the cost of a bad prompt. In a loop, an error can accumulate with every iteration, the result can gradually drift away from the goal, and a lot of tokens can be just wasted.
That's why verification becomes one of the most critical parts of Loop Engineering. If we want autonomous loops to produce stable results, they need a reliable way to verify their own work.
For a deeper dive:
- Getting Started with Loops Claude blog
- Loop Engineering: The Anthropic Playbook for Designing Systems That Prompt Your Agents
- Loop Engineering by Addy Osmani
- Practical Loop Engineering by Addy Osmani
#engineering #ai
Over the past year, AI has been constantly bringing new terms and practices into our work. Yet another XX Engineering probably doesn't already surprise anyone.
Let's take a look at Loop Engineering, which is hyped right now.
The main idea is that an engineer no longer prompts an agent directly. Instead, they design a system that does it on its own. A loop here can be thought of as a recursive goal: you define the purpose, and the AI iterates until complete.
A typical loop consists of the following steps:
๐ธ Discovery: Define what should be done during the iteration. The key is letting the agent find its own work rather than providing it manually.
๐ธ Handoff. Move the task from the scheduling system into the hands of the agent that does the work.
๐ธ Verification. Check whether the result satisfies the goal. This is what prevents the loop from blindly moving forward.
๐ธ Persistence. Save the result somewhere it can survive between sessions and agents, e.g., PR, issue tracker, database, etc.
๐ธ Scheduling. Define when the loop should run. For example, an automated morning CI triage.
Building blocks for the loop:
๐ธ Automations. Automated procedures or triggers that start the loop.
๐ธ Worktrees. Built-in git mechanism for multiple independent working directories in one repo to allow agents to work independently.
๐ธ Skills. Project knowledge and instructions.
๐ธ Plugins\Connectors. Tools (usually MCPs) that give agents access to the outside world: issue trackers, databases, APIs, and other systems.
๐ธ Subagents. Separation of tasks and responsibilities. For example, one agent does the work while another verifies it.
๐ธ Memory. Long-term storage for task state, execution results, and lessons learned.
The overall concept is pretty cool, but it requires very strong engineering discipline.
The cost of an incorrectly designed loop is much higher than the cost of a bad prompt. In a loop, an error can accumulate with every iteration, the result can gradually drift away from the goal, and a lot of tokens can be just wasted.
That's why verification becomes one of the most critical parts of Loop Engineering. If we want autonomous loops to produce stable results, they need a reliable way to verify their own work.
For a deeper dive:
- Getting Started with Loops Claude blog
- Loop Engineering: The Anthropic Playbook for Designing Systems That Prompt Your Agents
- Loop Engineering by Addy Osmani
- Practical Loop Engineering by Addy Osmani
#engineering #ai
๐2๐ฅ2โค1
Loop Engineering from First Principles
Continuing the topic of Loop Engineering, I'd like to share the talk: Loop Engineering from First Principles.
The author criticizes the current trend of using "blind" agentic loops everywhere. They can lead to a huge volume of generated code that nobody really understands. Quality decreases, the number of bugs grows.
No, he doesn't reject the approach itself. Instead, he suggests looking at it from a more engineering perspective and applying a Control Theory Framework.
Sounds promising, right?
To apply it to agentic loops, the following elements should be defined:
๐ธ Sensor (Measurement). Define the desired state and use deterministic tools like tests and linters to detect violations.
๐ธ Controller (Prioritization). Use deterministic rules to prioritize tasks, for example, starting with the smallest unit of work.
๐ธ Actuator (Change Application). Use hand-written "golden patterns" to guide the agent and keep changes aligned with your team's standards.
๐ธ Feedback Loop. Verify the result after each change and feed it back into the next iteration:
- Run the loop in CI to detect regressions.
- Keep a human in the loop for the final review. They should understand the changes and own the code.
So the main point is not to let an agent generate a huge pile of code that a human can no longer understand. The point is to get controlled and verifiable results incrementally, in small pieces, keeping "human-in-the-loop".
I liked the talk. It's really great when we start moving from hype toward something more manageable and controllable.
#ai #engineering
Continuing the topic of Loop Engineering, I'd like to share the talk: Loop Engineering from First Principles.
The author criticizes the current trend of using "blind" agentic loops everywhere. They can lead to a huge volume of generated code that nobody really understands. Quality decreases, the number of bugs grows.
No, he doesn't reject the approach itself. Instead, he suggests looking at it from a more engineering perspective and applying a Control Theory Framework.
Control theory is a branch of engineering and applied mathematics that regulates the behavior of dynamical systems to achieve a desired output.
Sounds promising, right?
To apply it to agentic loops, the following elements should be defined:
๐ธ Sensor (Measurement). Define the desired state and use deterministic tools like tests and linters to detect violations.
๐ธ Controller (Prioritization). Use deterministic rules to prioritize tasks, for example, starting with the smallest unit of work.
๐ธ Actuator (Change Application). Use hand-written "golden patterns" to guide the agent and keep changes aligned with your team's standards.
๐ธ Feedback Loop. Verify the result after each change and feed it back into the next iteration:
- Run the loop in CI to detect regressions.
- Keep a human in the loop for the final review. They should understand the changes and own the code.
So the main point is not to let an agent generate a huge pile of code that a human can no longer understand. The point is to get controlled and verifiable results incrementally, in small pieces, keeping "human-in-the-loop".
I liked the talk. It's really great when we start moving from hype toward something more manageable and controllable.
#ai #engineering
YouTube
Loop Engineering from First Principles โ Kyle Mistele, HumanLayer
A coding agent will happily hand you a 40,000 line pull request that nobody can review and that quietly does the wrong thing. Kyle Mistele's argument is that the fix is not a better prompt but a better loop, borrowed from control theory: a thermostat sensesโฆ
๐3โค2๐ฅ2
The Culture Map
Have you ever worked in international distributed teams? Or collaborated with teams from other countries? Then you probably noticed that it can be quite challenging.
I didn't think much about that before reading The Culture Map by Erin Meyer. I saw some difficulties in my international teams and put a lot of effort into making it work. But after reading the book, I realized that some of my efforts were just fighting windmills.
So what is this book about?
It's about how cultural patterns impact our behavior and why sometimes it's so difficult to understand each other, even when we use the same language (like English for business). The language is the same, but the meanings can be very different.
The author defines 8 scales to measure the difference:
1. Communicating: low context vs high context. Low-context means to be explicit and clear; where high-context means to be implicit, expect to read between the lines.
2. Evaluating: direct negative feedback vs indirect negative feedback.
3. Persuading: principles first (theory before examples) vs application first (examples before theory).
4. Leading: egalitarian (small distance between boss and employee) vs hierarchical (strong vertical and status).
5. Deciding: consensual (based on group decision) vs top-down
6. Trusting: task-based (trust through work result and competence) vs relationship based.
7. Disagreeing: confrontational (open disagreement) vs avoid confrontations (indirect disagreement)
8. Scheduling: linear time (plans and strong deadlines) vs flexible time (no strict plans in changing circumstances).
A good example is the US and Japan. The US is a low-context culture, where things are usually described and explained explicitly. Japan is a high-context culture, where you are expected to read between the lines. This difference can easily lead to frustration and misunderstanding.
The general rule is to compare cultures relatively. What matters is not whether a culture is โdirectโ or โhierarchical,โ but whether it is more or less so than the culture you are comparing it with.
But what to do with all of that? First, this knowledge can help you better prepare for negotiations. Second, the author recommends that leaders be more flexible: understand what leadership style their team expects and adapt to it, taking cultural differences into account.
I really liked the book. Itโs clear and practical, with good examples and useful recommendations. I started to understand many things better. Iโd highly recommend it to anyone who works in an international company or with international partners.
#booknook #softskills #leadership
Have you ever worked in international distributed teams? Or collaborated with teams from other countries? Then you probably noticed that it can be quite challenging.
I didn't think much about that before reading The Culture Map by Erin Meyer. I saw some difficulties in my international teams and put a lot of effort into making it work. But after reading the book, I realized that some of my efforts were just fighting windmills.
So what is this book about?
It's about how cultural patterns impact our behavior and why sometimes it's so difficult to understand each other, even when we use the same language (like English for business). The language is the same, but the meanings can be very different.
The author defines 8 scales to measure the difference:
1. Communicating: low context vs high context. Low-context means to be explicit and clear; where high-context means to be implicit, expect to read between the lines.
2. Evaluating: direct negative feedback vs indirect negative feedback.
3. Persuading: principles first (theory before examples) vs application first (examples before theory).
4. Leading: egalitarian (small distance between boss and employee) vs hierarchical (strong vertical and status).
5. Deciding: consensual (based on group decision) vs top-down
6. Trusting: task-based (trust through work result and competence) vs relationship based.
7. Disagreeing: confrontational (open disagreement) vs avoid confrontations (indirect disagreement)
8. Scheduling: linear time (plans and strong deadlines) vs flexible time (no strict plans in changing circumstances).
A good example is the US and Japan. The US is a low-context culture, where things are usually described and explained explicitly. Japan is a high-context culture, where you are expected to read between the lines. This difference can easily lead to frustration and misunderstanding.
The general rule is to compare cultures relatively. What matters is not whether a culture is โdirectโ or โhierarchical,โ but whether it is more or less so than the culture you are comparing it with.
But what to do with all of that? First, this knowledge can help you better prepare for negotiations. Second, the author recommends that leaders be more flexible: understand what leadership style their team expects and adapt to it, taking cultural differences into account.
I really liked the book. Itโs clear and practical, with good examples and useful recommendations. I started to understand many things better. Iโd highly recommend it to anyone who works in an international company or with international partners.
#booknook #softskills #leadership
๐3๐ฅ2โค1
Illustrations from The Culture Map showing how different cultures compare on the scales.
#booknook #softskills #leadership
#booknook #softskills #leadership
โค2๐ฅ2๐1
Why Software Factories Fail
"Read the Code!" is one of the key ideas from Dex Horthy's talk "Why Software Factories Fail".
The author touches on a very hot topic right now: we are actively pushed to put AI-generated code into production, build software factories, and eliminate the human bottleneck from the process. And all of that would be great if, at the same time, the overall quality of software products wasn't going down and the number of incidents wasn't going up.
Dex explains this phenomenon by pointing to a fundamental limitation of current models: they can't maintain and improve the quality of a codebase over time. Models are trained to solve the issue and pass the tests. There is nothing there about good architecture, clean code, or future maintainability.
What to do with that? Dex suggests to do the following:
๐ธ Make Product Review: align goals, check mockups, etc.
๐ธ Do System Architecture: establish clear component boundaries and API contracts.
๐ธ Perform Program Design: define types and method signatures, test approach, program layout, and dependency stack.
๐ธ Implement by Vertical Slices: build features vertically, reducing the scope of what an agent can change in a single iteration.
So, it's all about shifting review left to earlier stages and focusing more on architecture and program design (I described similar ideas about code review here). And then let agents make small changes that engineers can actually understand and review.
And yes, we still need to review the changes.
#engineering #codereview #ai
"Read the Code!" is one of the key ideas from Dex Horthy's talk "Why Software Factories Fail".
The author touches on a very hot topic right now: we are actively pushed to put AI-generated code into production, build software factories, and eliminate the human bottleneck from the process. And all of that would be great if, at the same time, the overall quality of software products wasn't going down and the number of incidents wasn't going up.
Dex explains this phenomenon by pointing to a fundamental limitation of current models: they can't maintain and improve the quality of a codebase over time. Models are trained to solve the issue and pass the tests. There is nothing there about good architecture, clean code, or future maintainability.
If the model knew what good code looks like it would write it in the first place.
What to do with that? Dex suggests to do the following:
๐ธ Make Product Review: align goals, check mockups, etc.
๐ธ Do System Architecture: establish clear component boundaries and API contracts.
๐ธ Perform Program Design: define types and method signatures, test approach, program layout, and dependency stack.
๐ธ Implement by Vertical Slices: build features vertically, reducing the scope of what an agent can change in a single iteration.
So, it's all about shifting review left to earlier stages and focusing more on architecture and program design (I described similar ideas about code review here). And then let agents make small changes that engineers can actually understand and review.
And yes, we still need to review the changes.
#engineering #codereview #ai
YouTube
Why Software Factories Fail
In July 2025 Dex Horthy turned the lights off: an agent software factory where nobody read the code. It fell apart. An issue appeared that no amount of prompting could fix, the site was down, users were furious, and he was digging through a codebase he hadโฆ
๐ฅ3๐2โค1
Tracer Bullets
Continuing the topic from the previous post, let's talk in more detail about implementing vertical slices.
The original concept comes from The Pragmatic Programmer, where it was called tracer bullets. A tracer bullet is a small, end-to-end slice of functionality that touches all layers of the system at once.
Matt Pocock in his article adapts this concept to AI engineering:
- Build a small feature end-to-end
- Test it immediately
- Get feedback
- Move to the next slice in a fresh context window
- Repeat
Matt explained this in the following way: AI's natural inclination is to build big layers in isolation. But we need to do the opposite: build the whole scenario end to end across all layers.
The article also provides a prompt sample:
So this approach forces the agent to think by small chunks and produce more predictable results.
As they say, everything new is just well-forgotten old. As we can see, AI often doesn't bring really new concepts, but rather brings existing ones back to life.
#engineering #ai
Continuing the topic from the previous post, let's talk in more detail about implementing vertical slices.
The original concept comes from The Pragmatic Programmer, where it was called tracer bullets. A tracer bullet is a small, end-to-end slice of functionality that touches all layers of the system at once.
Matt Pocock in his article adapts this concept to AI engineering:
- Build a small feature end-to-end
- Test it immediately
- Get feedback
- Move to the next slice in a fresh context window
- Repeat
Matt explained this in the following way: AI's natural inclination is to build big layers in isolation. But we need to do the opposite: build the whole scenario end to end across all layers.
The article also provides a prompt sample:
When building features, build a tiny, end-to-end slice of the feature first, seek feedback, then expand out from there. When building systems, you want to write code that gets you feedback as quickly as possible. Tracer bullets are small slices of functionality that go through all layers of the system, allowing you to test and validate your approach early. This helps in identifying potential issues and ensures that the overall architecture is sound before investing significant time in development.
So this approach forces the agent to think by small chunks and produce more predictable results.
As they say, everything new is just well-forgotten old. As we can see, AI often doesn't bring really new concepts, but rather brings existing ones back to life.
#engineering #ai
๐3โค2๐ฅ1