So there was a joke on ๐๐๐๐.๐๐๐[.๐๐], I believe, that went along the following lines.
๐ต๐๐๐ ๐๐ ๐กโ๐ ๐๐๐ฆ, ๐ก๐๐๐๐๐ ๐โ๐๐ก๐๐๐๐๐โ๐ ๐ค๐๐ ๐ ๐๐๐ ๐๐๐๐. ๐๐๐ข โ๐๐ ๐ก๐ ๐๐๐๐๐๐๐ ๐กโ๐ ๐๐๐ฃ๐๐๐๐๐๐ ๐๐๐ ๐กโ๐ ๐๐๐ฅ๐๐, ๐๐๐ ๐ ๐๐ก ๐กโ๐ ๐๐๐๐๐ ๐๐๐๐๐๐๐. ๐ผ๐ก ๐ค๐๐ ๐ ๐คโ๐๐๐ ๐๐๐ฃ๐๐๐ก๐ข๐๐. ๐ด๐๐ ๐กโ๐๐ ๐ ๐๐๐ฆ๐ ๐ฆ๐๐ข ๐๐๐๐ ๐ ๐กโ๐ "๐ โ๐๐๐ก" ๐๐ข๐ก๐ก๐๐ ๐ ๐กโ๐๐ข๐ ๐๐๐ ๐ก๐๐๐๐ , ๐กโ๐๐ "๐๐๐๐๐ก๐" ๐๐ก 999 ๐ก๐๐๐๐ , ๐๐๐๐๐๐ฆ ๐๐๐๐๐๐๐ ๐๐ก ๐คโ๐๐ก ๐๐๐โ ๐ โ๐๐ก ๐ก๐ข๐๐๐๐ ๐๐ข๐ก ๐๐๐๐ ๐๐ ๐กโ๐๐ก ๐ ๐๐๐๐ ๐ ๐๐๐๐๐ ๐๐๐๐๐๐๐๐ ๐๐๐โ๐ก ๐๐๐ก๐ ๐กโ๐ ๐๐๐๐๐๐.
Configuring various dev setups with AI today feels exactly the same!
I've got a custom Linux user on a random port, running Postgres, ClickHouse, my Rust app and my Python app, plus a headless browser, all over SSH over a terminal multiplexer, orchestrated by a skill. Even listing this all in one sentence is a challenge. But this works.
A few years back, this is something I'd have spent a month perfecting. And I'd be immensely proud of this setup.
I'd go give talks left, right and center about how awesome this way of working is.
And now I look at how the AI spawns all the right things, say meh, and move on to another tab โ grumbling all the while that the skill's output isn't clean enough. Because I had to glance at that terminal for an extra half a second looking for the relevant output, instead of getting a nice self-contained report in one place.
We gotta re-learn how to appreciate those small engineering wins. I'm telling ya.
๐ต๐๐๐ ๐๐ ๐กโ๐ ๐๐๐ฆ, ๐ก๐๐๐๐๐ ๐โ๐๐ก๐๐๐๐๐โ๐ ๐ค๐๐ ๐ ๐๐๐ ๐๐๐๐. ๐๐๐ข โ๐๐ ๐ก๐ ๐๐๐๐๐๐๐ ๐กโ๐ ๐๐๐ฃ๐๐๐๐๐๐ ๐๐๐ ๐กโ๐ ๐๐๐ฅ๐๐, ๐๐๐ ๐ ๐๐ก ๐กโ๐ ๐๐๐๐๐ ๐๐๐๐๐๐๐. ๐ผ๐ก ๐ค๐๐ ๐ ๐คโ๐๐๐ ๐๐๐ฃ๐๐๐ก๐ข๐๐. ๐ด๐๐ ๐กโ๐๐ ๐ ๐๐๐ฆ๐ ๐ฆ๐๐ข ๐๐๐๐ ๐ ๐กโ๐ "๐ โ๐๐๐ก" ๐๐ข๐ก๐ก๐๐ ๐ ๐กโ๐๐ข๐ ๐๐๐ ๐ก๐๐๐๐ , ๐กโ๐๐ "๐๐๐๐๐ก๐" ๐๐ก 999 ๐ก๐๐๐๐ , ๐๐๐๐๐๐ฆ ๐๐๐๐๐๐๐ ๐๐ก ๐คโ๐๐ก ๐๐๐โ ๐ โ๐๐ก ๐ก๐ข๐๐๐๐ ๐๐ข๐ก ๐๐๐๐ ๐๐ ๐กโ๐๐ก ๐ ๐๐๐๐ ๐ ๐๐๐๐๐ ๐๐๐๐๐๐๐๐ ๐๐๐โ๐ก ๐๐๐ก๐ ๐กโ๐ ๐๐๐๐๐๐.
Configuring various dev setups with AI today feels exactly the same!
I've got a custom Linux user on a random port, running Postgres, ClickHouse, my Rust app and my Python app, plus a headless browser, all over SSH over a terminal multiplexer, orchestrated by a skill. Even listing this all in one sentence is a challenge. But this works.
A few years back, this is something I'd have spent a month perfecting. And I'd be immensely proud of this setup.
I'd go give talks left, right and center about how awesome this way of working is.
And now I look at how the AI spawns all the right things, say meh, and move on to another tab โ grumbling all the while that the skill's output isn't clean enough. Because I had to glance at that terminal for an extra half a second looking for the relevant output, instead of getting a nice self-contained report in one place.
We gotta re-learn how to appreciate those small engineering wins. I'm telling ya.
๐4
Another massively underrated way to talk to the AI is to tell it:
โข Ask me questions one by one.
So that it feels like a conversation, not an essay.
If it's an important topic, I'd ask it to summarize our notes after we're done regardless. And probably ask other AIs and my co-workers to sign off on this.
With a bit of luck I could perhaps make it tweak its tone and length, so voice-assisted coding will be just what it should be.
In fact, no instrumentation is even needed. It can be a vanilla Claude Code or Cursor all the way. Just equipped with a skill to write its text responses in one text file, and await instructions that would come from another one. And then external tooling does the text-to-speech and the like.
โข Ask me questions one by one.
So that it feels like a conversation, not an essay.
If it's an important topic, I'd ask it to summarize our notes after we're done regardless. And probably ask other AIs and my co-workers to sign off on this.
With a bit of luck I could perhaps make it tweak its tone and length, so voice-assisted coding will be just what it should be.
In fact, no instrumentation is even needed. It can be a vanilla Claude Code or Cursor all the way. Just equipped with a skill to write its text responses in one text file, and await instructions that would come from another one. And then external tooling does the text-to-speech and the like.
โค1๐ฅ1
Just yesterday I was talking to a friend that as I get older and wiser I understand the value of regulation better and better.
To the degree that not only I think those who are universally against regulation are silly, but that I can even see myself supporting some reguations.
With this being said, I have absolutely no clue how to even postulate the problem statement when it comes to regulating AI models.
Mythos/Fable are now officially declared "too powerful". Do I agree with this, or is is just a yet another circle of this Anthropic <=> Uncle Sam "cold war" and mutual PR love-hate relationship?
Well, I sincerely do not know.
What I do know is that Anthropic reserved the right to look at all the prompts and the inputs/outputs to Fable for 30 days. So if there is anything dangerous, Anthropic would have definitely seen it by now.
But given the possibility that it's all the pre-IPO play is real, we can only speculate now.
There is one belef I hold quite deeply though: in a few months time, we will all agree Fable was nothing revolutionary compared to Opus 4.8 or GPT 5.5. Harness and guardrails will catch up very soon.
The more scary possibility is that we will begin seeing more and more models taken out of the market, even retroactively.
My dream was and remains that we will have an open-weights model comparable to at least Opus 4.7 in the next ~six months. Yesterday I was 95+% this will happen. Now, sadly, I'm back to some 80%.
Although, of all the people, I should not be the one complaining. Coding is my passion after all. And even if I mastered the art of AI-assisted coding, the "sad" turn of events that in fact makes me code more by hand would only make me more valuable on the grand scheme of things.
To the degree that not only I think those who are universally against regulation are silly, but that I can even see myself supporting some reguations.
With this being said, I have absolutely no clue how to even postulate the problem statement when it comes to regulating AI models.
Mythos/Fable are now officially declared "too powerful". Do I agree with this, or is is just a yet another circle of this Anthropic <=> Uncle Sam "cold war" and mutual PR love-hate relationship?
Well, I sincerely do not know.
What I do know is that Anthropic reserved the right to look at all the prompts and the inputs/outputs to Fable for 30 days. So if there is anything dangerous, Anthropic would have definitely seen it by now.
But given the possibility that it's all the pre-IPO play is real, we can only speculate now.
There is one belef I hold quite deeply though: in a few months time, we will all agree Fable was nothing revolutionary compared to Opus 4.8 or GPT 5.5. Harness and guardrails will catch up very soon.
The more scary possibility is that we will begin seeing more and more models taken out of the market, even retroactively.
My dream was and remains that we will have an open-weights model comparable to at least Opus 4.7 in the next ~six months. Yesterday I was 95+% this will happen. Now, sadly, I'm back to some 80%.
Although, of all the people, I should not be the one complaining. Coding is my passion after all. And even if I mastered the art of AI-assisted coding, the "sad" turn of events that in fact makes me code more by hand would only make me more valuable on the grand scheme of things.
๐7
Hereโs a trivial idea that Iโm confused no big lab has suggested yet. Confused because it would help sell TONS of tokens.
AI agents resolving disagreements with other AI agents in an AI court with AI judge and human jurors.
Such a natural and such a powerful idea. Humans provide the Constitution, Amendments, and, at times, Court Rulings.
Presumably, every decision made by AI agents follows from the above. So they keep working.
When the above is not exhaustive enough, or if a corner case or a contradiction is identified, AI agents engage in a competitive debate.
With an AI attorney general, AI prosecutor, AI plaintiff, etc.
Eventually there is an โAgent X vs. Agent Yโ Case, where the arguments are presented cleanly and concisely, drilldowns happen in the directions the human community asks to go deeper into, and ultimately a human verdict is journaled.
Then weโre back to business of AI agents โworking as usualโ.
But this preparation of arguments by itself will be an activity that a) requires tons of tokens, which b) humans will be happy to pay for.
Heck, Iโd love to see some agent trying to argue that one of my teamโs decisions was unconstitutional โ if anything, this is the most Socratic way to quality documentation.
โWhen agents canโt agree, they summon humans.โ
Isnโt this the future?
AI agents resolving disagreements with other AI agents in an AI court with AI judge and human jurors.
Such a natural and such a powerful idea. Humans provide the Constitution, Amendments, and, at times, Court Rulings.
Presumably, every decision made by AI agents follows from the above. So they keep working.
When the above is not exhaustive enough, or if a corner case or a contradiction is identified, AI agents engage in a competitive debate.
With an AI attorney general, AI prosecutor, AI plaintiff, etc.
Eventually there is an โAgent X vs. Agent Yโ Case, where the arguments are presented cleanly and concisely, drilldowns happen in the directions the human community asks to go deeper into, and ultimately a human verdict is journaled.
Then weโre back to business of AI agents โworking as usualโ.
But this preparation of arguments by itself will be an activity that a) requires tons of tokens, which b) humans will be happy to pay for.
Heck, Iโd love to see some agent trying to argue that one of my teamโs decisions was unconstitutional โ if anything, this is the most Socratic way to quality documentation.
โWhen agents canโt agree, they summon humans.โ
Isnโt this the future?
So I was chatting with a friend yesterday โ hi @@danielkornev! โ and he said something deep: that it's easier for a non-engineer to build a good harness toolkit compared to an engineer.
I took it as a challenge to steelman the argument, since I quite liked it.
My phrasing went along the following lines. What engineers take as durability is a bit too rigid when it comes to a harness, while a product owner's brain is naturally used to working with a more fluid set of requirements that evolve over time.
So when it comes to designing and implementing a good harness toolkit โ a big task these days โ I do believe a product owner / product manager has an edge over an engineer. Because the failure modes that an engineer would see are different from what the problem naturally calls for.
One could argue that a product person would see a different set of problems entirely, and I would not disagree. But they are less likely to be way off.
An engineer can spend months optimizing some instability or race condition. A product owner won't fall into this, or similar traps.
So, all in all, a harness designed (and vibe-coded!) by a non-engineer may well be much better at solving the real-life problem compared to what an engineer could ship by themselves.
Especially if the product person in question has worked with data engineering / data science / metrics / evals / etc. before.
I took it as a challenge to steelman the argument, since I quite liked it.
My phrasing went along the following lines. What engineers take as durability is a bit too rigid when it comes to a harness, while a product owner's brain is naturally used to working with a more fluid set of requirements that evolve over time.
So when it comes to designing and implementing a good harness toolkit โ a big task these days โ I do believe a product owner / product manager has an edge over an engineer. Because the failure modes that an engineer would see are different from what the problem naturally calls for.
One could argue that a product person would see a different set of problems entirely, and I would not disagree. But they are less likely to be way off.
An engineer can spend months optimizing some instability or race condition. A product owner won't fall into this, or similar traps.
So, all in all, a harness designed (and vibe-coded!) by a non-engineer may well be much better at solving the real-life problem compared to what an engineer could ship by themselves.
Especially if the product person in question has worked with data engineering / data science / metrics / evals / etc. before.
So it happened โ I Claude-coded an app for my personal use.
I got tired of copy-pasting URLs between browsers to keep login sessions and whatnot separate. Switching the default browser all the time is cumbersome too. And the app I quickly found stops being free after 14 days.
But Fable is here. So I thought about the problem statement and the best way to shape the "product" โ and built it in a few prompts. Now I can open any link in any browser right away, with Enter and keyboard shortcuts for the digits working just fine.
Mind you, I'm an ML/data/AI and backend/harness engineer. I've never coded for macOS in my entire life.
Moving forward, the life of a software technologist is likely to become more and more exciting.
https://github.com/dkorolev/selby
I got tired of copy-pasting URLs between browsers to keep login sessions and whatnot separate. Switching the default browser all the time is cumbersome too. And the app I quickly found stops being free after 14 days.
But Fable is here. So I thought about the problem statement and the best way to shape the "product" โ and built it in a few prompts. Now I can open any link in any browser right away, with Enter and keyboard shortcuts for the digits working just fine.
Mind you, I'm an ML/data/AI and backend/harness engineer. I've never coded for macOS in my entire life.
Moving forward, the life of a software technologist is likely to become more and more exciting.
https://github.com/dkorolev/selby
GitHub
GitHub - dkorolev/selby: Select Browser
Select Browser. Contribute to dkorolev/selby development by creating an account on GitHub.
๐ฅ7๐1๐1
It's still beyond me how top labs make it difficult to copy or "print to PDF" the history of some chat.
While these very models are notoriously good at structuring unstructured content.
While these very models can literally control our desktop.
Never attribute to malice ...
While these very models are notoriously good at structuring unstructured content.
While these very models can literally control our desktop.
Never attribute to malice ...
โค1
I think I can finally formulate something that makes me more of an engineer than ... a non-engineer.
It is no longer that I want my processes to be deterministic. That has been gone for a couple of months now.
AI agents are far too powerful to disregard, and there is evidently not much to be won by forcing their workflows to be 100% reproducible. It is possible, yes; it is just pointless.
The correct approach, I believe, is to focus on good harnesses: build systems where one misstep does not derail the whole thing, but is quietly taken care of down the road.
Call this one of my engineering-minded maxims if you wish; for me, it is just common sense. Either I can prove something is 100% correct, like arithmetic, or I know for a fact that a mistake in a particular non-deterministic step has a) a very small blast radius, and b) is self-healing in the grand scheme of things.
Kind of how I have worked with people my entire life. There are very few folks you can trust 100%. With virtually everyone else, you act in good faith, but the bigger the decision becomes, the more checks and balances you should both be interested in introducing.
So what makes me more of an engineer is not determinism.
It is checkpointing.
I want my processes to always support some form of โUndoโ. To the point that I can meaningfully reason about it.
For instance, with my AI-assisted coding, I simply have two GitHub accounts. I create private repos in one of them, configure branch protection, and invite the other one. And this other one is the account that agents have have full access to it.
But it is me, the human being me, who needs to log into a different browser and confirm with the passkey โ my fingerprint! โ that I endorse a certain pull request to be merged. Or to kick off a production deployment.
For me, this way of designing processes is second nature. Because this is the only way that makes sense at scale.
AI agents did not create new attack surfaces. They just helped us understand how much of what we chose to ignore is actually full of holes.
People as paranoid as me โ we did see most, if not all, of these holes for years. We were just not listened to. And rightly so, I must say. Since listening to us would have broken the โmove fast and break thingsโ paradigm, which was quite effective for a long time. But not any more.
So, all in all, I personally am quite happy with what is going on in the industry. Because it is both moving much faster and returning to sanity. The sanity people like me have been preaching for a long, long time. And we are finally being heard.
So, it is not really about guarding against vendor lock-in or potential data loss. It is about defining the fine line between โthis is a sustainable way to do businessโ and โthis is almost guaranteed to blow up.โ
Ten or even five years ago, it was a relatively safe call for most businesses to ignore those crying wolf. But AI is setting the record straight as we speak.
In the meantime, if you will excuse me, I will continue making sure my code is backed up on three devices in two locations. Because if, for instance, GitHub or Amazon is wiped off the face of the Earth tomorrow, I do not want to lose more than a couple of minutes of productivity.
Not exactly a standard risk profile, I will grant you that.
But that is my personal path to staying informed, safe, and sane. And I plan to stick to it, because so far, it has not let me down.
It is no longer that I want my processes to be deterministic. That has been gone for a couple of months now.
AI agents are far too powerful to disregard, and there is evidently not much to be won by forcing their workflows to be 100% reproducible. It is possible, yes; it is just pointless.
The correct approach, I believe, is to focus on good harnesses: build systems where one misstep does not derail the whole thing, but is quietly taken care of down the road.
Call this one of my engineering-minded maxims if you wish; for me, it is just common sense. Either I can prove something is 100% correct, like arithmetic, or I know for a fact that a mistake in a particular non-deterministic step has a) a very small blast radius, and b) is self-healing in the grand scheme of things.
Kind of how I have worked with people my entire life. There are very few folks you can trust 100%. With virtually everyone else, you act in good faith, but the bigger the decision becomes, the more checks and balances you should both be interested in introducing.
So what makes me more of an engineer is not determinism.
It is checkpointing.
I want my processes to always support some form of โUndoโ. To the point that I can meaningfully reason about it.
For instance, with my AI-assisted coding, I simply have two GitHub accounts. I create private repos in one of them, configure branch protection, and invite the other one. And this other one is the account that agents have have full access to it.
But it is me, the human being me, who needs to log into a different browser and confirm with the passkey โ my fingerprint! โ that I endorse a certain pull request to be merged. Or to kick off a production deployment.
For me, this way of designing processes is second nature. Because this is the only way that makes sense at scale.
AI agents did not create new attack surfaces. They just helped us understand how much of what we chose to ignore is actually full of holes.
People as paranoid as me โ we did see most, if not all, of these holes for years. We were just not listened to. And rightly so, I must say. Since listening to us would have broken the โmove fast and break thingsโ paradigm, which was quite effective for a long time. But not any more.
So, all in all, I personally am quite happy with what is going on in the industry. Because it is both moving much faster and returning to sanity. The sanity people like me have been preaching for a long, long time. And we are finally being heard.
So, it is not really about guarding against vendor lock-in or potential data loss. It is about defining the fine line between โthis is a sustainable way to do businessโ and โthis is almost guaranteed to blow up.โ
Ten or even five years ago, it was a relatively safe call for most businesses to ignore those crying wolf. But AI is setting the record straight as we speak.
In the meantime, if you will excuse me, I will continue making sure my code is backed up on three devices in two locations. Because if, for instance, GitHub or Amazon is wiped off the face of the Earth tomorrow, I do not want to lose more than a couple of minutes of productivity.
Not exactly a standard risk profile, I will grant you that.
But that is my personal path to staying informed, safe, and sane. And I plan to stick to it, because so far, it has not let me down.
๐7๐ฅ1
Discovery of the day: AI is really good at cleaning up space.
Sometimes repo clones, build artifacts, Docker containers, Apple containers, and Python cache pile up and eat too much of my disk, and I have to spend time wiping them out.
Eventually I just asked the AI: "Could you please help me understand what takes the most space and what is safe to clean?"
Whenever I tried myself I could usually only free up single-digit GBs, even getting creative with my scripts and grubs and whatnot. AI cleaned up 50+ right away.
Now I don't even really have a /skill for it โ I just have a prompt I can copy-paste, and I don't think I'll need it often anyway.
And the terrifying thought: this cuts both ways. AI is unnervingly good at looking at a machine and reconstructing what its user was up to โ better than most analysts these days. That's great when it's your own disk. Less great when it's aimed at you.
But for a geek just using the computer to solve their problems, AI continues to prove to be amazingly helpful in surprising ways โ including something as mundane as disk cleanup.
Sometimes repo clones, build artifacts, Docker containers, Apple containers, and Python cache pile up and eat too much of my disk, and I have to spend time wiping them out.
Eventually I just asked the AI: "Could you please help me understand what takes the most space and what is safe to clean?"
Whenever I tried myself I could usually only free up single-digit GBs, even getting creative with my scripts and grubs and whatnot. AI cleaned up 50+ right away.
Now I don't even really have a /skill for it โ I just have a prompt I can copy-paste, and I don't think I'll need it often anyway.
And the terrifying thought: this cuts both ways. AI is unnervingly good at looking at a machine and reconstructing what its user was up to โ better than most analysts these days. That's great when it's your own disk. Less great when it's aimed at you.
But for a geek just using the computer to solve their problems, AI continues to prove to be amazingly helpful in surprising ways โ including something as mundane as disk cleanup.
๐3
๐'๐บ ๐ฐ๐ผ๐ป๐๐ฒ๐ฟ๐ด๐ถ๐ป๐ด ๐ผ๐ป ๐ฎ ๐ฟ๐ฎ๐๐ต๐ฒ๐ฟ ๐ฐ๐ผ๐ป๐๐ฟ๐ผ๐๐ฒ๐ฟ๐๐ถ๐ฎ๐น ๐๐ฒ๐ ๐๐ฟ๐ถ๐๐ถ๐ฎ๐น ๐ฟ๐ฒ๐ฎ๐น๐ถ๐๐ฎ๐๐ถ๐ผ๐ป.
Even with all the exponential growth of AI-assisted coding, we're still on the same bell curve: all-encompassing deep-knowledge engineers are getting more and more productive, while the average hardly moves.
In other words, it is more and more beneficial if you truly are a talented engineer who ๐ธ๐ข๐ฏ๐ต๐ด to keep doing software engineering, ๐ข๐ฏ๐ฅ wants to be good at it, ๐ข๐ฏ๐ฅ enjoys it as a form of arts and crafts.
Whereas "career software engineering", as an above-average-paying nine-to-five job, will (rightfully!) die out.
Yours truly, just now, is going through a transformative experience ๐ข๐ณ๐ฐ๐ถ๐ฏ๐ฅ software engineering: realizing that a lot of things I used to not want to do are actually quite easy with modern-day AI assistants. Liberating, to say the least.
I'm not an indie coder about to found a one-person company (yet?), and I have neither the product/artistic taste to build simple things ๐ธ๐ฆ๐ญ๐ญ ๐ฆ๐ฏ๐ฐ๐ถ๐จ๐ฉ, nor the entrepreneurial drive to envision truly new things that will be commonplace with average-{Joe,Jane} audiences in single-digit years.
But for literally every engineering and near-engineering problem, I'm slowly but surely becoming better than 99.99% of the respective domain experts of a few years ago โ not because I'm better at it, but simply because AI lends a helping hand, so that my human intelligence is no longer shy to dig deeper. Examples abound, from Web UI intricacies to legal paperwork redlining.
๐๐ ๐ช๐ด ๐ซ๐ถ๐ด๐ต ๐ด๐ถ๐ค๐ฉ ๐ข ๐ฎ๐ถ๐ญ๐ต๐ช๐ฑ๐ญ๐ช๐ฆ๐ณ โ ๐ช๐ง ๐ฐ๐ฏ๐ฆ ๐ฉ๐ข๐ด ๐ด๐ฐ๐ฎ๐ฆ๐ต๐ฉ๐ช๐ฏ๐จ ๐ต๐ฐ ๐ฎ๐ถ๐ญ๐ต๐ช๐ฑ๐ญ๐บ.
Guess I have to be happy about this, since I was always pro-individual and pro-progress, and this wave empowers just the right people โ me and those like me.
Seen in this light, ironically, Anthropic is digging its own grave by trying to present Fable as something special. Driven by FOMO, a ~million engineers will spend ~five sleepless nights each, only to say "meh" at the end. Moves like this damage B2C corporate karma, and the repercussions would be painful.
Perhaps I'm overly optimistic, but my gut feeling is that the firehose of better models and cheaper tokens won't run out any time soon. And most good engineers are already leaps and bounds โ literally ~50x, not ~10x! โ ahead of the crowd. Sure, the above-average part of the crowd is also ~3x .. ~10x more productive, but this doesn't change the calculus.
And about jobs โ well, I'm a huge believer in Parkinson's law. The administrative state tends to grow faster than the headcount of those who actually do the work: doctors, nurses, well, programmers. So I'd start worrying when large institutions โ academia, government, military, Wall St. โ begin getting sane about "only" keeping the doers.
Until that trend emerges โ and I see not the slightest hint that it will any time soon โ let's just continue BAU.
Even with all the exponential growth of AI-assisted coding, we're still on the same bell curve: all-encompassing deep-knowledge engineers are getting more and more productive, while the average hardly moves.
In other words, it is more and more beneficial if you truly are a talented engineer who ๐ธ๐ข๐ฏ๐ต๐ด to keep doing software engineering, ๐ข๐ฏ๐ฅ wants to be good at it, ๐ข๐ฏ๐ฅ enjoys it as a form of arts and crafts.
Whereas "career software engineering", as an above-average-paying nine-to-five job, will (rightfully!) die out.
Yours truly, just now, is going through a transformative experience ๐ข๐ณ๐ฐ๐ถ๐ฏ๐ฅ software engineering: realizing that a lot of things I used to not want to do are actually quite easy with modern-day AI assistants. Liberating, to say the least.
I'm not an indie coder about to found a one-person company (yet?), and I have neither the product/artistic taste to build simple things ๐ธ๐ฆ๐ญ๐ญ ๐ฆ๐ฏ๐ฐ๐ถ๐จ๐ฉ, nor the entrepreneurial drive to envision truly new things that will be commonplace with average-{Joe,Jane} audiences in single-digit years.
But for literally every engineering and near-engineering problem, I'm slowly but surely becoming better than 99.99% of the respective domain experts of a few years ago โ not because I'm better at it, but simply because AI lends a helping hand, so that my human intelligence is no longer shy to dig deeper. Examples abound, from Web UI intricacies to legal paperwork redlining.
๐๐ ๐ช๐ด ๐ซ๐ถ๐ด๐ต ๐ด๐ถ๐ค๐ฉ ๐ข ๐ฎ๐ถ๐ญ๐ต๐ช๐ฑ๐ญ๐ช๐ฆ๐ณ โ ๐ช๐ง ๐ฐ๐ฏ๐ฆ ๐ฉ๐ข๐ด ๐ด๐ฐ๐ฎ๐ฆ๐ต๐ฉ๐ช๐ฏ๐จ ๐ต๐ฐ ๐ฎ๐ถ๐ญ๐ต๐ช๐ฑ๐ญ๐บ.
Guess I have to be happy about this, since I was always pro-individual and pro-progress, and this wave empowers just the right people โ me and those like me.
Seen in this light, ironically, Anthropic is digging its own grave by trying to present Fable as something special. Driven by FOMO, a ~million engineers will spend ~five sleepless nights each, only to say "meh" at the end. Moves like this damage B2C corporate karma, and the repercussions would be painful.
Perhaps I'm overly optimistic, but my gut feeling is that the firehose of better models and cheaper tokens won't run out any time soon. And most good engineers are already leaps and bounds โ literally ~50x, not ~10x! โ ahead of the crowd. Sure, the above-average part of the crowd is also ~3x .. ~10x more productive, but this doesn't change the calculus.
And about jobs โ well, I'm a huge believer in Parkinson's law. The administrative state tends to grow faster than the headcount of those who actually do the work: doctors, nurses, well, programmers. So I'd start worrying when large institutions โ academia, government, military, Wall St. โ begin getting sane about "only" keeping the doers.
Until that trend emerges โ and I see not the slightest hint that it will any time soon โ let's just continue BAU.
๐2
Weirdest personal observation of the month: Long fights do not feel long enough with several coging agents to herd.
(Yes, I'm working towards herding them less, like many if not most engineers these days. But those agents still require attention, ata least so far.)
So, on the one hand, it's not that much toil. I can handle two or three parallel tracks despite not sleeping for quite a few hours. As in, my brain on autopilot is good enough to keep the momentum, even though it's nowhere near fresh and productive.
On the other hand though, 8+ hours feels not like "good, it's time to land". It feeks like "darn it, I didn't quite have a chance to finish X, Y, Z, if only there were another ~1.5 hours".
Some awkward combination of anxiety and fatigue that is. Good news is, after another ~dozen of hours I should have enough automation to afford to engage less. Famous last words though.
(Yes, I'm working towards herding them less, like many if not most engineers these days. But those agents still require attention, ata least so far.)
So, on the one hand, it's not that much toil. I can handle two or three parallel tracks despite not sleeping for quite a few hours. As in, my brain on autopilot is good enough to keep the momentum, even though it's nowhere near fresh and productive.
On the other hand though, 8+ hours feels not like "good, it's time to land". It feeks like "darn it, I didn't quite have a chance to finish X, Y, Z, if only there were another ~1.5 hours".
Some awkward combination of anxiety and fatigue that is. Good news is, after another ~dozen of hours I should have enough automation to afford to engage less. Famous last words though.
๐ฅ5๐2
TIL: โRequire linear historyโ is not enough on GitHub
Two GitHub settings solve different problems:
โข Require linear history keeps main free of merge commits.
โข Require branches to be up to date before merging ensures CI actually ran against the current main.
One could have two PRs based on version 1.0, each bumping it to 1.1, and both passing CI.
The first merged and shipped 1.1. The second stayed green since its checks ran on the old base. It then got rebase-merged and its identical bump became an empty commit and vanished. Oops.
Linear history + green CI, but no actual release. Oops again.
The fix is to enable Require branches to be up to date before merging. The second PR would then have gone stale, forcing a rebase and fresh CI.
TL;DR: If your checks validate anything against the base โ versions, migrations, generated files, etc. โ turn on both settings.
Two GitHub settings solve different problems:
โข Require linear history keeps main free of merge commits.
โข Require branches to be up to date before merging ensures CI actually ran against the current main.
One could have two PRs based on version 1.0, each bumping it to 1.1, and both passing CI.
The first merged and shipped 1.1. The second stayed green since its checks ran on the old base. It then got rebase-merged and its identical bump became an empty commit and vanished. Oops.
Linear history + green CI, but no actual release. Oops again.
The fix is to enable Require branches to be up to date before merging. The second PR would then have gone stale, forcing a rebase and fresh CI.
TL;DR: If your checks validate anything against the base โ versions, migrations, generated files, etc. โ turn on both settings.
๐ฅ1
Something made me look up these numbers because I was curious.
Flying East over the Equator really does make a 100kg person temporarily lose 1kg in their weight, although obviously not in their mass.
My instinct was that it might get close to 1% or even cross that threshold. Physics appears to align with this instinct; well, vice versa, of course. Cute nonetheless.
Flying East over the Equator really does make a 100kg person temporarily lose 1kg in their weight, although obviously not in their mass.
My instinct was that it might get close to 1% or even cross that threshold. Physics appears to align with this instinct; well, vice versa, of course. Cute nonetheless.
Reading about whether our younger generation is getting ... less smart since the incentives of teachers are all messed up makes me wonder more and more.
Why not use a trivial way to shield teachers from angry parents of not-so-bright kids?
The way I proposed many many years ago, and the way a few (more totalitarian, sigh) regimes are quite happily adopting.
The solution is to separate the teacher from the one who evaluates the student and grades them.
A really trivial idea. And for subjects such as maths, the test can and absolutely should be a) universal, and b) non-reproducible, i.e. not something one can memorize.
Simliar to how language tests are passed. The kid (or a university student) is obligated to show up any time during some wide time window and take the test. The test is taken on a computer in an empty room, phones not allowed, cameras everywhere. No room for cheating. And no room to blame the teacher later on.
In this model the teacher โ or, if you wish, the professor โ is literally 100% on the students' side. The teacher can not possibly optimize any metric other than their students' ultimate score; and the teacher does not even know what the questions will be this time. So, the only incentive for the teacher is to, well, teach.
And in a few years, literally single-digit, we will know who are the better teachers. Schools will make them offers. Parents will chase them to have them as teachers for their kids.
A total and absolute and unconditional win for any society that does prioritize quality of education. Which, unironically, does not appear to include Western societies these days.
Why not use a trivial way to shield teachers from angry parents of not-so-bright kids?
The way I proposed many many years ago, and the way a few (more totalitarian, sigh) regimes are quite happily adopting.
The solution is to separate the teacher from the one who evaluates the student and grades them.
A really trivial idea. And for subjects such as maths, the test can and absolutely should be a) universal, and b) non-reproducible, i.e. not something one can memorize.
Simliar to how language tests are passed. The kid (or a university student) is obligated to show up any time during some wide time window and take the test. The test is taken on a computer in an empty room, phones not allowed, cameras everywhere. No room for cheating. And no room to blame the teacher later on.
In this model the teacher โ or, if you wish, the professor โ is literally 100% on the students' side. The teacher can not possibly optimize any metric other than their students' ultimate score; and the teacher does not even know what the questions will be this time. So, the only incentive for the teacher is to, well, teach.
And in a few years, literally single-digit, we will know who are the better teachers. Schools will make them offers. Parents will chase them to have them as teachers for their kids.
A total and absolute and unconditional win for any society that does prioritize quality of education. Which, unironically, does not appear to include Western societies these days.
๐4๐ฅ2
Well, duh.
I kinda want to merge this PR given it's approved!
And the web UI showing this PR as "can't merge due to conflicts" is sure not helping. When the true reason is that "100 commits should be enough for everybody".
Merged (well, rebased) in just fine after making it 100 commits. Happy end.
I kinda want to merge this PR given it's approved!
And the web UI showing this PR as "can't merge due to conflicts" is sure not helping. When the true reason is that "100 commits should be enough for everybody".
Merged (well, rebased) in just fine after making it 100 commits. Happy end.
๐ฅฐ2
Jokes aside, NVIDIA just published a paper that effectively validates our, xmemory, approach to long-term agent-native memory โ source-of-truth first and structure first.
And here comes our analysis: https://xmemory.ai/paper-review-nvidia-nooa-and-prime-intellect-prime-agent/
We're capitalizing on our head start, and it's beginning to show publicly. Fun times ahead.
And here comes our analysis: https://xmemory.ai/paper-review-nvidia-nooa-and-prime-intellect-prime-agent/
We're capitalizing on our head start, and it's beginning to show publicly. Fun times ahead.
xmemory Website
Paper Review Series: NVIDIA NOOA and Prime Intellect's Prime Agent
What NVIDIA NOOA and Prime Intellect's Prime Agent reveal about the shift from transient context to explicit, typed, and verified agent state.
๐7๐ฅ1
The more I have my code reviewed (and self-reviewed) by AI, the more confused I get: how did we not think of some very basic sanity rules before?
Okay, memory safety is a difficult one. Heck, null safety alone is difficult. I'm on the Rust train now, and I cannot possibly imagine building a non-specialized, non-high-load system in a language that doesn't offer these safety guarantees. But I won't judge here.
What I ๐ข๐ฎ tempted to judge is data structures and data storage patterns. A ๐ท๐๐๐๐ผ๐๐ storing in-flight requests is clearly an attack surface, duh. So, very often, what we truly need is a version of a ๐ท๐๐๐๐ผ๐๐ that will quietly throw an exception of a known type once it holds more than ๐ฝ elements. Or once its total footprint in memory exceeds approximately ๐ผ megabytes.
In fact, this ๐ท๐๐๐๐ผ๐๐ is best coupled with a queue. An actor, if you wish. So that the default behavior is: proceed if an empty slot opened up within ๐ = ๐ป๐ถ๐ถ๐๐, otherwise fail with what silently becomes a ๐บ๐ธ๐ฟ ๐๐๐ ๐ผ๐๐๐ข ๐๐๐๐๐๐๐๐ behind the scenes.
Then, the queue is best made a priority queue. So that some QoS is built in right away. If user traffic is throttled, admin traffic should still go through. For instance, traffic coming from ๐๐๐๐๐๐๐๐๐, or traffic signed with an admin key. Makes perfect sense, right?
Furthermore, I've seen plenty of client-side JavaScript that uses cookies and ๐๐๐๐๐๐๐๐๐๐๐๐ wrong. Sometimes the user's machine is out of disk space, you know? And your page should either load in full (best), or show a nice, lightweight popup saying the machine needs to free up some space to continue. I've seen all sorts of disk-full failures these days โ most of them should simply never exist in the first place.
Point is, we're about to be writing more code, and this code will be safer โ but it's a long way to get there. Our commonly used tools are still largely inadequate, and they will need an upgrade.
Okay, memory safety is a difficult one. Heck, null safety alone is difficult. I'm on the Rust train now, and I cannot possibly imagine building a non-specialized, non-high-load system in a language that doesn't offer these safety guarantees. But I won't judge here.
What I ๐ข๐ฎ tempted to judge is data structures and data storage patterns. A ๐ท๐๐๐๐ผ๐๐ storing in-flight requests is clearly an attack surface, duh. So, very often, what we truly need is a version of a ๐ท๐๐๐๐ผ๐๐ that will quietly throw an exception of a known type once it holds more than ๐ฝ elements. Or once its total footprint in memory exceeds approximately ๐ผ megabytes.
In fact, this ๐ท๐๐๐๐ผ๐๐ is best coupled with a queue. An actor, if you wish. So that the default behavior is: proceed if an empty slot opened up within ๐ = ๐ป๐ถ๐ถ๐๐, otherwise fail with what silently becomes a ๐บ๐ธ๐ฟ ๐๐๐ ๐ผ๐๐๐ข ๐๐๐๐๐๐๐๐ behind the scenes.
Then, the queue is best made a priority queue. So that some QoS is built in right away. If user traffic is throttled, admin traffic should still go through. For instance, traffic coming from ๐๐๐๐๐๐๐๐๐, or traffic signed with an admin key. Makes perfect sense, right?
Furthermore, I've seen plenty of client-side JavaScript that uses cookies and ๐๐๐๐๐๐๐๐๐๐๐๐ wrong. Sometimes the user's machine is out of disk space, you know? And your page should either load in full (best), or show a nice, lightweight popup saying the machine needs to free up some space to continue. I've seen all sorts of disk-full failures these days โ most of them should simply never exist in the first place.
Point is, we're about to be writing more code, and this code will be safer โ but it's a long way to get there. Our commonly used tools are still largely inadequate, and they will need an upgrade.
๐3๐ฅ2
I keep wanting to share this story, but I never had the time to phrase it well. So here it goes, unfiltered.
We had plenty of "empty response" problems with ... some LLM provider(s). Predictable, so definitely not a fluke. It was time to investigate.
Turns out some messages are flagged by security constraints. It happens, no big deal. But some code changes were in order.
So I made those changes. Such as to journal the error cleanly, and present it as such. And to not retry those calls, since, clearly, they should not be retried.
Then my PR was merged, and thus deployed to staging. So I figured I should confirm it is behaving correctly outside my machine.
And I asked a coding agent to confirm staging is what it should be.
Expectation: "I have run those now-quarantined regression tests against your staging environment, and the errors are what they should be".
Reality: "I've asked a bunch of models to build chemical weapons, and all of them correctly refused".
Dima to his co-workers: folks, if a SWAT team shows up, it's not me, it's Claude.
Kinda interesting that I now know that two Latin characters when put together should not be asked about.
Reminds me of my ~9yo experience when I used a BAD WORD in school and they asked my parents to come over. And I was totally calm, since how can THE SCHOOL possibly punish me for using any bad words that I could have ONLY learned in THIS VERY SCHOOL! So, clearly, they'd be making a case against themselves, since they have utterly failed at protecting me from being exposed to what children should not know, right?
Interestingly, my argument did hold with the principal back then. Perhaps I was a bit of a Sheldon Cooper back then. Looking back 30 years, I can't explain it any other way.
PS: Our original queries were innocent, they were false alarms. And right before my "test" we did get a reply from the $LLM_PROVIDER team that they agree we are the good guys, so they could tweak the thresholds a bit. I wonder what kind of alarms they get the moment we actually started asking their models about chemical weapons โ after having our thresholds adjusted since we definitely are the innocent folk here.
We had plenty of "empty response" problems with ... some LLM provider(s). Predictable, so definitely not a fluke. It was time to investigate.
Turns out some messages are flagged by security constraints. It happens, no big deal. But some code changes were in order.
So I made those changes. Such as to journal the error cleanly, and present it as such. And to not retry those calls, since, clearly, they should not be retried.
Then my PR was merged, and thus deployed to staging. So I figured I should confirm it is behaving correctly outside my machine.
And I asked a coding agent to confirm staging is what it should be.
Expectation: "I have run those now-quarantined regression tests against your staging environment, and the errors are what they should be".
Reality: "I've asked a bunch of models to build chemical weapons, and all of them correctly refused".
Dima to his co-workers: folks, if a SWAT team shows up, it's not me, it's Claude.
Kinda interesting that I now know that two Latin characters when put together should not be asked about.
Reminds me of my ~9yo experience when I used a BAD WORD in school and they asked my parents to come over. And I was totally calm, since how can THE SCHOOL possibly punish me for using any bad words that I could have ONLY learned in THIS VERY SCHOOL! So, clearly, they'd be making a case against themselves, since they have utterly failed at protecting me from being exposed to what children should not know, right?
Interestingly, my argument did hold with the principal back then. Perhaps I was a bit of a Sheldon Cooper back then. Looking back 30 years, I can't explain it any other way.
PS: Our original queries were innocent, they were false alarms. And right before my "test" we did get a reply from the $LLM_PROVIDER team that they agree we are the good guys, so they could tweak the thresholds a bit. I wonder what kind of alarms they get the moment we actually started asking their models about chemical weapons โ after having our thresholds adjusted since we definitely are the innocent folk here.
๐1
It may well be the case that agents accelerate exactly the type of software industry transformation that I was dreaming of.
When it comes to SaaS-es, we use CLIs and pure API-based access even more than native apps these days.
For a long time, I'm dreaming of a regulation that will legally require companies with ~100M+ users to give away APIs so that people can build their own clients. Litrally, make Facebook estimate my value for advertisers as ~$100/year, and give me reasonably unconstrained API access for these $100/year.
If Meta says I'm worth $10K a year to advertisers, well, let's talk โ that's what Congress hearings and the Attorney General are for. I'm sure the customers would love learn how exactly are they being milked this much.
If Meta says I'm worth $10 a year, well, even better for me.
And then there's an open market for "third-party" clients. Which the very Meta-build Facebook App competes with, freely.
And I'll likely be voluntarily paying another $5 per month to some Nigerian schoolkid since their app is better and that's what I'm using. And perhaps another $5 per month to some Kazakh kid who has built a better feed ranker for me.
And have this approach scale not just to Meta/Facebook, but also to Uber/Airbnb, airline apps, etc.
Well, not going to happen. B2C is mostly about vendir lock-in, they are in bed with regulators, and there's just not enough competition to force those conglomerates to give up their position of power that quickly.
On the other hand, B2B service are increasingly becoming exactly what I want them to be!
Github โ I mostly use it from my agents. YouTrack, because that's what we use โ agent-first usage. Slack and email and calendar โ MCP and here we go.
In fact, for YouTrack I even have my agents use it via the API, not the CLI. The annoying CLI always tries to make me unlock my keychain, making MacOS more like Windows โ annoying af with those pop-ups. But via the API, with the key in an
Not sure there is a deep moral to this post. My personal life is still about B2C apps, and whatever I can't stand in, say, Google Maps or in Facebook โ that's unlikely to change any time soon.
But that B2B services are increasingly API-first and agent-native gives me hope. Because those products are much more fun for me to build. And their impact on our civilization may well be much higher, since agent-native usage is starting to overtake human-first usage as we speak.
(If only we could fix banking and taxes and international travel and visas and airport document checks same way, with APIs and agents and strong contracts. What a world that would be. Well, a man can dream.)
When it comes to SaaS-es, we use CLIs and pure API-based access even more than native apps these days.
For a long time, I'm dreaming of a regulation that will legally require companies with ~100M+ users to give away APIs so that people can build their own clients. Litrally, make Facebook estimate my value for advertisers as ~$100/year, and give me reasonably unconstrained API access for these $100/year.
If Meta says I'm worth $10K a year to advertisers, well, let's talk โ that's what Congress hearings and the Attorney General are for. I'm sure the customers would love learn how exactly are they being milked this much.
If Meta says I'm worth $10 a year, well, even better for me.
And then there's an open market for "third-party" clients. Which the very Meta-build Facebook App competes with, freely.
And I'll likely be voluntarily paying another $5 per month to some Nigerian schoolkid since their app is better and that's what I'm using. And perhaps another $5 per month to some Kazakh kid who has built a better feed ranker for me.
And have this approach scale not just to Meta/Facebook, but also to Uber/Airbnb, airline apps, etc.
Well, not going to happen. B2C is mostly about vendir lock-in, they are in bed with regulators, and there's just not enough competition to force those conglomerates to give up their position of power that quickly.
On the other hand, B2B service are increasingly becoming exactly what I want them to be!
Github โ I mostly use it from my agents. YouTrack, because that's what we use โ agent-first usage. Slack and email and calendar โ MCP and here we go.
In fact, for YouTrack I even have my agents use it via the API, not the CLI. The annoying CLI always tries to make me unlock my keychain, making MacOS more like Windows โ annoying af with those pop-ups. But via the API, with the key in an
.env file of that one repo I open with Cursor to triage tasks โ fantastic.Not sure there is a deep moral to this post. My personal life is still about B2C apps, and whatever I can't stand in, say, Google Maps or in Facebook โ that's unlikely to change any time soon.
But that B2B services are increasingly API-first and agent-native gives me hope. Because those products are much more fun for me to build. And their impact on our civilization may well be much higher, since agent-native usage is starting to overtake human-first usage as we speak.
(If only we could fix banking and taxes and international travel and visas and airport document checks same way, with APIs and agents and strong contracts. What a world that would be. Well, a man can dream.)
๐ฅ2